Swimming pool robot positioning and trajectory prediction method and system fusing sonar and vision
By combining terahertz sonar with visual sensors, a trajectory model and environmental model of the swimming pool robot were constructed, solving the problems of inaccurate ranging and high maintenance difficulty of the swimming pool robot, and realizing high-precision positioning and intelligent management.
Patent Information
- Application Number
- CN202510700015.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-27
- Publication Date
- 2025-12-16
- Estimated Expiration
- 2045-05-27
AI Technical Summary
Existing pool robots have low ranging accuracy and are difficult to maintain and manage. They are also susceptible to interference from factors such as water quality, water temperature, and impurities in the water, leading to inaccurate positioning and frequent robot malfunctions.
By combining terahertz sonar equipment with visual sensors, an activity trajectory model is constructed using terahertz sonar data. Multi-source image data is acquired by combining multispectral and event cameras. Data processing and model fusion are performed using deep learning and quantum machine learning algorithms to achieve real-time positioning and trajectory prediction. Real-time simulation and visualization are then achieved through a digital twin model.
It improves the positioning accuracy and trajectory prediction accuracy of pool robots in complex environments, reduces maintenance difficulty, and enhances autonomous operation capabilities and intelligent management levels.
Smart Images

Figure CN120489100B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of data prediction, in particular to a pool robot positioning and trajectory prediction method and system fusing sonar and vision. BACKGROUND
[0002] At present, a pool robot is an intelligent device specially designed for a swimming pool environment and capable of autonomously completing a series of tasks.
[0003] In the related art, although an ultrasonic sensor is a commonly used ranging sensor for a pool robot, its accuracy is easily affected by factors such as water quality, water temperature, and impurities in the water. For example, in the case of high water temperature or turbid water quality, the propagation speed of ultrasonic waves will change, resulting in an increase in ranging error. Moreover, the beam angle of the ultrasonic sensor is large and is easily disturbed by the surrounding environment, and for some smaller or closer objects, the distance may not be accurately measured. In addition, the pool environment is complex, and impurities and microorganisms in the water are also easily attached to the surface of the robot, affecting its appearance and heat dissipation performance, and even possibly blocking the sensor and water inlet, causing the robot to malfunction, and the robot maintenance is difficult. Therefore, it is urgent to design a technical solution to overcome at least one technical problem in the related art. SUMMARY
[0004] The main purpose of the embodiments of the present application is to provide a pool robot positioning and trajectory prediction method and system fusing sonar and vision, aiming to solve the technical problems of low ranging accuracy of the pool robot and high difficulty in maintenance and management of the pool robot.
[0005] In a first aspect, the embodiments of the present application provide a pool robot positioning and trajectory prediction method fusing sonar and vision, comprising:
[0006] By deploying a terahertz sonar device in the robot, terahertz sonar data for a target area are collected to detect distance information of the pool robot and surrounding objects in the target area; wherein the target area is a water environment in which the pool robot is located;
[0007] Based on the terahertz sonar data, an activity trajectory model of the pool robot is constructed;
[0008] By deploying a vision sensor around the pool and in the water environment, a multispectral and event camera cooperative acquisition strategy is adopted to obtain multi-source image data containing changes in underwater, water surface, and pool environment in the target area;
[0009] The dynamic SLAM algorithm based on deep learning is adopted to identify and process dynamic objects in the pool environment in real time based on the multi-source image data, and the pool environment model is constructed by using the environment image data from which the dynamic objects are removed.
[0010] The activity trajectory model and the pool environment model are deeply fused through a knowledge graph network, and a quantum machine learning algorithm is used to predict the running state of the pool robot to obtain a panoramic model containing the predicted activity trajectory and the predicted running state of the pool robot.
[0011] The panoramic model corresponding to the digital twin model is created and updated in real time, realizing real-time simulation and visualization of the running state of the pool robot and the change of the environment, facilitating the maintenance and management decision of the pool robot.
[0012] In a second aspect, the embodiments of the present application provide a pool robot positioning and trajectory prediction method and system fusing sonar and vision, comprising:
[0013] A first acquisition module is deployed in a terahertz sonar device in the robot, configured to acquire terahertz sonar data for a target area to detect distance information of the pool robot and surrounding objects in the target area; wherein the target area is a water environment where the pool robot is located.
[0014] A first construction module is configured to construct an activity trajectory model of the pool robot based on the terahertz sonar data.
[0015] A second acquisition module is deployed in a vision sensor around the pool and in the water environment, configured to acquire multi-source image data containing changes in the underwater, water surface and pool environment in the target area by using a multi-spectral and event camera cooperative acquisition strategy.
[0016] A second construction module is configured to use a dynamic SLAM algorithm based on deep learning to identify and process dynamic objects in the pool environment in real time based on the multi-source image data, and to construct a pool environment model by using environment image data from which the dynamic objects are removed; the dynamic objects include swimmers or active objects in the environment.
[0017] A fusion module is configured to deeply fuse the activity trajectory model and the pool environment model through a knowledge graph network, and to use a quantum machine learning algorithm to predict the running state of the pool robot to obtain a panoramic model containing the predicted activity trajectory and the predicted running state of the pool robot.
[0018] The display module is used for creating and updating a digital twin model corresponding to the panoramic model in real time, realizing real-time simulation and visualization of the running state of the pool robot and the change of the environment, and facilitating maintenance and management decision of the pool robot.
[0019] In a third aspect, the embodiments of the present application further provide a terminal device, which comprises a processor, a memory for storing a computer program, and the processor is configured to execute the computer program and implement the fusion sonar and vision pool robot positioning and trajectory prediction method of the first aspect or any of the embodiments of the present application when executing the computer program.
[0020] The embodiments of the present application provide a fusion sonar and vision pool robot positioning and trajectory prediction method and system. In the method, a terahertz sonar device deployed in the robot is used to collect terahertz sonar data for a target area to detect distance information of the pool robot and surrounding objects in the target area; the target area is a water environment where the pool robot is located; based on the terahertz sonar data, an activity trajectory model of the pool robot is constructed; a vision sensor is deployed around the pool and in the water environment, and a multispectral and event camera cooperative acquisition strategy is used to acquire multi-source image data containing changes of underwater, water surface and pool environment in the target area; a dynamic SLAM algorithm based on deep learning is used to identify and process dynamic objects in the pool environment in real time based on the multi-source image data, and an environment image data without the dynamic objects is used to construct a pool environment model; the dynamic objects include swimmers or active objects in the environment; the activity trajectory model and the pool environment model are deeply fused through a knowledge graph network, and a quantum machine learning algorithm is used to predict the running state of the pool robot to obtain a panoramic model containing a predicted activity trajectory and a predicted running state of the pool robot; a digital twin model corresponding to the panoramic model is created and updated in real time to realize real-time simulation and visualization of the running state of the pool robot and the change of the environment, and facilitate maintenance and management decision of the pool robot.
[0021] The embodiments of the present application can realize intelligent upgrading of the pool robot ranging and positioning process from environmental perception, data processing, model construction to running prediction and visualization management. In terms of positioning accuracy, terahertz sonar and multispectral vision data fusion, combined with advanced algorithms, improve the accurate determination of the position of the robot in a complex environment. Secondly, in terms of trajectory prediction, quantum machine learning and multi-model fusion technology significantly improve the accuracy and real-time performance of the prediction. In terms of environmental modeling capability, the three-dimensional map constructed by the dynamic SLAM technology can reflect the change of the pool environment in real time. Overall, the embodiments of the present application effectively improve the autonomous operation capability, environmental adaptability and intelligent management level of the pool robot, and provide reliable and efficient technical support for pool cleaning, monitoring and other tasks. BRIEF DESCRIPTION OF DRAWINGS
[0022] Figure 1 A flowchart of a pool robot positioning and trajectory prediction method fusing sonar and vision provided for an embodiment of the present application is shown in the figure;
[0023] Figure 2 A module structure diagram of a pool robot positioning and trajectory prediction system fusing sonar and vision provided for an embodiment of the present application is shown in the figure;
[0024] Figure 3 A structure schematic block diagram of a terminal device provided for an embodiment of the present application is shown in the figure. DETAILED DESCRIPTION
[0025] In view of the technical problems in the related art, the present application provides a pool robot positioning and trajectory prediction method and system fusing sonar and vision.
[0026] Specifically, data is collected by a terahertz sonar device deployed in the robot. Terahertz waves have wideband and low photon energy characteristics, and can penetrate interference materials such as water surface foam, enabling super-resolution distance detection in low signal-to-noise ratio environments. Compared with traditional sonar, this technology can more accurately obtain the distance information of the pool robot and the surrounding objects, especially the detection ability of small foreign objects or underwater structural details is significantly improved, providing a high-precision data foundation for subsequent trajectory model construction. Second, based on the terahertz sonar data, an activity trajectory model is constructed, combined with the inertial measurement unit (IMU) data built-in the robot, and the spatio-temporal Transformer network is used to mine the correlation of data in time and space dimensions. At the same time, through the dynamic optimization of model parameters by reinforcement learning algorithm, the model can adapt to different pool environments and robot motion modes, accurately depicting the motion trajectory of the robot in the pool, and providing reliable motion information for the running state prediction. Third, a multi-spectral and event camera cooperative acquisition strategy is adopted to obtain multi-source image data. The multi-spectral camera can obtain images of different wavebands, enhancing the recognition ability of complex environmental entities such as transparent floating objects and underwater pipelines; the event camera is sensitive to dynamic changes, and the combination of the two can comprehensively capture underwater, water surface and pool environment change information. In addition, the intelligent triggering system and dynamic sampling strategy based on environment perception can reduce redundant data acquisition, improve data acquisition efficiency and quality. Then, based on the deep learning dynamic SLAM algorithm, the spatio-temporal graph convolution network (ST-GCN) and the Transformer hybrid architecture are used to realize the accurate identification and real-time tracking of dynamic objects combined with multi-source image data. Through the conditional random field (CRF) algorithm and event camera data assistance, the dynamic objects are accurately segmented and removed, and the pure static environment data is obtained for map construction. Then, with the help of sparse bundle adjustment (SBA) algorithm and deep reinforcement learning optimization parameters, as well as incremental point cloud fusion technology, the map is updated in real time, generating an accurate and dynamic three-dimensional pool environment model, providing a reliable environmental reference for path planning and trajectory prediction. Further, a knowledge graph network containing semantic information of the pool environment and motion rules of the robot is constructed through knowledge graph technology, integrating multi-source knowledge. The graph attention mechanism GAT is used to calculate the weight of different model information, realizing the intelligent fusion of the activity trajectory model and the pool environment model, and fully mining the correlation information between models. Based on quantum machine learning algorithm, the data is mapped to quantum feature space, and the superposition and entanglement characteristics of quantum states are used to efficiently process high-dimensional complex data. Quantum support vector machine (QSVM) or quantum neural network (QNN) is selected for training and prediction, which can more accurately predict the running state of the pool robot, generate a panoramic model containing predicted activity trajectory and running situation, and improve the accuracy and foresight of prediction. Finally, the digital twin model corresponding to the panoramic model is created and updated in real time, combined with augmented reality (AR) and mixed reality (MR) technology, through edge computing and 5G communication to realize the real-time interaction between virtual model and physical robot.The robot operation data and environment perception data are mapped into the virtual model in real time, so that the real-time simulation and visualization of the operation state of the pool robot and the environment change are realized. Based on the virtual model, reinforcement learning simulation is carried out, optimal solutions are provided for robot maintenance and management decisions, intelligent operation and maintenance are realized, maintenance cost is reduced, and management efficiency is improved.
[0027] The embodiment of the application provides a kind of fusion sonar and vision's pool robot positioning and trajectory prediction method and system. Wherein, the fusion sonar and vision's pool robot positioning and trajectory prediction method can be applied in terminal equipment, such as mobile phone, virtual reality equipment, tablet computer, notebook computer, desktop computer, wearable device and other electronic devices. The terminal equipment can be the server connected with tower crane equipment, can also be server cluster. The connection mode can be realized by hardware circuit, can also be realized by communication module.
[0028] Some embodiments of the application will be described in detail below with reference to the accompanying drawings. In the case of no conflict, the embodiments described below and the features in the embodiments can be combined with each other. Please refer to Figure 1 , Figure 1 The flowchart of the fusion sonar and vision's pool robot positioning and trajectory prediction method provided by the embodiment of the application is shown.
[0029] As shown in Figure 1 , the fusion sonar and vision's pool robot positioning and trajectory prediction method includes the following steps S101 to S106.
[0030] Step S101, through the terahertz sonar equipment deployed in the robot, terahertz sonar data for the target area is collected to detect the distance information of the pool robot and the surrounding objects in the target area.
[0031] Step S102, based on the terahertz sonar data, the activity trajectory model of the pool robot is constructed.
[0032] In the embodiment of the application, the target area is the water environment where the pool robot is located.
[0033] In the embodiment of the application, the terahertz wave has the characteristics of wide band and low photon energy, and the combination of terahertz technology and sonar can realize higher resolution distance measurement, penetrate interference materials such as water surface foam, and accurately detect small foreign matters or underwater structure details in the pool.
[0034] The terahertz sonar device mainly realizes distance detection based on the characteristics of terahertz waves. The principle is as follows: terahertz waves are electromagnetic waves between microwaves and infrared, with the characteristics of wide band and low photon energy. The transmitting module of the device generates and transmits terahertz waves to the target area. When the terahertz waves encounter objects in the pool (such as pool walls, floating objects, underwater pipes, etc.), they will be reflected. The receiving module captures these reflected terahertz wave signals and converts them into electrical or digital signals. By calculating the time delay of the terahertz waves from transmission to reception, combined with the propagation speed of terahertz waves in water, the distance between the pool robot and the surrounding objects can be accurately calculated.
[0035] The terahertz sonar data mainly includes the following contents: distance information, which is the most core data, clearly records the straight-line distance between the pool robot and each detected object, which is the key basis for judging the relative position of the robot and the surrounding environment. Signal strength, reflecting the strength of the reflected terahertz wave signal. Signal strength is related to factors such as object material, surface roughness, and distance. For example, the surface of the object reflects a strong signal, and the farther the distance, the weaker the signal strength. These information can assist in judging the material and size of the object. Timestamp, recording the specific time of terahertz wave transmission and reception, used to calculate the time delay, and also provides a time dimension reference for subsequent data processing and analysis, facilitating understanding of the detection situation of the robot at different times. Waveform characteristics, including waveform shape, frequency, and other characteristic information of the reflected signal. The waveforms of terahertz waves reflected by different objects are different. By analyzing these characteristics, the type and attribute of the object can be further identified.
[0036] For example, in step S101, when the pool robot is performing an underwater cleaning task, it needs to detect in real time whether there are obstacles (such as pool walls, floating toys, underwater pipes, etc.) in front to avoid collision and plan a cleaning path. The terahertz sonar device on the bottom of the pool robot transmits a beam of terahertz waves (for example, electromagnetic waves with a frequency of 0.1 THz) to the front, which propagates in water at a speed of about 2.25 x 10 8 m / s (slightly lower than the speed of light in air). When the terahertz wave encounters the pool wall 2 meters in front, part of the wave is reflected back to the robot. The receiving module captures the reflected signal and records the signal transmission time (t1) and reception time (t2). For example, the time delay Δt = t2-t1 = about 1.78 x 10 -8 seconds (calculated according to the distance of 2 meters: Δt = 2 x 2 / 2.25 x 10 8 ≈1.78 x 10 -8 seconds, the round trip distance needs to be multiplied by 2). According to the formula, distance = (propagation speed x time delay) / 2, the distance between the robot and the pool wall is calculated to be 2 meters.
[0037] Meanwhile, the device records the intensity of the reflected signal (e.g., -30 dBm, reflecting the smoothness of the pool wall surface) and the waveform characteristics (used to distinguish the material of the object later, e.g., the waveform of the concrete material of the pool wall is different from that of the plastic floating object). According to the detected distance of 2 meters, the robot determines that it does not need to turn immediately and continues to clean along the original path. If the distance is less than a safety threshold (e.g., 0.5 meters), a turning instruction is triggered. The distance data collected multiple times (e.g., the continuous position of the pool wall) can construct a pool contour map to assist in path planning.
[0038] Further optionally, in step S101, after collecting the terahertz sonar data for the target area by the terahertz sonar device deployed in the robot, pool environment data can also be obtained; the pool environment data at least includes: venue type, water turbidity, pool lighting conditions changing over time, weather information of the area where the pool is located. Then, the dynamic changes of the pool environment are predicted according to the pool environment data, and the sonar noise reduction parameters are adjusted based on the prediction results; wherein the dynamic changes of the pool environment include: water visibility change trend, lighting condition change trend, physical environmental disturbance caused by weather changes. Finally, the terahertz sonar data is real-time denoised using the sonar noise reduction parameters adjusted by optimization, to improve the quality of the sonar data.
[0039] Specifically, when the terahertz sonar device collects data for the target area, pool environment factors will interfere with the data quality. Different venue types have different spatial structures and echo characteristics, such as the closed space of an indoor swimming pool and the open environment of an outdoor pool, the propagation path and interference sources of the sonar reflected signal are different. Water turbidity affects the propagation of sound waves, the higher the turbidity, the more serious the sound wave scattering and absorption. Changes in pool lighting conditions and weather factors will change the water surface fluctuations, environmental temperature, etc., indirectly affecting the sonar data. First, obtain various pool environment data such as venue type and water turbidity, use data analysis and prediction algorithms to analyze the correlation between environmental data and predict the dynamic change trends of water visibility, lighting, and physical environmental disturbance. For example, according to the weather information, it is predicted that strong winds will cause the water surface to fluctuate violently, and according to the change trend of the lighting conditions, it is estimated that the period when the reflection interference is enhanced. Then, according to the prediction results, adjust the sonar noise reduction parameters according to the preset rules or machine learning model, such as adjusting the filtering strength, frequency range, etc. Finally, the terahertz sonar data is real-time denoised using the optimized parameters to remove the noise caused by environmental interference and improve the data quality, providing reliable data support for subsequent robot positioning, target detection, etc.
[0040] It can be understood that in practical application, in an open swimming pool, a sudden change in weather blows a strong wind, the water surface fluctuates violently, and a large amount of interference sound waves are generated. At the same time, the light condition changes due to cloud cover. The system predicts that the increase in physical environment disturbance and the change in light condition may cause reflection interference, and adjusts the sonar noise reduction parameter in advance to enhance the filtering capability of high-frequency noise. When the terahertz sonar equipment collects data, the noise generated by the water surface fluctuation is effectively suppressed, the information of the swimming pool bottom and the environment around the robot is clearly presented, the positioning deviation and target misjudgment caused by noise interference are avoided, the robot can accurately perform the cleaning task, and obstacles can be efficiently avoided. For example, in a swimming pool with turbid water quality, the system predicts that the visibility of the water body is reduced and the sound wave scattering is enhanced according to the turbidity data of the water quality, adjusts the sonar noise reduction parameter to optimize the signal receiving and processing mode. After the terahertz sonar data is reduced, the swimming pool wall, underwater obstacles and other targets can still be accurately identified, the robot can provide accurate environmental perception, ensure stable operation in a complex water quality environment, and greatly improve the adaptability and working reliability of the robot in different swimming pool environments.
[0041] Further optionally, in the above steps, the swimming pool environment dynamic change is predicted according to the swimming pool environment data, and the sonar noise reduction parameter is adjusted based on the prediction result, comprising:
[0042] The venue type, water quality turbidity, swimming pool light condition, and weather information are input into the adaptive deep learning model in the form of spatiotemporal graph nodes combined with sequence data. The dynamic capture module in the adaptive deep learning model captures the correlation between the environment data in the spatial dimension to obtain spatiotemporal environment correlation features, and the long-range dependence relationship between the environment data in the time dimension is established by the spatiotemporal correlation module to obtain long-range dependence features. The multi-scale feature fusion module is used to cross-dimensionally fuse the high-frequency detail features, low-frequency structure features in the terahertz sonar data, the spatiotemporal environment correlation features used to represent the swimming pool environment dynamic change, and the long-range dependence features to obtain environment fusion features. The signal-to-noise ratio improvement value of the noise-reduced sonar data and the positioning accuracy improvement rate are used as a reward function, and the reinforcement learning module is used to optimize the sonar noise reduction parameter adjustment strategy. The sonar noise reduction parameter adjustment strategy obtained by reinforcement learning is combined to dynamically optimize and adjust the convolution kernel size and the hidden layer parameters of the recurrent neural network used in the noise reduction process, and the optimized sonar noise reduction parameter is used to perform real-time noise reduction on the environment fusion features to filter the noise data in the terahertz sonar data.
[0043] Specifically, environmental data such as venue type, water turbidity, pool lighting conditions, and weather information are converted into a combination of spatio-temporal graph nodes and sequence data, which are used as inputs for an adaptive deep learning model. The dynamic capture module analyzes spatio-temporal graph nodes using graph neural networks (GNN) to mine the spatial dimension of different environmental data, such as how the closed structure of an indoor venue affects sound wave reflection, and the relationship between water turbidity and venue ventilation conditions, to extract spatio-temporal environmental correlation features. The spatio-temporal correlation module uses sequence modeling structures such as long short-term memory networks or Transformers to process sequence data and establish long-range dependencies in the time dimension of environmental data, such as predicting the subsequent impact of weather changes on pool lighting and water surface fluctuations, to obtain long-range dependency features.
[0044] The multi-scale feature fusion module decomposes terahertz sonar data to obtain high-frequency detail features (such as reflection signals of small obstacles) and low-frequency structure features (such as the overall profile signal of the pool wall), and cross-dimensionally fuses them with spatio-temporal environmental correlation features and long-range dependency features to form environmental fusion features that include dynamic environmental changes and sonar data features. The reinforcement learning module uses the signal-to-noise ratio improvement value of the denoised sonar data and the positioning accuracy improvement rate as a reward function, and through continuous trial and error and learning, explores the optimal sonar denoising parameter adjustment strategy. For example, different combinations of filter strength, frequency range, and other parameters are tried, and the strategy is optimized based on reward feedback. Finally, combining the strategy obtained by reinforcement learning, the size of the convolution kernel and the parameters of the recurrent neural network hidden layer in the denoising process are dynamically adjusted to perform real-time denoising on the environmental fusion features, specifically filtering out noise in the terahertz sonar data and retaining valid signals.
[0045] For example, in an open-air swimming pool scenario, the weather information for the day shows that there will be strong winds in the afternoon, the current water turbidity is high, and it is currently in the intense sunlight period. These environmental data are converted into spatio-temporal graph nodes and sequence data, which are input into the adaptive deep learning model. The dynamic capture module analyzes and finds that strong winds will exacerbate water surface fluctuations, and strong light will enhance water surface reflections, both of which will interfere with sonar data in space. The spatio-temporal correlation module predicts that as the strong winds approach, water surface fluctuations will continue to increase, and changes in light angle will also cause changes in reflection areas.
[0046] The multi-scale feature fusion module processes the terahertz sonar data, fuses the high-frequency water surface ripple reflection detail features, low-frequency pool bottom contour features, and the above-mentioned environmental features. The reinforcement learning module continuously tries different noise reduction parameters, such as adjusting the filter convolution kernel size and the number of hidden layer neurons of the recurrent neural network, finds that increasing the filtering strength of a specific frequency band and adjusting the convolution kernel size can effectively reduce the noise interference caused by water surface fluctuations and reflections, improve the signal-to-noise ratio and positioning accuracy, and thus determines the optimal noise reduction parameter adjustment strategy. Finally, the optimized parameters are used for real-time noise reduction of the sonar data, so that the robot can clearly identify the pool bottom and surrounding obstacles, and accurately plan the motion path even in harsh environments.
[0047] Therefore, through the above steps, the accuracy and adaptability of terahertz sonar data processing are significantly improved. In a complex and variable pool environment, through deep analysis and feature fusion of environmental data, the influence of environmental changes on sonar data can be predicted in advance, and noise reduction parameters can be optimized accordingly. Compared with traditional fixed parameter noise reduction methods, the signal-to-noise ratio of sonar data can be improved, the positioning error caused by noise can be effectively reduced, the positioning accuracy can be improved, the misjudgment and collision risk of the robot caused by data noise can be reduced, the working stability and reliability of the robot in different venue types, weather conditions, and water quality conditions can be improved, and the robot can efficiently complete cleaning, detection, and other tasks.
[0048] For example, when the visibility of the water body is predicted to decrease, the reinforcement learning module preferentially adjusts the filtering strength parameter of the de-noising algorithm. If the lighting conditions change drastically, the attention mechanism parameters of the deep learning-based de-noising model are optimized, making the de-noising process more targeted.
[0049] Further optionally, in step S102, based on the terahertz sonar data, an activity trajectory model of the pool robot is constructed, including:
[0050] A spatio-temporal Transformer model based on a hybrid architecture of a spatio-temporal graph convolution network (STGCN) and a Transformer network is constructed. An inertial measurement module (IMU) built-in the pool robot is obtained to obtain inertial measurement data. The spatio-temporal Transformer model is used to simulate the activity trajectory of the pool robot based on the inertial measurement data, to construct an initial activity trajectory model of the pool robot. The distance relationship between the pool robot and surrounding objects is simulated based on the terahertz sonar data, to optimize the initial activity trajectory model, so that the initial activity trajectory model is adaptive to different pool environments and robot motion modes, and an finally output activity trajectory model is obtained.
[0051] In the embodiments of the present application, the activity trajectory model is a mathematical model used to depict the movement trajectory of the pool robot in the pool environment, and plays a key role in the whole positioning and trajectory prediction method. The activity trajectory model construction method realizes accurate trajectory modeling through network architecture and multi-source data fusion. Specifically, it is constructed based on terahertz sonar data and robot built-in inertial measurement unit (IMU) data. The terahertz sonar data provides distance information between the robot and the surrounding objects, and clearly defines the positional relationship of the robot in the environment; the IMU data reflects the motion state changes such as acceleration and angular velocity of the robot itself. A spatio-temporal Transformer network, i.e., a hybrid architecture of spatio-temporal graph convolution network (STGCN) and Transformer network, is adopted. The STGCN mines the spatial correlation and temporal dynamic change rule of the data, and the Transformer network utilizes the attention mechanism to globally model the long sequence data and capture complex dependency relationships, and the combination of the two effectively processes the spatio-temporal characteristics and long-term information of the robot motion data. First, the spatio-temporal Transformer model simulates the activity trajectory of the robot based on the inertial measurement data to obtain an initial activity trajectory model; then, the distance relationship between the robot and the surrounding objects is simulated based on the terahertz sonar data to optimize the initial model and adjust the position parameters, motion direction, etc. of the robot, so that the model is self-adaptive to different pool environments and robot motion modes, and finally an accurate activity trajectory model is output.
[0052] Obviously, the model accurately depicts the movement trajectory of the robot in the pool, provides a reliable motion information basis for subsequent operation state prediction, is also helpful for the robot to plan the path, and at the same time provides key data for the digital twin model, helping to realize intelligent operation and maintenance.
[0053] Among them, the spatio-temporal graph convolution network (STGCN) is good at processing data with spatial topological structure and time sequence characteristics, and can mine the spatial correlation (such as the relative position relationship between the robot and the surrounding objects) and the temporal dynamic change rule of the pool robot motion data. The Transformer network can globally model long sequence data and capture complex dependency relationships in the data by virtue of its powerful attention mechanism. The combination of the hybrid architecture of the two can process the spatio-temporal characteristics of the robot motion data and effectively analyze long-term motion information, providing a powerful tool for trajectory modeling.
[0054] The inertial measurement module (IMU) can obtain real-time inertial measurement data such as acceleration and angular velocity of the robot, reflecting the change of the motion state of the robot itself; the terahertz sonar data provides the distance information between the robot and the surrounding objects, and clearly shows the position relationship of the robot in the pool environment. Through the space-time Transformer model, the initial activity trajectory model is first simulated by using the inertial measurement data, and then the initial model is optimized by combining the terahertz sonar data, so that the model comprehensively considers the motion state of the robot itself and external environmental factors, thereby adapting to different pool environments and motion modes.
[0055] When optimizing the initial activity trajectory model based on terahertz sonar data, the position parameters and motion direction of the robot in the model are adjusted by calculating the distance relationship between the robot and the surrounding objects. For example, if the terahertz sonar detects an obstacle in front, the model will correct the original motion trajectory of the robot according to the distance information, so that the trajectory bypasses the obstacle, thereby continuously optimizing the model and improving the accuracy of trajectory modeling.
[0056] Exemplarily, it is assumed that in a standard pool, the pool robot starts to perform a cleaning task. In the initial stage, the inertial measurement module (IMU) collects the acceleration data of the robot as 0.5 m / s 2 , and the angular velocity as 10° / s, and inputs these inertial measurement data into the space-time Transformer model to preliminarily simulate the motion trajectory of the robot in the pool, obtain an initial activity trajectory model, and show that the robot moves in a straight line in a certain direction of the pool at the current state. As the robot moves, the terahertz sonar device continuously collects data and detects a floating object 2 meters in front. At this time, the initial activity trajectory model is optimized based on the distance relationship between the robot and the floating object in the terahertz sonar data. The model adjusts the motion direction of the robot to make the trajectory bypass the floating object, and re-plans a new trajectory, thereby obtaining the final output activity trajectory model, ensuring that the robot can safely and efficiently perform the task in the complex pool environment.
[0057] Therefore, by combining inertial measurement data and terahertz sonar data, a comprehensive model is established from two levels of robot's own motion state and external environment, which can more accurately depict the robot's activity trajectory compared to a single data source, providing a reliable foundation for subsequent operating state prediction, path planning, etc. By optimizing the model with terahertz sonar data, the activity trajectory model can automatically adjust the trajectory according to changes in the pool environment (such as the appearance of new obstacles, different pool shapes) and changes in the robot's motion mode (such as acceleration, turning), enhancing the model's adaptability in different scenarios and improving the robot's autonomous operation capability in complex pool environments. The hybrid architecture of the spatio-temporal Transformer model fully leverages the advantages of STGCN and Transformer networks, enabling fast processing of large amounts of spatio-temporal data, real-time modeling and updating of the robot's activity trajectory, meeting the timeliness requirements of trajectory modeling for pool robots in dynamic environments, and improving the overall system's operational efficiency.
[0058] In step S103, multi-spectral and event camera cooperative acquisition strategy is adopted to obtain multi-source image data containing underwater, water surface and pool environment changes in the target area through visual sensors deployed around the pool and in the water environment.
[0059] In the embodiments of the present application, the visual sensors around the pool can be installed in the cameras at the edge of the pool, the ceiling or the wall, etc. These sensors are used to monitor the overall condition of the pool from different angles, including the water surface situation, the environment around the pool and possible abnormal situations, etc. For example, it can monitor whether there are people accidentally falling into the water, whether there are objects falling around the pool, etc. They usually have a wide field of view and can cover a large range of pool area, but their monitoring ability for underwater conditions is relatively limited.
[0060] The visual sensors in the water environment mainly refer to underwater cameras or other optical sensors. These sensors can directly obtain underwater image information, including the condition of the pool bottom, the pool wall, underwater facilities and objects in the water, etc. For example, they can be used to detect stains or cracks on the pool bottom, or identify the movements of underwater swimmers, etc. In order to adapt to the underwater environment, they usually need to have waterproof, pressure-resistant and other characteristics.
[0061] For example, in the multi-spectral and event camera cooperative acquisition strategy, a multi-spectral camera can capture light in different wavelength ranges, thereby obtaining image information in multiple spectral channels. In the pool environment, different substances have different absorption and reflection characteristics for different wavelengths of light, and through multi-spectral imaging, these characteristics can be used to distinguish different objects or scene features. For example, through light of specific wavelengths, the recognition ability for water quality, underwater organisms, stains, etc. can be enhanced, and the reflection differences under different spectra can also be used to better detect the fluctuations of the water surface, the changes of light and shadow, and the outlines of underwater objects, etc.
[0062] Event cameras are a new type of visual sensor that do not capture images at a fixed frame rate like traditional cameras, but instead record changes in brightness in the scene based on pixel-level event triggers. In a pool environment, event cameras can quickly capture dynamic changes in the scene, such as the rapid movements of swimmers, the transient fluctuations of the water surface, the sudden movements of objects, etc., with extremely high temporal resolution. Since it only records brightness change events, it can greatly reduce the amount of data when processing dynamic scenes, while also more accurately capturing rapidly occurring events.
[0063] The cooperative acquisition strategy is to combine the use of multispectral cameras and event cameras, giving full play to their respective advantages. Multispectral cameras provide rich spectral information, which helps to analyze and identify scenes in detail, but may have frame rate limitations when capturing fast dynamics. Event cameras can well capture rapidly changing events, but lack spectral information and complete description of static scenes. Through cooperative acquisition, when the event camera detects dynamic events in the scene, it triggers the multispectral camera to capture hyperspectral images at the corresponding time, or uses the information of the event camera to guide the adjustment of the shooting parameters of the multispectral camera to better capture multispectral information of dynamic scenes. At the same time, the images captured by the multispectral camera can provide static background and more rich scene information for the event camera to assist the understanding and positioning of events by the event camera. In this way, the two complement each other and can obtain more comprehensive and accurate multi-source image data containing underwater, water surface and pool environment changes in the target area.
[0064] Further optionally, the visual sensor further comprises a multispectral camera for acquiring image information of different wavebands. Based on the foregoing assumption, by deploying the visual sensor around the pool and in the water environment, using the multispectral and event camera cooperative acquisition strategy, multi-source image data containing underwater, water surface and pool environment changes in the target area are obtained, including:
[0065] Environmental changes are monitored in real time by light sensors and water flow velocity sensors deployed in the swimming pool to assess the complexity of the pool environment. A reinforcement learning algorithm is used to train a dynamic sampling strategy, dynamically adjusting the band acquisition combination of the multispectral camera, the event trigger threshold of the event camera, and the sensor collaboration mode according to the complexity of the pool environment. Specifically, when a sudden change in light or increased water flow disturbance is detected, the multispectral camera and event camera are automatically triggered to enter a high frame rate collaborative acquisition mode. Using the multispectral camera, based on the event trigger threshold and sensor collaboration mode, visual image data in different bands is acquired for the target area. Based on the visual image data in different bands, complex environmental entities in the pool environment are identified, and a dynamic filter bank is used to perform image enhancement processing on the identified complex environmental entities to obtain optimized visual image data. The complex environmental entities include at least: transparent floating objects and underwater pipes.
[0066] Specifically, light sensors can monitor changes in light intensity and color in real time, while water flow velocity sensors can monitor changes in the magnitude and direction of water flow. This information reflects the dynamics of the pool environment and, when combined, assesses its complexity. For example, a sudden change in light may indicate the passage of an object or a change in lighting, while increased water flow disturbance may be caused by swimmer activity or equipment. A reinforcement learning algorithm is used to train a dynamic sampling strategy, adjusting the band acquisition combination of the multispectral camera, the event trigger threshold of the event camera, and the sensor collaboration mode based on environmental complexity. Through continuous trial and error, reinforcement learning learns the optimal sampling strategy under different environmental conditions to adapt to environmental changes and improve the effectiveness of data acquisition. When a sudden change in light or increased water flow disturbance is detected, indicating rapid changes in the environment, the multispectral camera and event camera are automatically triggered to enter a high-frame-rate collaborative acquisition mode to more accurately capture these dynamic changes and avoid missing important information. The multispectral camera acquires visual image data in different bands based on the set event trigger threshold and sensor collaboration mode. Different bands reflect different object and scene features differently, helping to identify various complex environmental entities. Then, the identified complex environmental entities are enhanced using a dynamic filter bank to highlight target features, suppress noise and interference, and improve image quality.
[0067] For example, suppose there are transparent plastic floats in a large swimming pool, and a complex underwater pipe system. When a swimmer passes by quickly, it causes a sudden change in water flow velocity. A water flow velocity sensor detects this change, and a light sensor detects changes in light reflection caused by the swimmer's movement. Based on these detected environmental changes, a reinforcement learning algorithm adjusts the multispectral camera to select specific combinations of wavelengths sensitive to transparent objects and underwater pipes for acquisition, while lowering the event trigger threshold of the event camera to more sensitively capture events related to these objects. After the multispectral camera acquires image data in different wavelengths, an image recognition algorithm identifies complex environmental entities such as transparent floats and underwater pipes. Then, a dynamic filter bank is used to enhance these entities, such as increasing the edge contrast of transparent floats and highlighting the outline of underwater pipes, making them more clearly visible in the image.
[0068] For example, a recognition system based on a multimodal large model (such as a CLIP-like architecture) can be constructed by jointly inputting the spectral features of multispectral images with the spatiotemporal features of event cameras into the model. Through contrastive learning, different modal data are aligned in the same semantic space, and the global attention mechanism of the Transformer is used to capture the subtle differences of complex environmental entities such as transparent floating objects and underwater pipes under multimodal conditions. Combined with meta-learning algorithms, the model can quickly adapt to the feature changes of entities in different swimming pool environments, thereby improving recognition accuracy.
[0069] For example, for identified complex environmental entities, an image enhancement method based on Conditional Generative Adversarial Network (cGAN) and dynamic filter bank is employed. cGAN generates enhancement conditions based on entity type (transparent floating objects, underwater pipes, etc.), while the dynamic filter bank adaptively adjusts filtering parameters based on the band characteristics of the multispectral image to perform targeted enhancement. Simultaneously, a reinforcement learning module is introduced, using the edge sharpness and semantic consistency of the enhanced image as reward functions to optimize the enhancement process, resulting in high-quality optimized visual image data.
[0070] In this way, by dynamically adjusting the acquisition strategy according to the complexity of the environment, more targeted and useful data can be collected, avoiding over-collection in simple environments or data loss in complex environments, thus improving the quality and effectiveness of image data. The acquisition of different bands by the multispectral camera and image enhancement processing help to more accurately identify various complex environmental entities in the pool environment, such as transparent floating objects and underwater pipes. Even if these objects are difficult to distinguish clearly under normal vision, they can be clearly presented in the processed image, providing a good foundation for subsequent analysis and processing. The system can adapt to dynamic changes in the pool environment in real time, such as sudden changes in light and water flow, and adjust the acquisition mode and parameters in a timely manner to ensure that the acquired data reflects the true situation of the current environment, improving the robustness and adaptability of the system.
[0071] Further optionally, in the above steps, a dynamic sampling strategy is trained using a reinforcement learning algorithm, dynamically adjusting the band acquisition combination of the multispectral camera, the event trigger threshold of the event camera, and the sensor cooperation mode according to the complexity of the pool environment, including:
[0072] The multispectral camera employs an adaptive band switching method. Based on the complexity of the pool environment, including real-time light, water turbidity, and weather conditions, it acquires the spatial and spectral correlations between images of different bands. Based on these correlations, it automatically adjusts the combination of acquisition bands, event trigger thresholds, and sensor collaboration modes to ensure that visual image data adapted to different pool environments is acquired under different operating conditions, reducing redundant data acquisition and improving data effectiveness.
[0073] In principle, multispectral cameras and related sensors monitor the pool environment in real time, acquiring information such as real-time light (e.g., light intensity, light direction), water turbidity (reflecting water clarity), and weather conditions (e.g., sunny, cloudy, rainy). Multispectral cameras collect images in different wavelengths; different wavelengths react differently to objects of different materials and states. By analyzing the spatial and spectral correlations between images of different wavelengths, the characteristics and distribution of objects in the pool environment can be understood. For example, some wavelengths are more sensitive to transparent objects, while others better reveal details of underwater structures. Correlation analysis identifies the relationships between these wavelengths and their connection to environmental factors in the pool.
[0074] Reinforcement learning algorithms take the complexity of the swimming pool environment as input and aim to find the optimal dynamic sampling strategy through continuous trial and error and learning. The algorithm uses the combination of sampling bands, event trigger thresholds, and sensor cooperation modes as adjustable actions, and the effectiveness of the acquired visual image data (such as the useful information contained in the data and the accuracy of object recognition in the swimming pool environment) as rewards. During training, the algorithm selects actions based on the current environmental state (adjusting the sampling strategy), and then evaluates the quality of the actions based on the ability of the acquired data to describe the swimming pool environment (reward), continuously adjusting the strategy to maximize the reward. For example, when the water turbidity is high, the algorithm tries different band combinations and adjusts subsequent sampling strategies based on the recognizability of underwater objects in the acquired images to improve data effectiveness.
[0075] Based on a strategy derived from a reinforcement learning algorithm, the multispectral camera adaptively switches its acquisition band combinations and adjusts the event trigger threshold of the event camera and the sensor collaboration mode. For example, in low-light environments, it increases the event trigger threshold of the event camera to reduce false triggers caused by light interference; simultaneously, the multispectral camera selects band combinations more suitable for low-light environments to obtain clearer image data. Through this adaptive adjustment, it ensures that visual image data adapted to the pool environment can be acquired under different operating conditions, reducing the amount of redundant data acquisition and improving data effectiveness and the ability to describe the pool environment.
[0076] For example, the swimming pool has strong real-time lighting and low water turbidity. Multispectral cameras analyze the spatial and spectral correlations between images of different bands, discovering that certain bands (such as the visible light band) are effective at imaging floating objects and pool facilities. Reinforcement learning algorithms select appropriate band combinations (e.g., enhancing the acquisition of the visible light band) based on the current environmental conditions, while simultaneously adjusting the event trigger threshold of the event camera (appropriately increasing it due to the relatively stable environment) and the sensor collaboration mode (e.g., increasing the synchronous acquisition frequency of the multispectral camera and the event camera). The resulting image data is clear, accurately identifies objects in the pool, has low redundancy, and high data validity.
[0077] For example, in low light and with high water turbidity, after acquiring images of various bands using a multispectral camera, correlation analysis revealed that the infrared band has strong penetrating power for underwater objects and can better capture underwater structures. The reinforcement learning algorithm adjusted the sampling strategy, switching to a combination of acquisition bands primarily based on the infrared band, lowering the event trigger threshold of the event camera (because changes in light and water turbidity may make dynamic changes of objects difficult to capture), and adjusting the sensor collaboration mode (e.g., changing the acquisition time interval between the multispectral camera and the event camera). Through these adjustments, the acquired image data can more clearly present underwater objects. Although the environment is complex, the adaptive sampling strategy reduces redundant data and improves the data's ability to describe the pool environment, such as accurately identifying the location and shape of underwater pipes.
[0078] Based on the above principles, training dynamic sampling strategies using reinforcement learning algorithms can enable multispectral cameras and event cameras to better adapt to different swimming pool environments, thereby improving the quality and efficiency of data acquisition.
[0079] Step S104: Using a deep learning-based dynamic SLAM algorithm, dynamic objects in the pool environment are identified and processed in real time based on the multi-source image data, and a pool environment model is constructed by removing the environmental image data containing the dynamic objects.
[0080] In this embodiment, the deep learning-based dynamic SLAM algorithm, using a deep learning-based dynamic SLAM (Simultaneous Localization and Mapping) algorithm, can identify and process dynamic objects (such as swimmers) in the pool environment in real time, avoids them being mistakenly included in the map construction, generates a more accurate and dynamic 3D map, and provides a more reliable environmental model for path planning and trajectory prediction.
[0081] Image data is acquired using visual sensors (such as multispectral cameras and event cameras mentioned in the previous steps), and deep learning algorithms are used to extract and match features in the images. For example, convolutional neural networks (CNNs) are used to extract feature points such as corners and edges in the images, and then feature matching algorithms are used to find the correspondence between different frames, thereby calculating the camera pose changes and achieving localization.
[0082] For example, based on localization, environmental information observed by the camera is fused to construct a map. For dynamic SLAM, the key is to distinguish between dynamic objects and the static environment. Deep learning object detection algorithms, such as YOLO (You Only Look Once) or SSD (Single Shot MultiBox Detector), are used to identify dynamic objects (such as swimmers) in the images. Then, when constructing the map, information about these dynamic objects is excluded, and only image data of the static environment is used to build the map, resulting in a more accurate model of the pool environment.
[0083] Deep learning object detection is used to identify moving objects. Taking the YOLO algorithm as an example, the input image is divided into multiple grids, each responsible for detecting the center position of the object. Convolutional and fully connected layers are used to extract features and classify the image, predicting the object's category, location, and confidence level. In a swimming pool environment, it can quickly detect moving objects such as swimmers. Deep learning models are used for feature extraction, such as deep learning versions of traditional feature extraction methods like SIFT (Scale-Invariant Feature Transform) or ORB (Oriented Fast and Rotated BRIEF). These models can learn more robust features, unaffected by factors such as lighting and viewing angle. Feature matching determines the correspondence between different frames by calculating the similarity between features; commonly used methods include matching algorithms based on Euclidean distance or cosine similarity. Optionally, the camera pose change can be calculated based on the feature matching results. Methods such as PnP (Perspective-n-Point) can be used to solve for the camera's rotation and translation parameters using known feature point coordinates and camera intrinsic parameters, thus obtaining the camera's pose at different times.
[0084] Besides using object detection algorithms to identify dynamic objects, other methods can be employed to track their movement. For example, filtering algorithms such as the Kalman filter can be used to predict and update the position of dynamic objects, allowing for better exclusion of their influence when building maps. Meanwhile, for static objects that might be misidentified as dynamic objects (such as floating objects in water), further semantic analysis or contextual information can be used for judgment and correction.
[0085] In this way, deep learning can learn rich image features, making it more accurate in recognizing dynamic objects and modeling static environments. Compared to traditional SLAM algorithms, it reduces map errors caused by interference from dynamic objects. It also exhibits good adaptability to complex environments such as varying lighting conditions and turbid water. During training, deep learning models can learn features from different environments, enabling stable localization and map building under various conditions. With the continuous improvement of hardware computing power, deep learning-based algorithms can complete calculations in a shorter time, meeting real-time requirements. For example, using high-performance GPUs can accelerate the inference process of deep learning models, allowing dynamic SLAM systems to process image data in real time and generate 3D maps of swimming pool environments.
[0086] In this embodiment, the pool environment model is a digital representation of the pool and its surrounding environment. It is constructed based on information such as collected multi-source image data and aims to accurately present the static structure and characteristics of the pool in the form of a 3D map, including the shape, size, depth, underwater facilities (such as underwater pipes), surrounding buildings or facilities, etc., to provide a basis for subsequent path planning, trajectory prediction and other analyses and applications related to the pool environment.
[0087] Specifically, the process begins by acquiring multi-source image data, including underwater, surface, and pool environment information, through visual sensors deployed around the pool and within the aquatic environment. Then, a deep learning-based dynamic SLAM algorithm is used to identify and process dynamic objects in the pool environment in real time. The environmental image data, after removing dynamic objects, serves as the foundation for a series of image processing, feature extraction, matching, and 3D reconstruction techniques to gradually construct a 3D model of the pool environment. During this process, multispectral camera images of different wavelengths are also incorporated to enhance the identification and modeling capabilities of different objects in the environment. For example, analyzing images in different wavelengths helps to more accurately determine the location and shape of underwater pipes and other facilities.
[0088] This provides accurate environmental information for pool cleaning robots and lifesaving equipment, helping them plan reasonable movement paths, avoid collisions with obstacles, and complete tasks efficiently. For example, a pool cleaning robot can plan a cleaning path covering the entire pool bottom based on the environmental model, while avoiding fixed facilities such as underwater pipes. This model can also serve as a reference model, compared with real-time monitored images, to promptly detect abnormal changes in the pool environment, such as whether new objects have entered the pool or whether pool facilities have been damaged, thus helping to ensure the safe operation of the pool. By combining the motion information of dynamic objects with the environmental model, the future trajectories of dynamic objects such as swimmers can be predicted, allowing for advance preparation. For instance, lifeguards can use the prediction results to promptly identify swimmers in potential danger and take rescue measures.
[0089] In this embodiment of the application, dynamic objects in the pool environment mainly refer to those objects whose position, posture or state changes over time, such as swimmers, or floating objects in the water (if their motion is unstable), pool cleaning equipment in operation, etc.
[0090] This is the most prominent feature of dynamic objects. In a swimming pool, they move at different speeds and in different directions, and their trajectories may be complex curves, influenced by factors such as the swimmer's movements and water currents. Taking a swimmer as an example, they will make various different postures during swimming, such as breaststroke, freestyle, and backstroke. The appearance of their body shape in the image will constantly change, posing a certain challenge to recognition and tracking.
[0091] The movement of dynamic objects can cause changes in the surrounding environment. For example, a swimmer's strokes create water currents and affect the surface ripples. This interaction information can serve as important clues for identifying and analyzing dynamic objects. A deep learning-based dynamic SLAM algorithm is used to identify dynamic objects. This algorithm typically first extracts features from multi-source image data, learning the characteristic patterns of dynamic objects under different postures and lighting conditions. Then, it uses object detection and tracking techniques to determine the position and range of the dynamic object in the image in real time. When constructing the pool environment model, dynamic objects are removed from the environmental image data to avoid interference with the static environment modeling. For example, the swimmer's body contour is identified and removed from the image, while the static background and facilities of the pool are retained for environmental model construction. Simultaneously, the motion information of dynamic objects is analyzed and recorded separately for subsequent trajectory prediction and other operations.
[0092] As an optional embodiment, in step S104, a deep learning-based dynamic SLAM algorithm is used to identify and process dynamic objects in the pool environment in real time based on the multi-source image data, including:
[0093] The process involves acquiring visual image data in different bands from a multispectral camera and spatiotemporal feature maps from an event camera triggered by an event threshold; performing spatiotemporal alignment between the visual image data in different bands and the spatiotemporal feature maps; establishing a correspondence between the visual image data in different bands and the spatiotemporal feature maps using a feature matching algorithm based on the color and spectral information of the multispectral images and the sensitivity of the event camera to dynamic changes; using a generative adversarial network (GAN) to enhance the visual image data in different bands, thereby strengthening the edge and texture details of each object in the visual image data in different bands; and employing a multi-head attention mechanism to globally model the visual image data in different bands, the spatiotemporal feature maps, and the correspondence to obtain high-dimensional fused features. The high-dimensional fusion features are used to capture the intrinsic connections of the dynamic objects in the spectral and spatiotemporal dimensions. Local regions and / or feature points in the high-dimensional fusion features are used to construct corresponding graph nodes. Graph convolution operations are used to learn the spatial structural relationships between graph nodes and their temporal motion patterns. Based on these motion patterns, dynamic objects in the pool environment are initially labeled. A conditional random field algorithm is used to perform pixel-level segmentation of visual image data in different bands based on the initial labeling information. The boundary features of the dynamic objects are obtained through the spatial positional relationships and feature similarities between pixels. Event stream data from an event camera is used to assist in the judgment of these boundary features. Using the boundary features obtained through this auxiliary judgment, the dynamic objects are deleted from the multi-source image data to obtain the environmental image data.
[0094] Specifically, firstly, visual image data in different bands is acquired from multispectral cameras. This data contains rich spectral information, which helps to identify objects with different materials and properties. Simultaneously, an event camera, triggered by an event threshold, acquires spatiotemporal feature maps. The event camera is highly sensitive to dynamic changes and can capture rapidly occurring events in the scene.
[0095] Furthermore, spatiotemporal alignment is performed on visual image data and spatiotemporal feature maps across different spectral bands to ensure temporal and spatial consistency. Based on the color and spectral information of multispectral images and the sensitivity of event cameras to dynamic changes, a feature matching algorithm is used to establish the correspondence between visual image data and spatiotemporal feature maps across different spectral bands. For example, finding the corresponding position of an object in a multispectral image in the event camera's spatiotemporal feature map, or vice versa.
[0096] Next, a Generative Adversarial Network (GAN) is used to enhance the visual image data in different bands. A GAN consists of a generator and a discriminator. The generator attempts to produce images similar to real images, while the discriminator distinguishes between real and generated images. Through this adversarial process, the edges and texture details of objects in the visual image data in different bands are enhanced, making it easier for subsequent recognition tasks to detect object features.
[0097] Then, a multi-head attention mechanism is used to globally model visual image data, spatiotemporal feature maps, and established correspondences across different spectral bands. This mechanism can simultaneously focus on information from different locations, comprehensively processing data from different modalities to obtain high-dimensional fusion features. These high-dimensional fusion features are used to capture the intrinsic connections between dynamic objects in the spectral and spatiotemporal dimensions. For example, through this fusion, the spectral features of objects in multispectral images (such as reflectivity in a specific band) can be combined with the spatiotemporal features of object motion recorded by an event camera (such as direction and velocity), forming a more comprehensive and representative feature vector for better identification and understanding of dynamic objects.
[0098] Local regions and / or feature points from high-dimensional fusion features are used to construct corresponding graph nodes. Graph convolution operations are then used to learn the spatial structural relationships and temporal motion patterns between these nodes. Based on the learned motion patterns, dynamic objects in the pool environment are initially labeled. For example, based on the connectivity and motion trends between graph nodes, it is determined which regions might be dynamic objects, and their approximate positions and orientations.
[0099] The Conditional Random Field (CRF) algorithm is used to perform pixel-level segmentation of visual image data in different bands based on preliminary annotation information. CRF considers the spatial relationships and feature similarities between pixels, enabling precise delineation of the boundaries of dynamic objects. The boundary features of dynamic objects are obtained through the spatial relationships and feature similarities between individual pixels. For example, the similarity of adjacent pixels in features such as color and texture, as well as their relative spatial positions, are used to determine the boundaries of dynamic objects.
[0100] By combining event stream data from event cameras, boundary features are used to assist in the judgment. If a certain area generates a large number of event responses in the event camera and matches the object regions segmented from the multispectral image, then the area is considered a dynamic object. Otherwise, falsely judged areas are removed. Using the boundary features determined through auxiliary judgment, dynamic objects are removed from the multi-source image data, resulting in environmental image data containing only static environmental information. This completes the identification and processing of dynamic objects, providing a foundation for subsequently building an accurate pool environment model.
[0101] Through the above steps, the deep learning-based dynamic SLAM algorithm can effectively identify and process dynamic objects in the pool environment, providing reliable data support for generating accurate 3D maps and subsequent tasks such as path planning and trajectory prediction.
[0102] Continuing with the aforementioned embodiments, in step S104, an environmental image data after removing the dynamic objects is used to construct a swimming pool environment model, including:
[0103] Using the Sparse Bundle Adjustment (SBA) algorithm, a 3D map of the swimming pool environment is constructed based on the feature point information in the environmental image data, resulting in the swimming pool environment model. Based on the incremental point cloud fusion algorithm, multi-source image data acquired in real time is fused into the swimming pool environment model. By adding new feature point information to the 3D map through feature matching and point cloud registration, the swimming pool environment model is updated in real time.
[0104] The Sparse Bundle Adjustment (SBA) algorithm is an optimization method used in multi-view geometry, aiming to simultaneously optimize the camera pose and the 3D coordinates of scene points by minimizing reprojection errors. In pool environment modeling, it works based on feature point information in environmental image data.
[0105] Specifically, the process begins by extracting feature points from the environmental image data. These feature points can be corners, edges, or other points with unique characteristics. For example, the edges of a pool wall or the corners of underwater facilities can be used as feature points. Commonly used feature point extraction algorithms include SIFT (Scale Invariant Feature Transform) and ORB (Accelerated Robust Feature Transform). By analyzing the correspondence between feature points in different images, triangulation and other methods are used to initially estimate the camera's pose (including rotation and translation). The camera pose describes the camera's position and orientation in space, which is crucial for constructing a 3D map. The estimated camera pose and the 3D coordinates of the scene points are projected back onto the image plane to obtain reprojected points. The error between the reprojected points and the actually observed feature points is calculated; this error is the reprojection error. A nonlinear optimization algorithm (such as the Levenberg-Marquardt algorithm) is used to minimize the reprojection error. During optimization, the camera pose and the 3D coordinates of the scene points are adjusted simultaneously to minimize the reprojection error. After multiple iterations of optimization, a more accurate camera pose and the 3D coordinates of the scene points are obtained, thus constructing a 3D map of the pool environment, i.e., the pool environment model.
[0106] The core idea of incremental point cloud fusion algorithms is to progressively fuse real-time acquired multi-source image data into an existing pool environment model. This is achieved by continuously adding new feature point information to update the model, enabling it to reflect real-time changes in the pool environment. Multi-source image data is acquired in real-time by various sensors deployed around the pool and in the water environment (such as multispectral cameras and event cameras). This data contains information about the pool environment at different times, potentially including newly emerging feature points or changes in existing feature points. Feature points from the real-time acquired multi-source image data are matched with feature points in the existing pool environment model. This is achieved by calculating the similarity between feature point descriptors (such as SIFT descriptors). Once matching feature point pairs are found, the correspondence between the newly acquired data and the model can be determined. Based on the feature matching results, a point cloud registration algorithm (such as the Iterative Closest Point Algorithm - ICP) is used to register the newly acquired point cloud data with the existing pool environment model. The purpose of registration is to accurately align the new point cloud data into the model, ensuring its position and orientation in 3D space are consistent with the model. After point cloud registration, the newly added feature point information is added to the 3D map to update the pool environment model. In this way, the model can continuously incorporate new information over time, achieving real-time updates and always maintaining an accurate description of the pool environment.
[0107] Through the above two steps, an initial pool environment model is first constructed using the SBA algorithm, and then the real-time collected data is fused into the model for real-time updates based on the incremental point cloud fusion algorithm, thereby obtaining an accurate and dynamic pool environment model, which provides a reliable foundation for subsequent tasks such as path planning and trajectory prediction.
[0108] Step S105: Through a knowledge graph network, the activity trajectory model and the pool environment model are deeply integrated, and a quantum machine learning algorithm is used to predict the operating status of the pool robot, thereby obtaining a panoramic model that includes the predicted activity trajectory and predicted operating status of the pool robot.
[0109] In this embodiment, the knowledge graph network is a graph-based data structure composed of nodes (entities) and edges (relationships) used to represent various types of knowledge and the relationships between them. In this application, the knowledge graph network is constructed based on various entities in the swimming pool environment (such as pool boundaries, facilities, swimmers, etc.) and the relationships between them (such as positional relationships, movement relationships, etc.). It can integrate and correlate information obtained from different data sources (such as multispectral cameras, event cameras, etc.) to form a structured knowledge network for better understanding and analysis of the swimming pool environment.
[0110] This approach integrates information from the activity trajectory model and the pool environment model—for example, combining the swimmer's trajectory with static pool environment information—allowing the system to understand the dynamic and static conditions of the pool from a global perspective. Relationships within the knowledge graph network enable reasoning and prediction. For instance, based on the swimmer's current position and movement trends, as well as obstacle information in the pool environment, potential collision risks can be inferred, providing a more comprehensive basis for the pool robot's path planning and operational status prediction.
[0111] Representing knowledge in an intuitive graphical structure facilitates computer understanding and processing, while also aiding humans in the management and maintenance of knowledge about the swimming pool environment.
[0112] In this application, quantum machine learning is a method that applies the principles and techniques of quantum computing to the field of machine learning. It utilizes the properties of quantum states such as superposition and entanglement to process and analyze data, potentially offering higher efficiency and performance compared to traditional machine learning algorithms when dealing with certain complex problems. For example, in this application, qubits may be used to represent the state information or environmental characteristics of a swimming pool robot, and these quantum states can be transformed and processed through quantum gate operations to predict the operating state of the swimming pool robot.
[0113] Therefore, the swimming pool environment contains a large amount of multi-source data, including images and motion trajectories. Quantum machine learning algorithms can process this complex data more efficiently, extract key features, and thus more accurately predict the operating status of the swimming pool robot. Leveraging the unique advantages of quantum computing, the prediction model can be optimized; for example, it can find the global optimum more quickly when searching for optimal parameters, improving the accuracy and reliability of predictions. This gives the model better generalization ability when facing different swimming pool environments and operating scenarios, enabling it to adapt to various complex situations and providing more stable and intelligent operational guidance for the swimming pool robot.
[0114] As an optional embodiment, in step S105, knowledge graph technology is used to construct a knowledge graph network containing semantic information of the pool environment and robot motion rules; the robot's historical motion trajectory data and robot motion state information in the activity trajectory model, as well as the static map data and real-time detected obstacle information in the pool environment model, are converted into node feature vectors in the knowledge graph network; the weights of the activity trajectory model and the pool environment model are calculated using the graph attention mechanism GAT to obtain the fusion result between the activity trajectory model and the pool environment model; wherein, the attention weights of the pool robot and its neighboring nodes in the fusion result are used to reflect the degree of influence of the neighboring nodes on the operating state of the pool robot; the neighboring nodes include: surrounding obstacle nodes and pool wall nodes; the fusion result is mapped to the quantum feature space, and using the quantum state superposition and entanglement properties, a quantum support vector machine or a quantum neural network is used as a quantum machine learning model to predict the operating state of the pool robot in the future; the prediction result is integrated with the activity trajectory model and the pool environment model to generate the panoramic model containing the predicted activity trajectory and predicted operating status, which is used to provide a decision basis for the path planning and task scheduling of the pool robot.
[0115] Knowledge graph technology was used to construct a network containing semantic information about the pool environment and rules governing robot movement. It functions like an intelligent knowledge database, organizing various information about the pool environment—such as the pool's shape, size, and facility locations—and the rules the robot must follow to move within the pool, such as speed and turning angle limits, in a structured manner. This knowledge graph network clearly represents the relationships between different pieces of information, providing a foundation for subsequent analysis and processing.
[0116] Specifically, in step S105, the robot's historical motion trajectory data and motion state information (such as current speed and direction) in the activity trajectory model, as well as the static map data (such as the layout of the pool and the location of fixed obstacles) and real-time detected obstacle information in the pool environment model, are all converted into node feature vectors in the knowledge graph network. This is equivalent to converting various types of information into digital vector forms that computers can understand and process, enabling the knowledge graph network to perform unified analysis and operations on this information. For example, the robot's historical motion trajectory can be converted into vectors representing features such as trajectory shape and speed changes through some feature extraction methods, while static map data can be converted into vectors representing the location and attributes of different areas.
[0117] Next, Graph Attention (GAT) is used to calculate the weights of the activity trajectory model and the pool environment model. GAT automatically learns the importance of each node relative to other nodes, i.e., the attention weight. In this scenario, GAT determines the importance of each part of the activity trajectory model and the pool environment model in describing the swimming robot's operating state, thus obtaining a fusion result between the two. For example, for the swimming robot, the attention weights of obstacle nodes and pool wall nodes around it are higher because these neighboring nodes have a greater impact on its operating state. In this way, the information from the two models can be fused more accurately, highlighting factors that have a significant impact on the robot's operating state.
[0118] The fusion result is then mapped to a quantum feature space, leveraging the properties of the quantum realm to further process information. Quantum states exhibit superposition and entanglement, enabling quantum systems to process multiple pieces of information simultaneously, with different quantum states exhibiting interrelationships. Quantum support vector machines or quantum neural networks are employed as quantum machine learning models, utilizing these properties to predict the future operating state of the pool robot. Quantum machine learning models can handle complex nonlinear problems more efficiently, and for complex motion systems like pool robots, they can more accurately predict their future behavior, such as predicting their position and velocity changes over a given period.
[0119] Finally, the prediction results are integrated with the activity trajectory model and the pool environment model to generate a panoramic model that includes the predicted activity trajectory and predicted operating conditions. This panoramic model integrates all relevant information, including the robot's past movement trajectory, current environmental information, and predictions of future operating states, providing a comprehensive and accurate decision-making basis for the pool robot's path planning and task scheduling. For example, based on the panoramic model, the robot can plan in advance to avoid obstacles, or rationally arrange the order and timing of cleaning tasks based on the predicted operating conditions.
[0120] Step S106: Create and update the digital twin model corresponding to the panoramic model in real time to realize the real-time simulation and visualization of the operating status and environmental changes of the pool robot, which facilitates the maintenance and management decisions of the pool robot.
[0121] In this embodiment of the application, the digital twin model is a virtual model corresponding to a real physical system. It accurately simulates and maps the state, behavior and performance of physical entities through data connection and real-time interaction, so as to better understand, predict and control physical entities.
[0122] Specifically, digital twin models utilize technologies such as the Internet of Things (IoT), big data, and artificial intelligence (AI) to transmit data from the physical world to a virtual model in real time, enabling data-driven modeling and simulation of physical entities. Simultaneously, the virtual model can optimize and control the physical entity based on simulation results, achieving bidirectional interaction between the physical and virtual worlds. Various data from the physical entity, such as position, speed, and temperature, are collected through sensors and other devices and transmitted to the digital twin model via network. Mathematical models and computer simulation techniques are used to simulate and predict the behavior and performance of the physical entity. AI and machine learning algorithms are used to analyze and process large amounts of data, enabling automatic model optimization and decision support. The simulation results of the digital twin model are displayed intuitively, such as through 3D graphics and animations, facilitating user understanding and analysis. Based on information provided by the panoramic model, the digital twin model can simulate the real-time operating status of a swimming pool robot, including its position, speed, and posture, as well as changes in the surrounding environment, such as the location of obstacles and water flow. The simulation results are presented to users in a visual manner, allowing them to intuitively see the swimming pool robot's operation through a graphical interface, promptly identify problems, and make adjustments. By analyzing digital twin models, potential malfunctions and problems of swimming pool robots can be predicted, allowing for proactive maintenance and upkeep. Simulation results can also be used to optimize robot task scheduling and path planning, improving work efficiency.
[0123] This application embodiment enables an intelligent upgrade of the swimming pool robot's ranging and positioning process, encompassing environmental perception, data processing, model building, operational prediction, and visual management. Regarding positioning accuracy, the fusion of terahertz sonar and multispectral visual data, combined with advanced algorithms, improves the precise determination of the robot's position in complex environments. Furthermore, in trajectory prediction, quantum machine learning and multi-model fusion technologies significantly enhance the accuracy and real-time performance of predictions. In terms of environmental modeling capabilities, the 3D map constructed using dynamic SLAM technology can reflect changes in the swimming pool environment in real time. Overall, this application embodiment effectively improves the autonomous operation capability, environmental adaptability, and intelligent management level of the swimming pool robot, providing reliable and efficient technical support for tasks such as swimming pool cleaning and monitoring.
[0124] Please see Figure 2 , Figure 2 This application provides a method and system for localization and trajectory prediction of a swimming pool robot that integrates sonar and vision. The method and system include the following modules:
[0125] The first acquisition module, a terahertz sonar device deployed in the robot, is used to acquire terahertz sonar data for a target area to detect the distance information between the pool robot and surrounding objects in the target area; where the target area is the aquatic environment in which the pool robot is located.
[0126] The first construction module is used to construct the activity trajectory model of the pool robot based on the terahertz sonar data;
[0127] The second acquisition module, a visual sensor deployed around the pool and in the aquatic environment, is used to acquire multi-source image data containing underwater, surface and pool environment changes in the target area using a multispectral and event camera collaborative acquisition strategy.
[0128] The second construction module is used to use a deep learning-based dynamic SLAM algorithm to identify and process dynamic objects in the pool environment in real time based on the multi-source image data, and to construct a pool environment model by removing the dynamic objects from the environmental image data; the dynamic objects include swimmers or moving objects in the environment.
[0129] The fusion module is used to deeply fuse the activity trajectory model with the pool environment model through a knowledge graph network, and to use a quantum machine learning algorithm to predict the operating status of the pool robot, thereby obtaining a panoramic model that includes the predicted activity trajectory and predicted operating status of the pool robot.
[0130] The display module is used to create and update the digital twin model corresponding to the panoramic model in real time, so as to realize the real-time simulation and visualization of the operating status and environmental changes of the pool robot, which facilitates the maintenance and management decisions of the pool robot.
[0131] In some implementations, the pool robot localization and trajectory prediction method and system integrating sonar and vision can be applied to terminal devices. It should be noted that, for the sake of convenience and brevity, the specific working process of the pool robot localization and trajectory prediction system integrating sonar and vision described above can be referred to the corresponding process in the aforementioned embodiments of the pool robot localization and trajectory prediction method integrating sonar and vision, and will not be repeated here.
[0132] Please see Figure 3 , Figure 3 This is a schematic block diagram illustrating the structure of a terminal device provided in an embodiment of this application. Figure 3As shown, the terminal device 300 includes a processor 301 and a memory 302, which are connected via a bus 303, such as an I2C bus. Specifically, the processor 301 provides computing and control capabilities to support the operation of the entire terminal device. The processor 301 can be a central processing unit, or it can be other general-purpose processors, digital signal processors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor. Specifically, the memory 302 can be a Flash chip, a read-only memory disk, an optical disc, a USB flash drive, or a portable hard drive, etc.
[0133] Those skilled in the art will understand that Figure 3 The structures shown are merely block diagrams of some structures related to the embodiments of this application and do not constitute a limitation on the terminal devices on which the embodiments of this application are applied. Specific servers may include more or fewer components than shown in the figures, or combine certain components, or have different component arrangements. The processor is used to run a computer program stored in the memory, and when executing the computer program, implements any of the sonar and vision-fusion pool robot localization and trajectory prediction methods provided in the embodiments of this application. It should be noted that those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the terminal device described above can be referred to the aforementioned embodiments of the sonar and vision-fusion pool robot localization and trajectory prediction method, and will not be repeated here.
Claims
1. A method for localization and trajectory prediction of a swimming pool robot integrating sonar and vision, characterized in that, include: Terahertz sonar equipment deployed in the robot collects terahertz sonar data for the target area to detect the distance information between the pool robot and surrounding objects in the target area; where the target area is the aquatic environment in which the pool robot is located. Based on the terahertz sonar data, a trajectory model of the pool robot is constructed. By deploying visual sensors around the pool and in the aquatic environment, and employing a multispectral and event camera collaborative acquisition strategy, multi-source image data containing underwater, surface, and pool environment changes in the target area are obtained. A deep learning-based dynamic SLAM algorithm is used to identify and process dynamic objects in the pool environment in real time based on the multi-source image data, and a pool environment model is constructed by removing the dynamic objects from the environmental image data; the dynamic objects include swimmers or moving objects in the environment. The activity trajectory model and the pool environment model are deeply integrated through a knowledge graph network, and the operation status of the pool robot is predicted by a quantum machine learning algorithm to obtain a panoramic model that includes the predicted activity trajectory and operation status of the pool robot. A digital twin model corresponding to the panoramic model is created and updated in real time to realize the real-time simulation and visualization of the operating status and environmental changes of the pool robot, which facilitates the maintenance and management decisions of the pool robot.
2. The method according to claim 1, characterized in that, After acquiring terahertz sonar data for the target area using a terahertz sonar device deployed in the robot, the process further includes: Acquire swimming pool environmental data; the swimming pool environmental data shall include at least: venue type, water turbidity, swimming pool lighting conditions changing over time, and weather information for the area where the swimming pool is located; The dynamic changes in the pool environment are predicted based on the pool environment data, and the sonar noise reduction parameters are adjusted based on the prediction results. The dynamic changes in the pool environment include: the trend of changes in water visibility, the trend of changes in lighting conditions, and physical environmental disturbances caused by weather changes. The terahertz sonar data is denoised in real time using optimized sonar denoising parameters to improve the sonar data quality.
3. The method according to claim 2, characterized in that, The step of predicting dynamic changes in the pool environment based on the pool environment data and adjusting sonar noise reduction parameters based on the prediction results includes: The environmental data, including venue type, water turbidity, pool lighting conditions, and weather information, are combined with spatiotemporal graph nodes and sequence data and then input into the adaptive deep learning model. The dynamic capture module in the adaptive deep learning model captures the spatial relationship of various environmental data to obtain spatiotemporal environmental correlation features, and the spatiotemporal correlation module establishes the long-range dependency relationship of various environmental data in the time dimension to obtain long-range dependency features. A multi-scale feature fusion module is used to fuse the high-frequency detail features, low-frequency structural features, spatiotemporal environmental correlation features used to represent the dynamic changes of the pool environment, and long-range dependency features in the terahertz sonar data across dimensions to obtain environmental fusion features. Using the improvement in signal-to-noise ratio and the improvement in positioning accuracy of the denoised sonar data as reward functions, the sonar denoising parameter adjustment strategy is optimized through a reinforcement learning module. By combining the sonar noise reduction parameter adjustment strategy obtained from reinforcement learning, the convolution kernel size and the hidden layer parameters of the recurrent neural network are dynamically optimized and adjusted during the noise reduction process. The optimized sonar noise reduction parameters are then used to perform real-time noise reduction on the environmental fusion features to filter out noise data in the terahertz sonar data.
4. The method according to claim 1, characterized in that, The process of constructing a trajectory model for the pool robot based on the terahertz sonar data includes: Construct a spatiotemporal Transformer model based on a hybrid architecture of spatiotemporal graph convolutional network STGCN and Transformer network; Acquire inertial measurement data from the inertial measurement unit (IMU) built into the pool robot; The spatiotemporal Transformer model is used to simulate the movement trajectory of the pool robot based on the inertial measurement data, so as to construct the initial movement trajectory model of the pool robot; Based on the terahertz sonar data, the distance relationship between the pool robot and surrounding objects is simulated to optimize the initial activity trajectory model, so that the initial activity trajectory model can adapt to different pool environments and robot movement modes, and the final output activity trajectory model is obtained.
5. The method according to claim 1, characterized in that, The visual sensor also includes a multispectral camera for acquiring image information in different bands; The method utilizes visual sensors deployed around the pool and in the aquatic environment, employing a multispectral and event camera collaborative acquisition strategy to acquire multi-source image data containing underwater, surface, and pool environment changes in the target area, including: Environmental changes are monitored in real time by deploying light sensors and water flow velocity sensors in the pool to assess the complexity of the pool environment; A dynamic sampling strategy is trained using reinforcement learning algorithms. The band acquisition combination of the multispectral camera, the event trigger threshold of the event camera, and the sensor collaboration mode are dynamically adjusted according to the complexity of the pool environment. Specifically, when a sudden change in light or an increase in water flow disturbance is detected, the multispectral camera and the event camera are automatically triggered to enter a high frame rate collaborative acquisition mode. Using a multispectral camera, visual image data in different bands is collected for the target area based on event trigger thresholds and sensor collaboration mode. Based on visual image data in different bands, complex environmental entities in the swimming pool environment are identified, and the identified complex environmental entities are image-enhanced through a dynamic filter bank to obtain optimized visual image data; wherein, the complex environmental entities include at least: transparent floating objects and underwater pipes.
6. The method according to claim 5, characterized in that, The method of training a dynamic sampling strategy using reinforcement learning algorithms, and dynamically adjusting the band acquisition combination of the multispectral camera, the event trigger threshold of the event camera, and the sensor cooperation mode according to the complexity of the pool environment, includes: The multispectral camera employs an adaptive band switching method. Based on the complexity of the pool environment, including real-time light, water turbidity, and weather conditions, it acquires the spatial and spectral correlations between images of different bands. Based on these correlations, it automatically adjusts the combination of acquisition bands, event trigger thresholds, and sensor collaboration modes to ensure that visual image data adapted to different pool environments is acquired under different operating conditions, reducing redundant data acquisition and improving data effectiveness.
7. The method according to claim 1, characterized in that, The method employs a deep learning-based dynamic SLAM algorithm to identify and process dynamic objects in the pool environment in real time based on the multi-source image data, including: Acquire visual image data in different bands from a multispectral camera, as well as spatiotemporal feature maps acquired by an event camera triggered by an event threshold; Spatiotemporal alignment is performed between visual image data in different bands and the spatiotemporal feature map; Based on the color and spectral information of multispectral images and the sensitivity of event cameras to dynamic changes, a correspondence between visual image data in different bands and the spatiotemporal feature map is established through a feature matching algorithm. Generative Adversarial Network (GAN) is used to enhance visual image data in different bands to improve the edge and texture details of objects in the visual image data in different bands. A multi-head attention mechanism is employed to globally model visual image data from different spectral bands, the spatiotemporal feature maps, and the corresponding relationships to obtain high-dimensional fusion features. These high-dimensional fusion features are used to capture the intrinsic connections between the dynamic objects in the spectral and spatiotemporal dimensions. Local regions and / or feature points in the high-dimensional fusion features are constructed as corresponding graph nodes. The spatial structural relationships and temporal motion patterns between graph nodes are learned through graph convolution operations. Based on the motion patterns, dynamic objects in the pool environment are initially labeled. The Conditional Random Field algorithm is used to perform pixel-level segmentation of visual image data under different bands based on preliminary annotation information. The boundary features of the dynamic object are obtained by the spatial position relationship and feature similarity between each pixel. The boundary features are further evaluated by combining event stream data from the event camera. By using the boundary features determined through auxiliary judgment, the dynamic object is deleted from the multi-source image data to obtain the environmental image data.
8. The method according to claim 7, characterized in that, The process of constructing a swimming pool environment model by removing the dynamic objects from the environmental image data includes: Using the Sparse Bundle Adjustment (SBA) algorithm, a 3D map of the pool environment is constructed based on the feature point information in the environmental image data, thus obtaining the pool environment model. Based on the incremental point cloud fusion algorithm, multi-source image data collected in real time is fused into the swimming pool environment model. The newly added feature point information is added to the 3D map through feature matching and point cloud registration, so as to realize the real-time update of the swimming pool environment model.
9. The method according to claim 1, characterized in that, The process involves deeply fusing the activity trajectory model with the pool environment model using a knowledge graph network, and employing quantum machine learning algorithms to predict the pool robot's operating state, thereby obtaining a panoramic model that includes the predicted activity trajectory and predicted operating status of the pool robot. Using knowledge graph technology, a knowledge graph network containing semantic information about the pool environment and robot movement rules is constructed. The robot's historical motion trajectory data and robot motion state information in the activity trajectory model, as well as the static map data and real-time detected obstacle information in the pool environment model, are converted into node feature vectors in a knowledge graph network. The weights of the activity trajectory model and the pool environment model are calculated using the graph attention mechanism (GAT) to obtain a fusion result between the activity trajectory model and the pool environment model. The attention weights of the pool robot and its neighboring nodes in the fusion result are used to reflect the degree of influence of the neighboring nodes on the operating state of the pool robot. The neighboring nodes include: surrounding obstacle nodes and pool wall nodes. The fusion results are mapped to the quantum feature space, and the quantum state superposition and entanglement properties are utilized. Quantum support vector machines or quantum neural networks are used as quantum machine learning algorithms to predict the operating state of the pool robot in the future. The prediction results are integrated with the activity trajectory model and the pool environment model to generate the panoramic model containing the predicted activity trajectory and predicted operation status, which is used to provide a decision-making basis for the path planning and task scheduling of the pool robot.
10. A pool robot localization and trajectory prediction system integrating sonar and vision, characterized in that, The system includes: The first acquisition module, a terahertz sonar device deployed in the robot, is used to collect terahertz sonar data for a target area to detect the distance information between the pool robot and surrounding objects in the target area; where the target area is the aquatic environment in which the pool robot is located. The first construction module is used to construct the activity trajectory model of the pool robot based on the terahertz sonar data; The second acquisition module, consisting of visual sensors deployed around the pool and in the aquatic environment, is used to acquire multi-source image data containing underwater, surface, and pool environment changes in the target area using a multispectral and event camera collaborative acquisition strategy. The second construction module is used to use a deep learning-based dynamic SLAM algorithm to identify and process dynamic objects in the pool environment in real time based on the multi-source image data, and to construct a pool environment model by removing the dynamic objects from the environmental image data; the dynamic objects include swimmers or moving objects in the environment. The fusion module is used to deeply fuse the activity trajectory model with the pool environment model through a knowledge graph network, and to use a quantum machine learning algorithm to predict the operating status of the pool robot, thereby obtaining a panoramic model that includes the predicted activity trajectory and predicted operating status of the pool robot. The display module is used to create and update the digital twin model corresponding to the panoramic model in real time, so as to realize the real-time simulation and visualization of the operating status and environmental changes of the pool robot, which facilitates the maintenance and management decisions of the pool robot.
Citation Information
Patent Citations
Multi-sensor fusion underwater robot positioning method and system
CN119178430A
Movement control method and device for swimming pool cleaning robot and swimming pool map construction method and device
CN119937553A