Intelligent cleaning control method, system and equipment for unmanned aerial vehicle
By constructing a collaborative control model and a deep reinforcement learning agent, the flight control and cleaning parameters of the UAV cleaning system were optimized, solving the problems of low operating efficiency and high energy consumption of tethered UAVs in dynamic environments, and achieving efficient and safe cleaning results.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- JIANGSU XINLIT TECHNOLOGY (GROUP) CO LTD
- Filing Date
- 2026-01-26
- Publication Date
- 2026-05-05
AI Technical Summary
Existing tethered drone cleaning systems suffer from low operating efficiency, high energy consumption, and difficulty in coordination when facing dynamic and complex environments. They are unable to respond in real time to interference factors such as sudden changes in wind speed and direction, changes in facade geometry, and uneven distribution of stains, resulting in safety hazards and inconsistent cleaning effects.
A collaborative control model is constructed, which is combined with a deep reinforcement learning agent. By using real-time environmental state and equipment state variables, the flight control, cleaning execution and mooring cable management of the UAV are optimized. A reward function is designed with a comprehensive optimization objective to achieve adaptive decision-making for action space commands.
It realizes global adaptive collaborative control of the drone cleaning system in dynamic environments, reduces energy consumption, improves cleaning effect and operation safety, and ensures high efficiency and stability of the cleaning process.
Smart Images

Figure CN121979259A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of cleaning control technology, and more specifically to intelligent cleaning control methods, systems and equipment for unmanned aerial vehicles (UAVs). Background Technology
[0002] With the acceleration of urbanization, the demand for cleaning the exterior walls of high-rise buildings is increasing. Traditional manual methods, such as suspended platforms or spider-man operations, are not only inefficient and costly but also pose significant safety risks. The development of drone technology provides an innovative solution for high-altitude operations. In particular, cleaning drones using tethered power supply technology, which continuously deliver power and cleaning media through ground base stations, can theoretically achieve long-term, large-scale continuous operations, demonstrating great potential in terms of efficiency and safety.
[0003] However, existing technologies still face significant challenges in practical applications. Current mainstream tethered drone cleaning systems typically employ separate or simple serial control strategies for flight control, high-pressure cleaning execution, and tether cable management. For example, flight paths are often based on preset static trajectory planning, while cleaning parameters (such as water pressure and flow rate) are set to fixed values based on experience or subject to simple segmented adjustments. This control mode appears rigid and lacks adaptability in dynamic and complex high-altitude facade work environments. It cannot respond in real-time and collaboratively to interference factors such as sudden changes in wind speed and direction, changes in facade geometry, and uneven dirt distribution, leading to a series of problems such as high energy consumption, potential safety hazards due to tether cable tension fluctuations, inconsistent cleaning results, and overall operational efficiency falling short of expectations. Consequently, it struggles to cope with unforeseen complex working conditions. Summary of the Invention
[0004] This application provides an intelligent cleaning control method, system, and equipment for unmanned aerial vehicles (UAVs), which solves the technical problems of low operating efficiency, high energy consumption, and difficulty in coordination caused by the use of static or segmented control strategies when existing cleaning UAVs face dynamic and complex environments.
[0005] The first aspect of this application provides an intelligent cleaning control method for unmanned aerial vehicles (UAVs), the method comprising: A collaborative control model is constructed based on a comprehensive optimization objective, which includes at least minimizing total operational energy consumption and optimizing cleaning effect per unit area. The input variables of the collaborative control model include real-time environmental state variables and operational equipment state variables. Based on the collaborative control model, a deep reinforcement learning agent is constructed and trained. The state space of the deep reinforcement learning agent maps to the input variables, and its action space includes UAV flight control parameters, cleaning actuator parameters, and tether cable management parameters. Its reward function is designed based on the comprehensive optimization objective and preset constraints. During the cleaning operation, the deep reinforcement learning agent performs decision analysis based on the real-time state space information and outputs corresponding action space commands. Based on the action space commands, the UAV and its cleaning actuator are controlled to perform collaborative operations, and updated state space information and reward information are obtained after the actions are executed. The decision-making strategy of the deep reinforcement learning agent is continuously fine-tuned.
[0006] A second aspect of this application provides an intelligent cleaning control system for unmanned aerial vehicles (UAVs), the system comprising: Model Building Module: Constructs a collaborative control model based on the comprehensive optimization objective, which includes at least minimizing total operational energy consumption and optimizing cleaning effect per unit area. The input variables of the collaborative control model include real-time environmental state variables and operational equipment state variables. Agent Training Module: Based on the collaborative control model, constructs and trains a deep reinforcement learning agent. The state space of the deep reinforcement learning agent maps to the input variables, and its action space includes UAV flight control parameters, cleaning actuator parameters, and tether cable management parameters. Its reward function is designed based on the comprehensive optimization objective and preset constraints. Decision Analysis Module: During the cleaning operation, the deep reinforcement learning agent performs decision analysis based on the real-time state space information and outputs corresponding action space commands. Collaborative Operation Module: Based on the action space commands, controls the UAV and its cleaning actuator to perform collaborative operations, and obtains updated state space information and reward information after action execution, continuously fine-tuning the decision strategy of the deep reinforcement learning agent.
[0007] A third aspect of this application provides an electronic device, comprising: a memory for storing executable instructions; and a processor for implementing the intelligent cleaning control method for unmanned aerial vehicles provided in this application when executing the executable instructions stored in the memory.
[0008] One or more technical solutions provided in this application have at least the following technical effects or advantages: A collaborative control model is constructed based on the comprehensive optimization objectives, which include at least minimizing total operational energy consumption and optimizing cleaning effect per unit area. The input variables of the collaborative control model include real-time environmental state variables and operational equipment state variables. Next, based on the collaborative control model, a deep reinforcement learning agent is constructed and trained. The state space of the deep reinforcement learning agent maps to the input variables, and its action space includes UAV flight control parameters, cleaning actuator parameters, and tether cable management parameters. Its reward function is designed based on the comprehensive optimization objectives and preset constraints. Furthermore, during the cleaning operation, the deep reinforcement learning agent performs decision analysis based on the real-time state space information and outputs corresponding action space commands. Finally, based on the action space commands, the UAV and its cleaning actuator are controlled to perform collaborative operations, and updated state space information and reward information are obtained after the actions are executed, allowing for continuous fine-tuning of the deep reinforcement learning agent's decision-making strategy. This invention addresses the technical challenges of low operational efficiency, high energy consumption, and difficulties in coordination among existing cleaning drones operating in dynamic and complex environments due to the use of static or segmented control strategies. It achieves global, adaptive, and collaborative control of drone flight control, cleaning parameters, and tether cable management through deep reinforcement learning, thereby reducing energy consumption and improving cleaning effectiveness and operational safety. Attached Figure Description
[0009] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0010] Figure 1 This is a schematic flowchart of an intelligent cleaning control method for unmanned aerial vehicles provided in an embodiment of this application.
[0011] Figure 2 This is a schematic diagram of the intelligent cleaning control system for drones provided in an embodiment of this application.
[0012] Figure 3 This is a schematic diagram of the structure of an exemplary electronic device of this application.
[0013] Figure labeling: Model building module 11, Agent training module 12, Decision analysis module 13, Collaborative operation module 14, Processor 21, Memory 22, Input device 23, Output device 24. Detailed Implementation
[0014] To further illustrate the technical means and effects of the present invention in achieving its intended purpose, the following detailed description of the specific implementation methods, structures, features, and effects of the present invention, in conjunction with the accompanying drawings and preferred embodiments, is provided below.
[0015] Example 1, as Figure 1 As shown in the embodiment of this application, an intelligent cleaning control method for unmanned aerial vehicles (UAVs) is provided, wherein the method includes: A collaborative control model is constructed based on the comprehensive optimization objectives, which include at least minimizing total operating energy consumption and optimizing cleaning effect per unit area. The input variables of the collaborative control model include real-time environmental state variables and operating equipment state variables.
[0016] In one embodiment, a computable and solvable collaborative control model for UAV cleaning operations is established, with the comprehensive optimization objectives of minimizing total operational energy consumption and optimizing cleaning effect per unit area. Specifically, the cleaning task is defined as a continuous control process of completing coverage cleaning within a given target facade area, and energy consumption and cleaning effect are quantified uniformly. Total operational energy consumption characterizes the sum of the electrical energy consumption of the UAV and the power consumption of the cleaning actuators throughout the entire cleaning cycle, including at least the energy consumption of the flight propulsion system, attitude stabilization control, and high-pressure water pump / spray gun system. Cleaning effect per unit area characterizes the effective cleanliness of a unit cleaning area, including at least the effective cleaned area per unit time and the proportion of repeated cleaning. During the model construction, real-time environmental state variables and operational equipment state variables are defined as input variables. Real-time environmental state variables include wind speed, wind direction, and three-dimensional surface geometric features; operational equipment state variables include at least the aircraft battery level, power supply mode, and input voltage. Subsequently, the aforementioned input variables are synchronously collected, filtered, and normalized at a unified time step to form a state vector. Within the Markov decision process framework, the coupling relationship between each input variable and the comprehensive optimization objective is modeled and quantified in each control cycle, thus constructing a collaborative control model for UAV intelligent cleaning. This model is used for subsequent flight control, cleaning parameter adjustment, and mooring cable deployment and tension control, thereby providing clear and executable optimization objectives, state definitions, and constraint boundaries for the training and online decision-making of the deep reinforcement learning agent.
[0017] Furthermore, the collaborative control model for intelligent cleaning using drones includes: The collaborative control model is modeled, and a Markov decision process framework guided by the comprehensive optimization objective is established. Under the Markov decision process framework, the coupling relationship between each input variable and the comprehensive optimization objective in the collaborative control model is initialized and modeled using historical operation data and simulation data.
[0018] Preferably, during model construction, the entire cleaning operation process is first discretized according to a fixed control period Δt, which can be set to 50ms to 200ms; for example, 100ms can be taken as a decision time step. At each time step t, the system is abstracted as a closed-loop decision process of state-action-reward-state transition. Subsequently, the state space S is defined. The state vector S... t The state vector is composed of real-time environmental state variables and operational equipment state variables, and is represented in the form of a fixed-length vector. For example, the environmental state variables include real-time wind speed and direction information (2D), facade normal vector and local curvature features (3-6D), stain coverage and pollution level (2D), and mooring cable tension and deployment length (2D). The operational equipment state variables include the drone's remaining battery power and instantaneous power consumption (2D), motor temperature (4D), and high-pressure water pump operating pressure and flow rate (2D). After normalization, the above dimensions form a fixed-dimensional state vector S. t (25-40 dimensions) are used to fully characterize the overall operational status of the UAV cleaning system at the current moment. Next, the action space A and action vector A' are defined. t The continuous action space represents the combination of control parameters that the system can execute within the current decision cycle, such as the UAV's forward / lateral velocity, pitch and roll angle adjustments, water pump pressure setpoint, nozzle opening, tether cable tension, and cable deployment / retraction speed. Each action component is limited within a preset safety physical boundary to ensure the executability and safety of the action output. Next, the state transition process is defined, with the state transition function describing the state in the current state S. t Next, execute action A t Then, the system evolves to the next state S. t+1 The dynamic process of this transition implicitly includes the flight dynamics of the UAV, the effects of aerodynamic disturbances, the mechanical response of the tethered cable, and the time-varying nature of the cleaning effect. Under the Markov assumption, the next state S... t+1 Depends only on the current state S t With current action A t This is independent of the state at earlier times. Next, the control effect of a single decision cycle is quantified with the comprehensive optimization objective as the core, and a reward function R is constructed. t The reward function employs a weighted combination, including a cleaning efficiency reward, an energy efficiency reward, and a safety constraint penalty. This approach ensures that the reward function directly reflects the degree to which the overall optimization objective is achieved, thereby guiding the decision-making strategy towards minimizing total operational energy consumption and optimizing cleaning performance per unit area. This completes the construction of the Markov decision process framework oriented towards the overall optimization objective.
[0019] After establishing the Markov decision process framework, to avoid ineffective exploration or unstable convergence of the deep reinforcement learning agent in the initial training phase, the system uses historical operation data and simulation data to initialize the coupling relationship between each input variable and the comprehensive optimization objective in the cooperative control model. Specifically, historical UAV cleaning operation data is first collected and organized, which includes at least the state sequences S recorded in multiple actual cleaning tasks. t The corresponding action sequence A t The data includes energy consumption statistics and cleaning effect evaluation results for each time step. These data are sliced in chronological order to form state-action-result sequence samples of length T, for example, 100-300 time steps. Simultaneously, a 3D facade model, wind field model, and UAV-tethered cable-cleaning device dynamic model consistent with the real operation scenario are constructed in a high-fidelity digital twin simulation environment. By changing wind speed, facade curvature, dirt distribution density, and cleaning parameter combinations in the simulation environment, simulation sample data covering a wider range of operating conditions is generated to compensate for the deficiencies of real historical data under extreme or rare conditions. Subsequently, a unified data preprocessing workflow is applied to the historical operation data and simulation data, including time alignment, outlier removal, and feature normalization. The processed data is then used to fit and model the initial mapping relationship between state-action-result. For example, a multi-layer feedforward neural network or a function approximation model guided by physical constraints can be used for a given state S. t And Action A t At the next time step, the system predicts the changes in energy consumption and cleaning effect, thus obtaining an initial approximate state transition model and reward function parameter estimates. During this initialization modeling process, the model parameters are initialized with small-amplitude randomization, allowing the model to reflect fundamental physical and empirical laws such as "increased wind speed leads to increased energy consumption," "increased jet pressure improves cleaning efficiency within a certain range but also increases energy consumption," and "changes in tether tension affect flight stability" from the early stages of training. This approach provides a reasonable prior structure and stable optimization direction for the subsequent policy learning of the deep reinforcement learning agent, significantly improving the training convergence speed and the feasibility and safety of the policy in real cleaning operations.
[0020] Furthermore, the real-time environmental state variables include at least real-time wind speed and direction information obtained through airborne sensors, three-dimensional geometric features and stain distribution information of the work surface identified through visual sensors, and real-time tension information of the mooring cable obtained through tension sensors; the work equipment state variables include at least the UAV battery power and motor temperature, water tank remaining capacity, and high-pressure water pump working pressure and flow rate.
[0021] Optionally, real-time environmental state variables include at least real-time wind speed and direction information, three-dimensional geometric features and dirt distribution information of the work surface, and real-time tension information of the mooring cable. The real-time wind speed and direction information is obtained by using an airborne wind speed and direction sensor at a fixed sampling frequency, such as 10Hz to 50Hz, to collect the wind speed and direction near the drone's location in real time. The three-dimensional geometric features of the work surface are obtained by using a vision sensor mounted on the drone, such as a monocular camera, binocular camera, or depth camera, to continuously image and perceive the depth of the target work surface, combined with structured light, stereo matching, or visual- The inertial fusion algorithm reconstructs the working surface, and these three-dimensional geometric features include parameters such as surface normal vectors, tilt angles, curvature changes, and boundary contours. Stain distribution information is extracted by using a camera-pre-defined convolutional neural network to identify and classify stain areas on the working surface, including descriptive features such as stain coverage ratio and contamination level. Real-time tension information of the tether cable is obtained by tension sensors deployed on the tether cable deployment / retraction mechanism or the tether cable itself, which collects data on the stress on the tether cable under current flight attitude and cleaning load conditions, characterizing whether the tether cable is within a safe working range. The operating equipment status variables include at least the UAV battery level and motor temperature, water tank remaining capacity, and high-pressure water pump operating pressure and flow rate. The UAV battery level is read in real-time by the onboard power management module; the motor temperature is collected by temperature sensors installed on each power motor or ESC, used to assess motor load level and thermal safety status; the water tank remaining capacity is obtained by a level sensor; and the high-pressure water pump operating pressure and flow rate are collected by pressure and flow sensors installed at the high-pressure water pump outlet or in the pipeline, characterizing the current cleaning intensity and energy consumption level. These data enable the system to obtain a complete, accurate, and physically meaningful state description in each decision cycle, providing a reliable data foundation for subsequent decision analysis and collaborative control of deep reinforcement learning agents.
[0022] Based on the aforementioned collaborative control model, a deep reinforcement learning agent is constructed and trained. The state space of the deep reinforcement learning agent maps to the input variables, and its action space includes UAV flight control parameters, cleaning actuator parameters, and tether cable management parameters. Its reward function is designed based on the aforementioned comprehensive optimization objective and preset constraints.
[0023] In one embodiment, a deep reinforcement learning agent is constructed and trained based on a cooperative control model aimed at minimizing total operational energy consumption and optimizing cleaning effect per unit area. This agent enables real-time cooperative optimization of flight control, cleaning execution, and mooring cable management during cleaning operations. Specifically, the input variables of the cooperative control model are directly mapped to the agent's state space S. tIn each decision cycle, the real-time environmental state variables, obtained from airborne and ground stations and processed through time alignment and normalization, together with the operational equipment state variables, constitute the current state vector, which is then input into the policy network. The agent outputs its action space A based on this state vector. t A set of continuous control commands includes UAV flight control parameters, cleaning actuator parameters, and tether cable management parameters. The UAV flight control parameters include at least the distance adjustment Δd relative to the normal direction of the work surface and the heading speed fine-tuning Δv parallel to the work surface. y With lateral velocity fine adjustment Δv x The cleaning actuator parameters should include at least the pressure setpoint P of the high-pressure water jet. set With the flow rate setpoint Q set And the mixing ratio of cleaning agents C mix When a rotary brush is used, the output brush rotation speed ω is... set With downward pressure F set The tethered cable management parameters should at least include the cable deployment / retraction speed request command sent to the tethered base station. tether Continuous values or discrete values ranging from -2 to 2 m / s are used to coordinate cable release / reeling during changes in flight position or increased wind disturbance to maintain tension within a safe range and reduce the impact on flight stability. To guide the agent in learning strategies that satisfy the comprehensive optimization objective and preset constraints, the reward function is calculated in each decision cycle according to the structure of "cleaning efficiency gain - energy consumption cost - constraint penalty". The training phase can be conducted in a digital twin simulation environment. Through interactive exploration between the agent and the environment, the reward function for each decision cycle (S) is calculated. t A t ,R t ,S t+1 The data is stored as empirical data, and the policy network parameters are iteratively updated using policy gradient or Actor-Critic algorithms. In actual cleaning operations, the agent receives the latest status online at a set frequency and outputs control commands. At the same time, it continuously fine-tunes the policy using real-time feedback rewards and updated status, thereby achieving adaptive and coordinated control of factors such as wind speed changes, differences in stain distribution, facade geometric complexity, and changes in mooring forces. This ensures that low energy consumption and high-quality cleaning results are achieved while meeting safety constraints.
[0024] Furthermore, based on the comprehensive optimization objective and preset safety and performance constraints, the reward function is designed and initialized, including: Positive reward weights are assigned to cleaning efficiency indicators, including effective cleaning area per unit time and stain removal rate based on visual feedback; positive reward weights are also assigned to energy efficiency indicators, including flight trajectory smoothness, motor torque output stability, and high-pressure water pump operating point efficiency; negative rewards are assigned to behaviors that violate preset safety and efficiency constraints, including predicted collision risk, exceeding tether tension limits, cleaning effect failing to reach thresholds, and excessive redundancy in operation path planning.
[0025] Preferably, the reward function is constructed using a component-based calculation and weighted summation method, and is calculated in real time within each decision cycle Δt. This is used to provide quantitative feedback on the effect of the current action decision of the deep reinforcement learning agent. This reward function can be expressed as: ,in, This indicates a cleaning performance bonus. This indicates an energy efficiency bonus item. This indicates safety and performance constraints and penalties. The positive reward weight for cleaning efficiency incentive items. The positive reward weight for energy efficiency incentive items. The negative reward weights for safety and efficiency constraints can be preset and adjusted according to the job type or stage. Within each decision cycle, for cleaning efficiency rewards, the effective cleaning area per unit time is first calculated based on the actual trajectory of the drone on the work surface and the effective range of the cleaning actuator. That is, based on the projected displacement Δs of the drone along the work surface between two adjacent time steps t-1 and t. t Combined with the current spray width or brush effective contact width W t The system calculates the new cleaning area through multiplication. When an area is detected as already sufficiently cleaned or exhibiting significant overlapping coverage, this area is reduced or zeroed out. The processed new effective cleaning area is then normalized to a preset baseline cleaning efficiency and its minimum value is calculated by taking the minimum value from 1 to obtain the area bonus component. Simultaneously, based on before-and-after cleaning images captured by a visual sensor, the stain removal rate of the current cleaning area is evaluated. This is achieved by calculating the difference between the proportion of stain pixels in the target area in the before-and-after image and the corresponding proportion in the after-and-after image, and then dividing this difference by the proportion of stain pixels in the target area in the before-and-after image to obtain the stain removal rate. If the stain removal rate is higher than a preset minimum effective stain removal threshold, a stain removal bonus component is calculated proportionally; otherwise, it is normalized to the preset minimum effective stain removal threshold and its minimum value is calculated by taking the minimum value from 1 to obtain the stain removal bonus component. Finally, the area bonus component and the stain removal bonus component are weighted and summed to form the cleaning efficiency index for the current decision cycle, which serves as the result of the cleaning efficiency bonus item.
[0026] For energy efficiency rewards, three dimensions are quantified: flight trajectory smoothness, motor output stability, and cleaning actuator efficiency. Flight trajectory smoothness is evaluated by calculating the rate of change of the UAV's speed and acceleration within adjacent time steps. Specifically, the absolute values of the speed and acceleration changes are added together, divided by a preset maximum smoothness, and then the quotient is subtracted from 1 to obtain the trajectory smoothness. When the speed or acceleration change exceeds a preset threshold, the reward gradually decreases until it reaches zero. Motor torque output stability is evaluated by monitoring the time series variance of the motor torque. When the motor torque variance is lower than a preset stability threshold within a window of length N, a positive reward is given. Specifically, the ratio of the motor torque variance to the preset motor torque variance is calculated, and then this ratio is subtracted from 1 to obtain the positive reward. High-pressure water pump operating point efficiency is evaluated by comparing the current pressure-flow combination with the high-efficiency range of the water pump efficiency curve. When the current operating point is within the high-efficiency range, a full reward is given; otherwise, the reward decreases linearly according to the degree of deviation. Finally, the three calculation results are weighted to obtain the energy efficiency index, which serves as the result of the energy efficiency reward item.
[0027] For safety and efficiency constraint penalties, a short-term prediction (e.g., 0.5–1 second) of the current trajectory is performed using a reserved long short-term memory network. When the predicted minimum distance between the drone and the work surface is less than the safe distance threshold, a collision risk penalty is triggered. This is achieved by taking the maximum value of 0 and the difference between the safe distance threshold and the current minimum distance, and then multiplying it by the corresponding preset proportional coefficient. Similarly, when the real-time tension of the tethering cable exceeds the safe upper limit, a tension penalty is applied. This is also achieved by taking the maximum value of 0 and the difference between the real-time tension and the safe upper limit, and then multiplying it by the corresponding preset proportional coefficient. Furthermore, when the cleaning effect in a certain area is lower than the preset minimum decontamination threshold, or the flight distance required per unit effective cleaning area is significantly higher than the reference value, indicating redundancy in path planning, corresponding insufficient cleaning penalties and path redundancy penalties are applied. Finally, all penalties are summed to form the total constraint penalty, which serves as the result of the safety and efficiency constraint penalties. Through the above process, without introducing any manual rule switching, cleaning quality, energy efficiency and safety constraints are uniformly incorporated into the same reward function framework, enabling the deep reinforcement learning agent to autonomously learn the optimal cooperative control strategy under complex environments and multiple constraints during long-term interaction and training.
[0028] Furthermore, the training of the deep reinforcement learning agent is completed in a high-fidelity digital twin simulation environment, including: A digital twin environment is constructed, which includes a 3D model of the target operation scenario, a fluid dynamics wind field model, and an equipment dynamics model. In the digital twin environment, the deep reinforcement learning agent interacts with the simulation environment to explore and experiment, and its policy network is pre-trained offline until the policy converges to obtain an initial agent model.
[0029] Preferably, environmental data of the work area is first collected using LiDAR and multi-view cameras. These sensors can capture detailed geometric features of the target building or cleaning object, such as surface morphology, curvature, local slope, and the location of potential obstacles. The collected point cloud data is then converted into a high-precision 3D model using software tools such as AutoCAD and Rhino, and further meshed to form a 3D model of the target work scene suitable for simulation calculations. In high-altitude cleaning operations, wind speed, wind direction, and wind turbulence characteristics significantly affect the flight stability and cleaning effect of the UAV. Therefore, computational fluid dynamics (CFD) technology is used to simulate wind speed, wind direction, and turbulence effects within the work area. In this process, a fluid computational domain is constructed based on the building morphology and wind speed data in the environment. This domain is typically discretized using the finite volume method or finite element method, and then simulated using CFD simulation software such as ANSYS Fluent and OpenFOAM. During the calculations, the Navier-Stokes equations and turbulence models are used to describe the motion and turbulence characteristics of the fluid, simulating the interaction between wind and buildings under different meteorological conditions. Through these calculations and simulations, a final fluid dynamics wind field model is constructed. In addition, equipment dynamics models corresponding to the UAV, cleaning execution equipment, and tethering cable are constructed. These models primarily describe the flight dynamics of the UAV, the operating state of the cleaning execution equipment, and the physical behavior of the tethering cable. The UAV's flight dynamics model includes motion equations based on the Newton-Euler equations, involving calculations of thrust, aerodynamic drag, and attitude control. Real-time flight attitude information is obtained through onboard sensor data (such as accelerometers and gyroscopes), and the UAV's position changes in three-dimensional space are calculated using this data. The dynamics model of the cleaning execution mechanism focuses on the hydrodynamic characteristics of the water pump, changes in jet pressure and flow rate, and the impact of these changes on the cleaning effect. The dynamics model of the tethering cable considers the cable tension, length changes, and the interaction between the cable and the environment, simulating their impact on aircraft stability to ensure the aircraft maintains a safe flight attitude under different operating conditions. After constructing the 3D model of the target operation scenario, the fluid dynamics wind field model, and the equipment dynamics model, the 3D model of the target operation scenario serves as the spatial foundation layer of the digital twin environment. The world coordinate system, scale units, and gravity direction are uniformly defined in the simulation engine, ensuring that the drone, the object being cleaned, and the mooring cable are all within the same 3D spatial reference frame. Based on this, the fluid dynamics wind field model is embedded as the environmental disturbance layer in the 3D space, and the equipment dynamics model serves as the behavior execution layer, achieving closed-loop coupling and forming a complete high-fidelity digital twin simulation environment.
[0030] After constructing the digital twin simulation environment, the system trains a deep reinforcement learning agent. Specifically, in the digital twin environment, the initial state of the environment is first passed to the agent as the basis for its first interaction with the environment. The initial state vector of the environment contains all environmental state variables and device state variables. Based on this state information, the agent selects an appropriate action through a policy network. The action space includes the UAV's flight control parameters, the cleaning actuator's operating parameters, and the tether cable management parameters. When the agent selects and executes an action, the digital twin environment updates its current state based on the result of that action and calculates the reward for that decision. This reward signal provides feedback to the agent on the quality of the current decision; a larger reward indicates a more favorable action decision, while a smaller reward indicates a poorer decision that may require adjustment. Next, the agent stores the current state, action, reward, and the new environmental state as empirical data (S). t A t ,R t ,S t+1These experiential data are stored in an experience pool. Each new accumulation of experiential data provides the agent with more environmental feedback, helping it accumulate more policy experience through continuous interaction. This experiential data forms the foundation for the agent's learning; the agent repeatedly optimizes its policy using these samples. The agent replays samples from the experience pool and utilizes deep reinforcement learning algorithms, such as Q-learning and DQN, to optimize the parameters of the policy network using gradient descent. During this process, the agent calculates the error between the action value predicted by the current policy network and the target value calculated from the experiential data, then propagates the error back into the network through backpropagation, gradually updating the network parameters so that the policy network can output more accurate action choices in the next decision. This process is repeated in each training iteration, gradually causing the agent's policy network to converge over time, eventually reaching a stable policy model. During training, the agent's policy gradually improves, and its decision-making ability continuously strengthens. As training progresses, the agent gradually becomes better able to adapt to complex environments and make efficient decisions, ensuring optimized cleaning results, reduced energy consumption, and safe operation. The training process typically involves several iterations. Each iteration allows the agent to gain more experience in the environment and further optimize its policy until the agent's policy network converges in the digital twin environment, enabling it to stably execute tasks and achieve optimal cleaning results and energy efficiency. After training, the agent is tested on unseen simulation data to verify its generalization ability. If the agent can maintain stable cleaning results and ensure system safety and energy consumption control during testing, the training is considered converged, and the agent's policy network has strong adaptive capabilities, making it usable in practical applications. Ultimately, the resulting agent can serve as the initial decision model for actual operations, incorporated into the real-time decision control of UAV cleaning tasks to ensure high-quality cleaning and high-energy-efficiency operation of the UAV.
[0031] During the cleaning operation, the deep reinforcement learning agent performs decision analysis based on the real-time state space information and outputs corresponding action space instructions.
[0032] In one embodiment, during the cleaning operation, a deep reinforcement learning agent operates online with a fixed control cycle to output directly executable action space commands based on real-time state information, enabling collaborative decision-making for flight control, cleaning execution, and mooring cable management. Specifically, at the beginning of each decision cycle Δt, the current state space information is collected and extracted based on the initial coverage path, and timestamp alignment and fusion processing are completed. Subsequently, the obtained current state space information is input into the deep reinforcement learning agent, which performs coupling relationship deduction and multi-step prediction based on the trained knowledge, and outputs a corresponding set of action space commands. This set of action space commands includes at least parameter commands for flight control, parameter commands for cleaning execution, and parameter commands for mooring cable management. The parameter commands for flight control include a distance adjustment Δd along the normal to the working surface and a heading speed fine-tuning Δv parallel to the working surface. y With lateral velocity fine adjustment Δv x The parameter commands used for cleaning execution include the pressure setpoint P of the high-pressure water jet. set With the flow rate setpoint Q set And the mixing ratio of cleaning agents C mix Brush rotation speed ω set With downward pressure F set ; Parameter commands used for mooring cable management include cable reel-in / deel-out speed requests u tether and target tension setpoint T set To ensure the execution of commands and compliance with safety boundaries, the output phase performs constraint projection and amplitude limiting processing on the motion commands. Specifically, each motion component is clipped to a preset physical range and undergoes secondary verification based on safety constraints such as predicted collision distance, tension upper limit, and attitude angle upper limit. When a risk state is detected, such as when the predicted collision distance is less than a threshold or the tension is close to the upper limit, a safety-first motion correction strategy is triggered, correcting the output commands towards safer directions such as moving away from obstacles, reducing speed, minimizing jet reaction, or adjusting the cable deployment / retraction direction. Finally, the verified motion space commands are sent as control commands to the flight control and cleaning actuator controller and the tethering base station controller. These controllers execute the corresponding control actions within the current decision cycle and continue to collect updated status information in the next decision cycle for the next decision. This enables real-time adaptive collaborative control of complex factors such as wind disturbance, surface geometric changes, differences in dirt distribution, and changes in tethering forces, ensuring a continuous, stable, safe, and reliable cleaning operation.
[0033] Furthermore, the method also includes an initial coverage path generation step: Before the operation begins, based on the 3D point cloud model of the target operation area and combined with the preset standard cleaning width, an initial coverage path is generated that can completely cover the target operation area without collision; the initial coverage path is converted into an initial state space sequence, which serves as the reference state input for the first decision of the deep reinforcement learning agent.
[0034] Preferably, before the cleaning operation begins, a 3D point cloud model of the target work area is first acquired. This 3D point cloud model can be acquired using a lidar, depth camera, or multi-view vision reconstruction system mounted on a drone. The point cloud must contain at least the spatial coordinate information (x, y, z) of each sampling point, as well as the normal vector information (n). x ,n y ,n z The data includes reflection intensity and texture features. The raw point cloud data is then preprocessed, including removing outlier noise points using statistical filtering or radius filtering algorithms, unifying the point cloud density to a preset resolution (e.g., 5mm-20mm) using voxel mesh downsampling to reduce subsequent computational complexity, and mapping the point cloud coordinates to a local coordinate system referenced to the work facade, giving the normal, vertical, and parallel directions clear physical meaning. Next, the preprocessed 3D point cloud model undergoes work area segmentation and boundary extraction. During this process, based on the consistency and spatial continuity of the point cloud normal vectors, region growing or surface segmentation is performed to extract continuously cleanable sub-regions. Within each sub-region, the boundary contour is extracted using 2D projection and convex / concave hull calculation methods. Simultaneously, non-cleanable areas or obstacle areas, such as window frames, protruding structures, and equipment mounting parts, are identified and labeled, and preset safety expansion distances are set for them for subsequent path collision detection. Based on this, and considering the physical parameters of the cleaning actuator, a standard cleaning width W is set. This standard cleaning width is determined according to the effective coverage width of the nozzle spray or the diameter of the rotating brush, and can be set to a fixed value, such as 0.4m, based on experience or experimental calibration results. Then, using this standard cleaning width as a path spacing constraint, the work sub-area is decomposed into strips. That is, a set of equidistant reference scan lines is generated in the main direction of the work surface, such as the vertical or horizontal direction, ensuring that the distance between adjacent scan lines is no greater than the standard cleaning width W. This theoretically guarantees that any point is effectively cleaned at least once.
[0035] Next, a specific 3D coverage path point sequence is generated along each scan line. Specifically, the path is discretely sampled along the scan line direction at a fixed step size Δs, such as 5cm to 20cm. Each sampling point corresponds to a path node, and each path node is assigned 3D spatial coordinates, the corresponding surface normal direction, and the desired working normal distance d. refThe path generation process involves collision detection and safety assessment for each path node. This involves calculating the minimum distance from the node's location to the obstacle region and boundary contour. If this distance is less than a preset safety threshold, the node is deemed infeasible, and the scan line is locally truncated, offset, or rearranged. If local adjustments still fail to meet safety constraints, a transition node is inserted into the path or the region is skipped, ensuring the generated initial coverage path is spatially collision-free and continuously executable. After these steps, an initial coverage path arranged in the job order is obtained. This initial coverage path is stored as an ordered sequence of path nodes, each containing spatial position, normal information, reference job distance, and desired motion direction information. The initial coverage path is then further converted into an initial state space sequence, used as the reference state input for the deep reinforcement learning agent's initial decision. Specifically, the path node sequence is mapped to a reference state sequence in job time order. Each reference state corresponds to a theoretical job time and includes at least a reference position vector, a reference heading vector, a reference normal distance, and the desired displacement direction and step size information between the node and its predecessor and successor nodes. The aforementioned reference state sequence is stored as static reference data in the airborne computing unit or ground control station. When the cleaning operation begins, the deep reinforcement learning agent aligns the current state space information, acquired in real-time from this initial state space sequence, with the current reference state during the initial decision cycle and subsequent decision cycles. It calculates the deviation information between the current position and the reference path nodes, including distance error along the normal direction, lateral offset error parallel to the work surface, and heading angle deviation. This deviation information is added to the collected state space information as input, enabling the agent to gradually introduce adaptive adjustment capabilities based on real-time environmental changes while ensuring coverage integrity. This improves operational stability, coverage consistency, and cleaning efficiency.
[0036] Furthermore, the deep reinforcement learning agent performs decision analysis based on the real-time acquired state space information and outputs corresponding action space instructions, including: The deep reinforcement learning agent receives and processes the fused current state space information, which includes environmental state information, device state information, and deviation information relative to the reference state input. Based on the current state space information, the deep reinforcement learning agent performs coupling relationship deduction and multi-step prediction, and autonomously generates a global operation strategy with the goal of minimizing the current overall cost. The global operation strategy includes a set of real-time UAV flight control parameters, cleaning actuator parameters, and tether cable management parameters. The global operation strategy is output as the action space command.
[0037] Optionally, at the beginning of each decision cycle, the deep reinforcement learning agent receives and processes the fused state space information of the current moment. This state space information is divided into environmental state information, equipment state information vectors, and deviation information relative to the reference state input according to a preset structure. These state information are normalized and concatenated before entering the agent to form the state space information of the current moment. Subsequently, the deep reinforcement learning agent performs forward reasoning on this state space information of the current moment. During the reasoning process, not only is the current state instantaneously mapped, but the coupling relationship between environmental disturbances, flight dynamics response, cleaning execution effect, and changes in tether cable force is also modeled and deduced through the internal hidden layer. Specifically, the agent internally performs multi-step prediction of the state evolution trend for several future time steps, such as 2 to 10 decision cycles. The prediction content includes at least the impact of different flight direction and speed choices on attitude stability and tether tension under the current wind field conditions, the changing trend of different cleaning parameter combinations on cleaning efficiency and energy consumption, and the impact of path selection on future coverage efficiency and safety margin. Based on the above predictions, the agent, with the goal of minimizing the current overall cost, implicitly or explicitly evaluates multiple candidate strategies. This overall cost includes at least total operational energy consumption, cleaning effect per unit area, and safety risk costs. On this basis, the agent autonomously generates a global operational strategy. This global strategy is no longer limited to a preset fixed coverage path but can make overall optimization decisions based on complex coupling relationships perceived in real time. For example, when the current state space information reflects a persistent strong wind in the area ahead, and this wind direction will cause the tether cable tension to rise rapidly, the agent will not only reduce its heading speed or adjust its attitude angle to passively resist the wind, but also actively plan its flight path by comprehensively utilizing the three-dimensional geometric features of the work surface. This allows the drone to fly close to building protrusions, corners, or facade edges, thereby providing temporary guidance or shielding for the tether cable in space, reducing the direct impact of wind load on the drone and tethering system. Simultaneously, the agent will also adjust the cleaning actuator parameters in conjunction with the system. For example, in areas of reduced stability, it may appropriately reduce the water pump pressure or flow rate to reduce jet reaction force and further reduce energy consumption and tension fluctuations. After completing the above deduction and optimization, the agent structures the generated global operation strategy into action space instructions for output. During the action output phase, limiting and safety checks are performed on each parameter to ensure they meet preset physical and safety constraints. Finally, the checked action space instructions are sent. Through this process, the deep reinforcement learning agent can achieve global, adaptive, and forward-looking collaborative control of flight, cleaning, and tethered management in complex, multi-perturbation, and multi-constraint cleaning operation environments.
[0038] Based on the action space instructions, the drone and its cleaning execution mechanism are controlled to perform collaborative operations, and updated state space information and reward information are obtained after the action is executed, so as to continuously fine-tune the decision-making strategy of the deep reinforcement learning agent.
[0039] In one embodiment, after the deep reinforcement learning agent outputs action space commands, the system enters a closed-loop online collaborative control process of command execution, state feedback, reward calculation, and policy fine-tuning to achieve continuous optimization of UAV flight control, cleaning execution, and tether cable management. Specifically, within each decision cycle Δt, the action space commands are parsed into three types of executable control quantities. The first is the UAV flight control command, including distance adjustment along the normal direction of the work surface and heading / lateral velocity fine-tuning parallel to the work surface. This command is input to the flight control system through the position-velocity-attitude control link. The flight control system generates motor outputs and attitude control quantities under the premise of satisfying physical constraints such as attitude angle, angular velocity, and acceleration, so that the UAV stably approaches the work surface along the target trajectory and maintains an effective cleaning distance. The second is the cleaning execution mechanism command, including high pressure... The system sends commands to the cleaning actuator, including water jet pressure setpoints, flow rate setpoints, cleaning agent mixing ratios, brush rotation speed, and downforce. The pump frequency converter, valve opening adjustment module, and mixing metering module adjust the working point in real time to match the cleaning intensity with the current level of contamination and flight stability. Thirdly, the system sends tether cable management commands, including requests for cable reeling / unloading speeds or target tension setpoints to the tether base station. This command is executed by the tether base station, which uses a closed-loop control of the cable reel motor to adjust the cable length and tension, maintaining the tether tension within a safe range and reducing disturbances to the UAV's attitude. During the execution of the action, the system synchronously collects execution feedback, then performs time alignment, filtering, and normalization on this feedback data to construct the updated state space information S after the action is executed. t+1, Simultaneously, a reward calculation is performed on the control effect of this cycle based on a preset reward function. Finally, the experience data generated in this cycle is written into the online experience cache, and the policy network is updated incrementally with small steps at a preset frequency. This allows for continuous adaptation to changes in wind disturbance, differences in dirt, and fluctuations in equipment operating conditions while ensuring safety constraints, thereby achieving stable, efficient, and low-energy collaborative control of drone cleaning operations.
[0040] Furthermore, acquiring updated state space information and reward information after action execution, and continuously fine-tuning the decision-making strategy of the deep reinforcement learning agent, includes: After an action is executed, the deep reinforcement learning agent collects and updates state space information, and generates immediate reward information according to the reward function. The state space information, action space instructions, immediate reward information, and updated state space information for the next decision period are stored together as a set of empirical data. Based on the empirical data, the policy network parameters of the deep reinforcement learning agent are updated in a gradient to optimize the decision strategy for the next round of cleaning.
[0041] Preferably, firstly, at any decision time t, real-time data is collected synchronously, and the current state space information S is constructed. t Subsequently, the deep reinforcement learning agent calculates the current state space information S through forward propagation. t The optimal action combination A t Before issuing the action, a safety constraint projection process is performed on the action combination. This involves tailoring each parameter to a preset physical and safety range to avoid generating control commands that are unexecutable or pose significant risks. Next, the constraint-processed action combination A is... t Synchronous distribution. Within a complete decision cycle Δt, each controller completes the corresponding control action according to the action command and records key feedback data in real time during the execution process, including attitude error, speed error, motor and water pump power consumption, tether tension changes, and image feedback of the cleaning area. After the action is completed, a new state is entered. At this time, various feedback data are collected and processed again to construct and update the state space information S. t+1 and the state S before execution t The results are compared, and the performance of the tasks within the current decision-making cycle is evaluated based on the reward function, generating immediate reward information R. t This instant reward R t It consists of multiple sub-rewards, including cleaning efficiency rewards, energy efficiency rewards, and negative penalties. Then, a complete empirical data point generated within this decision-making cycle is recorded as (S). t A t ,R t ,S t+1These empirical data can be accumulated chronologically and labeled with priority weights for sampling selection during subsequent policy updates. In online update mode, the agent selects the latest or high-priority empirical data from the experience cache after each decision cycle or every few decision cycles to perform a small-step parameter update on the policy network. In offline update mode, the agent samples multiple empirical data points in batches from the experience replay pool when the task ends or when computing resources are sufficient, and performs multiple rounds of iterative training on the policy network. Finally, based on the sampled empirical data, the gradient of the policy network parameters with respect to the objective function is calculated using the backpropagation algorithm, and the network parameters are updated in combination with a preset learning rate and gradient pruning strategy. To prevent online learning from affecting the stability of actual tasks, the update process can be configured with an upper limit on the update magnitude, a policy change threshold, and a rollback mechanism. That is, when the new policy causes a significant decrease in reward or triggers a security risk in the short term, it automatically rolls back to the previous stable policy parameters. Through the continuous accumulation of experience and fine-tuning of strategy network parameters, the deep reinforcement learning agent can continuously correct its decision-making bias in real cleaning operations, and gradually learn better control strategies under different wind fields, different facade structures and different stain distribution conditions, thereby achieving continuous improvement in cleaning operation efficiency, safety and energy consumption performance.
[0042] In summary, the embodiments of this application have at least the following technical effects: A collaborative control model is constructed based on a comprehensive optimization objective, which includes at least minimizing total operational energy consumption and optimizing cleaning effect per unit area. The input variables of the collaborative control model include real-time environmental state variables and operational equipment state variables. Next, based on the collaborative control model, a deep reinforcement learning agent is constructed and trained. The state space of the deep reinforcement learning agent maps to the input variables, and its action space includes UAV flight control parameters, cleaning actuator parameters, and tether cable management parameters. Its reward function is designed based on the comprehensive optimization objective and preset constraints. Then, during the cleaning operation, the deep reinforcement learning agent performs decision analysis based on the real-time state space information and outputs corresponding action space commands. Finally, based on the action space commands, the UAV and its cleaning actuator are controlled to perform collaborative operations, and updated state space information and reward information after action execution are obtained to continuously fine-tune the decision-making strategy of the deep reinforcement learning agent. This invention addresses the technical challenges of low operational efficiency, high energy consumption, and difficulties in coordination among existing cleaning drones operating in dynamic and complex environments due to the use of static or segmented control strategies. It achieves global, adaptive, and collaborative control of drone flight control, cleaning parameters, and tether cable management through deep reinforcement learning, thereby reducing energy consumption and improving cleaning effectiveness and operational safety.
[0043] Example 2, based on the same inventive concept as the intelligent cleaning control method for drones in the foregoing examples, such as... Figure 2 As shown, this application provides an intelligent cleaning control system for drones, wherein the system includes: Model building module 11: Constructs a collaborative control model based on the comprehensive optimization objective, which includes at least minimizing total operational energy consumption and optimizing cleaning effect per unit area. The input variables of the collaborative control model include real-time environmental state variables and operational equipment state variables. Agent training module 12: Constructs and trains a deep reinforcement learning agent based on the collaborative control model. The state space of the deep reinforcement learning agent maps to the input variables, and its action space includes UAV flight control parameters, cleaning actuator parameters, and tether cable management parameters. Its reward function is designed based on the comprehensive optimization objective and preset constraints. Decision analysis module 13: During the cleaning operation, the deep reinforcement learning agent performs decision analysis based on the real-time state space information and outputs corresponding action space instructions. Collaborative operation module 14: Based on the action space instructions, controls the UAV and its cleaning actuator to perform collaborative operations, and obtains updated state space information and reward information after the action execution, continuously fine-tuning the decision strategy of the deep reinforcement learning agent.
[0044] Furthermore, the model building module 11 is used to perform the following methods: The collaborative control model is modeled, and a Markov decision process framework guided by the comprehensive optimization objective is established. Under the Markov decision process framework, the coupling relationship between each input variable and the comprehensive optimization objective in the collaborative control model is initialized and modeled using historical operation data and simulation data.
[0045] Furthermore, the model building module 11 is used to perform the following methods: The real-time environmental state variables include at least real-time wind speed and direction information obtained through airborne sensors, three-dimensional geometric features and dirt distribution information of the work surface identified through visual sensors, and real-time tension information of the mooring cable obtained through tension sensors; the work equipment state variables include at least the UAV battery power and motor temperature, water tank remaining capacity, and high-pressure water pump working pressure and flow rate.
[0046] Furthermore, the agent training module 12 is used to perform the following methods: Positive reward weights are assigned to cleaning efficiency indicators, including effective cleaning area per unit time and stain removal rate based on visual feedback; positive reward weights are also assigned to energy efficiency indicators, including flight trajectory smoothness, motor torque output stability, and high-pressure water pump operating point efficiency; negative rewards are assigned to behaviors that violate preset safety and efficiency constraints, including predicted collision risk, exceeding tether tension limits, cleaning effect failing to reach thresholds, and excessive redundancy in operation path planning.
[0047] Furthermore, the agent training module 12 is used to perform the following methods: A digital twin environment is constructed, which includes a 3D model of the target operation scenario, a fluid dynamics wind field model, and an equipment dynamics model. In the digital twin environment, the deep reinforcement learning agent interacts with the simulation environment to explore and experiment, and its policy network is pre-trained offline until the policy converges to obtain an initial agent model.
[0048] Furthermore, the decision analysis module 13 is used to perform the following methods: Before the operation begins, based on the 3D point cloud model of the target operation area and combined with the preset standard cleaning width, an initial coverage path is generated that can completely cover the target operation area without collision; the initial coverage path is converted into an initial state space sequence, which serves as the reference state input for the first decision of the deep reinforcement learning agent.
[0049] Furthermore, the decision analysis module 13 is used to perform the following methods: The deep reinforcement learning agent receives and processes the fused current state space information, which includes environmental state information, device state information, and deviation information relative to the reference state input. Based on the current state space information, the deep reinforcement learning agent performs coupling relationship deduction and multi-step prediction, and autonomously generates a global operation strategy with the goal of minimizing the current overall cost. The global operation strategy includes a set of real-time UAV flight control parameters, cleaning actuator parameters, and tether cable management parameters. The global operation strategy is output as the action space command.
[0050] Furthermore, the collaborative operation module 14 is used to perform the following methods: After an action is executed, the deep reinforcement learning agent collects and updates state space information, and generates immediate reward information according to the reward function. The state space information, action space instructions, immediate reward information, and updated state space information for the next decision period are stored together as a set of empirical data. Based on the empirical data, the policy network parameters of the deep reinforcement learning agent are updated in a gradient to optimize the decision strategy for the next round of cleaning.
[0051] Example 3, Figure 3 This is a schematic diagram of the structure of an electronic device provided in Embodiment 3 of the present invention, showing a block diagram of an exemplary electronic device suitable for implementing the embodiments of the present invention. Figure 3 The electronic device shown is merely an example and should not be construed as limiting the functionality or scope of the embodiments of the present invention. Figure 3 As shown, the electronic device includes a processor 21, a memory 22, an input device 23, and an output device 24; the number of processors 21 in the electronic device can be one or more. Figure 3 Taking a processor 21 as an example, the processor 21, memory 22, input device 23, and output device 24 in an electronic device can be connected via a bus or other means. Figure 3 Taking the example of a connection between China and Israel via a bus.
[0052] The memory 22, as a computer-readable storage medium, can be used to store software programs, computer-executable programs, and modules, such as the program instructions / modules corresponding to the intelligent cleaning control method for drones in this embodiment of the invention. The processor 21 executes various functional applications and data processing of the electronic device by running the software programs, instructions, and modules stored in the memory 22, thereby realizing the aforementioned intelligent cleaning control method for drones.
[0053] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention in any way. Although the present invention has been disclosed above with reference to preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art can make some modifications or alterations to the above-disclosed technical content to create equivalent embodiments without departing from the scope of the present invention. Any modifications, equivalent changes, and alterations made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the scope of the present invention.
Claims
1. A smart cleaning control method for unmanned aerial vehicles (UAVs), characterized in that, The method includes: A collaborative control model is constructed based on the comprehensive optimization objectives, which include at least minimizing total operating energy consumption and optimizing cleaning effect per unit area. The input variables of the collaborative control model include real-time environmental state variables and operating equipment state variables. Based on the cooperative control model, a deep reinforcement learning agent is constructed and trained. The state space of the deep reinforcement learning agent maps the input variables, and its action space includes UAV flight control parameters, cleaning actuator parameters, and tether cable management parameters. Its reward function is designed based on the comprehensive optimization objective and preset constraints. During the cleaning operation, the deep reinforcement learning agent performs decision analysis based on the real-time state space information and outputs corresponding action space instructions. Based on the action space instructions, the drone and its cleaning execution mechanism are controlled to perform collaborative operations, and updated state space information and reward information are obtained after the action is executed, so as to continuously fine-tune the decision-making strategy of the deep reinforcement learning agent.
2. The intelligent cleaning control method for unmanned aerial vehicles as described in claim 1, characterized in that, Construct a collaborative control model for intelligent cleaning using drones, including: The collaborative control model is modeled, and a Markov decision process framework guided by the comprehensive optimization objective is established. Within the framework of the Markov decision process, historical operational data and simulation data are used to initialize and model the coupling relationship between each input variable and the comprehensive optimization objective in the collaborative control model.
3. The intelligent cleaning control method for unmanned aerial vehicles as described in claim 1, characterized in that, The real-time environmental state variables include at least the real-time wind speed and direction information obtained by airborne sensors, the three-dimensional geometric features and stain distribution information of the work surface identified by visual sensors, and the real-time tension information of the mooring cable obtained by tension sensors. The status variables of the operating equipment include at least the drone battery charge and motor temperature, water tank remaining capacity, and high-pressure water pump operating pressure and flow rate.
4. The intelligent cleaning control method for unmanned aerial vehicles as described in claim 1, characterized in that, Based on the comprehensive optimization objective and preset safety and performance constraints, the reward function is designed and initialized, including: Positive reward weights are assigned to cleaning efficiency indicators, which include the effective cleaning area per unit time and the stain removal rate based on visual feedback. Positive reward weights are assigned to energy efficiency indicators, including flight trajectory smoothness, motor torque output stability, and high-pressure water pump operating point efficiency. Negative rewards are assigned to behaviors that violate preset safety and performance constraints, including predicted collision risks, exceeding tether tension limits, failing to achieve cleaning thresholds, and excessive redundancy in work path planning.
5. The intelligent cleaning control method for unmanned aerial vehicles as described in claim 1, characterized in that, The training of the deep reinforcement learning agent is completed in a high-fidelity digital twin simulation environment, including: Construct a digital twin environment that includes a 3D model of the target operation scenario, a fluid dynamics wind field model, and a equipment dynamics model; In the digital twin environment, the deep reinforcement learning agent interacts with the simulation environment to explore and experiment, and its policy network is pre-trained offline until the policy converges, thus obtaining the initial agent model.
6. The intelligent cleaning control method for unmanned aerial vehicles as described in claim 1, characterized in that, The method further includes an initial coverage path generation step: Before the operation begins, based on the 3D point cloud model of the target operation area and combined with the preset standard cleaning width, an initial coverage path is generated that can completely cover the target operation area without collision. The initial coverage path is converted into an initial state space sequence, which serves as the reference state input for the first decision of the deep reinforcement learning agent.
7. The intelligent cleaning control method for unmanned aerial vehicles as described in claim 6, characterized in that, During the cleaning operation, the deep reinforcement learning agent performs decision analysis based on the real-time acquired state space information and outputs corresponding action space instructions, including: The deep reinforcement learning agent receives and processes the fused current state space information, which includes environmental state information, device state information, and deviation information relative to the reference state input. The deep reinforcement learning agent performs coupling relationship deduction and multi-step prediction based on the current state space information. With the goal of minimizing the current comprehensive cost, it autonomously makes decisions to generate a global operation strategy. The global operation strategy includes a set of real-time UAV flight control parameters, cleaning actuator parameters, and tether cable management parameters. The global job strategy is output as the action space instruction.
8. The intelligent cleaning control method for unmanned aerial vehicles as described in claim 1, characterized in that, Acquire updated state space information and reward information after action execution, and continuously fine-tune the decision-making strategy of the deep reinforcement learning agent, including: After the deep reinforcement learning agent performs an action, it collects and updates the state space information and generates instant reward information according to the reward function. The state space information, action space instructions, immediate reward information obtained in the current decision cycle, and the updated state space information for the next decision cycle are stored together as a single piece of empirical data. Based on the empirical data, the policy network parameters of the deep reinforcement learning agent are updated using gradients to optimize the decision-making strategy for the next round of cleaning.
9. An intelligent cleaning control system for unmanned aerial vehicles (UAVs), characterized in that, The system is used to implement the intelligent cleaning control method for unmanned aerial vehicles according to any one of claims 1-8, the system comprising: Model building module: Constructs a collaborative control model based on the comprehensive optimization objectives, which include at least minimizing total operating energy consumption and optimizing cleaning effect per unit area. The input variables of the collaborative control model include real-time environmental state variables and operating equipment state variables. Intelligent agent training module: Based on the cooperative control model, a deep reinforcement learning intelligent agent is constructed and trained. The state space of the deep reinforcement learning intelligent agent maps the input variables. Its action space includes UAV flight control parameters, cleaning actuator parameters and tether cable management parameters. Its reward function is designed based on the comprehensive optimization objective and preset constraints. Decision analysis module: During the cleaning operation, the deep reinforcement learning agent performs decision analysis based on the real-time state space information and outputs corresponding action space instructions; Collaborative Operation Module: Based on the action space instructions, the module controls the UAV and its cleaning execution mechanism to perform collaborative operations, and obtains updated state space information and reward information after the action is executed, and continuously fine-tunes the decision-making strategy of the deep reinforcement learning agent.
10. An electronic device, characterized in that, The electronic device includes: Memory, used to store executable instructions; The processor, when executing executable instructions stored in the memory, implements the intelligent cleaning control method for a drone as described in any one of claims 1-8.