An unmanned sanitation vehicle self-adaptive cleaning method and device based on multi-sensor fusion

CN122653283APending Publication Date: 2026-08-28XUZHOU XUGONG ENVIRONMENTAL TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610893540.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-18
Publication Date
2026-08-28

AI Technical Summary

Technical Problem

当面临地面条件变化或环境障碍增加等情况时,系统无法实时调整清洁方式,进一步降低了清洁效果

Benefits of technology

本发明通过卡尔曼滤波算法融合数据,有效减少多传感器之间同一信息的差异与冲突,获得高置信度的车辆位姿信息,基于卷积神经网络准确识别污渍类型与污渍等级,克服了传统方法对复杂地面环境识别能力不足的缺陷。通过将位姿、识别结果与障碍物信息共同构建状态向量,并采用深度Q网络进行自适应清洁策略学习,实现了清洁力度与作业方式的智能决策,避免了固定清洁模式在环境变化时效果下降的问题。采用多通道并行PID闭环控制将决策输出的目标转速与流量转化为精确的电机驱动信号,确保扫刷、风机与喷水系统的执行精度与响应速度,通过清扫覆盖率与能耗构建奖励函数,结合目标网络硬更新机制与经验回放,使深度Q网络的权重不断优化,逐步逼近全局最优清扫策略。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122653283A_ABST
    Figure CN122653283A_ABST
Patent Text Reader

Abstract

The application discloses an environment perception and adaptive cleaning algorithm field, and discloses an unmanned environmental sanitation vehicle adaptive cleaning method and device based on multi-sensor fusion, aiming at solving the technical problem of insufficient ground material, stain and obstacle recognition in a complex environment. The method comprises the following steps: acquiring multi-source sensor environment data and fusing the multi-source sensor environment data to obtain optimal pose state estimation; constructing an environment state matrix based on the optimal pose state estimation, and fusing visual features to obtain environment pollution structured information; using a deep Q network to learn and decide an adaptive cleaning strategy, and outputting a continuous control quantity; converting the control quantity into a driving signal to adjust a cleaning action; determining a cleaning coverage rate according to an actual pose and environment information, and planning a rescan path; constructing a reward function based on the coverage rate change and actual energy consumption, updating deep Q network weights, and generating an optimal cleaning strategy. The application can realize precise perception of the environment and intelligent and efficient cleaning function of the unmanned environmental sanitation vehicle.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to an adaptive cleaning method and device for unmanned sanitation vehicles based on multi-sensor fusion, belonging to the field of environmental perception and adaptive cleaning algorithm technology. Background Technology

[0002] Unmanned sanitation vehicles and sweepers, as representatives of modern intelligent cleaning equipment, have significantly improved the efficiency and effectiveness of floor cleaning through automation and intelligent technologies. However, current technologies still have some significant limitations, mainly reflected in the inadequacy of environmental perception and recognition technologies, as well as the immaturity of adaptive cleaning algorithms and control methods.

[0003] Specifically, existing technologies primarily rely on pre-set environmental models or simple sensor data for perception, resulting in limited recognition capabilities when faced with complex and varied ground materials, different types of stains, and diverse obstacles. Because they cannot accurately identify ground materials, stain types, and their distribution, precise adjustment of cleaning intensity is difficult, impacting the cleaning effectiveness of unmanned sanitation vehicles in actual use. Furthermore, current cleaning algorithms and control methods are mostly based on fixed cleaning patterns and path designs, lacking the ability to automatically adjust cleaning strategies and intensity according to actual ground conditions and environmental changes. When faced with changes in ground conditions or increased environmental obstacles, the system cannot adjust its cleaning methods in real time, further reducing cleaning effectiveness. Summary of the Invention

[0004] The purpose of this invention is to provide an adaptive cleaning method and device for unmanned sanitation vehicles based on multi-sensor fusion, which can accurately identify ground material, stain type and distribution, and obstacle location to achieve accurate perception, intelligent decision-making and efficient cleaning of unmanned sanitation vehicles in complex environments.

[0005] To achieve the above-mentioned technical effects, the present invention is implemented using the following technical solution.

[0006] In a first aspect, the present invention provides an adaptive cleaning method for unmanned sanitation vehicles based on multi-sensor fusion, comprising: Acquire environmental data from multiple sensor sources; The multi-source sensor environmental data are fused to obtain the optimal pose state estimate of the unmanned sanitation vehicle at the current moment. Based on the optimal pose state estimation, an environmental state matrix is ​​constructed, and the RGB color image features and disparity depth data extracted from the multi-source sensor data are identified and classified and then fused into the environmental state matrix to obtain structured information on environmental pollution. Based on the structured information of environmental pollution, a deep Q-network is used to learn and make decisions on adaptive cleaning strategies, and output continuous control variables. The continuous control quantity is converted into a motor drive signal through a multi-channel parallel PID regulator to regulate the cleaning action of the unmanned sanitation vehicle. Based on the current cleaning actions of unmanned sanitation vehicles, their actual position and posture information and actual energy consumption are obtained, and the cleaning coverage rate is determined according to the actual position and posture information and the structured information of environmental pollution. Plan a re-sweeping path based on the distribution of the sweeping coverage, and update the sweeping coverage after sweeping along the re-sweeping path is completed; A reward function is constructed based on the difference in cleaning coverage before and after the update and the actual execution energy consumption, and the weights of the deep Q network are updated based on the reward function to generate the optimal cleaning strategy.

[0007] In conjunction with the first aspect, the multi-source sensor environmental data further includes three-dimensional point cloud depth information generated by lidar, ranging and velocity measurement data from millimeter-wave radar, RGB color image data and parallax depth data collected by high-definition cameras, and satellite positioning data and inertial measurement offset from the integrated navigation system. The 3D point cloud data, ranging and velocity measurement data, and RGB color image data are fused to obtain the pre-processed multi-source fusion data. ; Based on the satellite positioning data and inertial measurement offset, calculate the pose information after integrated navigation. ; Based on the aforementioned multi-source fusion data pose information after integrated navigation Construct a state-space model that includes a state transition matrix and an observation matrix; Using the state space model, state prediction is performed based on the pose state estimate of the previous time step, and the state prediction value of the unmanned sanitation vehicle at the current time step and the prediction error covariance matrix are output. The pre-processed multi-source fusion data pose information after integrated navigation Using the observed value at the current moment, the Kalman gain is calculated by combining the prediction error covariance matrix and the preset measurement noise covariance matrix. The Kalman gain is used to weight and fuse the predicted state value of the unmanned sanitation vehicle at the current moment with the observed value to obtain the optimal pose state estimate of the unmanned sanitation vehicle at the current moment.

[0008] In conjunction with the first aspect, the expression for the optimal pose state estimation of the unmanned sanitation vehicle at the current moment is further as follows: ; in, This represents the optimal pose state estimate after fusion at time k; This represents the optimal pose estimate after fusion at the previous time step; Indicates Kalman gain; This represents the actual observations from multiple sensors at time k; This represents the observation matrix.

[0009] In conjunction with the first aspect, the obtained structured environmental pollution information further includes: Based on the optimal pose state estimation at the current moment and the optimal pose state estimation at the previous moment, the position coordinate components and attitude angle components are extracted respectively, and the relative pose transformation matrix between adjacent frames is calculated. Based on the relative pose transformation matrix, from the preceding multi-source fusion data Extract the 3D point cloud depth information from the previous frame, and project the 3D point cloud depth information onto the current frame coordinates to generate an environment state matrix. ; A pre-trained convolutional neural network is used to extract the feature vector of the next frame of RGB color image for recognition and classification, and the feature vector containing the ground material type and the stain intensity level is extracted. The feature vectors are mapped onto the spatial structure of the environment state matrix using an attention mechanism, and then fused with the environment state matrix to obtain fused state information. ; Calculate the fusion state information Compared with the preset reference state value The residual between them yields the degree of environmental staining. ; Based on the degree of environmental stains Constructing structured information on environmental pollution .

[0010] In conjunction with the first aspect, the expression for the fused state information is further as follows: ; in, Indicates fusion status information; Indicates the weights for visual feature fusion; This represents the feature vector of the RGB color image in frame t+1. Indicates the weighting of environmental spatial features; Represents the environment state matrix; This indicates matrix vectorization operations.

[0011] In conjunction with the first aspect, the continuously controlled output quantity further includes: Based on the pollution structure information Stain levels Perform a linear mapping to obtain the real-time stain index. ; The real-time stain index and fusion status information As an observation value, the remaining battery capacity of the unmanned sanitation vehicle is also included. and current regional coverage completion rate After zero-mean standardization, the time-series state input is formed. ; Input the timing state The input is fed into the deep Q-network, which outputs the action value corresponding to each candidate cleaning strategy in the discrete action space A. And select the optimal discrete action index according to the ε-greedy principle. ; The optimal discrete action index is obtained through a preset parameter mapping table. Convert to continuous control quantity .

[0012] In conjunction with the first aspect, further, the continuous control quantity is converted into a motor drive signal through a multi-channel parallel PID controller, including: Based on the continuous control quantity A multi-channel parallel PID controller is constructed; the multi-channel parallel PID controller corresponds to the brush speed loop, the energy consumption deviation fan pressure loop, and the water pump flow loop, respectively. Based on the brush speed circuit, energy consumption deviation fan pressure circuit, and water pump flow circuit, calculate the tracking error of each circuit; Based on the tracking error, the PID control output of each loop is calculated, and its expression is as follows: ; in, Indicates the first Each loop in The PID control output at any given time; Indicates the first The proportional gain coefficient of each loop; Indicates the first Each loop in Tracking error at any given moment; Indicates the first Integral gain coefficient of each loop; Indicates the current moment; In the process of integration Tracking error; Represents the integral variable; Indicates the first The differential gain coefficient of each loop; PWM drive signals are generated based on the PID control values ​​of each loop.

[0013] In conjunction with the first aspect, the generation of the optimal cleaning strategy further includes: Mark the cleaned areas; Based on the current cleaning actions of unmanned sanitation vehicles, and according to the actual position information and structured environmental pollution information, residual features in uncleaned areas are removed, and the cleaning coverage rate is calculated. ; Obtain the theoretical energy consumption value calculated based on the current cleaning actions and load status of the unmanned sanitation vehicle, compare the actual execution energy consumption with the decision estimate, and calculate the energy consumption deviation. ; Areas that have not been cleaned due to slipping, missed areas, or stubborn stains are marked as a set of paths to be re-cleaned. Based on the set of paths to be re-cleaned and the remaining battery capacity of the unmanned sanitation vehicle, a re-cleaning path is planned. Control the unmanned sanitation vehicle to clean along the re-sweeping path and update the cleaning coverage rate; Calculate the sweep coverage increase based on the sweep coverage before and after the update. ; Based on the increase in cleaning coverage and energy consumption deviation Construct a reward function ; Using the reward function Update the weights of the deep Q network to gradually approach the optimal sweeping strategy.

[0014] In conjunction with the first aspect, the updating of the weights of the deep Q-network further includes: Set the playback buffer; A small batch of transfer samples is randomly sampled from the replay buffer using an empirical replay mechanism. ; The target value is calculated using the target network in the deep Q-network. And the value of the current action is calculated using the main network in the deep Q-network; Based on the difference between the target value and the current action value, a loss function is constructed, and the parameters of the main network are updated by minimizing the loss function. The expression for the loss function is: ; ; in, The mean squared error loss function representing the parameters of the main network; This indicates the number of small-batch transfer samples sampled in the experience replay buffer; Indicates the index of the currently processed mini-batch of samples; Indicates the first The target value corresponding to each transferred sample; This indicates the main network's view of the current state. Take action below Q-value evaluation; Indicates the first Instant reward for each transferred sample; Indicates the discount factor; Indicate the next state The action that maximizes the Q-value of the target network evaluation; Indicates the target network's next state and actions Q-value evaluation; Represents the network parameters of the target network; The parameters of the main network are changed every fixed number of steps. Copy to the target network to update the parameters of the target network. .

[0015] Secondly, an adaptive cleaning device for unmanned sanitation vehicles based on multi-sensor fusion includes: The data acquisition module is used to acquire environmental data from multiple sensor sources. The data fusion module is used to fuse the environmental data from the multi-source sensors to obtain the pose estimation value of the unmanned sanitation vehicle. The environment construction and fusion module is used to construct an environment state matrix based on the optimal pose state estimation, and to fuse the RGB color images and parallax depth data extracted from the multi-source sensor data into the environment state matrix after identification and classification to obtain structured information on environmental pollution. An adaptive learning module is used to learn and make decisions on adaptive cleaning strategies based on the structured information of environmental pollution using a deep Q-network, and output continuous control variables. The data conversion module is used to convert the continuous control quantity into a motor drive signal through a multi-channel parallel PID regulator in order to regulate the cleaning action of the unmanned sanitation vehicle. The cleaning area determination module is used to obtain the actual position and posture information and actual energy consumption of the unmanned sanitation vehicle based on the current cleaning action of the unmanned sanitation vehicle, and determine the cleaning coverage rate according to the actual position and posture information and the structured information of environmental pollution. The replanning path module is used to plan a resweeping path based on the distribution of the sweeping coverage, and update the sweeping coverage after sweeping along the resweeping path is completed. The optimal strategy generation module is used to construct a reward function based on the difference in cleaning coverage before and after the update and the actual execution energy consumption, and to update the weights of the deep Q network based on the reward function to generate the optimal cleaning strategy.

[0016] Compared with the prior art, the beneficial effects achieved by the present invention are as follows: This invention fuses data using a Kalman filter algorithm, effectively reducing discrepancies and conflicts in the same information from multiple sensors to obtain high-confidence vehicle pose information. Based on a convolutional neural network, it accurately identifies stain types and levels, overcoming the shortcomings of traditional methods in recognizing complex ground environments. By constructing a state vector from pose, recognition results, and obstacle information, and employing a deep Q-network for adaptive cleaning strategy learning, it achieves intelligent decision-making regarding cleaning intensity and operation methods, avoiding the problem of decreased effectiveness of fixed cleaning modes when the environment changes. Multi-channel parallel PID closed-loop control converts the target speed and flow rate output into precise motor drive signals, ensuring the execution accuracy and response speed of the sweeping brushes, fans, and water spraying systems. A reward function is constructed using sweeping coverage and energy consumption, combined with a target network hard update mechanism and experience replay, continuously optimizing the weights of the deep Q-network to gradually approach the globally optimal cleaning strategy.

[0017] Furthermore, by dynamically correcting the energy consumption penalty coefficient in the reward function based on the deviation between actual energy consumption and theoretically estimated energy consumption, and by combining the erasure of residual features in uncleaned areas to achieve accurate estimation of coverage, cleaning efficiency can be continuously improved, energy consumption reduced, and full coverage cleaning achieved in long-term operation. Ultimately, this enables unmanned sanitation vehicles to achieve accurate perception, intelligent decision-making, and efficient cleaning in complex environments. Attached Figure Description

[0018] Figure 1 The diagram shown is a flowchart of an adaptive cleaning method for unmanned sanitation vehicles based on multi-sensor fusion provided in an embodiment of the present invention. Figure 2 The diagram shown is a flowchart of the Kalman filtering operation provided in an embodiment of the present invention. Figure 3 The diagram shown is a flowchart of a convolutional neural network provided in an embodiment of the present invention. Detailed Implementation

[0019] The technical solution of the present invention will be described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the embodiments of the present invention and the specific features in the embodiments are detailed descriptions of the technical solution of the present invention, rather than limitations thereof. In the absence of conflict, the embodiments of the present invention and the technical features in the embodiments can be combined with each other. Example 1

[0020] See Figure 1This embodiment introduces an adaptive cleaning method for unmanned sanitation vehicles based on multi-sensor fusion. The method first acquires environmental data from multiple sensor sources during the multi-sensor data collection stage. The data sources include lidar, high-definition cameras, millimeter-wave radar, and integrated navigation systems.

[0021] Secondly, in order to reduce the differences and conflicts between the same information collected by different sensors and improve the accuracy and precision of environmental perception, the Kalman filter algorithm is used to fuse environmental data from multiple sources to obtain the pose estimate of the unmanned sanitation vehicle.

[0022] Next, based on the pose estimation values, an environmental state matrix is ​​constructed, and the RGB color image features extracted from the high-definition camera are fused into the environmental state matrix to obtain structured information on environmental pollution.

[0023] Then, based on the structured information of environmental pollution, a deep Q-network is used to construct a state vector for learning and decision-making of adaptive cleaning strategies. The optimal action combination is selected by maximizing the expected return, and the continuous control quantity is output. The continuous control quantity includes the brush speed, water flow rate and fan speed.

[0024] Finally, the continuous control quantity is converted into a motor drive signal through a multi-channel parallel PID controller to regulate the cleaning action of the unmanned sanitation vehicle. Based on the current cleaning action, the actual pose information and actual execution energy consumption of the unmanned sanitation vehicle are obtained. The cleaning coverage rate is determined according to the actual pose information and the structured information of environmental pollution. The re-sweeping path is planned and the cleaning coverage rate is updated according to the distribution of the cleaning coverage rate. The reward function is constructed based on the change in cleaning coverage rate and the actual execution energy consumption to update the weights of the deep Q network, thus completing the intelligent cleaning operation of the unmanned sanitation vehicle. Example 2

[0025] This embodiment, based on Embodiment 1, further describes an adaptive cleaning method for unmanned sanitation vehicles based on multi-sensor fusion, including the following steps: S1. Acquire environmental data from multiple sensor sources; The system utilizes a collaborative approach involving lidar, high-definition cameras, millimeter-wave radar, and a combined navigation system to acquire perception data of the environment surrounding the unmanned sanitation vehicle and its own position and pose information. Specifically, the lidar emits laser beams and receives reflected signals to obtain 3D point cloud depth information; the high-definition camera acquires RGB color image data and parallax depth data; the millimeter-wave radar emits millimeter-wave signals and analyzes the echoes to obtain ranging and velocity data of target objects (such as pedestrians and vehicles) in the environment surrounding the unmanned sanitation vehicle; and the combined navigation system integrates data from GPS, IMU (Inertial Measurement Unit), and wheel speedometers. Furthermore, the S1 incorporates a front-mounted long-range sensing module, which includes a lidar, millimeter-wave radar, and a high-definition camera.

[0026] 3D modeling and spatial positioning are performed based on 3D point cloud data generated by LiDAR. Simultaneously, ranging and velocity measurement data from millimeter-wave radar are integrated to improve the accuracy of velocity and ranging measurements and penetration capability of target objects. This data is then weighted and fused with RGB color image data and parallax depth data from a high-definition camera to obtain pre-processed multi-source fusion data. Its expression is: (1) in, This indicates front-end multi-source fusion data; This represents the weighting coefficient of the lidar. Represents 3D point cloud data; The weighting coefficients represent the weighting coefficients for millimeter-wave radar. This represents the ranging and velocity measurement data from millimeter-wave radar. Indicates the weighting coefficient of high-definition cameras; This represents the parallax depth data captured by the high-definition camera.

[0027] The S1 also includes a side and blind spot sensing module, which includes millimeter-wave radar.

[0028] The ranging and velocity measurement data from the millimeter-wave radar are fused. A monitoring window centered on the vehicle is selected. Within this window, time synchronization and noise filtering are performed on each sensor group in the integrated navigation system. The pose information after integrated navigation is obtained through integration calculations, expressed as: (2) in, This represents the pose information after the integrated navigation system is activated; t represents time. Indicates the initial time; Represents satellite positioning data; This indicates the inertial measurement offset.

[0029] Furthermore, time synchronization processing for each group of sensors in the integrated navigation system includes: A synchronization method combining hardware triggering and software interpolation is adopted. The PPS (pixel per second) signal of the RTK module in the integrated navigation system is used as the global time reference. The data of each sensor is timestamped to obtain the initial synchronization timestamp, the expression of which is as follows: (3) in, This represents the timestamp after the 1st sensor is synchronized; Indicates global reference time; Indicates the original sampling time; Indicates the sensor startup time; Indicates the sampling frequency; This indicates a fixed delay offset.

[0030] After completing the initial timestamp calibration, the data from each sensor is interpolated or extrapolated using the maximum sampling interval as the reference period. The asynchronous data is then mapped onto a unified time grid to generate a unified timestamp data frame. The data frame alignment expression is as follows: (4) in, This represents the k-th frame of the unified timestamp data frame; Indicates that the i-th sensor is in Interpolated data at any given time; N is the total number of sensors; This represents the time corresponding to the kth unified time grid point; i represents the sensor index value; and k represents the frame index.

[0031] S2. The Kalman filter algorithm is used to fuse environmental data from multiple sources to obtain the pose estimate of the unmanned sanitation vehicle. See Figure 2 To eliminate data redundancy and conflicts arising from different sensors observing the same environmental features, and to obtain more accurate and reliable environmental state estimates, the following measures are taken: S21. Based on the pre-processed multi-source fusion data pose information after integrated navigation Construct a state-space model that includes a state transition matrix and an observation matrix; S22. Using the state-space model, based on the pose state estimate of the previous time step, perform state prediction, and output the predicted state value of the unmanned sanitation vehicle at the current time step and the prediction error covariance matrix. The expression for the predicted state value of the unmanned sanitation vehicle at the current time step is: (5) in, This represents the predicted state value at time k; Represents the state transition matrix; This represents the optimal state estimate at time k-1; Represents the control matrix; This indicates a control input.

[0032] S23, merging multi-source data from the front end pose information after integrated navigation Using the observed value at the current moment, combined with the prediction error covariance matrix and the preset measurement noise covariance matrix, the Kalman gain is calculated, and its expression is: (6) in, Indicates Kalman gain; Represents the prediction error covariance matrix; Represents the observation matrix; Represents the measurement noise covariance matrix; This represents the transpose of the observation matrix.

[0033] S24. Using Kalman gain, the predicted state value and the observed value of the unmanned sanitation vehicle at the current moment are weighted and fused to obtain the optimal pose estimate of the unmanned sanitation vehicle at the current moment, the expression of which is: (7) in, This represents the optimal pose estimate after fusion at time k; This represents the optimal pose estimate after fusion at the previous time step; This represents the actual observations from multiple sensors at time k; This represents the observation matrix.

[0034] S3. Based on the pose estimation value, construct an environmental state matrix, and fuse the visual features extracted from the multi-source sensor data into the environmental state matrix to obtain structured information on environmental pollution. See Figure 3 Specifically, we obtain structured information about environmental pollution, including: S31. Based on the optimal pose state estimation at the current moment and the optimal pose state estimation at the previous moment, extract their position coordinate components and attitude angle components respectively, and calculate the relative pose transformation matrix between adjacent frames. Its expression is: (8) in, Represents the relative pose transformation matrix; Indicates the pose of the current frame; This indicates the pose of the previous frame.

[0035] S32. Based on the relative pose transformation matrix, from the pre-processed multi-source fusion data Extract the 3D point cloud depth information from the previous frame, and project the 3D point cloud depth information onto the current frame coordinates to generate an environment state matrix. Its expression is: (9) in, Represents the environment state matrix. Represents the spatial projection operator. This indicates the environmental depth information of the previous frame. This represents the geometric alignment residual.

[0036] S33. Use a pre-trained convolutional neural network to extract the feature vector of the next frame of RGB color image for recognition and classification, and extract the feature vector containing the ground material type and the stain intensity level. S34. The feature vectors are mapped onto the spatial structure of the environment state matrix through an attention mechanism, and then fused with the environment state matrix to obtain fused state information. Its expression is: (10) in, Indicates fusion status information; Indicates the weights for visual feature fusion; This represents the feature vector of the RGB color image in frame t+1. Indicates the weighting of environmental spatial features; Represents the environment state matrix; This indicates matrix vectorization operations.

[0037] S35. Calculate fusion status information Compared with the preset reference state value The residual between them yields the degree of environmental staining. Its expression is: (11) in, Indicates the baseline state value; Indicates the degree of environmental soiling.

[0038] S36. Based on the degree of environmental stains Constructing structured information on environmental pollution ; in, Indicates the location of the stained area; Indicate the type of stain; Indicates the stain level.

[0039] S4. Based on structured information about environmental pollution, a deep Q-network is used to learn and make decisions on adaptive cleaning strategies, and output continuous control variables. Specifically, outputting continuous control quantities includes the following steps: S41, Based on structured pollution information Stain levels Perform a linear mapping to obtain the real-time stain index. ; S42. Real-time Stain Index and fusion status information As an observation value, the remaining battery capacity of the unmanned sanitation vehicle is also included. and current regional coverage completion rate After zero-mean standardization, the time-series state input is formed. Its expression is: (12) (13) in, Indicates the scaling factor for the grade; Indicates timing state input; Indicates the real-time stain index; This represents the mean of the stain index; The standard deviation of the stain index; Indicates timing state input; Indicates the remaining battery capacity; This indicates the current coverage completion rate of the area; This indicates the vector transpose symbol; This indicates standardized operations.

[0040] S43. Input the timing status. The input is fed into a deep Q-network, which employs a three-layer fully connected hidden layer structure with 128, 64, and 32 neurons in each layer, respectively. The activation function used is ReLU, and the network output layer corresponds to the action value of each candidate cleaning strategy in the discrete action space A. And select the optimal discrete action index according to the ε-greedy principle. The optimal discrete action index can be represented as: (14) in, Indicates the optimal action index; This represents the k-th candidate cleaning strategy; Indicates the main network parameters; Action value function.

[0041] S44. Index the optimal discrete action using a preset parameter mapping table. Convert to continuous control quantity .

[0042] In this embodiment of the invention, the action selection and parameter output follow the ε-greedy principle. The initial exploration rate ε = 0.9 decays exponentially to 0.01 with each training round, and the network outputs the optimal discrete action index. Converted into continuous control quantity through a preset parameter mapping table Its expression is: (15) in, Indicates the brush rotation speed; Indicates the fan power; Indicates the speed of travel; This indicates a continuous control quantity.

[0043] S5, Continuous control quantity The multi-channel parallel PID controller converts the signals into motor drive signals to regulate the cleaning actions of the unmanned sanitation vehicle. Specifically, based on continuous control quantities, a low-level execution closed loop is constructed to convert control commands into motor drive signals, as follows: A multi-channel parallel PID controller is established, corresponding to the brush speed, energy consumption deviation, fan pressure, and water pump flow loops; among them, the brush speed is one of the continuous control variables. Fan power speed of travel These are respectively used as the input setpoints for the corresponding loops; Based on the brush rotation speed, energy consumption deviation, fan pressure, and water pump flow rate loop, the tracking error of each loop is calculated, and its expression is as follows: (16) in, express The appropriate amount; This indicates the actual value fed back by the sensor; Let represent the tracking error of the j-th loop.

[0044] Based on the tracking error, the PID control output of each loop is calculated, and its expression is as follows: (17) in, Indicates the first Each loop in The PID control output at any given time; Indicates the first The proportional gain coefficient of each loop; Indicates the first Each loop in Tracking error at any given moment; Indicates the first The integral gain coefficient of each loop; t represents the current time. In the process of integration Tracking error; Represents the integral variable; Indicates the first The differential gain coefficient of each loop; Indicates the circuit number ( Represents the brush rotation speed circuit. Represents the fan pressure circuit. (Represents the water pump flow circuit).

[0045] An error weighting mechanism is introduced to assign the absolute value of the tracking error of the j-th loop. The loop with the largest absolute deviation is given priority in allocating computing resources to ensure synchronized dynamic response of multiple actuators, and PWM drive signals are generated based on the PID control values ​​of each loop.

[0046] Based on the current cleaning actions of the unmanned sanitation vehicle, the actual pose information and actual execution energy consumption of the unmanned sanitation vehicle are obtained, and the cleaning coverage rate is determined according to the actual pose information and the structured information of environmental pollution. The re-sweeping path is planned according to the distribution of the cleaning coverage rate, and the cleaning coverage rate is updated after cleaning is completed along the re-sweeping path. A reward function is constructed based on the difference between the cleaning coverage rate before and after the update and the actual execution energy consumption, and the weights of the deep Q network are updated based on the reward function to generate the optimal cleaning strategy.

[0047] Specifically, in order to identify coverage blind spots and plan rescanning paths in the next decision cycle, thereby optimizing the strategy throughout the entire operation cycle, the cleaned areas are first marked; Based on the current cleaning actions of unmanned sanitation vehicles, and according to the actual position information and structured environmental pollution information, residual features in uncleaned areas are removed, and the cleaning coverage rate is calculated. ; Obtain the theoretical energy consumption value (decision estimate) calculated based on the current cleaning actions and load status of the unmanned sanitation vehicle. Compare the actual energy consumption with the decision estimate to calculate the energy consumption deviation. Its expression is: (18) in, Indicates energy consumption deviation; Indicates actual energy consumption during execution; This represents a theoretically estimated energy consumption.

[0048] Uncleaned areas due to slippage, missed areas, or stubborn stains are marked as a set of path points to be re-sweeped. Based on this set of path points and the remaining battery capacity of the unmanned sanitation vehicle, a re-sweeping path is planned for the next decision cycle. The expression for this path is: (19) in, To rescan the path; For path planning functions; Current battery status; This represents the set of path points to be rescanned in the next decision cycle.

[0049] The system controls unmanned sanitation vehicles to sweep along the re-sweeping path, updates the sweeping coverage rate, and calculates the sweeping coverage increment based on the sweeping coverage rate before and after the update. Its expression is: (20) in, Indicates the increase in coverage; Indicates cleaning coverage rate; This indicates the cleaning coverage rate in the previous decision-making cycle.

[0050] Based on the increase in cleaning coverage and energy consumption deviation Construct a reward function This is used to quantify the merits of a strategy, and its expression is: ;(twenty one) in, The first point indicates that the efficiency of stain removal is encouraged; This indicates that the second item is a power safety constraint; Indicates the battery threshold; This indicates that the third item excludes motor power consumption; Indicates action The corresponding instantaneous power of the motor; Indicates the amount of stain at the next moment.

[0051] Using the reward function Update the weights of the deep Q-network to gradually approach the optimal sweeping strategy. Updating the weights of the deep Q-network includes: Set the playback buffer; An experience replay mechanism is used to break the temporal correlation between data, and small batches of transfer samples are randomly sampled from the replay buffer. ; The target value is calculated using the target network in the deep Q-network. And the value of the current action is calculated using the main network in the deep Q-network; A loss function is constructed based on the difference between the target value and the current action value. The parameters of the main network are updated by minimizing this loss function. The expression for the loss function is: ;(twenty two) ;(twenty three) in, The mean squared error loss function representing the parameters of the main network; This indicates the number of small-batch transfer samples sampled in the experience replay buffer; Indicates the index of the currently processed mini-batch of samples; Indicates the first The target value corresponding to each transition sample is used to measure the accuracy of the main network's long-term cumulative reward estimation for the current state-action pair. Its calculation includes the immediate reward and the discounted future maximum Q value. This indicates the main network's view of the current state. Take action below The Q-value assessment represents the expected long-term cumulative return of taking this action in this state; Indicates the first The immediate reward for each transferred sample is obtained by a reward function constructed from the cleaning coverage increment and energy consumption deviation; This represents the discount factor, with a value range of (0, 1), used to balance the weights of immediate rewards and future rewards; Indicate the next state The action that maximizes the Q-value of the target network evaluation; Indicates the target network's next state and actions The Q-value evaluation represents the action to be taken in this state. Long-term cumulative return estimates; The network parameters of the target network are represented, and the stability of the estimation is maintained by periodically replicating the main network parameters.

[0052] At fixed intervals, the parameters of the main network are... Copy to the target network to update the parameters of the target network. .

[0053] Furthermore, the weight update of the deep Q-network also incorporates the self-evolution mechanism of the sweeping strategy, iteratively correcting the weight coefficients in the reward function through the following rules: ;(twenty four) in, This represents the updated value of the i-th weight coefficient at the next time step; This represents the value of the i-th weight coefficient at the current moment; Indicates the learning rate; This indicates the overall optimization goal; This represents the i-th weight coefficient.

[0054] This update rule enables the system to dynamically adjust the weight coefficients of cleaning efficiency, safety constraints, and energy consumption penalties in the reward function based on feedback information such as cleaning coverage and energy consumption deviation. This allows the system to continuously correct state estimation and strategy preferences, gradually approaching the globally optimal cleaning strategy, and ultimately achieving efficient, energy-saving, and fully covered intelligent sanitation operations. Example 3

[0055] An adaptive cleaning device for unmanned sanitation vehicles based on multi-sensor fusion includes: The data acquisition module is used to acquire environmental data from multiple sensor sources. The data fusion module is used to fuse the environmental data from the multi-source sensors to obtain the pose estimation value of the unmanned sanitation vehicle. The environment construction and fusion module is used to construct an environment state matrix based on the optimal pose state estimation, and to fuse the RGB color images and parallax depth data extracted from the multi-source sensor data into the environment state matrix after identification and classification to obtain structured information on environmental pollution. An adaptive learning module is used to learn and make decisions on adaptive cleaning strategies based on the structured information of environmental pollution using a deep Q-network, and output continuous control variables. The data conversion module is used to convert the continuous control quantity into a motor drive signal through a multi-channel parallel PID regulator in order to regulate the cleaning action of the unmanned sanitation vehicle. The cleaning area determination module is used to obtain the actual position and posture information and actual energy consumption of the unmanned sanitation vehicle based on the current cleaning action of the unmanned sanitation vehicle, and determine the cleaning coverage rate according to the actual position and posture information and the structured information of environmental pollution. The replanning path module is used to plan a resweeping path based on the distribution of the sweeping coverage, and update the sweeping coverage after sweeping along the resweeping path is completed. The optimal strategy generation module is used to construct a reward function based on the difference in cleaning coverage before and after the update and the actual execution energy consumption, and to update the weights of the deep Q network based on the reward function to generate the optimal cleaning strategy.

[0056] The embodiments of the present invention have been described above with reference to the accompanying drawings. However, the present invention is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of the present invention without departing from the spirit and scope of the claims. All of these forms are within the protection scope of the present invention.

Claims

1. An adaptive cleaning method for unmanned sanitation vehicles based on multi-sensor fusion, characterized in that, include: Acquire environmental data from multiple sensor sources; The multi-source sensor environmental data are fused to obtain the optimal pose state estimate of the unmanned sanitation vehicle at the current moment. Based on the optimal pose state estimation, an environmental state matrix is ​​constructed, and the RGB color image features and disparity depth data extracted from the multi-source sensor data are identified and classified and then fused into the environmental state matrix to obtain structured information on environmental pollution. Based on the structured information of environmental pollution, a deep Q-network is used to learn and make decisions on adaptive cleaning strategies, and output continuous control variables. The continuous control quantity is converted into a motor drive signal through a multi-channel parallel PID regulator to regulate the cleaning action of the unmanned sanitation vehicle. Based on the current cleaning actions of unmanned sanitation vehicles, their actual position and posture information and actual energy consumption are obtained, and the cleaning coverage rate is determined according to the actual position and posture information and the structured information of environmental pollution. Plan a re-sweeping path based on the distribution of the sweeping coverage, and update the sweeping coverage after sweeping along the re-sweeping path is completed; A reward function is constructed based on the difference in cleaning coverage before and after the update and the actual execution energy consumption, and the weights of the deep Q network are updated based on the reward function to generate the optimal cleaning strategy.

2. The adaptive cleaning method for unmanned sanitation vehicles based on multi-sensor fusion according to claim 1, characterized in that, The multi-source sensor environmental data includes three-dimensional point cloud depth information generated by lidar, ranging and velocity measurement data from millimeter-wave radar, RGB color image data and parallax depth data collected by high-definition cameras, and satellite positioning data and inertial measurement offset from the integrated navigation system. The 3D point cloud data, ranging and velocity measurement data, and RGB color image data are fused to obtain the pre-processed multi-source fusion data. ; Based on the satellite positioning data and inertial measurement offset, calculate the pose information after integrated navigation. ; Based on the aforementioned multi-source fusion data pose information after integrated navigation Construct a state-space model that includes a state transition matrix and an observation matrix; Using the state space model, state prediction is performed based on the pose state estimate of the previous time step, and the state prediction value of the unmanned sanitation vehicle at the current time step and the prediction error covariance matrix are output. The pre-processed multi-source fusion data pose information after integrated navigation Using the observed value at the current moment, the Kalman gain is calculated by combining the prediction error covariance matrix and the preset measurement noise covariance matrix. The Kalman gain is used to weight and fuse the predicted state value of the unmanned sanitation vehicle at the current moment with the observed value to obtain the optimal pose state estimate of the unmanned sanitation vehicle at the current moment.

3. The adaptive cleaning method for unmanned sanitation vehicles based on multi-sensor fusion according to claim 2, characterized in that, The expression for the optimal pose state estimation of the unmanned sanitation vehicle at the current moment is: ; in, This represents the optimal pose state estimate after fusion at time k; This represents the optimal pose estimate after fusion at the previous time step; Indicates Kalman gain; This represents the actual observations from multiple sensors at time k; This represents the observation matrix.

4. The adaptive cleaning method for unmanned sanitation vehicles based on multi-sensor fusion according to claim 2, characterized in that, The obtained structured information on environmental pollution includes: Based on the optimal pose state estimation at the current moment and the optimal pose state estimation at the previous moment, the position coordinate components and attitude angle components are extracted respectively, and the relative pose transformation matrix between adjacent frames is calculated. Based on the relative pose transformation matrix, from the preceding multi-source fusion data Extract the 3D point cloud depth information from the previous frame, and project the 3D point cloud depth information onto the current frame coordinates to generate an environment state matrix. ; A pre-trained convolutional neural network is used to extract the feature vector of the next frame of RGB color image for recognition and classification, and the feature vector containing the ground material type and the stain intensity level is extracted. The feature vectors are mapped onto the spatial structure of the environment state matrix using an attention mechanism, and then fused with the environment state matrix to obtain fused state information. ; Calculate the fusion state information Compared with the preset reference state value The residual between them yields the degree of environmental staining. ; Based on the degree of environmental stains Constructing structured information on environmental pollution .

5. The adaptive cleaning method for unmanned sanitation vehicles based on multi-sensor fusion according to claim 4, characterized in that, The expression for the fusion state information is: ; in, Indicates fusion status information; Indicates the weights for visual feature fusion; This represents the feature vector of the RGB color image in frame t+1. Indicates the weighting of environmental spatial features; Represents the environment state matrix; This indicates matrix vectorization operations.

6. The adaptive cleaning method for unmanned sanitation vehicles based on multi-sensor fusion according to claim 1, characterized in that, The output continuous control quantity includes: Based on the pollution structure information Stain levels Perform a linear mapping to obtain the real-time stain index. ; The real-time stain index and fusion status information As an observation value, the remaining battery capacity of the unmanned sanitation vehicle is also included. and current regional coverage completion rate After zero-mean standardization, the time-series state input is formed. ; Input the timing state The input is fed into the deep Q-network, which outputs the action value corresponding to each candidate cleaning strategy in the discrete action space A. And select the optimal discrete action index according to the ε-greedy principle. ; The optimal discrete action index is obtained through a preset parameter mapping table. Convert to continuous control quantity .

7. The adaptive cleaning method for unmanned sanitation vehicles based on multi-sensor fusion according to claim 6, characterized in that, The continuous control quantity is converted into a motor drive signal through a multi-channel parallel PID controller, including: Based on the continuous control quantity A multi-channel parallel PID controller is constructed; the multi-channel parallel PID controller corresponds to the brush speed loop, the energy consumption deviation fan pressure loop, and the water pump flow loop, respectively. Based on the brush speed circuit, energy consumption deviation fan pressure circuit, and water pump flow circuit, calculate the tracking error of each circuit; Based on the tracking error, the PID control output of each loop is calculated, and its expression is as follows: ; in, Indicates the first Each loop in The PID control output at any given time; Indicates the first The proportional gain coefficient of each loop; Indicates the first Each loop in Tracking error at any given moment; Indicates the first Integral gain coefficient of each loop; Indicates the current moment; In the process of integration Tracking error; Represents the integral variable; Indicates the first The differential gain coefficient of each loop; PWM drive signals are generated based on the PID control values ​​of each loop.

8. The adaptive cleaning method for unmanned sanitation vehicles based on multi-sensor fusion according to claim 1, characterized in that, The generation of the optimal cleaning strategy includes: Mark the cleaned areas; Based on the current cleaning actions of unmanned sanitation vehicles, and according to the actual position information and structured environmental pollution information, residual features in uncleaned areas are removed, and the cleaning coverage rate is calculated. ; Obtain the theoretical energy consumption value calculated based on the current cleaning actions and load status of the unmanned sanitation vehicle, compare the actual execution energy consumption with the decision estimate, and calculate the energy consumption deviation. ; Areas that have not been cleaned due to slipping, missed areas, or stubborn stains are marked as a set of paths to be re-cleaned. Based on the set of paths to be re-cleaned and the remaining battery capacity of the unmanned sanitation vehicle, a re-cleaning path is planned. Control the unmanned sanitation vehicle to clean along the re-sweeping path and update the cleaning coverage rate; Calculate the sweep coverage increase based on the sweep coverage before and after the update. ; Based on the increase in cleaning coverage and energy consumption deviation Construct a reward function ; Using the reward function Update the weights of the deep Q network to gradually approach the optimal sweeping strategy.

9. The adaptive cleaning method for unmanned sanitation vehicles based on multi-sensor fusion according to claim 8, characterized in that, The updated weights of the deep Q-network include: Set the playback buffer; A small batch of transfer samples is randomly sampled from the replay buffer using an empirical replay mechanism. ; The target value is calculated using the target network in the deep Q-network. And the value of the current action is calculated using the main network in the deep Q-network; Based on the difference between the target value and the current action value, a loss function is constructed, and the parameters of the main network are updated by minimizing the loss function. The expression for the loss function is: ; ; in, The mean squared error loss function representing the parameters of the main network; This indicates the number of small-batch transfer samples sampled in the experience replay buffer; Indicates the index of the currently processed mini-batch of samples; Indicates the first The target value corresponding to each transferred sample; This indicates the main network's view of the current state. Take action below Q-value evaluation; Indicates the first Instant reward for each transferred sample; Indicates the discount factor; Indicate the next state The action that maximizes the Q-value of the target network evaluation; Indicates the target network's next state and actions Q-value evaluation; Represents the network parameters of the target network; The parameters of the main network are changed every fixed number of steps. Copy to the target network to update the parameters of the target network. .

10. An adaptive cleaning device for unmanned sanitation vehicles based on multi-sensor fusion, characterized in that, include: The data acquisition module is used to acquire environmental data from multiple sensor sources. The data fusion module is used to fuse the environmental data from the multi-source sensors to obtain the pose estimation value of the unmanned sanitation vehicle. The environment construction and fusion module is used to construct an environment state matrix based on the optimal pose state estimation, and to fuse the RGB color images and parallax depth data extracted from the multi-source sensor data into the environment state matrix after identification and classification to obtain structured information on environmental pollution. An adaptive learning module is used to learn and make decisions on adaptive cleaning strategies based on the structured information of environmental pollution using a deep Q-network, and output continuous control variables. The data conversion module is used to convert the continuous control quantity into a motor drive signal through a multi-channel parallel PID regulator in order to regulate the cleaning action of the unmanned sanitation vehicle. The cleaning area determination module is used to obtain the actual position and posture information and actual energy consumption of the unmanned sanitation vehicle based on the current cleaning action of the unmanned sanitation vehicle, and determine the cleaning coverage rate according to the actual position and posture information and the structured information of environmental pollution. The replanning path module is used to plan a resweeping path based on the distribution of the sweeping coverage, and update the sweeping coverage after sweeping along the resweeping path is completed. The optimal strategy generation module is used to construct a reward function based on the difference in cleaning coverage before and after the update and the actual execution energy consumption, and to update the weights of the deep Q network based on the reward function to generate the optimal cleaning strategy.