Smart city environmental sanitation unmanned vehicle path planning and real-time monitoring method
By utilizing multi-source environmental perception data and advanced algorithm technology, efficient path planning and real-time monitoring of smart city sanitation unmanned vehicles are realized, solving the problems of low efficiency, high cost and uneven quality in traditional sanitation operation methods, and achieving efficient, safe and intelligent sanitation operations.
Patent Information
- Application Number
- CN202510571193.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-06
- Publication Date
- 2025-06-03
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Traditional sanitation operations have problems such as low efficiency, high cost and uneven quality, and lack the ability to adapt to real-time road conditions and environmental changes, making it difficult to achieve efficient and safe sanitation operations.
Multi-source environment perception data (lidar point cloud data, camera image sequence and real-time traffic flow information) are used for path planning and real-time monitoring. Through dynamic road network modeling, timing target tracking, path optimization model and reinforcement learning algorithm, a dynamic adjustment strategy network is generated to realize efficient path planning and real-time monitoring of unmanned vehicles.
It has improved the efficiency and quality of sanitation operations, enhanced driving safety, reduced labor costs, realized intelligent and informatized management of sanitation operations, and provided strong support for the construction and development of smart cities.
Smart Images

Figure CN120087580A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of unmanned vehicle control, and specifically to a method for path planning and real-time monitoring of a smart city sanitation unmanned vehicle. Background Art
[0002] Under the background of the construction of a smart city, the intellectualization and high efficiency of sanitation operations have become important development directions. The traditional sanitation operation mode mainly relies on manual cleaning and simple mechanical assistance, with many drawbacks and being difficult to meet the needs of modern urban development. The efficiency of manual cleaning is relatively low. When facing the cleaning tasks of large-area urban roads, a large amount of labor costs need to be invested, and the cleaning speed is slow, unable to quickly respond to sudden environmental sanitation problems. For example, after a garbage spill occurs on the main urban road, it often takes a long time for manual cleaning to restore the road to cleanliness, affecting traffic flow and the city image. At the same time, the quality of manual cleaning also varies, being greatly affected by factors such as the working state and skill level of sanitation workers, and it is difficult to ensure a consistent cleaning effect.
[0003] Although traditional sanitation mechanical vehicles have improved the operation efficiency to a certain extent, there are also many problems. On the one hand, their path planning usually relies on pre-set fixed routes and lacks the ability to adapt to real-time road conditions and environmental changes. When encountering road construction, traffic control, sudden obstacles, etc., the route cannot be adjusted in time, resulting in operation interruption or a significant reduction in efficiency. On the other hand, traditional sanitation vehicles lack effective real-time monitoring means during driving, and it is difficult to accurately grasp the vehicle's operating state, operation progress, and whether there are abnormal situations. Once the vehicle breaks down or deviates from the normal operation range, it cannot be discovered and processed in time, resulting in waste of resources and delays in sanitation operations.
[0004] With the rapid development of the city, the traffic flow is becoming increasingly complex, the road environment is constantly changing, and the requirements for sanitation operations are getting higher and higher. It is not only necessary to improve the cleaning efficiency and quality but also ensure the safety and reliability of the operation process and avoid interfering with the normal traffic order. At the same time, urban managers also hope to be able to obtain relevant data of sanitation operations in real time for scientific scheduling and management. However, the existing sanitation operation methods and technical means are difficult to meet these diverse needs.
[0005] In such a situation, it is urgent to study a method for path planning and real-time monitoring of a smart city sanitation unmanned vehicle that can adapt to complex urban environments, achieve efficient path planning, and conduct real-time monitoring. It will help to improve the intellectualization level of sanitation operations, reduce labor costs, improve operation efficiency and quality, and provide strong support for the construction and development of smart cities. Summary of the Invention
[0006] The purpose of the present invention is to provide a method for path planning and real-time monitoring of a smart city sanitation unmanned vehicle to solve the problems raised in the above-mentioned background technology.
[0007] To achieve the above object, the present invention provides the following technical solution: A method for path planning and real-time monitoring of a smart city sanitation unmanned vehicle, the method comprising: Obtain a multi-source environmental perception data set; the multi-source environmental perception data includes lidar point cloud data, camera image sequences, and real-time traffic flow information; the lidar point cloud data includes obstacle contour features and road surface flatness information, and the camera image sequences include dynamic target trajectories and road sign recognition results; Based on the lidar point cloud data, extract local road network features through a dynamic road network modeling algorithm, and the road network features include obstacle distribution density, boundary coordinates of the passing area, and road topological connection relationships; According to the camera image sequences, generate a dynamic target motion prediction map through a time-series target tracking algorithm, and the map includes a target trajectory prediction path and a collision risk probability value; Perform spatio-temporal slicing processing on the real-time traffic flow information to generate an evolution curve of regional traffic efficiency; Input the local road network features, the dynamic target motion prediction map, and the evolution curve of regional traffic efficiency into a path optimization model to generate an initial path sequence; Based on the initial path sequence, divide the cleaning task priorities through a regional segmentation algorithm, and output a cleaning area weight distribution map; Fuse the cleaning area weight distribution map with the initial path sequence, and construct a dynamic adjustment policy network through a reinforcement learning algorithm to generate a final path planning scheme.
[0008] Preferably, the extraction of local road network features through the dynamic road network modeling algorithm includes: Perform voxel downsampling processing on the lidar point cloud data to generate a sparse point cloud map; Based on the grid division rule, divide the sparse point cloud map into several sub-regions, and calculate the proportion of the obstacle coverage area of each sub-region; Adopt a curvature fitting algorithm to extract the road surface boundary contour, and generate a road topological connection relationship in combination with the boundary coordinates of the passing area; Encode the obstacle distribution density, boundary coordinates of the passing area, and road topological connection relationships as local road network features.
[0009] Preferably, the generation of the dynamic target motion prediction map through the time-series target tracking algorithm includes: Perform inter-frame difference processing on the camera image sequences to extract dynamic target motion vectors; Predict the future position of the target based on the Kalman filtering algorithm and calculate the confidence of the trajectory prediction path; Generate a risk probability model according to the historical collision data distribution and output the collision risk probability value in combination with the confidence; Map the trajectory prediction path and the collision risk probability value into a dynamic target motion prediction atlas.
[0010] Preferably, the path optimization model includes a feature fusion module and a sequence generation module, and the feature fusion module includes: Normalize the obstacle distribution density in the local road network features to obtain a first fusion vector; Perform Gaussian smoothing on the collision risk probability value in the dynamic target motion prediction atlas to generate a second fusion vector; Perform piecewise linear interpolation on the regional traffic efficiency evolution curve, extract the traffic efficiency gradient feature, and obtain a third fusion vector; Combine the first fusion vector, the second fusion vector, and the third fusion vector into an optimized input sequence through a multi-layer perceptron.
[0011] Preferably, the sequence generation module includes: Perform spatio-temporal alignment on the optimized input sequence to generate a path correlation tensor; Extract the dependency relationship between nodes through a graph convolutional network to generate a path node feature matrix; Perform matrix decomposition operation on the path correlation tensor and the path node feature matrix to generate a candidate path set; Select an initial path sequence from the candidate path set through a greedy search algorithm.
[0012] Preferably, constructing a dynamic adjustment policy network through a reinforcement learning algorithm includes: Initialize network nodes according to the clean area weight distribution map and generate a connection weight matrix based on the priority; Use the initial path sequence as the node state, and the connection weight matrix consists of the cleaning task time consumption and the priority weight; Update the policy scores of each node through the value iteration algorithm and adjust the connection weight matrix; Generate an optimal path planning scheme covering all clean areas according to the adjusted connection weight matrix.
[0013] Preferably, the parameter optimization method of the grid division rule includes: Calculate the initial grid size and the minimum obstacle density threshold according to the historical point cloud data distribution; Traverse the parameter combinations through a grid search algorithm and select the parameters with the highest matching degree between the division result and the manually annotated road network; Dynamically adjust the grid size and the minimum obstacle density threshold according to the matching degree, and optimize the division accuracy of the road topology connection relationship.
[0014] Preferably, the method for constructing the risk probability model includes: Collect historical collision event data in multiple scenarios, and extract spatio-temporal distribution characteristics and environmental impact factors; Perform kernel density estimation on the spatio-temporal distribution characteristics to generate a basic risk probability distribution map; Perform weighted correction on the basic risk probability distribution map according to the environmental impact factors, and output the risk probability model.
[0015] Preferably, the method for setting parameters of the value iteration algorithm includes: Define the value score between nodes as the balance coefficient of the cleaning task time-consuming and the priority weight; Initialize the policy scores of each node to zero, and the starting point value score to a preset reference value; Calculate the maximum cumulative value score of each node based on the previous nodes through the dynamic programming algorithm, and record the optimal policy path; Generate a complete path planning scheme by backtracking reversely according to the optimal policy path.
[0016] Preferably, the method further includes: Collect the operation status data of the sanitation unmanned vehicle in real time, generate a path correction instruction set through the anomaly detection model, and update the dynamic adjustment policy network; the method for constructing the anomaly detection model includes: Collect normal mode samples in the historical operation status data, and extract the speed fluctuation characteristics and the steering angle change sequence; Construct a normal behavior benchmark model based on the isolation forest algorithm, and calculate the deviation degree between the real-time data and the benchmark model; Statistically calculate the cumulative anomaly score of the deviation degree through a sliding window, and generate a path correction instruction set when it exceeds the preset threshold; Update the node connection weights of the dynamic adjustment policy network in combination with the correction instruction set.
[0017] Compared with the prior art, the beneficial effects of the present invention are: In terms of path planning, by obtaining a multi-source environmental perception data set, covering lidar point cloud data, camera image sequences, and real-time traffic flow information, various situations of the operating environment of the sanitation unmanned vehicle can be comprehensively and accurately grasped. Based on the local road network features extracted from the lidar point cloud data, such as the distribution density of obstacles, the boundary coordinates of the passing area, and the road topology connection relationship, precise basic road network information is provided for path planning, enabling the unmanned vehicle to know in advance the passable areas and the distribution of potential obstacles, and effectively avoiding dangerous areas. According to the dynamic target motion prediction map generated from the camera image sequence, including the predicted path of the target trajectory and the collision risk probability value, it can help the unmanned vehicle anticipate in advance the possibility of colliding with dynamic targets (such as pedestrians, other vehicles), and thus adjust the path in advance, greatly improving the driving safety. By performing spatio-temporal slicing on the real-time traffic flow information to generate the evolution curve of regional traffic efficiency, the unmanned vehicle can select the optimal passing route according to the real-time road conditions, avoid congested sections, reduce the driving time, and improve the operation efficiency. After generating the initial path sequence through a series of algorithms, combined with the regional segmentation algorithm to divide the priority of cleaning tasks, output the weight distribution map of the cleaning area, and fuse it with the initial path sequence, and use the reinforcement learning algorithm to construct a dynamic adjustment strategy network to generate the final path planning scheme. This method comprehensively considers the urgency of cleaning tasks and the importance of regions, ensuring that the unmanned vehicle can efficiently complete the cleaning tasks and avoid resource waste and time delays.
[0018] In terms of real-time monitoring, the operating state data of the sanitation unmanned vehicle is collected in real time, and a path correction instruction set is generated through an anomaly detection model, which can timely detect abnormal situations during the operation of the unmanned vehicle, such as abnormal speed fluctuations, abnormal steering angles, etc. Once an anomaly is detected, a path correction instruction set is immediately generated to update the dynamic adjustment strategy network, enabling the unmanned vehicle to quickly adjust the driving path, ensuring the continuity and safety of the operation. This not only reduces the losses caused by vehicle failures but also avoids potential threats to the surrounding environment and personnel caused by the abnormal driving of the unmanned vehicle.
[0019] From the perspective of overall sanitation operation management, this method helps to realize the intelligent and information-based management of sanitation operations. Urban managers can obtain information such as the location, operation progress, and operating state of the unmanned vehicle in real time through relevant systems, which is convenient for unified scheduling and resource allocation, improving the scientific and refined level of urban sanitation management. Moreover, the application of this method can effectively reduce labor costs, reduce the risks of sanitation workers operating in complex and dangerous environments, enhance the overall image and development level of the sanitation industry, and provide a strong guarantee for the sustainable development of smart cities. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Figure 1 It is the working principle diagram of the path planning and real-time monitoring method for the smart city sanitation unmanned vehicle described in the present invention; Figure 2 Flowchart for extracting local road network features by a dynamic road network modeling algorithm; Figure 3 Flowchart for generating a dynamic target motion prediction map by a time-series target tracking algorithm; Figure 4 Flowchart for a path optimization model sequence generation module. Specific implementation manners
[0021] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0022] Please refer to Figures 1-4 , the path planning and real-time monitoring method for a smart city sanitation unmanned vehicle provided by the present invention is specifically implemented as follows: Through sensors such as lidar and cameras mounted on the sanitation unmanned vehicle, multi-source environmental perception data is obtained. This data set includes lidar point cloud data, camera image sequences, and real-time traffic flow information. Among them, lidar point cloud data can accurately obtain the contour features of obstacles and the road surface flatness information, camera image sequences can identify dynamic target trajectories and road signs, and real-time traffic flow information reflects the current road traffic conditions.
[0023] The lidar point cloud data is processed using a dynamic road network modeling algorithm to extract local road network features, including obstacle distribution density, passing area boundary coordinates, and road topological connection relationships, providing basic road network information for subsequent path planning.
[0024] The camera image sequences are analyzed using a time-series target tracking algorithm to generate a dynamic target motion prediction map, which includes target trajectory prediction paths and collision risk probability values, used to evaluate the possibility of collision between the unmanned vehicle and dynamic targets during driving.
[0025] The obtained real-time traffic flow information is subjected to spatio-temporal slicing processing, and through specific algorithms and analyses, a regional traffic efficiency evolution curve is generated to intuitively reflect the changes in traffic efficiency in different regions at different times.
[0026] The extracted local road network features, the generated dynamic target motion prediction map, and the regional traffic efficiency evolution curve are input into the path optimization model. After a series of operations and processing, an initial path sequence is generated to preliminarily plan the driving path of the unmanned vehicle.
[0027] Based on the generated initial path sequence, the regional segmentation algorithm is used to divide the cleaning task priorities according to factors such as the urgency of the cleaning task and the area of the region, and then output the cleaning area weight distribution map to clarify the importance of each cleaning area.
[0028] Fuse the cleaning area weight distribution map with the initial path sequence, and construct a dynamic adjustment policy network through the reinforcement learning algorithm. After continuous learning and optimization, this network generates the final path planning scheme, enabling the sanitation unmanned vehicle to complete the cleaning task efficiently and safely.
[0029] Example 1: After obtaining the multi-source environmental perception data set, when extracting the local road network features from the lidar point cloud data, the following specific operations are adopted: Perform voxel downsampling on the lidar point cloud data. During the operation of the lidar, a large amount of point cloud data will be generated. This data volume is huge and contains a lot of redundant information. Through voxel downsampling, the three-dimensional space is divided into small voxel units. Within each voxel unit, multiple points are merged into a representative point, thereby generating a sparse point cloud map, which can not only reduce the data volume but also retain the key geometric features.
[0030] Divide the sparse point cloud map into several sub-regions based on the grid division rule. Here, a suitable grid size is set, for example, the side length of each grid is a meters. After the division, calculate the proportion of the obstacle coverage area in each sub-region. The specific calculation method is to count the number of grids occupied by obstacles n in each sub-region, and then divide it by the total number of grids N in this sub-region to obtain the proportion of the obstacle coverage area This proportion can intuitively reflect the density of obstacles in each sub-region and is an important basis for subsequent analysis.
[0031] Adopt the curvature fitting algorithm to extract the road surface boundary contour. Since the road surface boundary presents certain geometric features in the point cloud data, the curvature fitting algorithm can fit these points to obtain an accurate road surface boundary curve. Combining the previously determined boundary coordinates of the passing area, use relevant methods such as graph theory to generate the road topology connection relationship. The road topology connection relationship describes the connectivity between roads, such as which roads are connected to each other, the connection methods and sequences, etc.
[0032] Encode the obstacle distribution density (i.e., the proportion of the obstacle coverage area in each sub-region calculated above), the boundary coordinates of the passing area, and the road topology connection relationship to form local road network features. This encoding method can convert complex road network information into a data structure that is easy for the computer to process, facilitating the use of subsequent path planning algorithms.
[0033] Embodiment 2: When generating a dynamic target motion prediction map based on a camera image sequence, the following specific steps are taken: Perform inter-frame difference processing on the camera image sequence. In the image sequence continuously captured by the camera, there are certain differences between adjacent two frames of images, and these differences are mainly caused by the motion of dynamic targets. By subtracting each pixel of adjacent two frames of images, a difference image is obtained. In the difference image, the positions of dynamic targets will show obvious changes. By setting an appropriate threshold, these changed regions are extracted, thereby obtaining the motion vectors of dynamic targets. The motion vectors contain the motion direction and speed information of dynamic targets.
[0034] Predict the future position of the target based on the Kalman filter algorithm. The Kalman filter algorithm is an efficient recursive filter, which can predict the future state of the system according to the current state of the system and the observation data. In this embodiment, the motion vectors of dynamic targets are used as the observation data and input into the Kalman filter algorithm model. Assume that the state of the dynamic target at the current moment is , including information such as position and speed, and the observation data is . According to the prediction formula of the Kalman filter algorithm (where is the state at the current moment predicted based on the state at the previous moment, is the state transition matrix, which describes the change law of the system state over time; is the optimal estimated state at the previous moment; is the control input matrix, which is generally 0 in this scenario; is the control input at the previous moment, and there is no control input in this scenario, so it is also 0) and the update formula (where is the optimal estimated state at the current moment, is the Kalman gain, which is used to balance the weights of the predicted value and the observed value; is the observation matrix, which is used to convert the system state into the form of observation data), the future position of the target can be predicted. At the same time, by calculating parameters such as the covariance of the prediction result, the confidence of the trajectory prediction path is obtained. The confidence reflects the reliability of the prediction result.
[0035] Generate a risk probability model according to the historical collision data distribution. Collect a large amount of historical collision event data in multiple scenarios, analyze these data, and extract the spatio-temporal distribution characteristics and environmental impact factors. The spatio-temporal distribution characteristics include information such as the time and location of the collision, and the environmental impact factors such as weather conditions and road types. Perform kernel density estimation on the spatio-temporal distribution characteristics to generate a basic risk probability distribution map. Kernel density estimation is a non-parametric estimation method, which obtains the probability density estimation by placing kernel functions around the data points and then performing weighted summation on these kernel functions. Assume that the historical collision data points are , the kernel function is , and the bandwidth is , then the estimation formula for the basic risk probability distribution map is , where is the number of data points). Then, the basic risk probability distribution map is weighted and corrected according to environmental impact factors. For example, if the probability of a collision is higher on rainy days, the weight of the environmental impact factor on rainy days will increase accordingly, and finally, a risk probability model is output.
[0036] Map the trajectory prediction path to the collision risk probability value to generate a dynamic target motion prediction map. In this way, the motion trend of the dynamic target and the risk degree of collision with the unmanned vehicle can be intuitively displayed, providing an important reference for path planning.
[0037] Embodiment 3: The feature fusion module in the path optimization model operates as follows: For the obstacle distribution density in the local road network features, since its value range may be different, in order to facilitate subsequent unified processing and comparison, normalization processing is required. Assume that the original value of the obstacle distribution density is , and its value range is , through the normalization formula , map it to the interval to obtain the first fusion vector. After normalization processing, the obstacle distribution densities in different regions have a unified measurement standard.
[0038] For the collision risk probability value in the dynamic target motion prediction map, in order to make it smoother and reduce the influence of noise, Gaussian smoothing processing is adopted. When performing Gaussian smoothing processing on the collision risk probability value, at each point , calculate the weighted average value of the points in its neighborhood to generate the second fusion vector. This can make the change of the collision risk probability value more continuous and stable.
[0039] For the regional traffic efficiency evolution curve, in order to extract its traffic efficiency gradient feature, a piecewise linear interpolation method is adopted. The regional traffic efficiency evolution curve is segmented according to time or other appropriate dimensions. Within each segment, through the known two endpoint values, use the linear interpolation formula (where and are the two endpoints of the line segment, is the point to be interpolated, is the interpolation result) to calculate the value of the intermediate point, so as to obtain more dense traffic efficiency data. Then, perform a derivative operation on these data to obtain the traffic efficiency gradient feature, forming the third fusion vector. The traffic efficiency gradient feature reflects the speed of change of traffic efficiency over time or space.
[0040] Finally, the first fusion vector, the second fusion vector, and the third fusion vector are combined into an optimized input sequence through a multi-layer perceptron. A multi-layer perceptron is a neural network structure containing multiple neurons, which can perform complex non-linear transformations on the input data. Taking the three fusion vectors as the input of the multi-layer perceptron, after the calculation of the hidden layer and the processing of the activation function, a unified optimized input sequence is finally output, providing comprehensive feature data for subsequent path generation.
[0041] Example 4: The sequence generation module in the path optimization model is specifically implemented as follows: Perform spatio-temporal alignment on the optimized input sequence. The optimized input sequence contains information from different data sources, and these information may vary in time and space. Through spatio-temporal alignment operations, this information is adjusted according to a unified time and space benchmark, enabling them to match and correlate with each other. For example, the road network features, dynamic target motion prediction maps, and regional traffic efficiency evolution curve data collected at different times are aligned according to the driving time and position of the unmanned vehicle to generate a path correlation tensor. A path correlation tensor is a mathematical structure that can simultaneously describe the spatio-temporal relationships between multiple data.
[0042] Extract the dependencies between nodes through a graph convolutional network. A graph convolutional network is a neural network specifically designed for processing graph-structured data. Consider the path correlation tensor as a graph structure, where each node represents the information of a location or a time period, and the edges between nodes represent the association relationships between them. The graph convolutional network extracts the dependencies between nodes through convolutional operations on the features of nodes and edges. Assume the input of the graph convolutional network is (node feature matrix), and the convolutional kernel is , then the output after graph convolutional operation is (where is the adjacency matrix of the graph, representing the connection relationships between nodes, is the activation function, used to introduce non-linearity), generating a path node feature matrix. The path node feature matrix contains the dependency relationship information between each node and other nodes.
[0043] Perform matrix factorization operations on the path correlation tensor and the path node feature matrix. Matrix factorization is a method of decomposing a matrix into the product of multiple low-rank matrices. Common matrix factorization methods include singular value decomposition (SVD), etc. Through matrix factorization operations, the high-dimensional path correlation tensor and path node feature matrix can be decomposed into multiple low-dimensional matrices, and these low-dimensional matrices contain the key information of the original matrix. Through this decomposition, a candidate path set is generated. The candidate path set contains multiple possible path planning schemes.
[0044] Select an initial path sequence from the candidate path set through a greedy search algorithm. The greedy search algorithm is a heuristic search algorithm based on local optimal selection. In the candidate path set, calculate a certain evaluation index (such as path length, travel time, etc.) for each path in turn, and select the path with the optimal current evaluation index as the initial path sequence. For example, if the path length is used as the evaluation index, then select the path with the shortest path length as the initial path sequence. In this way, a relatively optimal initial path sequence can be quickly selected from numerous candidate paths, providing a basis for subsequent path optimization.
[0045] Embodiment 5: When constructing a dynamic adjustment policy network through a reinforcement learning algorithm, the following specific steps are adopted: Initialize network nodes according to the cleaning area weight distribution map. The cleaning area weight distribution map reflects the importance of each cleaning area, and each cleaning area is regarded as a node in the dynamic adjustment policy network. Generate a connection weight matrix based on priorities. The elements in the connection weight matrix represent the connection strength between different nodes. Suppose the cleaning area and the cleaning area The connection weight between them is , if the cleaning area has a higher priority and is closer to the cleaning area , then The value of will be relatively large. The connection weight matrix consists of the cleaning task time consumption and the priority weight. The cleaning task time consumption can be estimated based on historical data or experience, and the priority weight is determined according to the cleaning area weight distribution map.
[0046] Take the initial path sequence as the node state. The initial path sequence contains the driving path information preliminarily planned by the unmanned vehicle. Take it as the initial state of the network node for subsequent policy updates. In this state, the unmanned vehicle is at a certain position on the initial path, and the network needs to make decisions based on the current state and environmental information to adjust the path.
[0047] Update the policy scores of each node through the value iteration algorithm. The value iteration algorithm is a classic algorithm for solving Markov decision processes. Define the value score between nodes as the balance coefficient of the cleaning task time consumption and the priority weight. Let the balance coefficient be , the cleaning task time consumption is , the priority weight is , then the value score . Initialize the policy scores of each node to zero, and the starting point value score is the preset reference value . Calculate the maximum cumulative value score of each node based on the previous nodes through the dynamic programming algorithm. Suppose from node to node The transition probability is , the immediate reward is , then for node , the maximum cumulative value score (where is the discount factor, used to balance the importance of current and future rewards), and record the optimal policy path. During the calculation process, continuously update the value scores of each node until convergence.
[0048] Generate an optimal path planning scheme covering all cleaning areas according to the adjusted connection weight matrix. When the value iteration algorithm converges, the obtained connection weight matrix reflects the optimal path selection strategy. According to this matrix, starting from the starting point, sequentially select the next node according to the optimal policy path to generate an optimal path planning scheme covering all cleaning areas. In this way, through the dynamically adjusted policy network constructed by the reinforcement learning algorithm, it is possible to optimize the initial path sequence according to the priority and actual situation of the cleaning area, and obtain a more efficient path planning scheme.
[0049] Example 6: The method of the present invention further includes collecting the operation state data of the sanitation unmanned vehicle in real time, and generating a path correction instruction set through an anomaly detection model to update the dynamically adjusted policy network. The specific implementation is as follows: Collect normal mode samples from historical operation state data. During the daily operation of the sanitation unmanned vehicle, collect a large amount of operation state data, which includes information such as speed, steering angle, acceleration, etc. Screen out the data samples under normal operation conditions from these data as the basis for subsequent analysis.
[0050] Extract the speed fluctuation characteristics and the steering angle change sequence. For the speed fluctuation characteristics, calculate the change rate of the speed within a period of time. For example, within the time interval , the speed changes from to , and the speed change rate is . By statistically calculating multiple such speed change rates, obtain the speed fluctuation characteristics. For the steering angle change sequence, record the change of the steering angle of the unmanned vehicle over time during driving to form a steering angle change sequence.
[0051] Construct a normal behavior benchmark model based on the Isolation Forest algorithm. The Isolation Forest algorithm is an anomaly detection algorithm based on the isolation idea. It constructs a binary tree and uses the path length of each data point in the tree as an index to measure its degree of abnormality. For the speed fluctuation characteristics and steering angle change sequence data in the normal mode samples, use the Isolation Forest algorithm to construct a normal behavior benchmark model. During the construction process, the algorithm will automatically learn the distribution characteristics of normal data. Calculate the deviation degree between the real-time data and the benchmark model. Assume the real-time data point is , the path length in the isolation forest is , the average path length of normal data points is , then the deviation degree .
[0052] The cumulative anomaly score of the deviation degree is statistically calculated through a sliding window. Set a sliding window size , within each time window, calculate the cumulative sum of the deviation degree as the cumulative anomaly score , where is the deviation degree of the -th data point within the window. When the cumulative anomaly score exceeds the preset threshold , it is determined that the running state of the driverless vehicle is abnormal, and a path correction instruction set is generated. The path correction instruction set contains information for adjusting the driving path of the driverless vehicle, such as changing the driving direction, adjusting the speed, etc.
[0053] Update the node connection weights of the dynamic adjustment policy network in combination with the correction instruction set. Input the path correction instruction set into the dynamic adjustment policy network, and adjust the node connection weights of the network according to the content of the instruction. For example, if the instruction requires the driverless vehicle to avoid a certain obstacle, then the node connection weights in the area related to the obstacle are adjusted accordingly, so that the driverless vehicle can avoid the obstacle during subsequent driving, thereby ensuring the safe and efficient operation of the driverless vehicle.
[0054] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device.
[0055] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principle and spirit of the present invention, and the scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A method for path planning and real-time monitoring of a smart city sanitation unmanned vehicle, characterized in that: include: Acquire a multi-source environmental perception data set; the multi-source environmental perception data includes laser radar point cloud data, camera image sequence and real-time traffic flow information; the laser radar point cloud data includes obstacle contour features and road surface flatness information, and the camera image sequence includes dynamic target trajectory and road sign recognition results; Based on the laser radar point cloud data, extract local road network features through a dynamic road network modeling algorithm, wherein the road network features include obstacle distribution density, boundary coordinates of the pass area, and road topological connection relationships; According to the camera image sequence, a dynamic target motion prediction map is generated by a time-series target tracking algorithm, wherein the map includes a target trajectory prediction path and a collision risk probability value; The real-time traffic flow information is processed by time-space slicing to generate a regional traffic efficiency evolution curve; Inputting the local road network characteristics, dynamic target motion prediction map and regional traffic efficiency evolution curve into the path optimization model to generate an initial path sequence; Based on the initial path sequence, the cleaning task priorities are divided by a region segmentation algorithm, and a cleaning region weight distribution map is output; The cleaning area weight distribution map is fused with the initial path sequence, and a dynamic adjustment strategy network is constructed through a reinforcement learning algorithm to generate a final path planning solution.
2. A method for path planning and real-time monitoring of a smart city sanitation unmanned vehicle according to claim 1, characterized in that: The extracting of local road network features by a dynamic road network modeling algorithm includes: Performing voxel downsampling processing on the laser radar point cloud data to generate a sparse point cloud map; Divide the sparse point cloud map into several sub-areas based on the grid division rules, and calculate the obstacle coverage area ratio of each sub-area; The curvature fitting algorithm is used to extract the road boundary contour, and the road topological connection relationship is generated by combining the boundary coordinates of the traffic area; The obstacle distribution density, the boundary coordinates of the traffic area and the road topological connection relationship are encoded as local road network features.
3. A method for path planning and real-time monitoring of a smart city sanitation unmanned vehicle according to claim 1, characterized in that: The method of generating a dynamic target motion prediction map by using a time series target tracking algorithm includes: Performing inter-frame difference processing on the camera image sequence to extract the dynamic target motion vector; Predict the target's future position based on the Kalman filter algorithm and calculate the confidence of the trajectory prediction path; Generate a risk probability model based on the distribution of historical collision data, and output a collision risk probability value based on the confidence level; The trajectory prediction path and collision risk probability value are mapped into a dynamic target motion prediction map.
4. A method for path planning and real-time monitoring of a smart city sanitation unmanned vehicle according to claim 1, characterized in that: The path optimization model includes a feature fusion module and a sequence generation module, and the feature fusion module includes: Normalizing the obstacle distribution density in the local road network feature to obtain a first fusion vector; Performing Gaussian smoothing on the collision risk probability value in the dynamic target motion prediction map to generate a second fusion vector; Performing piecewise linear interpolation on the regional traffic efficiency evolution curve, extracting traffic efficiency gradient features, and obtaining a third fusion vector; The first fusion vector, the second fusion vector and the third fusion vector are combined into an optimized input sequence through a multi-layer perceptron.
5. A method for path planning and real-time monitoring of a smart city sanitation unmanned vehicle according to claim 4, characterized in that: The sequence generation module comprises: Perform spatiotemporal alignment on the optimized input sequence to generate a path association tensor; The dependency relationship between nodes is extracted through the graph convolutional network to generate the path node feature matrix; Perform matrix decomposition operation on the path association tensor and the path node feature matrix to generate a set of candidate paths; The initial path sequence is selected from the candidate path set through a greedy search algorithm.
6. A method for path planning and real-time monitoring of a smart city sanitation unmanned vehicle according to claim 1, characterized in that: The method of constructing a dynamic adjustment strategy network through a reinforcement learning algorithm includes: Initialize network nodes according to the clean area weight distribution map and generate a connection weight matrix based on the priority; The initial path sequence is used as the node state, and the connection weight matrix is composed of the cleaning task time consumption and priority weight; Update the strategy score of each node through the value iteration algorithm and adjust the connection weight matrix; The optimal path planning scheme covering all cleaning areas is generated based on the adjusted connection weight matrix.
7. A method for path planning and real-time monitoring of a smart city sanitation unmanned vehicle according to claim 2, characterized in that: The parameter optimization method of the grid division rule includes: Calculate the initial grid size and minimum obstacle density threshold based on the historical point cloud data distribution; The grid search algorithm is used to traverse the parameter combinations and select the parameters with the highest matching degree between the partitioning result and the manually annotated road network; The grid size and minimum obstacle density threshold are dynamically adjusted according to the matching degree to optimize the accuracy of road topology connection relationship division.
8. A method for path planning and real-time monitoring of a smart city sanitation unmanned vehicle according to claim 3, characterized in that: The method for constructing the risk probability model includes: Collect historical collision event data in multiple scenarios and extract spatiotemporal distribution characteristics and environmental influencing factors; Conduct kernel density estimation on spatiotemporal distribution characteristics to generate a basic risk probability distribution map; The basic risk probability distribution map is weighted and modified according to environmental influencing factors to output the risk probability model.
9. A method for path planning and real-time monitoring of a smart city sanitation unmanned vehicle according to claim 6, characterized in that: The parameter setting method of the value iteration algorithm includes: The value score between nodes is defined as the balance coefficient between the cleaning task time consumption and the priority weight; Initialize the strategy score of each node to zero, and the starting value score to the preset benchmark value; Calculate the maximum cumulative value score of each node based on the previous node through the dynamic programming algorithm, and record the optimal strategy path; Generate a complete path planning solution based on reverse backtracking of the optimal strategy path.
10. A method for path planning and real-time monitoring of a smart city sanitation unmanned vehicle according to claim 1, characterized in that: The method further comprises: The operating status data of the sanitation unmanned vehicle is collected in real time, a path correction instruction set is generated through an anomaly detection model, and a dynamic adjustment strategy network is updated; the construction method of the anomaly detection model includes: Collect normal mode samples from historical operating status data and extract speed fluctuation characteristics and steering angle change sequences; A normal behavior benchmark model is constructed based on the isolation forest algorithm, and the deviation between the real-time data and the benchmark model is calculated; The accumulated anomaly score of the deviation is calculated through a sliding window, and a path correction instruction set is generated when it exceeds a preset threshold. The node connection weights of the policy network are dynamically adjusted by updating the revised instruction set.
Citation Information
Patent Citations
Sanitation robot vehicle scheduling method and system based on vehicle infrastructure cooperation and reinforcement learning
CN116611635A
A method and device for safe navigation of unmanned vehicles in controlled road sections
CN116804560A
Complex scene path planning generation method and system based on multiple models
CN119687955A
Unmanned aerial vehicle intelligent inspection path generation method for interior of building based on BIM model feature extraction
CN119849724A
Unmanned sanitation vehicle path planning method based on artificial intelligence
CN119901306A
Cited By
Data preset learning-based dynamic material allocation system for multiple stations
CN120317643A
Intelligent inspection robot path optimization method and system based on edge reasoning model
CN120335455A
Street cleanliness real-time evaluation method based on multi-modal data fusion
CN120449111A
Real-time evaluation method of street cleanliness based on multimodal data fusion
CN120449111B
Map dynamic construction method and system fused with laser vision
CN120506938A