Snow sweeping and deicing path control method based on vehicle track dynamic feedback

By combining Kafka and WebSocket technologies to access snowplow GPS data in real time and using the Q-learning algorithm to optimize snowplow routes, the problems of inflexible route planning and vehicle scheduling delays in traditional snow removal and de-icing operations have been solved, enabling real-time, efficient, and intelligent de-icing operations.

CN120998051APending Publication Date: 2025-11-21浪潮智慧城市科技有限公司 +1
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202511152743.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-18
Publication Date
2025-11-21

AI Technical Summary

Technical Problem

Traditional snow removal and ice clearing operations suffer from static path planning, leading to blind spots or repetitive work, and vehicle dispatch response is delayed, making it difficult to cope with sudden snowfall.

Method used

By combining Kafka technology with WebSocket communication to access GPS data from snowplows in real time, and constructing reward and state-action value functions through Q-learning reinforcement learning algorithms, the operation path of snowplows is dynamically optimized. Combined with road priority and energy consumption parameters, intelligent allocation of path strategies is achieved.

Benefits of technology

It achieves real-time, intelligent, and energy-efficient optimization of snow removal operations, reduces redundant and invalid routes, ensures a balance between prioritizing key road sections and covering the entire road network, and possesses fault tolerance and efficient route update capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120998051A_ABST
    Figure CN120998051A_ABST
Patent Text Reader

Abstract

The invention provides a snow sweeping and deicing path control method based on vehicle track dynamic feedback, and belongs to the technical field of smart city traffic management.The snow sweeping and deicing path control method comprises the steps that a kafka technical component and a websocket communication technology are combined, GPS data of a snow removing vehicle are accessed in real time, the vehicle driving track is dynamically displayed, and urban area operation progress is visited; a Q-learning reinforcement learning algorithm is introduced, parameters such as road priority and operation vehicle energy consumption are combined, dynamic allocation of the optimal operation path strategy of the snow removal truck is achieved by constructing a reward function and a continuous iteration state-action value function (Q value function), and the effect of dynamic balance of snow removal operation efficiency and operation vehicle energy consumption is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of smart city traffic management technology, and involves technologies such as Apache Kafka real-time data access, WebSocket communication technology, Q-learning reinforcement learning algorithm, Python, Java Springboot backend architecture, etc. Specifically, it is a snow removal and de-icing path control method based on dynamic feedback of vehicle trajectory. Background Technology

[0002] Snow and ice accumulation on roads in winter pose a significant threat to traffic safety. Traditional snow removal and de-icing operations typically rely on planned fixed routes or manual dispatching, which has obvious limitations: static route planning easily leads to blind spots or repetitive work, lacks a dynamic weight allocation mechanism based on spatiotemporal characteristics, and traditional vehicle dispatching methods have long response delays, making it difficult to cope with sudden changes in conditions such as sudden snowfall. Therefore, it is necessary to propose a more flexible and efficient technical solution to address these problems. Summary of the Invention

[0003] To address the above technical problems, this invention provides a snow removal and de-icing path control method based on dynamic feedback of vehicle trajectory.

[0004] The technical solution of this invention is:

[0005] A snow removal and de-icing path control method based on dynamic feedback of vehicle trajectory is proposed. This method combines Kafka technology with WebSocket communication technology to access GPS data of snowplows in real time and dynamically display vehicle driving trajectories, providing an overview of urban operation progress. A Q-learning reinforcement learning algorithm is introduced, combining parameters such as road priority and vehicle energy consumption. By constructing a reward function and a continuously iterative state-action value function (Q-value function), the optimal operation path strategy for snowplows is dynamically allocated, achieving a dynamic balance between snow removal efficiency and vehicle energy consumption.

[0006] Specifically, it includes:

[0007] Real-time vehicle trajectory rendering: The system efficiently consumes and persistently stores GPS data streams from the snowplow using a Kafka distributed message queue. Simultaneously, a two-way communication channel is established using the WebSocket protocol to push real-time vehicle location information to the cockpit system within seconds. The cockpit dynamically renders the vehicle's real-time location based on a geographic information system and supports time-axis-driven historical trajectory playback.

[0008] Q-Learning Model Construction: Based on the task objective and dynamic changes in the environment, the state space and action space are defined, and the reward function is designed by combining parameters such as snow thickness, vehicle energy consumption and road priority with task constraints.

[0009] Dynamic feedback mechanism: Design dynamic parameters such as discount factors, based on the Q-value update formula, and through continuous iteration to update the Q-value, the agent can eventually learn the optimal strategy to take different actions in different states, generate work instructions and push them synchronously to the cockpit system and mobile terminal.

[0010] Furthermore,

[0011] Kafka data access and storage, along with a WebSocket bidirectional communication mechanism, enable real-time rendering of vehicle location data and driving trajectories.

[0012] The location data of the vehicle monitored in the cockpit is loaded and displayed on the GIS map, and the location changes dynamically in real time; based on the full GPS data of the vehicle on the same day, the trajectory is rendered and the driving trajectory of the day is played according to the data access time.

[0013] Furthermore,

[0014] Define the state space and action space based on the actual environment; plan reward and punishment conditions and task constraints, and design the reward function.

[0015] Furthermore,

[0016] Based on the real-time status of the vehicle, the Q value is periodically and continuously updated to generate the optimal path strategy, which is then persistently stored and pushed to the cockpit system or mobile terminal.

[0017] Dynamic feedback mechanism: Real-time recording of vehicle position and operation parameters, updating the status-action value function (Q-value function) every five minutes, generating path correction instructions and simultaneously pushing them to the cockpit system and mobile terminal, supporting manual confirmation or adjustment by the driver.

[0018] The system periodically receives the current status of the snowplow. When a command conflict is detected, it immediately triggers the Q-table to be recalculated and updated, and a correction command is issued.

[0019] When communication is interrupted, switch to the locally stored Q table or predefined rules to perform the job.

[0020] The beneficial effects of this invention are

[0021] This invention discloses a snow removal and de-icing path control method based on dynamic feedback of vehicle trajectory. This technical solution can dynamically and efficiently update snow removal operation paths, achieving a balance between prioritizing key road sections and ensuring full network coverage. Specific beneficial effects are as follows:

[0022] 1. Real-time performance: A real-time data pipeline is built using the Kafka distributed message queue to realize the collection and transmission of GPS data from snowplows. Through the processing and rendering of GPS data, the system can respond to the progress of snow removal operations in seconds.

[0023] 2. Intelligent: The Q-learning reinforcement learning algorithm is used to build an adaptive decision-making model. By defining a multi-dimensional state space such as road priority and snow thickness, and combining a dynamic reward function, the operation path is optimized in real time, which greatly reduces the rate of invalid path repetition and realizes dynamic priority skipping of emergency road sections.

[0024] 3. Energy saving: The optimized path generated by the algorithm can effectively shorten the total mileage of a single vehicle operation, reduce the invalid turnaround distance of snow removal vehicles, and reduce the total energy consumption of vehicles.

[0025] 4. Fault Tolerance: When a command conflict is detected, such as overlapping vehicle paths, the algorithm can be immediately triggered to recalculate and update, and a correction command can be issued. When a vehicle signal is lost, the locally stored Q table can be switched to call the historical best path strategy, maintaining a path planning reliability of over 80%. Attached Figure Description

[0026] Figure 1 This is a schematic diagram of the workflow of the present invention. Detailed Implementation

[0027] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are some embodiments of the present invention, but not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.

[0028] This invention discloses a snow removal and de-icing path control method based on dynamic feedback of vehicle trajectory, including...

[0029] (1) Real-time vehicle trajectory rendering: The snowplow GPS data is consumed and stored from the traffic management platform in real time through the Kafka component, and the latest location data is pushed to the cockpit system through the WebSocket protocol. The cockpit dynamically loads the vehicle location and renders the vehicle's driving trajectory playback for the day.

[0030] (2) Q-Learning model construction: Based on the task objectives and dynamic changes in the environment, the state space and action space are defined, and the reward function is designed by combining parameters such as snow thickness, vehicle energy consumption, and road priority.

[0031] (3) Dynamic feedback mechanism: Real-time recording of vehicle position and operation parameters, updating the status-action value function (Q value function) every five minutes, generating path correction instructions and simultaneously pushing them to the cockpit system and mobile terminal, supporting manual confirmation or adjustment by the driver.

[0032] The specific implementation scheme of the present invention is as follows:

[0033] Figure 1 This is a schematic diagram of the workflow of the present invention. This method is applicable to dynamically loading vehicle locations in the cockpit and rendering playback of the vehicle's daily driving trajectory.

[0034] like Figure 1 As shown, the present invention includes the following steps:

[0035] 1. Consume and store real-time GPS data of snowplows from the traffic management platform using the Kafka component. Key information includes latitude and longitude, speed, vehicle type, and vehicle identification.

[0036] 2. The latest location data of the snowplow is polled and pushed to the cockpit system via WebSocket communication technology, with a push frequency of once every 5 seconds.

[0037] 3. The location data of the snowplows monitored in the cockpit is loaded and displayed on the GIS map, with real-time dynamic changes in location. Based on the full GPS data of the snowplows for the day, trajectory rendering is performed, and the driving trajectory for the day is played back according to the data access time.

[0038] The Q-Learning model construction provided by this invention is suitable for introducing the Q-Learning reinforcement learning algorithm and designing a reward function based on environmental state parameters. It includes the following steps:

[0039] 1. Example of State Space Definition: The state of a snowplow consists of its environmental perception and its own state.

[0040] S = (x, y, s, e)

[0041]

[0042] 2. Example of Action Space Definition:

[0043] A={North,South,East,West,Clean,Charge}

[0044] Vehicle movement: The target grid is free of obstacles and the path is reachable.

[0045] North:(x,y)->(x-1,y)

[0046] South:(x,y)->(x+1,y)

[0047] East:(x,y)->(x,y+1)

[0048] West:(x,y)->(x,y-1)

[0049] Energy consumption model: Each move consumes 0.5% of the electricity.

[0050] Clean action: Current grid s>=1 and e>0

[0051] Charging action (Charge): Currently located in the charging station grid.

[0052] 3. Example of reward function (R) definition:

[0053] (1) Definition of the basic reward function (Rbase):

[0054] Rbase=λ1s–λ2(d+s 2 )+λ3(1-e)-λ4t

[0055] Snow removal reward: λ1 = 15, the thicker the snow, the higher the reward;

[0056] Energy consumption penalty: λ2=0.5, d is the movement distance within the grid, calculate the movement energy consumption and snow removal energy consumption;

[0057] Charging reward: λ3 = 20, reward for successfully returning to the charging station;

[0058] Delay penalty: λ4 = 0.1, t is the duration of the area without snow removal, and the penalty is +6 for each hour of unprocessed snow.

[0059] (2) Penalties for Abnormal Situations:

[0060] Rpenalty=-10: Repeat clearing areas that have no snow.

[0061] Rpenalty = -0.1s 3 Battery depleted

[0062] Rpenalty = -5: Collision with obstacles

[0063] (3) Road priority reward:

[0064] Rlevel=50: Key areas such as hospitals, fire exits, and schools.

[0065] Rlevel=30: Main Road

[0066] Rlevel=0: Other roads

[0067] (4) Total reward function:

[0068] R = Rbase + Rpenalty + Rlevel

[0069] (5) Example of a reward matrix: Taking a 3*3 grid as an example

[0070]

[0071]

[0072] The dynamic feedback mechanism of this invention is applicable to state-action value function (Q-value function) updates, generating path correction instructions and synchronously pushing them to the cockpit system and mobile terminal, supporting manual confirmation or adjustment by the driver. It includes the following steps:

[0073] 1. Q-value update formula:

[0074]

[0075] (1) Q(s,a) is the Q value under state s and action a;

[0076] (2) R represents instant reward;

[0077] (3) α is the learning rate, which determines the degree to which new information updates the Q value;

[0078] (4) γ is a discount factor used to represent the weight of future rewards;

[0079] (5) This represents the maximum Q value among all possible actions to take in the next state s'; by continuously iterating and updating the Q value, the agent can eventually learn the optimal strategy for taking different actions in different states, thereby achieving dynamic decision-making.

[0080] 2. Dynamic parameter configuration:

[0081] state Learning rate α Discount factor γ illustrate High snow cover s = 3 0.3 0.9 Accelerate learning emergency task Low battery e = 0 0.1 0.7 Conservative strategy to avoid running out of power td>2 0.25 0.9 Prioritize backlogged tasks Normal state 0.2 0.85 default

[0082] 3. Strategy Feedback:

[0083] By receiving real-time road conditions from snowplow equipment, the Q-table or action priority is dynamically adjusted to ensure the real-time performance of the strategy. The trained Q-table (convergence matrix) is used to output and store the optimal strategy based on the current vehicle state, and then distributed and pushed to the cockpit system and mobile terminal in real time. The driver then performs snow removal operations according to the issued strategy.

[0084] The system periodically receives the current status of snowplows (location, snow removal progress, etc.). When a command conflict is detected, such as overlapping vehicle paths, the system immediately triggers the recalculation and update of the Q table and issues a correction command.

[0085] When communication is interrupted, switch to the locally stored Q table or predefined rules to perform the job.

[0086] This invention processes real-time GPS data from snowplows, onboard sensor data, and road network information to form a multi-dimensional dynamic road perception network. Based on reinforcement learning algorithms, a composite reward function incorporating traffic flow weights and vehicle energy consumption is designed to enable the agent to autonomously learn optimal decision-making strategies during simulation training. Through a Markov decision process, dynamic path replanning is performed at 5-10 minute intervals to ensure adaptability to path optimization in the event of sudden snowfall or changes in traffic conditions. Compared to traditional fixed-route operations, this method significantly reduces energy consumption, improves snow removal efficiency, and achieves a balance between prioritizing key road sections and covering the entire road network. This technology can be widely applied to winter snow removal operations on urban roads, highways, and airport aprons.

[0087] The above description is merely a preferred embodiment of the present invention and is used only to illustrate the technical solution of the present invention, and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention are included within the scope of protection of the present invention.

Claims

1. A snow removal and de-icing path control method based on dynamic feedback of vehicle trajectory, characterized in that, include: By processing real-time GPS data from snowplows, onboard sensor data, and road network information, a multi-dimensional dynamic road perception network is formed. Based on reinforcement learning algorithms, a composite reward function is designed that includes traffic flow weights and vehicle energy consumption. This allows the agent to learn the optimal decision-making strategy autonomously during simulation training. Through Markov decision processes, dynamic path replanning is performed over a set period to ensure the adaptability of path optimization in the event of sudden snowfall or changes in traffic conditions.

2. The method according to claim 1, characterized in that, Specifically, it includes: Real-time vehicle trajectory rendering: Based on the Kafka distributed message queue, the system efficiently consumes the GPS data stream of the snowplow and performs persistent storage; at the same time, it uses the WebSocket protocol to build a two-way communication channel to push the real-time vehicle location information to the cockpit system in seconds; the cockpit side dynamically renders the real-time vehicle location based on the geographic information system and supports the historical trajectory backtracking function driven by the time axis. Q-Learning Model Construction: Based on the task objective and dynamic changes in the environment, the state space and action space are defined, and the reward function is designed by combining the parameters of snow thickness, vehicle energy consumption and road priority with the task constraints. Dynamic feedback mechanism: Design dynamic parameters such as discount factors, based on the Q-value update formula, and through continuous iteration to update the Q-value, the agent can eventually learn the optimal strategy to take different actions in different states, generate work instructions and push them synchronously to the cockpit system and mobile terminal.

3. The method according to claim 2, characterized in that, Kafka data access and storage, along with a WebSocket bidirectional communication mechanism, enable real-time rendering of vehicle location data and driving trajectories.

4. The method according to claim 3, characterized in that, The location data of the vehicle monitored in the cockpit is loaded and displayed on the GIS map, and the location changes dynamically in real time; based on the full GPS data of the vehicle on the same day, the trajectory is rendered and the driving trajectory of the day is played according to the data access time.

5. The method according to claim 2, characterized in that, Define the state space and action space based on the actual environment; plan reward and punishment conditions and task constraints, and design the reward function.

6. The method according to claim 2, characterized in that, Based on the real-time status of the vehicle, the Q value is periodically and continuously updated to generate the optimal path strategy, which is then persistently stored and pushed to the cockpit system or mobile terminal.

7. The method according to claim 6, characterized in that, Dynamic feedback mechanism: Real-time recording of vehicle position and operation parameters, updating the status-action value function (Q-value function) every five minutes, generating path correction instructions and simultaneously pushing them to the cockpit system and mobile terminal, supporting manual confirmation or adjustment by the driver.

8. The method according to claim 7, characterized in that, The system periodically receives the current status of the snowplow. When a command conflict is detected, it immediately triggers the Q-table to be recalculated and updated, and a correction command is issued. When communication is interrupted, switch to the locally stored Q table or predefined rules to perform the job.

Citation Information

Patent Citations

  • Sanitation vehicle real-time route planning method based on deep Q learning

    CN113420942A

  • Sanitation vehicle intelligent path planning method and system, scheduling method and storage medium

    CN119245682A

  • Deicing method of airport deicing vehicle

    CN119861745A

  • Snow sweeping robot path planning system and method based on optimized Astar algorithm

    CN119882734A

  • Unmanned sanitation vehicle path planning method based on artificial intelligence

    CN119901306A