Taxi dispatching method based on predictive reinforcement quantum reinforcement learning and application thereof

CN122529313APending Publication Date: 2026-08-07HEFEI UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HEFEI UNIV OF TECH
Filing Date
2026-05-14
Publication Date
2026-08-07

AI Technical Summary

Technical Problem

[0006]为了解决现有技术在无人网约车调度中安全约束适配不足、车路协同响应滞后、预测与决策脱节的问题,本发明提供一种基于预测增强型量子强化学习的网约车调度方法及其对应的存储介质、计算机程序产品、网约车调度设备

Benefits of technology

本发明针对无人网约车调度中安全约束适配不足、车路协同响应滞后等核心问题,通过构建三维量子态空间实现多因素动态耦合表征,结合预测结果进行量子相位补偿预判风险,依托量子强化学习算法并行求解最优调度方案并平衡多目标需求,最终通过滚动时域更新与安全触发机制构建决策-执行闭环。本发明的多环节协同创新可高效应对动态场景与安全风险,显著提升调度的安全性、实时性与运营效率,为无人网约车规模化商业化运营提供技术支撑,助力自动驾驶与智能交通协同融合。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122529313A_ABST
    Figure CN122529313A_ABST
Patent Text Reader

Abstract

The present application belongs to the field of intelligent transportation, and particularly relates to a ride-hailing scheduling method based on a prediction-enhanced quantum reinforcement learning and an application thereof. The scheme collects data in real time to construct a passenger demand parameter set, a vehicle state parameter set and a road network environment parameter set, maps them to a quantum superposition state and introduces a safety weight to serve as an initial quantum state; predicts a supply-demand matrix, a risk matrix and a cruising vector representing user travel demand, road network risk and vehicle endurance capacity at a future time, and fuses them into a prediction matrix; acquires vehicle real-time perception data and generates a state matrix, calculates the deviation between the state matrix and the prediction matrix and generates a phase compensation operator, and then obtains a corrected quantum state; presets a scheduling target and a safety constraint to construct a scheduling model based on QDQN; and generates a candidate scheduling scheme according to the corrected quantum state through the scheduling model. The present application solves the problems of insufficient safety constraint adaptation, response lag and disconnection between prediction and decision-making in existing unmanned ride-hailing scheduling.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of intelligent transportation, specifically relating to a ride-hailing scheduling method based on predictive reinforcement quantum reinforcement learning, and its corresponding storage medium, computer program product, and ride-hailing scheduling equipment. Background Technology

[0002] In the field of intelligent transportation and autonomous driving, the real-time dispatch efficiency and safety reliability of driverless ride-hailing vehicles are the core factors determining their large-scale deployment. With the upgrading of urban travel demands and the maturity of autonomous driving technology, driverless ride-hailing vehicles are gradually moving from pilot projects to commercial operation. An efficient and safe dispatch system is directly related to user experience, operating costs, and public transportation safety.

[0003] In traditional ride-hailing dispatch research, most methods revolve around a two-dimensional framework of supply-demand matching and route planning, focusing on the static matching of passenger orders and vehicles, or relying on traditional reinforcement learning for local optimization. However, they neglect the unique characteristics of autonomous ride-hailing vehicles, such as stringent safety constraints, strong real-time vehicle-to-infrastructure (V2X) capabilities, and complex supply-demand-environment coupling. When focusing on unmanned scenarios, the limitations of the aforementioned traditional dispatch methods become increasingly apparent: autonomous ride-hailing vehicles need to simultaneously meet the requirements of autonomous driving safety redundancy (such as sensor health and obstacle avoidance margin), millisecond-level response of V2X data, and dynamic adaptation under complex environments (congestion, weather, temporary traffic control). However, existing methods have almost never incorporated the multi-dimensional coupling of safety, vehicle-infrastructure, and supply-demand into the dispatch decision-making system.

[0004] Meanwhile, in the actual operation of driverless ride-hailing vehicles, the disconnect between prediction, decision-making, and execution of scheduling strategies is becoming increasingly prominent: traditional scheduling prediction modules are mostly based on a single dimension (such as supply and demand), lacking multi-dimensional fusion prediction of road network risks and vehicle range; decision-making modules rely on serial computation, making it difficult to cope with NP-hard problems in high-dimensional state spaces, resulting in insufficient real-time performance and optimality of scheduling schemes; the execution stage has not established a closed-loop feedback with decision-making, and when the environment changes suddenly (such as sudden congestion or sensor failure), the scheduling strategy cannot adapt quickly, greatly increasing the operational risks of driverless vehicles.

[0005] For the reasons mentioned above, it is difficult for operators of driverless ride-hailing services to balance the multiple objectives of passenger experience, vehicle efficiency, and safety redundancy in actual dispatching; autonomous driving technology suppliers lack dispatching algorithms adapted to driverless scenarios, making it difficult to implement related solutions in practice. Summary of the Invention

[0006] To address the problems of insufficient safety constraint adaptation, lagging vehicle-road cooperative response, and disconnect between prediction and decision-making in existing technologies for unmanned ride-hailing dispatch, this invention provides a ride-hailing dispatch method based on prediction-enhanced quantum reinforcement learning, along with its corresponding storage medium, computer program product, and ride-hailing dispatch equipment.

[0007] This invention is achieved using the following technical solution: A ride-hailing scheduling method based on prediction-enhanced quantum reinforcement learning, comprising: Collect data to construct a passenger demand parameter set D, a vehicle state parameter set V, and a road network environment parameter set R, and map them as quantum superposition states. And after introducing security weights, it is used as the initial quantum state. .

[0008] A demand forecasting model pre-trained using a spatiotemporal convolutional network is employed. Based on recent regional order data and population flow data, it predicts the future supply and demand intensity and passenger travel probability of the target area, outputting a supply and demand matrix. A risk forecasting model pre-trained using a long short-term memory network is used. Based on real-time data collected by roadside units, it predicts the future risk value of the road network within the target area and outputs a risk matrix. A battery degradation model is used to predict the remaining range and charging demand priority of each vehicle at future times, outputting a range vector. The supply and demand matrix, risk matrix, and range vector are then fused into a prediction matrix.

[0009] Real-time vehicle perception data is acquired and a state matrix is ​​generated. The deviation between the state matrix and the prediction matrix is ​​calculated. A phase compensation operator is generated based on the deviation and applied to the initial quantum state to obtain the corrected quantum state. .

[0010] A scheduling model based on a quantum deep Q-network is constructed with preset scheduling objectives and security constraints. Candidate scheduling schemes are generated by the scheduling model based on the input modified quantum state. When a security threshold is triggered, the quantum state is activated to rapidly collapse, and an emergency scheduling scheme is output.

[0011] As a further improvement of the present invention, the passenger demand parameter set D includes: passenger real-time location. d loc Expected travel time d time Destination location d dest Carpooling willingness d share Model preference d prefer .

[0012] As a further improvement of the present invention, the vehicle state parameter set V includes: the real-time position of the vehicle. vloc Remaining battery power v batt Sensor health v sensor Autonomous driving level v adlevel Recent 3-minute obstacle avoidance record v obstacle The success rate of dispatch execution in the past hour v history .

[0013] As a further improvement of the present invention, the road network environment parameter set R includes: road segment congestion index. r cong Real-time phase of traffic lights r signal Temporary traffic control area r control Real-time weather r weather Emergencies in the region r event .

[0014] As a further improvement of the present invention, the initial quantum state The expression is as follows: ; In the above formula, Represents a superposition quantum state of three parameter sets; , and These are the quantum states corresponding to parameter sets D, V, and R, respectively; , and To meet the normalization conditions , and quantum state amplitude, ; Represents the fundamental quantum state; Indicates safety weight; Indicates the safe distance margin between the vehicle and the obstacle. This refers to the emergency braking response time. To assess the reliability of sensor data.

[0015] As a further improvement to this invention, the demand forecasting model divides the target area into several 500m×500m grids and outputs the supply and demand intensity and passenger travel probability for each grid at various times within a specified future period. The supply and demand intensity is divided into 1-10 levels; the passenger probability is a probability value between 0 and 1. This results in a supply and demand matrix composed of supply and demand intensity and passenger travel probability under different time and spatial dimensions.

[0016] As a further improvement to this invention, the risk prediction model divides the target area into several 500m×500m grids and outputs the dynamic road network risk of each grid at various times within a specified future period. The types of dynamic road network risks include risk values ​​corresponding to sudden congestion sections, construction areas, and high-frequency areas of illegal lane changes; thus, a risk matrix is ​​obtained composed of various risk values ​​under different time and spatial dimensions.

[0017] As a further improvement of the present invention, the battery degradation model first predicts the remaining power at each time within a specified future period, and then generates charging demand priority based on the remaining power according to a preset mapping relationship; thereby obtaining a range vector composed of the remaining power and charging demand priority under different time dimensions.

[0018] The supply and demand matrix, risk matrix, and endurance vector are concatenated along the time dimension to obtain the required prediction matrix.

[0019] As a further improvement to this invention, the expression for the battery degradation model is as follows: ; In the above formula, v batt0 Indicates the initial battery level; t Indicates time, T Indicates ambient temperature; v Indicates vehicle speed; k , a and b These represent the influence coefficients of time, ambient temperature, and vehicle speed, respectively.

[0020] As a further improvement to the present invention, the modified quantum state The calculation formula is: ; In the above formula, Represents the quantum phase compensation operator; The standard 2×2 Hermitian unitary operator representing the spin z component of the Pauli z operator has eigenvalues ​​of +1 and -1, corresponding to the two orthogonal computational ground states of the qubit. X percept Represents the state matrix; X pred Represents the prediction matrix; This represents the deviation between the state matrix and the prediction matrix; i Represents the imaginary unit; Indicates the phase compensation amount; These are the proportional coefficient, derivative coefficient, and integral coefficient of PID control, respectively. This represents the time integral dummy variable (instantaneous time variable) of the PID integral element.

[0021] As a further improvement of the present invention, the scheduling model includes an input layer, a hidden layer, and an output layer; the input layer is used to convert the modified quantum state into a classical feature vector through quantum measurement and then perform quantum activation; the hidden layer realizes neuronal entanglement through CNOT quantum gates and introduces a quantum dropout layer to prevent overfitting; the output layer is used to output the Q value of each group of scheduling schemes and recommend candidate scheduling schemes.

[0022] The scheduling model minimizes the loss function using the quantum gradient descent algorithm. L To achieve network parameter updates; loss function L The expression is: ; In the above formula, Indicates the target Q value; This represents the actual Q value of the current prediction result; Representing quantum state The corresponding Hermitian conjugate left-hand vectors are used together to implement quantum inner product operations.

[0023] As a further improvement to this invention, the learning rate of the quantum gradient descent algorithm... Adaptive quantum learning rate strategy: Set initial learning rate When the scheduling success rate is above 95% for five consecutive rolling windows, the learning rate is reduced to [a lower threshold]. When the scheduling success rate falls below 80% or a safety threshold is triggered, the learning rate will be increased to [a higher percentage]. .

[0024] The present invention also includes a storage medium storing a computer program, which, when executed by a processor, implements the aforementioned ride-hailing scheduling method based on prediction-enhanced quantum reinforcement learning, and then generates candidate scheduling schemes for target vehicles based on dynamically updated data.

[0025] The present invention also includes a computer program product comprising a computer program that, when executed by a processor, implements the aforementioned ride-hailing scheduling method based on prediction-enhanced quantum reinforcement learning, and then generates candidate scheduling schemes for target vehicles based on dynamically updated data.

[0026] The present invention also includes a ride-hailing dispatching device, which includes a memory, a processor, and a computer program stored in the memory and running on the processor. When the processor executes the computer program, it implements the ride-hailing dispatching method based on prediction-enhanced quantum reinforcement learning as described above, and then generates candidate dispatching schemes for target vehicles based on dynamically updated data.

[0027] The technical solution provided by this invention has the following beneficial effects: This invention addresses core issues in autonomous ride-hailing dispatching, such as insufficient safety constraint adaptation and lagging vehicle-road cooperative response. It constructs a three-dimensional quantum state space to achieve dynamic coupling representation of multiple factors, combines prediction results with quantum phase compensation for risk assessment, and relies on quantum reinforcement learning algorithms to solve for the optimal dispatching scheme in parallel while balancing multiple objective requirements. Finally, it constructs a decision-making-execution closed loop through rolling time-domain updates and a safety triggering mechanism. This multi-stage collaborative innovation efficiently addresses dynamic scenarios and safety risks, significantly improving dispatching safety, real-time performance, and operational efficiency. It provides technical support for the large-scale commercial operation of autonomous ride-hailing vehicles and facilitates the synergistic integration of autonomous driving and intelligent transportation.

[0028] This invention can automatically output the optimal scheduling strategy in complex and dynamic scenarios, efficiently cope with sudden environmental changes and safety risks, significantly improve the safety, real-time performance and overall operational efficiency of unmanned ride-hailing scheduling, provide innovative technical solutions for the large-scale commercial operation of unmanned ride-hailing vehicles, and provide algorithmic support for the synergistic integration of autonomous driving technology and intelligent transportation systems. Attached Figure Description

[0029] Figure 1 This is a flowchart of the ride-hailing scheduling method based on prediction-enhanced quantum reinforcement learning provided in Embodiment 1 of the present invention.

[0030] Figure 2 This is a network architecture diagram of a demand prediction model based on a spatiotemporal convolutional network.

[0031] Figure 3 This is a schematic diagram of a scheduling model based on quantum deep Q-networks.

[0032] Figure 4 This is a module architecture diagram of the dispatching system used in the ride-hailing dispatching device in Embodiment 2 of the present invention.

[0033] Figure 5 The figure shows the reward convergence curve of the DQN network of the present invention during 2000 iterations in the simulation test.

[0034] Figure 6 This is a comparison chart of the supply and demand heat distribution in the early morning peak area output by the scheme of this invention during simulation testing.

[0035] Figure 7 This is a performance comparison chart of the present invention and the control group scheme in various indicators during simulation testing. Detailed Implementation

[0036] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.

[0037] Example 1 Traditional ride-hailing dispatching schemes typically focus on optimizing solutions from two dimensions: matching user travel demand and vehicle route planning. They lack safety constraints for autonomous ride-hailing vehicles and adaptability to real-time vehicle-road cooperative scenarios, thus easily leading to operational risks. To address this issue, this embodiment constructs a three-dimensional quantum state space of "supply and demand-vehicle-road-safety," mapping multi-dimensional factors to quantum superposition states and incorporating safety weight factors to achieve precise adaptation to autonomous scenarios, simultaneously considering safety redundancy and dispatching efficiency. For the collected multi-source real-time data, this embodiment generates a three-dimensional matrix of "supply and demand-risk-endurance" through three-stage prediction, combined with quantum phase compensation to correct errors, reducing prediction bias and anticipating risks in advance, thus gaining response time for strategy adjustments. At the dispatching decision-making level, this embodiment adopts a quantum reinforcement learning framework, overcoming the inefficiency of traditional serial computation algorithms. It achieves parallel policy exploration through quantum deep Q-networks, combining multi-objective optimization to generate dispatching schemes, thereby improving the optimization efficiency of dispatching strategies.

[0038] Specifically, such as Figure 1 As shown, the ride-hailing scheduling method based on prediction-enhanced quantum reinforcement learning provided in this embodiment includes the following process: I. Collect data to construct a passenger demand parameter set D, a vehicle state parameter set V, and a road network environment parameter set R, and map them as quantum superposition states. And after introducing security weights, it is used as the initial quantum state. .

[0039] In this embodiment, the passenger demand parameter set is defined as follows: ;in, d loc Indicates the passenger's real-time location, d time Indicates expected travel time, d dest Indicates the location of the destination. d share Indicates willingness to carpool (can be quantified as a value between 0 and 1). d prefer This indicates vehicle type preference (such as comfort, economy, etc.). In practical applications, the relevant location information can be represented by latitude and longitude, with an accuracy of up to 0.1m.

[0040] Define the autonomous vehicle state parameter set as ;in, v loc Indicates the real-time location of the vehicle. v batt This indicates the remaining battery level (accuracy 1%). v sensorIndicates sensor health (e.g., characterized by the probability of failure of devices such as LiDAR or cameras). v adlevel Indicates the level of autonomous driving (L2 to L4). v obstacle This indicates the obstacle avoidance record for the most recent 3 minutes. v history This indicates the success rate of scheduling execution over the past hour.

[0041] Define the road network environment parameter set as ;in, r cong This indicates the road congestion index (which can be divided into 1-10 levels). r signal Indicates the real-time phase of traffic lights, r control Indicates temporary traffic control area, r weather It indicates real-time weather (including relevant indicators such as temperature, precipitation, and visibility). r event It indicates an emergency in the area (such as a traffic accident, a short-term crowding caused by a concert, event or other activity).

[0042] Based on the data of the three types of parameter sets collected at each time point, this embodiment first maps them to a quantum superposition state. Then, a safety weighting factor is introduced. ,Will As weights, they are incorporated into the quantum state to obtain the initial quantum state. .

[0043] In summary, the initial quantum state The expression is as follows: ; In the above formula, Represents a superposition quantum state of three parameter sets; , and These are the quantum states corresponding to parameter sets D, V, and R, respectively; , and To meet the normalization conditions , and quantum state amplitude, ; Represents the fundamental quantum state; Indicates safety weight; Indicates the safe distance margin between the vehicle and the obstacle. This refers to the emergency braking response time. To assess the reliability of sensor data.

[0044] Second, the supply and demand matrix, risk matrix, and range vector, which represent user travel demand, road network risks, and vehicle range capabilities at future moments, are used to predict future moments. The supply and demand matrix, risk matrix, and range vector are then integrated into a prediction matrix.

[0045] In this embodiment, a demand forecasting model based on a pre-trained spatiotemporal convolutional network is employed. Based on recent regional order data and population flow data, it predicts the future supply and demand intensity and passenger travel probability of the target region, and outputs a supply and demand matrix. The architecture of the demand forecasting model is as follows: Figure 2 As shown, it is used to integrate order data. X order POI data X POI Historical data for the same period X hist and public transportation X trans Multi-source feature sets of data X demand For input, The model extracts spatiotemporal features through two SCTN-Conv Blocks: within each block, temporal-gated convolutions capture temporal dependencies, followed by spatial graph convolutions to mine region associations. A newly added spatial attention block (containing multi-head attention and layer normalization) strengthens the feature weights of high-demand regions. The model employs an adaptive spatial attention weight vector. and time attention weight vector The calculation formula is as follows: ; In the above formula, W s and W t These are the spatial feature weight matrix and the temporal feature weight matrix, respectively. F spatial and F temporal These are spatial feature maps and temporal feature maps, respectively. d k and b t These represent the spatial feature dimension and the bias term, respectively; softmax(·) and sigmoid(·) represent the softmax function and the sigmoid function, respectively.

[0046] The right-hand framework further breaks down the details of temporally gated convolution (1-D convolution + temporal attention + GLU gating), ultimately obtaining the supply and demand matrix through the output layer. This architecture achieves spatiotemporal fusion and accurate prediction of multi-dimensional data, providing crucial supply and demand inputs for subsequent quantum scheduling decisions.

[0047] Specifically, the demand forecasting model can divide the target area into several 500m×500m grids and output the supply and demand intensity and passenger travel probability for each grid at various times within a specified future period. The supply and demand intensity is divided into 1-10 levels; the passenger probability is a probability value between 0 and 1. This results in a supply and demand matrix composed of supply and demand intensity and passenger travel probability under different time and spatial dimensions.

[0048] Accordingly, this embodiment employs a risk prediction model pre-trained based on a Long Short-Term Memory (LSTM) network. Based on real-time data collected by roadside units (RSUs), it predicts the risk values ​​of the road network within the target area at future times and outputs a risk matrix. Specifically, the risk prediction model divides the target area into several 500m × 500m grids and outputs the dynamic road network risk for each grid at various times within a specified future time period. The types of dynamic road network risks include risk values ​​corresponding to sudden congestion sections, construction areas, and high-frequency areas of illegal lane changes; thus, a risk matrix is ​​obtained composed of various risk values ​​under different time and spatial dimensions.

[0049] In the risk matrix, the risk value of any i-th region at time j is... The calculation formula is: ; In the above formula, The value range is [0, 10], and the larger the value, the higher the risk. Let be the predicted average vehicle speed of region i at time j; be the design speed of the road segment. The severity of the traffic event in region i at time j is represented by the value [0, 5], where a value of 0 indicates no event, a value of 1 indicates minor congestion, and a value of 5 indicates a major accident. Let be the percentage of traffic light waiting time in region i at time j, with a value in the range [0,1]. A value of 0 indicates no waiting, and a value of 1 indicates waiting for the entire time.

[0050] Next, this embodiment predicts the remaining range and charging demand priority of each vehicle at future times based on the battery degradation model and outputs a range vector. The expression for the battery degradation model is: ; In the above formula, v batt0 Indicates the initial battery level; tIndicates time, T Indicates ambient temperature; v Indicates vehicle speed; k , a and b These represent the influence coefficients of time, ambient temperature, and vehicle speed, respectively.

[0051] The battery degradation model first predicts the remaining power at each time point within a specified future period, and then generates charging demand priorities based on the remaining power according to a preset mapping relationship; thus, it obtains a range vector composed of the remaining power and charging demand priorities under different time dimensions.

[0052] Finally, in this embodiment, the supply and demand matrix, risk matrix, and endurance vector are concatenated according to the time dimension to obtain the required prediction matrix.

[0053] Third, acquire real-time vehicle perception data and generate a state matrix, then calculate the deviation between the state matrix and the prediction matrix. Based on the deviation, generate a phase compensation operator and apply it to the initial quantum state to obtain the corrected quantum state. .

[0054] In this embodiment, the perception data used to generate the state matrix mainly includes data collected in real time by the autonomous vehicle and roadside units. The collected raw data includes LiDAR point clouds, camera image recognition results, etc. Based on this, deviations... The calculation formula is as follows: .

[0055] In the above formula, X percept Represents the state matrix; X pred This represents the prediction matrix.

[0056] Quantum phase compensation operator calculated based on the deviation between the state matrix and the prediction matrix. as follows: ; in, The standard 2×2 Hermitian unitary operator representing the spin z component of the Pauli z operator has eigenvalues ​​of +1 and -1, corresponding to the two orthogonal computational ground states of the qubit. i Represents the imaginary unit; The phase compensation amount is expressed by the following formula: ; In the above formula, These are the proportional coefficient, derivative coefficient, and integral coefficient of PID control, respectively. This represents the time integral dummy variable (instantaneous time variable) of the PID integral element.

[0057] Finally, the quantum phase compensation operator Applied to the initial quantum state Above, the corrected quantum state is obtained. as follows: .

[0058] IV. Preset scheduling objectives and security constraints, and construct a scheduling model based on quantum deep Q network (QDQN); generate candidate scheduling schemes based on the input modified quantum state through the scheduling model.

[0059] In this embodiment, a quantum deep Q-network is used as the scheduling model for scheduling strategy optimization, and its architecture is as follows: Figure 3 As shown, the scheduling model includes an input layer, a hidden layer, and an output layer. The input layer converts the modified quantum state into a classical feature vector via quantum measurement before quantum activation. The hidden layer implements neuron entanglement through CNOT quantum gates and introduces a quantum dropout layer to prevent overfitting. The output layer outputs the Q-values ​​of each scheduling scheme and recommends candidate scheduling schemes. The corresponding expressions are as follows: ; in, The measured classical feature vectors, with dimensions consistent with the quantum state dimensions, are used as input to the quantum deep Q-network; For the corrected quantum state, Representing quantum state The corresponding Hermitian conjugate left-hand vectors are used together to implement quantum inner product operations. The quantum measurement observable operator (quantum projective measurement operator) is the Hermitian operator corresponding to the observable physical quantity in quantum mechanics; x This represents the input to the quantum activation function; The output value of the quantum activation function. F l and F l-1 The first l Hidden layers and the first l -1 input feature map of hidden layer For the matrix representation of a controlled NOT gate, W l For the first l The weight matrix of the layer, b l For the first l The layer's bias term; ReLU(·) represents the ReLU activation function.

[0060] In practical applications, the scheduling model can avoid local optima through quantum entanglement exploration—utilizing a balancing mechanism combined with quantum annealing. Simultaneously, it relies on a quantum experience replay buffer to store interactive data, copying the current network parameters to the target quantum value network every N steps. Finally, it minimizes the QDQN loss function through quantum gradient descent, achieving a driverless ride-hailing scheduling decision that balances safety and efficiency. In practical applications, the innovative architecture of this embodiment solves the problem of inefficiency in traditional serial scheduling computation.

[0061] Specifically, this embodiment introduces quantum entanglement exploration into the scheduling model—utilizing a balance mechanism, at least five scheduling schemes are explored in parallel through quantum superposition states, and the Q-value of each scheduling scheme is output. The introduced quantum annealing strategy is used to optimize the scheduling strategy; during the solution process, an initial temperature is set. Annealing cycle T =100 iteration steps, using the formula Lower the temperature, with probability Accepting poorer solutions helps avoid getting trapped in local optima.

[0062] In the scheduling strategy optimization phase of the scheduling model, a scheduling target set is defined. O for: , in, T wait For average passenger waiting time, U vehicle For vehicle hourly utilization rate, S redundancy For safety redundancy, C energy Energy cost per order, T arrive This represents the average arrival time for passengers.

[0063] Based on the above five optimization objectives T wait , U vehicle , S redundancy , C energy and T arrive This embodiment further defines the quantum state weighting factors for each target quantization, and the corresponding weight vector is as follows: : ,satisfy .in, respectively T wait , U vehicle , S redundancy, C energy and T arrive The quantum state weighting factor.

[0064] Next, through quantum measurement operations ( , Calculate the multi-objective comprehensive benefit for each objective (using the observation operator corresponding to each objective) and output the optimal scheduling strategy. Among these, when the safety redundancy... ( S thresh When the safety threshold is reached, it will automatically be... Increase it to above 0.5 to prioritize the safe operation of unmanned vehicles.

[0065] In the application phase, design a quantum weighted reward function. r t : ; in, r t Let be the instantaneous reward value at time t. These represent passenger waiting time rewards, vehicle utilization rate rewards, safety redundancy rewards, energy consumption cost rewards, and passenger arrival time rewards, respectively. They represent The weights are determined by various factors. These weights can be dynamically adjusted based on the scenario; for example, in extreme weather scenarios, the weights will be adjusted accordingly. The initial value is set to 0.4. The initial value was reduced to 0.1 to adapt to the scheduling priority requirements of different scenarios.

[0066] The scheduling model minimizes the loss function using the quantum gradient descent algorithm. L To achieve network parameter updates; loss function L The expression is: ; In the above formula, Indicates the target Q value; This represents the actual Q value of the current prediction result; Representing quantum state The corresponding Hermitian conjugate left-hand vectors are used together to implement quantum inner product operations.

[0067] In this embodiment, the learning rate of the quantum gradient descent algorithm Adaptive quantum learning rate strategy: Set initial learning rate When the scheduling success rate is above 95% for five consecutive rolling windows, the learning rate is reduced to [a lower threshold]. When the scheduling success rate falls below 80% or a safety threshold is triggered, the learning rate will be increased to [a higher percentage]. .

[0068] In practical applications, this embodiment divides the time domain into 5-minute rolling time windows, collecting real-time execution feedback data of the autonomous vehicle within each window (including whether the dispatch was successfully completed, the deviation between the actual driving route and the planned route, and the number of obstacle avoidance operations). Then, the quantum state update quantity is calculated. ,in, To execute the quantum state corresponding to the feedback, (The quantum state in the previous time domain). Through quantum gate operations. ( Generate and update quantum states for the unit operator It then outputs new scheduling instructions.

[0069] Furthermore, this embodiment also includes a security failure threshold triggering mechanism. Specifically, a security constraint threshold set is set. T safe : ,in, T sensor The threshold for the probability of sensor failure (e.g., ≥0.3). T risk The threshold for road network risk level (e.g., ≥8 levels). T batt The remaining charge threshold (e.g., ≤10%) is used. When any safety constraint threshold is triggered, the rapid collapse mechanism of the quantum state is activated, and then measured via projection. The quantum state is collapsed into the emergency dispatch subspace, and emergency plans are output first, such as making safe stops nearby, changing to alternative routes, and dispatching backup vehicles for support.

[0070] Example 2 The ride-hailing scheduling method based on prediction-enhanced quantum reinforcement learning provided in Example 1 is essentially a data processing method. In order to better apply this scheme, this example further provides corresponding computer program products, storage media and ride-hailing scheduling equipment.

[0071] The storage medium provided in this embodiment stores a computer program. When the computer program is executed by the processor, it implements the ride-hailing scheduling method based on prediction-enhanced quantum reinforcement learning as described above, and then generates candidate scheduling schemes for target vehicles based on dynamically updated data.

[0072] The computer program product provided in this embodiment includes a computer program. When the computer program is executed by a processor, it implements the ride-hailing scheduling method based on prediction-enhanced quantum reinforcement learning as described above, and then generates candidate scheduling schemes for target vehicles based on dynamically updated data.

[0073] The ride-hailing dispatching device provided in this embodiment includes a memory, a processor, and a computer program stored in the memory and running on the processor. When the processor executes the computer program, it implements the aforementioned ride-hailing dispatching method based on predictive reinforcement learning, and then generates candidate dispatching schemes for target vehicles based on dynamically updated data. In practical applications, the ride-hailing dispatching device can adopt methods such as... Figure 4 The scheduling system shown implements the corresponding data processing logic. Specifically, the scheduling system includes a quantum state construction module, a prediction module, a quantum phase compensation module, a quantum deep Q-network, a rolling time-domain update module, and a secure triggering module.

[0074] The quantum state construction module constructs an initial quantum state based on the massive amount of collected data. The prediction module generates a prediction matrix using the prediction results from demand prediction, risk prediction, and battery degradation models. The quantum phase compensation module calculates the deviation between the prediction matrix and the state matrix and performs phase compensation on the initial quantum state to obtain a corrected quantum state. The quantum deep Q-network generates multiple candidate scheduling schemes based on the input corrected quantum state. The rolling time-domain update module continuously optimizes the scheduling strategy based on real-time updated collected data; the safety trigger module activates rapid quantum state collapse and outputs an emergency scheduling scheme when a safety threshold is triggered. Finally, corresponding scheduling instructions are generated and executed based on the scheduling scheme.

[0075] The ride-hailing dispatching device provided in this embodiment is essentially a computer device. In practical applications, this computer device can be an embedded device, integrated into autonomous vehicles or other higher-level management equipment. Alternatively, it can be a standalone computer device, such as a laptop, tablet, desktop computer, or a large-scale computer capable of executing computer programs, such as a rack server, blade server, tower server, or cabinet server (including standalone servers or server clusters composed of multiple servers), to perform backend dispatching decision processing on the multi-source real-time data collected from the front end.

[0076] The computer device in this embodiment includes, but is not limited to, a memory and a processor that can be interconnected via a system bus. In this embodiment, the memory (i.e., the readable storage medium) includes flash memory, hard disk, multimedia card, card-type memory (e.g., SD or DX memory), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disk, etc. In some embodiments, the memory can be an internal storage unit of the computer device, such as the hard disk or RAM of the computer device. In other embodiments, the memory can also be an external storage device of the computer device, such as a plug-in hard disk, smart media card (SMC), secure digital card (SD), flash card, etc. Of course, the memory can also include both internal storage units and external storage devices of the computer device. In this embodiment, the memory is typically used to store the operating system and various application software installed on the computer device. Furthermore, the memory can also be used to temporarily store various types of data that have been output or will be output.

[0077] In some embodiments, the processor may be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chip. The processor is typically used to control the overall operation of a computer device.

[0078] Simulation test To verify the effectiveness of the ride-hailing scheduling method based on prediction-enhanced quantum reinforcement learning provided by this invention, technicians conducted simulation and real-vehicle pilot tests on the relevant scheme. Multiple sets of comparative experiments were designed (the control schemes were the traditional deep Q-network scheduling method and the classical supply-demand matching scheduling method). An unmanned ride-hailing operation model was constructed through an urban traffic simulation platform (recreating actual scenarios such as morning and evening rush hours, extreme weather, and sudden traffic control). The core indicators of prediction accuracy, scheduling real-time performance, safety redundancy, operational efficiency, and scheduling cost of the scheme of this invention were quantitatively verified.

[0079] I. Experiment Content 1.1 Simulation Parameter Settings Prediction module parameters: The input to the spatiotemporal convolutional network is a multi-source spatiotemporal sequence of length M=10 (one time step every 5 minutes, covering nearly 50 minutes of data). Feature dimensions include 8 core features such as order density, capacity gap, POI popularity, and road condition level. The network contains two cascaded SCTN-Conv Blocks. Each Block has a 3×3 kernel size for the temporally gated convolutions and 64 output channels. The adjacency matrix of the spatial graph convolutions is dynamically updated based on road network connectivity. The spatiotemporal attention module employs an 8-head attention mechanism, and layer normalization parameters... The training batch size is 32, the learning rate is 0.001, and the number of iterations is 500.

[0080] Quantum reinforcement learning parameters: The quantum state dimension of the quantum DQN network is 64-dimensional (number of matching sub-regions), and the initial quantum state encoding adopts an amplitude-phase hybrid encoding scheme. The exploration rate linearly decays from 0.9 to 0.1 (decay period of 2000 iterations). The quantum annealing strategy has an initial temperature of 100, a decay period of 100, a discount factor of 0.9, a quantum experience replay buffer capacity of 10000 entries, a parameter synchronization interval of 50 time steps, and a learning rate of 0.0005 for quantum gradient descent.

[0081] Evaluation Index System: A multi-dimensional evaluation index system is constructed, covering 6 core indicators and 3 auxiliary indicators. Efficiency indicators include average passenger waiting time (minutes), average vehicle empty-run rate (%), order response rate (%), and vehicle utilization rate (%). Safety indicators include safety redundancy compliance rate (%, vehicle spacing ≥ 5 meters, risk section avoidance rate 100%), and vehicle-road cooperative response latency (milliseconds). Cost indicators include unit order energy consumption cost (yuan) and charging scheduling cost (yuan / day). Auxiliary indicators include prediction bias rate (%), reward function convergence speed (iterations), and extreme scenario adaptability rate (%).

[0082] 1.2 Control Group Setup Control group 1 (traditional DQN scheduling scheme): adopts the traditional deep Q network architecture, without quantum state encoding module, the input is classical spatiotemporal feature vector, without introducing spatiotemporal attention mechanism and quantum annealing strategy, and the reward function is the single order response rate weight.

[0083] Control Group 2 (Classic Supply and Demand Matching Scheduling Scheme): A greedy order dispatch strategy based on regional order density, which only considers real-time supply and demand relationships, without a prediction module or reinforcement learning optimization. The scheduling logic is to dispatch orders to the nearest location and plan the shortest path.

[0084] Experimental group (the scheme of this invention): adopts the patented core architecture, namely the improved ST-Conv Net multi-dimensional prediction module + quantum DQN decision module, which integrates key technologies such as supply and demand-vehicle-road-safety three-dimensional quantum state space, quantum parallel exploration, and quantum phase compensation correction.

[0085] II. Experimental Results and Analysis 2.1 Reward Convergence Analysis in Quantum Reinforcement Learning In the scheme of this invention, the reward convergence curve of the DQN network during 2000 iterations is as follows: Figure 5 As shown, Figure 5 The original reward curve (yellow) and the smoothed reward curve (orange) visually reflect the stability and convergence efficiency of the training process. Meanwhile, this embodiment further compares the reward convergence performance of the proposed solution and the control group 1 solution, as shown in Table 1: Table 1: Comparison of Reward Convergence between Reinforcement Learning Control Groups Analysis of the data in the table above shows that the traditional DQN scheme in control group 1 requires 1600 iterations to converge to a stable value (0.42±0.03), which is 37.5% slower than the experimental group, and the reward value after stabilization is 23.6% lower than that of the experimental group. This difference is due to the fact that the quantum gradient descent algorithm of this invention optimizes the parameter update efficiency, the quantum experience replay buffer reduces the training oscillations caused by data correlation, and the quantum annealing strategy effectively avoids the local optimum problem that traditional reinforcement learning is prone to get stuck in, which significantly improves the training stability and convergence speed.

[0086] 2.2 Multi-dimensional Prediction Accuracy Analysis This experiment further visualizes the predicted supply and demand dynamics and actual order dynamics in this invention, as shown below. Figure 6 The diagram shows a comparison of supply and demand distribution in the morning rush hour area. The prediction accuracy of the present invention and the control group was analyzed, and the experimental results are shown in Table 2. Table 2: Prediction Accuracy Analysis Analysis of the experimental data shows that the prediction results of this invention highly overlap with the spatial distribution of actual demand in high-demand areas (core business districts and areas around subway stations), with a calculated spatial matching degree of 92%. Specifically, the deviation between the predicted demand intensity and the actual demand intensity in the three core business districts is less than 8%, and the prediction deviation in the two transportation hubs is less than 10%. In contrast, the traditional DQN scheme in control group 1, lacking a spatiotemporal attention mechanism, exhibits a prediction deviation of 25%-30% in the core areas. This result verifies the effectiveness of the spatial graph convolution and spatial attention module in the improved ST-Conv Net—the spatial graph convolution captures the supply and demand relationship between adjacent areas based on road network connectivity, and the spatial attention module assigns weights to high-demand areas through a multi-head attention mechanism, enhancing feature extraction in key areas.

[0087] Based on time series data statistics, the experimental group's prediction deviation rate for supply and demand in the next 10 minutes was only 32%, a decrease of 44.8% compared to control group 1 (58%). Compared to control group 2 (without a prediction module), the experimental group's real-time response mode captured the migration trend of demand peaks earlier. This is attributed to the synergistic effect of temporally gated convolution and the temporal attention module: the temporally gated convolution uses the GLU gating mechanism to filter features of key time segments such as morning and evening peak hours and sudden orders; the temporal attention weight increases to 0.6-0.7 during peak periods, effectively strengthening the modeling of temporal dependencies; simultaneously, the quantum phase compensation operator dynamically adjusts the quantum state phase according to the prediction error, further correcting the prediction deviation caused by multi-source data noise.

[0088] In extreme scenarios, the prediction bias rate of the experimental group remained below 45%, while the bias rate of control group 1 soared to over 75%. For example, when large events ended, the experimental group identified areas of order surges in advance through long-term prediction (10-15 minutes), with a prediction bias rate of only 38%, reserving scheduling buffer time for the quantum DQN decision module. In contrast, control group 1, lacking accurate prediction, was unable to allocate transportation capacity in advance, resulting in a significant drop in order response rate.

[0089] 2.3 Comparative Analysis of Core Scheduling Indicators This experiment further analyzed and statistically analyzed the achievements of the present invention and the control group scheme in terms of core scheduling indicators. The results are shown in the table below: Table 3: Comparative Analysis of Core Scheduling Indicators Furthermore, based on the experimental data, plots were created as follows: Figure 7 The bar chart shows a comparison of different solutions across various indicators. The bar chart visually reflects the comprehensive advantages of the present invention in terms of efficiency, safety, cost, and other dimensions.

[0090] Analysis of the experimental data shows that, in terms of average passenger waiting time, the experimental group waited only 3.2 minutes, a 45% reduction compared to control group 1 (5.8 minutes) and a 57% reduction compared to control group 2 (7.5 minutes). The core reason is that the accurate prediction of the improved ST-Conv Net allows capacity to be deployed in advance to high-demand areas, while the quantum parallel exploration mechanism simultaneously optimizes order scheduling and route planning, reducing vehicle empty runs and detour time. In terms of vehicle utilization, the experimental group achieved 82.5%, a 26% increase compared to control group 1 (65.3%) and a 42% increase compared to control group 2 (58.1%). This is attributed to the global optimization of the three-dimensional quantum state space of supply and demand, vehicle-road, and safety, which avoids the drawbacks of traditional solutions that prioritize efficiency over safety or local factors over global ones, and achieves dynamic matching between transport capacity and demand.

[0091] In terms of order response rate, the experimental group reached 98.3%, which is 8.6% higher than control group 1 (89.7%) and 17.1% higher than control group 2 (81.2%), fully verifying the rapid response capability of the prediction-enhanced architecture to order demand.

[0092] Safety redundancy compliance rate: The experimental group reached 99.2%, significantly higher than control group 1 (92.7%) and control group 2 (85.4%). This is because a safety weight factor is incorporated into the quantum state space, and the objective function dynamically adjusts the weights of safety indicators such as vehicle spacing and avoidance of risky road sections, ensuring that the scheduling scheme meets safety constraints while optimizing efficiency.

[0093] In terms of vehicle-road cooperative response latency, the experimental group achieved only 12 milliseconds, a reduction of 65.7% compared to control group 1 (35 milliseconds) and a reduction of 79.3% compared to control group 2 (58 milliseconds). The characteristics of quantum parallel computing enabled the exploration and decision-making of the five scheduling schemes to be completed simultaneously, avoiding the latency accumulation of traditional serial computing and meeting the real-time requirements of unmanned ride-hailing vehicle-road cooperation.

[0094] In terms of unit order energy consumption cost, the experimental group cost only 2.1 yuan, which is 25% lower than control group 1 (2.8 yuan) and 40% lower than control group 2 (3.5 yuan). The quantum DQN network considers both energy consumption optimization and time optimization in path planning, reducing inefficient energy consumption scenarios such as rapid acceleration and traversing congested road sections.

[0095] In terms of charging scheduling costs, the daily average charging scheduling cost of the experimental group was 120 yuan, which was 33.3% lower than that of the control group 1 (180 yuan). This was because the multi-dimensional prediction module predicted the vehicle's range demand in advance, thus avoiding empty driving scheduling caused by temporary charging.

[0096] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A ride-hailing scheduling method based on predictive reinforcement quantum reinforcement learning, characterized in that, It includes: Collect data to construct a passenger demand parameter set D, a vehicle state parameter set V, and a road network environment parameter set R, and map them as quantum superposition states. And after introducing security weights, it is used as the initial quantum state. ; A demand prediction model based on spatiotemporal convolutional network pre-training is adopted to predict the future supply and demand heat and passenger travel probability of the target area based on recent regional order data and population flow data, and output a supply and demand matrix. A risk prediction model based on long short-term memory network pre-training is adopted to predict the risk value of the road network in the target area at future time based on real-time data collected by roadside units, and output a risk matrix. Based on the battery degradation model, the remaining range and charging demand priority of each vehicle at future time are predicted and the range vector is output. The supply and demand matrix, risk matrix, and endurance vector are integrated into a prediction matrix; Acquire real-time vehicle perception data and generate a state matrix, then calculate the deviation between the state matrix and the prediction matrix; A phase compensation operator is generated based on the deviation and applied to the initial quantum state to obtain the corrected quantum state. ; A scheduling model based on a quantum deep Q-network is constructed by pre-setting scheduling objectives and security constraints; candidate scheduling schemes are generated based on the input modified quantum state through the scheduling model; When the safety threshold is triggered, the quantum state is activated to rapidly collapse, and an emergency scheduling plan is output.

2. The ride-hailing scheduling method based on prediction-enhanced quantum reinforcement learning as described in claim 1, characterized in that: The passenger demand parameter set D includes: passenger real-time location d loc Expected travel time d time Destination location d dest Carpooling willingness d share Model preference d prefer ; The vehicle status parameter set V includes: real-time vehicle location. v loc Remaining battery power v batt Sensor health v sensor Autonomous driving level v adlevel Recent 3-minute obstacle avoidance record v obstacle The success rate of scheduling execution in the past hour v history ; The road network environment parameter set R includes: road segment congestion index. r cong Real-time phase of traffic lights r signal Temporary traffic control area r control Real-time weather r weather Emergencies in the region r event .

3. The ride-hailing scheduling method based on prediction-enhanced quantum reinforcement learning as described in claim 2, characterized in that: The initial quantum state The expression is as follows: ; In the above formula, Represents a superposition quantum state of three parameter sets; , and These are the quantum states corresponding to parameter sets D, V, and R, respectively; , and To meet the normalization conditions , and quantum state amplitude, ; Represents the fundamental quantum state; Indicates safety weight; Indicates the safe distance margin between the vehicle and the obstacle. This refers to the emergency braking response time. To assess the reliability of sensor data.

4. The ride-hailing scheduling method based on prediction-enhanced quantum reinforcement learning as described in claim 1, characterized in that: The demand forecasting model divides the target area into several 500m×500m grids and outputs the supply and demand intensity and passenger travel probability of each grid at various times in a specified future period. The supply and demand intensity is divided into 1-10 levels; the passenger probability is a probability value between 0 and 1; thus, a supply and demand matrix is ​​obtained, which is composed of the supply and demand intensity and passenger travel probability under different time and spatial dimensions. The risk prediction model divides the target area into several 500m×500m grids and outputs the dynamic road network risk of each grid at various times within a specified future period. The types of dynamic road network risk include risk values ​​corresponding to sudden congestion sections, construction areas, and high-frequency areas of illegal lane changes. This results in a risk matrix composed of various risk values ​​under different time and spatial dimensions. The battery degradation model first predicts the remaining power at each time point within a specified future period, and then generates charging demand priorities based on the remaining power according to a preset mapping relationship; thus obtaining a range vector composed of the remaining power and charging demand priorities under different time dimensions. The supply and demand matrix, risk matrix, and endurance vector are concatenated along the time dimension to obtain the required prediction matrix.

5. The ride-hailing scheduling method based on prediction-enhanced quantum reinforcement learning as described in claim 4, characterized in that: The expression for the battery degradation model is: ; In the above formula, v batt0 Indicates the initial battery level; t Indicates time, T Indicates ambient temperature; v Indicates vehicle speed; k , a and b These represent the influence coefficients of time, ambient temperature, and vehicle speed, respectively.

6. The ride-hailing scheduling method based on prediction-enhanced quantum reinforcement learning as described in claim 1, characterized in that: The corrected quantum state The calculation formula is: ; In the above formula, Represents the quantum phase compensation operator; The standard 2×2 Hermitian unitary operator representing the spin z component of the Pauli z operator has eigenvalues ​​of +1 and -1, corresponding to the two orthogonal computational ground states of the qubit. X percept Represents the state matrix; X pred Represents the prediction matrix; This represents the deviation between the state matrix and the prediction matrix; i Represents the imaginary unit; Indicates the phase compensation amount; These are the proportional coefficient, derivative coefficient, and integral coefficient of PID control, respectively. This represents the time integral dummy variable of the PID integral element.

7. The ride-hailing scheduling method based on prediction-enhanced quantum reinforcement learning as described in claim 1, characterized in that: The scheduling model includes an input layer, a hidden layer, and an output layer. The input layer is used to convert the modified quantum state into a classical feature vector through quantum measurement, and then perform quantum activation. The hidden layer realizes neuronal entanglement through CNOT quantum gates and introduces a quantum dropout layer to prevent overfitting. The output layer is used to output the Q value of each group of scheduling schemes and recommend candidate scheduling schemes. The scheduling model minimizes the loss function using the quantum gradient descent algorithm. L To achieve the updating of network parameters; the loss function L The expression is: ; In the above formula, Indicates the target Q value; This represents the actual Q value of the current prediction result; Representing quantum state The corresponding Hermitian conjugate left-hand state vectors are used together to realize quantum inner product operations; And / or, the learning rate of the quantum gradient descent algorithm Adaptive quantum learning rate strategy: Set initial learning rate When the scheduling success rate is above 95% for five consecutive rolling windows, the learning rate is reduced to [a lower threshold]. When the scheduling success rate falls below 80% or a safety threshold is triggered, the learning rate will be increased to [a higher percentage]. .

8. A storage medium storing a computer program, characterized in that: When the computer program is executed by the processor, it implements the ride-hailing scheduling method based on prediction-enhanced quantum reinforcement learning as described in any one of claims 1-7, and then generates candidate scheduling schemes for target vehicles based on dynamically updated data.

9. A computer program product comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the ride-hailing scheduling method based on prediction-enhanced quantum reinforcement learning as described in any one of claims 1-7, and then generates candidate scheduling schemes for target vehicles based on dynamically updated data.

10. A ride-hailing dispatching device, comprising a memory, a processor, and a computer program stored in the memory and running on the processor, wherein when the processor executes the computer program, it implements the ride-hailing dispatching method based on predictive reinforcement learning as described in any one of claims 1-7, and then generates candidate dispatching schemes for target vehicles based on dynamically updated data.