Vehicle formation method and system based on dobby machine
By using a two-layer learning framework based on multi-armed slot machines and blockchain technology, the problems of slow response latency, imperfect incentive mechanisms, and poor energy consumption fairness in vehicle platooning under highly dynamic traffic environments are solved, achieving efficient and stable vehicle platooning management.
Patent Information
- Application Number
- CN202511937075.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-22
- Publication Date
- 2026-03-31
AI Technical Summary
Existing vehicle platooning technologies face challenges such as slow response time, imperfect incentive mechanisms, and poor energy consumption fairness in highly dynamic traffic environments, making it difficult to maintain stability and synergy in complex environments.
A two-layer learning framework based on multi-armed slot machines is adopted, which combines UCB strategy and blockchain technology to design dynamic resource allocation and reward and punishment mechanisms. The multi-armed slot machine (MAB) learning framework is used to realize road screening and vehicle role matching. Combined with UCB strategy and compensation mechanism, blockchain technology is introduced to ensure data transparency and security.
It improves the fairness and incentive of the system, shortens the response time of formation tasks, optimizes energy consumption allocation, enhances the stability and efficiency of the system, and can maintain good robustness and real-time performance in complex traffic environments.
Smart Images

Figure CN121768183A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of vehicle platooning technology, and particularly relates to a vehicle platooning method and system based on a multi-armed slot machine. Background Technology
[0002] Vehicle platooning technology, an emerging vehicle-road cooperative driving mode developed with the advancement of wireless communication and autonomous driving technologies, has become an important research topic as a core technology for improving road utilization, reducing fuel consumption, and enhancing safety. Vehicle platooning (vehicle convoy driving) refers to a cooperative driving technology that organizes multiple vehicles to travel at specific intervals and sequences. The widespread application of vehicle platooning technology indicates its enormous potential for improving road operating efficiency in future transportation systems. With the development of vehicle-to-everything (V2X) and autonomous driving technologies, vehicle platooning has gradually moved from theory to practice, and it is of great significance for optimizing and intelligently developing modern transportation modes.
[0003] However, traditional vehicle platooning control technologies mainly include centralized optimization algorithms and static rule-driven approaches. While these address the stability and safety issues of platooning to some extent, they still have shortcomings. For example, traditional methods typically rely on rule-based control, lacking sufficient research on adapting to dynamic traffic environments and on vehicle benefit allocation and incentive mechanism design. Especially when facing high-speed movement, complex traffic flows, and unexpected situations, the stability and effectiveness of traditional methods are limited.
[0004] Current research attempts to address the problems of traditional methods through intelligent and distributed control strategies. Some studies have combined reinforcement learning, edge computing, and blockchain technologies to construct more flexible and efficient resource allocation models. Although these methods have shown certain advantages in theory and experiments, they still face challenges such as slow convergence speed, imperfect incentive mechanisms, and poor energy consumption fairness in highly dynamic traffic environments.
[0005] 1.1 Vehicle platooning technology optimization
[0006] Rebelo et al. elucidated the environmental impact of vehicle driving, analyzed the advantages of vehicle platooning in improving vehicle efficiency and air quality and health, and clarified the forward-looking research challenges in the field of vehicle platooning. Famularo et al. used a distributed multi-layer architecture for vehicles to achieve efficient management of autonomous vehicle logistics, which requires high communication stability. Therefore, Zhang et al. elucidated the impact of communication latency on vehicle platooning response lag, oscillation, and instability, and discussed the advantages and disadvantages of mitigation methods such as predictive control, lag compensation, and robust control. Zhu et al. designed a system to compensate for sensor failures to ensure platooning stability using adaptive fault-tolerant control and event-triggered communication, but this system has vehicle coordination problems in mixed traffic. Subsequently, D'Alfonso et al. proposed a distributed control framework based on deep reinforcement learning and MPC, but the computational complexity is still too high in large-scale scenarios. Wei et al. used multi-agent deep reinforcement learning to optimize global traffic efficiency and safety through collaborative learning, effectively alleviating the computational pressure of real-time decision-making for large-scale traffic flow.
[0007] Regarding platooning optimization, Pi et al. verified that close following and intelligent control can reduce wind resistance and fuel consumption; Dong et al. pointed out that vehicle cooperation in mixed platooning has a significant impact on energy consumption and traffic flow. Considering the dynamic changes in the environment, Yi et al. reviewed vehicle-road cooperative sensing technologies to find cooperative technologies adapted to complex environments. To further improve the theoretical foundation of platooning optimization, Nolan et al. analyzed the effects of different spacing and arrangement methods on airflow and resistance through experiments and simulations, providing a theoretical basis for further realizing reasonable vehicle platooning technology.
[0008] In terms of communication and computing collaboration, Zhu et al. proposed a dynamic task scheduling strategy based on real-time vehicle status and edge computing, which reduced system latency. Xiao Xiang et al. optimized safety performance by quantifying the impact of severe weather on vehicle lateral stability and spacing control. Xu et al.'s multi-objective resource allocation method based on DRL maintains formation safety while taking into account energy efficiency advantages, thus balancing energy efficiency and safety issues.
[0009] Intelligent optimization and learning algorithms provide powerful tools for complex decision-making in vehicle-to-everything (V2X) networks. Wang et al. used the MAB (Multi-Aspect Resource Allocation) algorithm to improve V2X data transmission, reduce latency and bandwidth consumption, but lacked consideration of the impact of dynamic environments. Therefore, Zhang et al. proposed a context-aware resource allocation algorithm for cellular V2X networks, which can adapt to environmental changes in complex traffic scenarios, significantly improving adaptability and overall network performance. Although the environmental adaptability of MAB was improved, its convergence speed was insufficient when vehicles were moving at high speeds. Therefore, Lei et al. proposed the Samp-DSGD federated learning algorithm to accelerate global convergence by filtering low-quality updates. Xiao et al. constructed a distributed collaborative MAB model that utilizes information sharing to reduce trial-and-error costs and improve deployment efficiency. Xiong et al. combined MAB with Q-learning to alleviate the curse of dimensionality in high-dimensional states and capture complex energy consumption patterns in vehicle platooning.
[0010] In research on optimization algorithms, Long et al. employed particle swarm optimization to accelerate the convergence of large-scale convoy tasks, addressing the problem of large-scale computational complexity. Xu et al. proposed a multi-objective resource allocation method based on deep reinforcement learning to balance communication latency, bandwidth utilization, and network reliability. Wen et al. improved system robustness through distributed trajectory optimization and sliding mode control, significantly reducing computational complexity and communication burden, making large-scale deployment possible. To address dynamic environmental challenges, He et al. proposed an adaptive trajectory control algorithm for rapid vehicle position adjustment. He et al. addressed the limitations of the fully observable assumption by proposing a vehicle coordination decision-making scheme based on partial information sharing. Zeng et al. designed a federated learning autonomous controller that optimizes convoy strategies while protecting privacy and maintaining near-centralized learning performance, mitigating privacy risks associated with data sharing.
[0011] 1.2 Incentive Mechanisms and Resource Allocation
[0012] In vehicle platooning, fair resource allocation and incentive mechanism design between PLs and PFs are crucial factors in maintaining system stability. Mohapatra et al. proposed a compensation strategy based on uniform crowdsourcing, ensuring task quality and efficient allocation in static environments. Wu et al. introduced an adaptive incentive mechanism, adjusting compensation amounts in real-time to improve its adaptability in dynamic environments. Wu et al. proposed a continuously online incentive framework, reducing computational complexity while maintaining incentive compatibility. Deng et al. addressed the potential trust risks of centralized mechanisms, utilizing blockchain technology to ensure transparency and immutability of the compensation process. Ying et al. emphasized long-term partnerships, designing a reputation mechanism that considers long-term relationships to address short-term vehicle behavior defects and strategic exit issues.
[0013] Most of the mechanisms described above assume that the vehicles are in an ideal state. Lesch et al. designed a heterogeneity compensation mechanism to address uncertainties such as differences in vehicle dynamics, sensor delays, and external disturbances. Acland et al. analyzed the trade-off between maximizing social welfare and individual interests using distributed weights. To address the problem of false vehicle reporting, Zhao et al. proposed a robust resource allocation mechanism under incomplete information conditions, compensating for the shortcomings of the complete information assumption.
[0014] In recent years, integrated learning and incentive mechanisms have attracted much attention. Gao et al. combined incentive models with MAB (Multi-Action Learning) to adjust strategies through dynamic learning, achieving long-term social welfare maximization. Gu et al. introduced a risk measure for energy consumption uncertainty, providing a theoretical basis for decision-making in dynamic environments. Shu et al. developed a multi-stage incentive mechanism to optimize system efficiency at different stages of vehicle entry, driving, and exit. Among the various incentive mechanisms mentioned above, multiple compensation mechanisms and reputation systems often focus on task efficiency and safety assurance. Multi-stage fairness models and learning methods are more suitable for dynamic resource optimization and balanced incentives. Hybrid mechanisms combining distributed learning and adaptive incentive strategies perform well in dynamic traffic scenarios. Yu et al.'s combined incentive method improves resource allocation efficiency and system reliability while enhancing vehicle participation.
[0015] In summary, in dynamic environments, the stability and synergy of vehicle platoons are significantly affected by external factors such as road conditions, traffic flow fluctuations, and severe weather. Furthermore, existing incentive mechanisms are insufficient in their incentive compatibility design, making it difficult to balance individual interests with overall benefits while preventing the potential risks of malicious vehicles. Therefore, there is an urgent need to develop an optimal platoon management scheme that can adapt to complex dynamic environments and achieve efficient and fair incentives.
[0016] Based on the above analysis, the problems and shortcomings of the existing technology are as follows:
[0017] Existing methods still face challenges in highly dynamic traffic environments, such as slow response time, imperfect incentive mechanisms, and poor energy consumption fairness. Summary of the Invention
[0018] To address the problems existing in the prior art, this invention provides a vehicle platooning method based on a multi-armed slot machine.
[0019] This invention is implemented as follows: a vehicle platooning method based on a multi-armed slot machine includes:
[0020] Step 1, Vehicle Utility Model;
[0021] Vehicles waiting to be platooned at any time The set of selectable mission paths is At time t, vehicle vi Take on a role The utility of time is:
[0022] U i,p (t)=v i,p (t)-c i,p (t)+r i,p (t)+s i,p (t),
[0023] U i,p (t) is v i The revenue and expenditure indicators when performing task M(n) assist in the MAB parameter update and decision-making process;
[0024] v i,p (t) 为基础估值 , is vehicle v i The expected value of the current task is related to contextual factors such as the current environment:
[0025]
[0026] Where, ξ p >0 represents the benefit coefficient corresponding to character p, d i,p (t) represents vehicle v i The distance to the mission destination, τ > 0 is a distance-sensitive parameter;
[0027] c i,p (t) represents energy consumption expenditure, which is the vehicle's energy consumption expenditure. i Energy consumption when playing role p in a formation mission, and vehicle speed v i (t) and the wind resistance experienced by different roles p can be expressed as:
[0028] c i,p (t)=α p ·f(v i (t)),
[0029] Where, α p >0 represents a specific energy consumption coefficient for the character, f(v) i (t) is the energy consumption function;
[0030] r i,p (t) As compensation revenue, to incentivize vehicles to assume high-energy-consuming positions, compensation revenue is provided to the lead vehicle PL:
[0031]
[0032] χ p >0 represents the corresponding compensation coefficient. This is an indicator function that takes the value 1 only when p = PL and 0 when p = PF;
[0033] s i,p (t) represents the synergistic effect, reflecting the positive benefits that vehicles gain from cooperating with each other. For example, following vehicles may benefit from aerodynamic wakes or safety may be improved when multiple vehicles cooperate.
[0034] It can be represented as:
[0035]
[0036] Where, N i For vehicle v i The set of neighboring vehicles, κ ij This represents the coordination strength coefficient. Numerous experiments have shown that the vehicle spacing remains at a fixed, relatively small distance d. ij (t) yields the best energy-saving effect, therefore this study uses h(d) ij The optimal cooperative distance is described by (t)). The closer the distance between vehicles is to the optimal cooperative distance, the stronger the cooperative effect; the farther the distance, the weaker the cooperative effect.
[0037] Step 2, MAB dynamic filtering;
[0038] Step 3: Reward and punishment mechanism and payment rules.
[0039] Furthermore, the MAB dynamic filtering:
[0040] 1) Road Filtering MAB
[0041] For the task path set When driving alone or in convoy, the vehicle selects the appropriate road. n After passage is completed, the system calculates the current average actual benefit R based on the comprehensive utility feedback from the vehicle under different conditions. r (t), while updating the vehicle's local ledger, recording road conditions under different environments. n Historical mission overall benefits μ r (t), the revenue update expression:
[0042] μ i,r (t)=φμ i,r (t-1)+(1-φ)R r (t),
[0043] Where φ∈(0,1) is the time decay factor, and μ is the vehicle's comprehensive gain at time t. i,r (t) The update is not based solely on the current average real return R. r(t) is also related to the gains at time t-1, which means that it takes into account two important influencing factors: historical data and the current performance of the vehicle. This ensures that the data information updates are highly timely and also improves the availability of information when implementing the MAB strategy for road screening.
[0044] Assuming novice vehicles with no "exploration" experience apply for platooning, they can be initialized using vehicles with similar driving experience, reducing the cold start for novice vehicles in the first layer of screening; For vehicle v i For road r n The initial estimate can be matched with v from historical data. i Similar features The estimated historical revenue from inheriting this similar vehicle is denoted as Right now:
[0045]
[0046] To address the issue of vehicles being unable to fully explore all roads, a strategy based on the Upper Confidence Bound (UCB) is adopted for road exploration.
[0047]
[0048] Where, n i,r (t) represents vehicle v i Driving on the road n The number of historical records, γ>0, β>0 are the exploration adjustment coefficients; μ i,r (t) represents the vehicle's comprehensive revenue at time t; UCB i,r (t) can reflect the vehicle's exploration level on the corresponding road; N r (t) represents road r n Historical selection count, via r * (t) can reflect the road r n Based on the historical best exploration value of UCB, through the vehicle's own UCB i,r (t) and the corresponding road requirements r * (t) Compare and filter to identify those that do not match the road r n Vehicles that meet the requirements will be included in the candidate vehicle set. among.
[0049] 2) Vehicle matching MAB.
[0050] Furthermore, the vehicle is equipped with MAB:
[0051] Use μ i,p (t) represents vehicle v i In the position of role The average utility estimate at time μ i,r (t) differs from μ i,p (t) More attention is paid to vehicles on the road. n Specific utility estimation when serving as PL or PF, μ i,r (t) More attention is paid to vehicles on the road. n The average utility estimate on, therefore μ i,p (t) can be greater than μ i,r (t) A more detailed description of the vehicle's position on the road. n Utility estimation when playing role p; using U i,p (t) represents vehicle v i If the actual utility observed at time t while playing role p is given, then its historical average utility update can be expressed as:
[0052] μ i,p (t)=φμ i,p (t-1)+(1-φ)U i,p (t),
[0053] To measure uncertainty, the UCB value of the vehicle matching MAB layer is defined as follows:
[0054]
[0055] Where, n i,p (t) represents vehicle v i The number of times a vehicle has been selected in the history of playing role p; the number of times a vehicle has played different roles reflects the vehicle's experience to some extent. Vehicles with more experience are better able to handle various unexpected situations that occur during formation.
[0056] After obtaining the candidate vehicle set After that, it is necessary to One PL and one |x were selected. * -1 PF, and maximize the profit W * At this point, the problem is transformed into an "Assignment Problem," and the Hungarian algorithm is used to find the optimal matching strategy x. * For ease of description, assume that in the current formation task, a total of N vehicles need to be selected (i.e., 1 PL and N-1 PF).
[0057] Before proceeding with the formal screening, a "vehicle-role" UCB value matrix A needs to be constructed. The size of the matrix depends on the set of candidate vehicles. Given the number of vehicles N required for the task, define a matrix row to store the corresponding... All vehicles, totaling The rows and columns correspond to the N role positions to be assigned in the formation. These N positions include: 1 PL and N-1 PF.
[0058] like If the value is greater than N, then there will be more vehicles than needed. However, the Hungarian algorithm requires that the number of rows and columns be equal, so matrix A needs to be expanded: if You can add at the end of the column. A virtual slot, making its UCB i,p (t) = 0, ensuring the number of columns equals the number of rows; after confirming the matrix size, corresponding values need to be filled in the corresponding positions for easy filtering, therefore, in vehicle v i At the intersection of the corresponding row and the navigator position column, enter the UCB of the vehicle when it served as the historical navigator. i,PL (t)(i=1,2,3···); At the intersection of the following position columns, fill in the UCB when the vehicle served as the historical following vehicle. i,PF (t); If there is a virtual slot, fill in 0 in the corresponding position; this will give you a A square matrix, the values in which represent the maximum profit W. * ;
[0059] Assume the set of candidate vehicles in a task is The task requires 3 vehicles N, meaning 1 PL and 2 PF vehicles need to be selected. The column slots are represented as {PL, PF1, PF2}, therefore two virtual slots need to be added, resulting in a 5×5 matrix A, represented as:
[0060]
[0061] Take a from the matrix ij (representing the element in the i-th row and j-th column of the "vehicle-role" UCB value matrix A) is transformed into max(a ij )-a ij , where max(a ij ) is the maximum value among all elements of the matrix, thus transforming the maximization problem into a minimization problem; for each row, find the minimum value of that row and subtract it from all elements of that row; for each column, find the minimum value of that column and subtract it from all elements of that column; this ensures that each row and each column has at least one 0 element;
[0062] In the processed matrix, try to mark some 0 elements so that each row and each column can be marked with at most one 0; try to cover all 0 elements with the minimum number of column lines; if the number of lines is equal to the matrix dimension, a feasible solution can be found; otherwise, it is necessary to continue to make further adjustments to the elements not covered by lines until all 0s can be covered with the same number of lines as the dimension.
[0063] A perfect match exists when all zeros can be covered by lines of the same order as the matrix. Based on the marked zero elements, a column index is assigned to each row sequentially, ultimately yielding the optimal match x. * ;
[0064] From the matching result x * It can be known that: vehicle v i Which position is it assigned to; if the column belongs to the navigator role, it indicates that vehicle v i Select the lead vehicle; if the column belongs to the follower role, i.e., vehicle v i Selected vehicle to follow; if matched with a "virtual slot", it means that the vehicle has not been selected into the formation and will continue to wait for the next round of selection.
[0065] Furthermore, the reward and punishment mechanism and payment rules are as follows:
[0066] (1) Reward and punishment mechanism;
[0067] (2) Payment rules.
[0068] Furthermore, the aforementioned reward and punishment mechanism:
[0069] The behavior of a vehicle during mission M(n) is monitored through collective verification by other vehicles in the formation; when a violation is detected, the other vehicles in the formation verify the behavior through multi-party consensus.
[0070]
[0071] Among them, O i (j) represents the detection value of the violation by vehicle i against vehicle j (1 indicates a confirmed violation, 0 indicates no violation was observed), and n represents the number of vehicles participating in the consensus. This is the consensus threshold; when C v If ≥Γ, then vehicle v is confirmed. j There were violations;
[0072] Violations are categorized into three types based on severity: T = {V} L V M V H}; where V L V M V H These represent different degrees of violation, corresponding to different levels of harm to the formation; when a violation occurs and causes economic losses to the formation mission, the violating vehicle v j The vehicles in the formation and the person who initiated the mission should be compensated according to the revenue generated when the mission is completed normally.
[0073]
[0074] Among them, P i The outstanding amount for vehicles involved in violations. The amount paid for marginal contribution loss, u i The amount of loss incurred by the task initiator; thereby further increasing the punishment for vehicles that violate regulations. When the cost of violation is greater than the benefits of violation, it can effectively curb the occurrence of violations.
[0075] Once the violation is confirmed through the above checks, the system records the violation information via the blockchain; however, to protect user privacy, the violation record uses blind signature and zero-knowledge proof technology, and the recording format is as follows:
[0076] R j (t)=(H(ID v ),V t ,T s ,L,Sig)
[0077] Among them, H(ID) v V is the hash value that uniquely identifies the vehicle. t For the type of violation, T s For timestamp, L is the geolocation hash, and Sig is the multi-signature verification information; by recording the result of the action but not the specific details, it prevents issues related to user privacy and security.
[0078] By leveraging the spread of violations to help more vehicles in a convoy receive early warnings, the propagation of blockchain records follows the mathematical model below:
[0079]
[0080] Among them, R j (t) represents vehicle v j The set of violation records possessed at time t, N i (t) represents vehicle v i The set of other vehicles encountered at time t; vehicle v at time t. i The existing early warning information is R i (t); vehicle v i By using set N i The vehicles in (t) communicate with each other to obtain their early warning information R. j (t); vehicle v i At the next time point t+1, it will interact with the records of its own vehicles and those of neighboring vehicles; in this way, violation records can be naturally spread among different vehicles, forming an "immune system"-like early warning mechanism.
[0081] Based on the violation history recorded by the blockchain, the system dynamically adjusts the vehicle's future platooning permissions; defining vehicle v jThe historical impact factors of violations are:
[0082]
[0083] Where, n L n M n H w represents the number of violations at different levels. L w M w H Let w be the weighting coefficient for each level of violation, and w L <w M <w H ; Let t be the time decay function. r Here, θ is the timestamp of the violation record, and t is the current time; the impact of the violation record gradually weakens over time; θ > 0 is the time decay coefficient, and r represents R. j (t) A single violation record in the set;
[0084] During the vehicle matching MAB phase, the UCB value of non-compliant vehicles is adjusted:
[0085] UCB′ i,p (t)=UCB i,p (t)·P(H v )
[0086] Among them, UCB′ i,p (t) represents the adjusted UCB value. The UCB value is a key factor affecting whether a vehicle can be selected into the formation. Therefore, a UCB value that is smaller may be directly rejected from joining the formation.
[0087] P(H v ) is the penalty function for violations:
[0088]
[0089] Where ω>0 is the penalty intensity coefficient;
[0090] When H v If a certain threshold is exceeded, the vehicle will be unable to participate in certain types of formation missions for a certain period of time.
[0091]
[0092] Where σ = 3 is the critical value for serious violations; T(v) is the prohibition periodic function, which is proportional to the violation history factor; within the T(v) period, vehicles are prohibited from gaining benefits by joining vehicle platoons. This restricts violating vehicles while ensuring that vehicles have the opportunity to regain platoon participation rights after a period of time.
[0093] In addition, the system allows vehicles to restore trust through behavioral proofs:
[0094]
[0095] When a vehicle successfully completes a platooning maneuver without violations m times, its historical violation impact factor K will decrease proportionally.
[0096] Furthermore, the payment rules are as follows:
[0097] Based on the actual service of all vehicles An optimal vehicle formation N can be calculated. * The total revenue can be expressed as:
[0098]
[0099] If vehicle v is removed from any vehicle platoon i The optimal profit that the system can achieve on the remaining vehicles is Then, the vehicle v is defined according to its marginal contribution. i The payment is defined as:
[0100]
[0101] In order to consider the contribution of vehicles to the overall formation system, it is necessary to calculate the true economic benefit of each vehicle, v. i The final net income is:
[0102]
[0103] Where π i The amount a vehicle needs to pay based on its marginal contribution can actually reduce its net revenue if the vehicle misreports, thus encouraging honest reporting. The rational choice for a vehicle is to report truthfully.
[0104] Marginal contribution payments ensure that honest reporting remains the dominant strategy in cases of vehicle violations; the distributed trust records on the blockchain constrain misconduct without the need for third-party escrow, reducing system complexity and trust requirements.
[0105] Another object of the present invention is to provide a vehicle platooning system based on a multi-armed slot machine, comprising:
[0106] Vehicle utility module, used for vehicle utility model;
[0107] Dynamic filtering module, used for dynamic filtering of MABs;
[0108] The reward and punishment module is used for reward and punishment mechanisms and payment rules.
[0109] Another object of the present invention is to provide a computer device including a memory and a processor, the memory storing a computer program that, when executed by the processor, causes the processor to perform the steps of the vehicle platooning method based on a multi-armed slot machine.
[0110] Another object of the present invention is to provide a computer-readable storage medium storing a computer program that, when executed by a processor, causes the processor to perform the steps of the vehicle platooning method based on a multi-armed slot machine.
[0111] Another objective of the present invention is to provide an information data processing terminal for implementing the vehicle platooning system based on the multi-armed slot machine.
[0112] Based on the above technical solutions and the technical problems solved, please analyze the advantages and positive effects of the technical solution to be protected by this invention from the following aspects:
[0113] First, addressing the technical problems existing in the prior art and the difficulty in solving them, this paper closely analyzes, in conjunction with the technical solution to be protected by this invention and the results and data obtained during the research and development process, how the technical solution of this invention solves the technical problems, and the inventive technical effects brought about by solving these problems. The specific description is as follows:
[0114] This invention proposes a vehicle platooning resource allocation method that integrates a Multi-Armed Bomber (MAB) learning framework and incentive mechanism. A two-layer MAB structure handles road selection and vehicle role matching, while the UCB strategy and compensation mechanism enhance the system's fairness and incentive. Furthermore, the introduction of blockchain technology effectively ensures data transparency and security, providing scalable and reliable technical support for the long-term development of vehicle platooning.
[0115] The main contributions of this invention are as follows:
[0116] (1) A dynamic vehicle platooning position optimization method based on the multi-armed slot machine learning framework is proposed to effectively address the resource allocation problem under uncertain traffic conditions.
[0117] (2) A compensation mechanism that balances fairness and incentives was designed, which optimized energy consumption allocation and improved the stability and efficiency of the system.
[0118] This invention addresses the platooning challenges involving multiple roads and vehicles, proposing a comprehensive solution based on a MAB (Multi-Area Learning) and incentive mechanism. By integrating "exploration" into the daily idle computing period of vehicles, the online computational pressure at the initiation of platooning tasks is reduced. The UCB (Unified Cost-Based Learning) strategy designs a two-layer dynamic matching system for roads and roles, balancing real-time performance with optimal matching. A reward-penalty payment rule ensures individual rationality and incentive compatibility among participating vehicles, effectively curbing false bids and violations. Experimental results show that this solution significantly improves social welfare and overall vehicle utility in scenarios of varying scales, while maintaining low network latency and computational overhead, verifying its feasibility and robustness in complex traffic environments. Even with malicious vehicles participating in platooning, the system maintains good robustness.
[0119] Secondly, as supplementary evidence of the inventive step of the claims of this invention, it is also reflected in the following important aspects:
[0120] (1) The expected benefits and commercial value of the technical solution of this invention after transformation are as follows:
[0121] With the rapid development and gradual maturation of vehicle wireless communication technology and intelligent driving technology, vehicle platooning technology is no longer confined to the theoretical research level, but is increasingly demonstrating its enormous potential for practical application. This invention addresses key bottlenecks in vehicle platooning, such as dynamic resource allocation, fair incentives, and trust building, providing a solid technical foundation for the commercialization and widespread adoption of vehicle platooning technology. It solves the problems of scalability and robustness in large-scale complex scenarios, laying a solid theoretical foundation for the successful commercialization of the technology. This solution demonstrates excellent performance in both scalability and robustness. Experimental verification shows that even in large-scale application scenarios involving a large number of vehicles and multiple roads, this solution can still maintain computational latency within an acceptable range. Simultaneously, the system exhibits good robustness; even with a certain proportion of malicious or non-cooperative vehicles participating in the platoon, the overall system performance and stability can still maintain a high level. These characteristics ensure that this solution has high practicality and reliability in real-world, complex commercial operating environments.
[0122] The commercial value of this solution extends beyond optimizing and improving the efficiency of existing transportation models; it has the potential to foster entirely new business models and mobility services. For small and medium-sized logistics companies or individual truck drivers: Based on the solution's dynamic and efficient platooning capabilities, a shared logistics platform can be built, enabling these companies or drivers to share the economies of scale of large fleets through dynamic platooning, such as uniformly optimized low-drag queues and potential right-of-way. For autonomous driving fleet operators: This technology enables higher-level fleet scheduling and management, optimizes vehicle utilization, and provides more cost-effective and competitive transportation services. Furthermore, the vehicle behavior and contribution data recorded using blockchain technology in the solution have inherent value. This trustworthy and tamper-proof data can be used for vehicle credit assessment, pricing of customized insurance products, or as qualification certificates for vehicles to participate in more complex collaborative tasks in the future (such as distributed energy trading and city-level collaborative perception), thereby generating entirely new data value-added services. It is expected that the widespread application of this technology will strongly promote the rapid development of multiple sub-markets such as intelligent logistics, autonomous driving transportation services, and vehicle-to-everything (V2X) value-added services, forming new economic growth points and industrial ecosystems.
[0123] (2) The technical solution of this invention fills a technical gap in the industry both domestically and internationally:
[0124] This invention systematically solves the core challenges that have long constrained the development of the vehicle platooning industry, such as dynamic resource allocation, fair incentives, and trust building. Addressing the shortcomings of traditional solutions in dynamic adaptability, a two-layer multi-armed platooning (MAB) framework is designed: the upper layer uses the UCB algorithm to dynamically screen the road environment (time complexity O(M)), while the lower layer uses an improved Hungarian algorithm to accurately match vehicle roles, compressing the decision-making latency of a 6-vehicle platoon to the millisecond level, significantly improving its response speed compared to traditional methods. Regarding fairness mechanisms, the solution establishes a PL energy consumption quantification model for the first time, accurately calculating the costs incurred due to additional wind resistance and communication coordination. A dynamic compensation function covers more than 50% of the additional expenditures, and combined with a payment rule based on marginal contribution, the utility of PL surpasses that of PF, solving the pain point of "insufficient motivation of the lead vehicle." To build a decentralized trust system, the solution creatively integrates blockchain technology, developing a closed-loop management system that includes multi-vehicle consensus verification, blind signature violation evidence storage, and UCB reputation linkage, improving the efficiency of malicious behavior identification while ensuring a reduction in the risk of privacy data leakage. At the engineering implementation level, a hybrid architecture of "offline pre-computation + RSU parallel processing" keeps the decision-making latency in city-level scenarios within an acceptable range, and the "multiple-selection secretary problem" optimization improves the matching efficiency of candidate vehicles. These innovations not only enable breakthroughs in key indicators such as dynamic adaptability, system stability, and participation enthusiasm in vehicle platooning technology, but more importantly, they construct a three-in-one technical paradigm of "intelligent algorithm - economic incentive - trustworthy mechanism." Its value extends beyond the transportation field, providing a reusable solution framework for various distributed collaborative systems, marking a significant turning point for intelligent collaboration technology from theoretical exploration to large-scale commercial application.
[0125] (3) Whether the technical solution of the present invention solves the technical problem that people have long wanted to solve but have never been able to solve successfully:
[0126] This invention innovatively solves the three core challenges of vehicle platooning technology in complex and dynamic traffic environments: dynamic adaptability, fair incentives, and trust building, achieving a key breakthrough from theoretical research to commercial application. Addressing the shortcomings of traditional solutions under static rule control, such as slow algorithm convergence, low PL (Power Line Provider) participation, and high risk of centralized system tampering, this solution constructs a three-in-one technical system of "intelligent decision-making - economic incentives - trustworthy mechanisms." First, it employs a two-layer multi-armed machine (MAB) learning framework, using the UCB algorithm to dynamically screen the road environment and the Hungarian algorithm to accurately match vehicle roles, compressing the vehicle platooning decision-making delay and improving dynamic adaptation efficiency. Second, it pioneers a role-differentiated compensation mechanism, accurately quantifying the additional air resistance and communication costs borne by the PL, designing a dynamic compensation function to cover more than 50% of the additional expenditure, and combining this with a payment rule based on marginal contribution, enabling the PL's utility to surpass the PF (Power Line Provider), thus improving system stability. Third, the innovative integration of blockchain technology constructs a distributed trust system, employing multi-vehicle consensus verification and blind signature for violation evidence storage. This reduces the false positive rate while mitigating the risk of privacy data leakage. Furthermore, it forms a closed loop of "behavior-reputation-reward" through UCB reputation linkage, improving the efficiency of malicious behavior identification. At the engineering implementation level, the "offline pre-computation + RSU parallel processing" architecture keeps city-level decision-making latency within an acceptable range, and the "multiple-selection secretary problem" optimization improves matching efficiency. This solution not only solves the technical bottlenecks in vehicle platooning but also provides a reusable technical paradigm for complex collaborative systems such as the sharing economy and distributed energy networks, marking a new stage in the large-scale application of distributed intelligent collaboration technology.
[0127] (4) Does the technical solution of the present invention overcome technical bias?
[0128] The technical solution of this invention significantly overcomes three major technical biases that have long existed in the field of vehicle platooning technology: First, by introducing a multi-armed machine (MAB) learning framework to replace traditional static rule control, it adopts an "exploration-exploitation" dynamic strategy to achieve data-driven adaptive decision-making, effectively solving the limitations of relying on idealized environment assumptions and deterministic models; Second, by using blockchain technology to build a decentralized trust system, it overcomes the inherent defects of traditional centralized management models, such as single point of failure, abuse of power, and information asymmetry, through distributed ledgers, multi-party consensus mechanisms, and transparent algorithm execution; Finally, addressing the problem of traditional designs prioritizing macro-efficiency over individual fairness, it establishes a refined utility model based on role differences, ensuring cost-benefit parity among all participants through mechanisms such as special compensation for the lead vehicle, fundamentally resolving the contradiction between system stability and participation enthusiasm. These innovations not only realize a paradigm shift from rule-driven to learning-driven, from centralized control to distributed collaboration, and from efficiency-first to a balance between fairness and efficiency, but also provide a complete technical solution for building a trustworthy, reliable, and sustainable intelligent platooning system adapted to real traffic environments. Attached Figure Description
[0129] Figure 1 This is a flowchart of a vehicle platooning method based on a multi-armed slot machine provided in an embodiment of the present invention.
[0130] Figure 2 This is a structural block diagram of a vehicle platooning system based on a multi-armed slot machine provided in an embodiment of the present invention.
[0131] Figure 3 This is a vehicle-road cooperative model diagram provided in an embodiment of the present invention.
[0132] Figure 4 This is a road screening MAB model diagram provided in an embodiment of the present invention.
[0133] Figure 5 This is a vehicle matching MAB model diagram provided in an embodiment of the present invention.
[0134] Figure 6 This is a comparison chart of vehicle social welfare provided in an embodiment of the present invention.
[0135] Figure 7 This is a comparison diagram of vehicle utility provided in an embodiment of the present invention.
[0136] Figure 8 This is a comparison chart of the overall effectiveness provided by the embodiments of the present invention.
[0137] Figure 9 This is a group test diagram provided in an embodiment of the present invention.
[0138] Figure 10 This is a large-scale test diagram provided in an embodiment of the present invention.
[0139] Figure 11 This is an elasticity index diagram provided in an embodiment of the present invention. Detailed Implementation
[0140] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.
[0141] like Figure 1 As shown, an embodiment of the present invention provides a vehicle platooning method based on a multi-armed slot machine, comprising the following steps:
[0142] S101, Vehicle Utility Model;
[0143] Vehicles waiting to be platooned at any time The set of selectable mission paths is At time t, vehicle v i Take on a role The utility of time is:
[0144] U i,p (t)=v i,p (t)-c i,p (t)+r i,p (t)+s i,p (t)
[0145] U i,p (t) is v i The revenue and expenditure indicators when performing task M(n) assist in the MAB parameter update and decision-making process;
[0146] v i,p (t) is the base valuation, which is the vehicle v i The expected value of the current task is related to contextual factors such as the current environment:
[0147]
[0148] Where, ξ p >0 represents the benefit coefficient corresponding to character p, d i,p (t) represents vehicle v i The distance to the mission destination, τ > 0 is a distance-sensitive parameter;
[0149] c i,p (t) represents energy consumption expenditure, which is the vehicle's energy consumption expenditure. i Energy consumption when playing role p in a formation mission, and vehicle speed v i (t) and the wind resistance experienced by different roles p can be expressed as:
[0150] ci,p (t)=α p ·f(v i (t))
[0151] Where, α p >0 represents a specific energy consumption coefficient for the character, f(v) i (t) is the energy consumption function;
[0152] r i,p (t) As compensation revenue, to incentivize vehicles to assume high-energy-consuming positions, compensation revenue is provided to the lead vehicle PL:
[0153]
[0154] χ p >0 represents the corresponding compensation coefficient. This is an indicator function that takes the value 1 only when p = PL and 0 when p = PF;
[0155] s i,p (t) represents the synergistic effect, reflecting the positive benefits that vehicles gain from cooperating with each other. For example, following vehicles may benefit from aerodynamic wakes or safety may be improved when multiple vehicles cooperate.
[0156] It can be represented as:
[0157]
[0158] Where, N i For vehicle v i The set of neighboring vehicles, κ ij This represents the coordination strength coefficient. Numerous experiments have shown that the vehicle spacing remains at a fixed, relatively small distance d. ij (t) yields the best energy-saving effect, therefore this study uses h(d) ij The optimal cooperative distance is described by (t)). The closer the distance between vehicles is to the optimal cooperative distance, the stronger the cooperative effect; the farther the distance, the weaker the cooperative effect.
[0159] S102, MAB dynamic filtering;
[0160] S103, Reward and Punishment Mechanism and Payment Rules.
[0161] MAB dynamic filtering provided in this embodiment of the invention:
[0162] 1) Road Filtering MAB
[0163] For the task path set When driving alone or in convoy, the vehicle selects the appropriate road. n After passage is completed, the system calculates the current average actual benefit R based on the comprehensive utility feedback from the vehicle under different conditions.r (t), while updating the vehicle's local ledger, recording road conditions under different environments. n Historical mission overall benefits μ r (t), the revenue update expression:
[0164] μ i,r (t)=φμ i,r (t-1)+(1-φ)R r (t)
[0165] Where φ∈(0,1) is the time decay factor, and μ is the vehicle's comprehensive gain at time t. i,r (t) The update is not based solely on the current average real return R. r (t) is also related to the gains at time t-1, which means that it takes into account two important influencing factors: historical data and the current performance of the vehicle. This ensures that the data information updates are highly timely and also improves the availability of information when implementing the MAB strategy for road screening.
[0166] Assuming novice vehicles with no "exploration" experience apply for platooning, they can be initialized using vehicles with similar driving experience, reducing the cold start for novice vehicles in the first layer of screening; For vehicle v i For road r n The initial estimate can be matched with v from historical data. i Similar features inherit Right now:
[0167]
[0168] To address the issue of vehicles being unable to fully explore all roads, a strategy based on the Upper Confidence Bound (UCB) is adopted for road exploration.
[0169]
[0170] Where, n i,r (t) represents vehicle v i Driving on the road n The number of historical records, γ>0, β>0 are the exploration adjustment coefficients; UCB i,r (t) can reflect the vehicle's exploration level on the corresponding road; N r (t) represents road r n Historical selection count, via r * (t) can reflect the road r n Based on the historical best exploration value of UCB, through the vehicle's own UCB i,r (t) and the corresponding road requirements r* (t) Compare and filter to identify those that do not match the road r n Vehicles that meet the requirements will be included in the candidate vehicle set. among.
[0171] 2) Vehicle matching MAB.
[0172] The vehicle matching MAB provided in this embodiment of the invention:
[0173] Use μ i,p (t) represents vehicle v i In the position of role The average utility estimate at time μ i,r (t) differs from μ i,p (t) More attention is paid to vehicles on the road. n Specific utility estimation when serving as PL or PF, μ i,r (t) More attention is paid to vehicles on the road. n The average utility estimate on, therefore μ i,p (t) can be greater than μ i,r (t) A more detailed description of the vehicle's position on the road. n Utility estimation when playing role p; using U i,p (t) represents vehicle v i If the actual utility observed at time t while playing role p is given, then its historical average utility update can be expressed as:
[0174] μ i,p (t)=φμ i,p (t-1)+(1-φ)U i,p (t)
[0175] To measure uncertainty, the UCB value of the vehicle matching MAB layer is defined as follows:
[0176]
[0177] Where, n i,p (t) represents vehicle v i The number of times a vehicle has been selected in the history of playing role p; the number of times a vehicle has played different roles reflects the vehicle's experience to some extent. Vehicles with more experience are better able to handle various unexpected situations that occur during formation.
[0178] After obtaining the candidate vehicle set After that, it is necessary to One PL and one |x were selected. * -1 PF, and maximize the profit W *At this point, the problem is transformed into an "Assignment Problem," and the Hungarian algorithm is used to find the optimal matching strategy x. * For ease of description, assume that in the current formation task, a total of N vehicles need to be selected (i.e., 1 PL and N-1 PF).
[0179] Before proceeding with the formal screening, a "vehicle-role" UCB value matrix A needs to be constructed. The size of the matrix depends on the set of candidate vehicles. Given the number of vehicles N required for the task, define a matrix row to store the corresponding... All vehicles, totaling The rows and columns correspond to the N role positions to be assigned in the formation. These N positions include: 1 PL and N-1 PF.
[0180] like If the value is greater than N, then there will be more vehicles than needed. However, the Hungarian algorithm requires that the number of rows and columns be equal, so matrix A needs to be expanded: if You can add at the end of the column. A virtual slot, making its UCB i,p (t) = 0, ensuring the number of columns equals the number of rows; after confirming the matrix size, corresponding values need to be filled in the corresponding positions for easy filtering, therefore, in vehicle v i At the intersection of the corresponding row and the navigator position column, enter the UCB of the vehicle when it served as the historical navigator. i,PL (t)(i=1,2,3···); At the intersection of the following position columns, fill in the UCB when the vehicle served as the historical following vehicle. i,PF (t); If there is a virtual slot, fill in 0 in the corresponding position; this will give you a A square matrix, the values in which represent the maximum profit W. * ;
[0181] Assume the set of candidate vehicles in a task is The task requires 3 vehicles N, meaning 1 PL and 2 PF vehicles need to be selected. The column slots are represented as {PL, PF1, PF2}, therefore two virtual slots need to be added, resulting in a 5×5 matrix A, represented as:
[0182]
[0183] Take a from the matrix ij (representing the element in the i-th row and j-th column of the "vehicle-role" UCB value matrix A) is transformed into max(a ij )-a ij , where max(a ij) is the maximum value among all elements of the matrix, thus transforming the maximization problem into a minimization problem; for each row, find the minimum value of that row and subtract it from all elements of that row; for each column, find the minimum value of that column and subtract it from all elements of that column; this ensures that each row and each column has at least one 0 element;
[0184] In the processed matrix, try to mark some 0 elements so that each row and each column can be marked with at most one 0; try to cover all 0 elements with the minimum number of column lines; if the number of lines is equal to the matrix dimension, a feasible solution can be found; otherwise, it is necessary to continue to make further adjustments to the elements not covered by lines until all 0s can be covered with the same number of lines as the dimension.
[0185] A perfect match exists when all zeros can be covered by lines of the same order as the matrix. Based on the marked zero elements, a column index is assigned to each row sequentially, ultimately yielding the optimal match x. * ;
[0186] From the matching result x * It can be known that: vehicle v i Which position is it assigned to; if the column belongs to the navigator role, it indicates that vehicle v i Select the lead vehicle; if the column belongs to the follower role, i.e., vehicle v i Selected vehicle to follow; if matched with a "virtual slot", it means that the vehicle has not been selected into the formation and will continue to wait for the next round of selection.
[0187] The reward and punishment mechanism and payment rules provided in this embodiment of the invention are as follows:
[0188] (1) Reward and punishment mechanism;
[0189] (2) Payment rules.
[0190] The reward and punishment mechanism provided in this embodiment of the invention:
[0191] The behavior of a vehicle during mission M(n) is monitored through collective verification by other vehicles in the formation; when a violation is detected, the other vehicles in the formation verify the behavior through multi-party consensus.
[0192]
[0193] Among them, O i (j) represents the detection value of the violation by vehicle i against vehicle j (1 indicates a confirmed violation, 0 indicates no violation was observed), and n represents the number of vehicles participating in the consensus. This is the consensus threshold; when C v If ≥Γ, then vehicle v is confirmed. j There were violations;
[0194] Violations are categorized into three types based on severity: T = {V} L V M V H}; where V L V M V H These represent different degrees of violation, corresponding to different levels of harm to the formation; when a violation occurs and causes economic losses to the formation mission, the violating vehicle v j The vehicles in the formation and the person who initiated the mission should be compensated according to the revenue generated when the mission is completed normally.
[0195]
[0196] Among them, P i The outstanding amount for vehicles involved in violations. The amount paid for marginal contribution loss, u i The amount of loss incurred by the task initiator; thereby further increasing the punishment for vehicles that violate regulations. When the cost of violation is greater than the benefits of violation, it can effectively curb the occurrence of violations.
[0197] Once the violation is confirmed through the above checks, the system records the violation information via the blockchain; however, to protect user privacy, the violation record uses blind signature and zero-knowledge proof technology, and the recording format is as follows:
[0198] R j (t)=(H(ID v ),V t ,T s ,L,Sig)
[0199] Among them, H(ID) v V is the hash value that uniquely identifies the vehicle. t For the type of violation, T s For timestamp, L is the geolocation hash, and Sig is the multi-signature verification information; by recording the result of the action but not the specific details, it prevents issues related to user privacy and security.
[0200] By leveraging the spread of violations to help more vehicles in a convoy receive early warnings, the propagation of blockchain records follows the mathematical model below:
[0201]
[0202] Among them, R j (t) represents vehicle v j The set of violation records possessed at time t, N i (t) represents vehicle v i The set of other vehicles encountered at time t; vehicle v at time t.i The existing early warning information is R i (t); vehicle v i By using set N i The vehicles in (t) communicate with each other to obtain their early warning information R. j (t); vehicle v i At the next time point t+1, it will interact with the records of its own vehicles and those of neighboring vehicles; in this way, violation records can be naturally spread among different vehicles, forming an "immune system"-like early warning mechanism.
[0203] Based on the violation history recorded by the blockchain, the system dynamically adjusts the vehicle's future platooning permissions; defining vehicle v j The historical impact factors of violations are:
[0204]
[0205] Where, n L n M n H w represents the number of violations at different levels. L w M w H Let w be the weighting coefficient for each level of violation, and w L <w M <w H ; Let t be the time decay function. r This is the timestamp of the violation record, where t is the current time; the impact of the violation record will gradually weaken over time.
[0206] During the vehicle matching MAB phase, the UCB value of non-compliant vehicles is adjusted:
[0207] UCB′ i,p (t)=UCB i,p (t)·P(H v )
[0208] Among them, UCB′ i,p (t) represents the adjusted UCB value. The UCB value is a key factor affecting whether a vehicle can be selected into the formation. Therefore, a UCB value that is smaller may be directly rejected from joining the formation.
[0209] P(H v ) is the penalty function for violations:
[0210]
[0211] When H v If a certain threshold is exceeded, the vehicle will be unable to participate in certain types of formation missions for a certain period of time.
[0212]
[0213] Where σ = 3 is the critical value for serious violations; T(v) is the prohibition periodic function, which is proportional to the violation history factor; within the T(v) period, vehicles are prohibited from gaining benefits by joining vehicle platoons. This restricts violating vehicles while ensuring that vehicles have the opportunity to regain platoon participation rights after a period of time.
[0214] In addition, the system allows vehicles to restore trust through behavioral proofs:
[0215]
[0216] When a vehicle successfully completes a platooning maneuver without violations m times, its historical violation impact factor K will decrease proportionally.
[0217] Payment rules provided in this embodiment of the invention:
[0218] Based on the actual service of all vehicles An optimal vehicle formation N can be calculated. * The total revenue can be expressed as:
[0219]
[0220] If vehicle v is removed from any vehicle platoon i The optimal profit that the system can achieve on the remaining vehicles is Then, the vehicle v is defined according to its marginal contribution. i The payment is defined as:
[0221]
[0222] In order to consider the contribution of vehicles to the overall formation system, it is necessary to calculate the true economic benefit of each vehicle, v. i The final net income is:
[0223]
[0224] Where π i The amount a vehicle needs to pay based on its marginal contribution can actually reduce its net revenue if it misreports, thus encouraging honest reporting. The rational choice for a vehicle is to report truthfully.
[0225] Marginal contribution payments ensure that honest reporting remains the dominant strategy in cases of vehicle violations; the distributed trust records on the blockchain constrain misconduct without the need for third-party escrow, reducing system complexity and trust requirements.
[0226] like Figure 2 As shown, an embodiment of the present invention provides a vehicle platooning system based on a multi-armed slot machine, comprising:
[0227] Vehicle utility module, used for vehicle utility model;
[0228] Dynamic filtering module, used for dynamic filtering of MABs;
[0229] The reward and punishment module is used for reward and punishment mechanisms and payment rules.
[0230] Another object of the present invention is to provide a computer device including a memory and a processor, the memory storing a computer program that, when executed by the processor, causes the processor to perform the steps of the vehicle platooning method based on a multi-armed slot machine.
[0231] Another object of the present invention is to provide a computer-readable storage medium storing a computer program that, when executed by a processor, causes the processor to perform the steps of the vehicle platooning method based on a multi-armed slot machine.
[0232] Another objective of the present invention is to provide an information data processing terminal for implementing the vehicle platooning system based on the multi-armed slot machine.
[0233] Specific implementation of the present invention:
[0234] This invention proposes a vehicle platooning resource allocation method that integrates a Multi-Armed Bomber (MAB) learning framework and incentive mechanism. A two-layer MAB structure handles road selection and vehicle role matching, while the UCB strategy and compensation mechanism enhance the system's fairness and incentive. Furthermore, the introduction of blockchain technology effectively ensures data transparency and security, providing scalable and reliable technical support for the long-term development of vehicle platooning.
[0235] 2. Preliminary Knowledge
[0236] 2.1 System Model
[0237] This invention aims to solve the problem of multi-road vehicle platooning allocation in intelligent transportation scenarios, based on a vehicle-road cooperative model, such as... Figure 3 As shown in the figure, this model mainly involves the following two types of entities: vehicles and roadside units (RSUs).
[0238] (1) Data preparation stage
[0239] Before the formation mission M(n) is initiated, vehicles v1, v2, ..., v3 that are not currently in formation but have the potential to join formation. i The vehicle is referred to as the candidate vehicle, vehicle vi Utilizing idle calculation period T (0) The computational resources within the framework are used for an "exploration" strategy within the Multi-armed Bandit (MAB) framework: exploring roads r1, r2, ..., r under different environments. n Vehicle v i An "exploration" strategy is used for evaluation, and a locally established encrypted ledger B(t) is used for data recording and updating. Vehicles to be platooned v i Accumulate experiential data in advance within a distributed environment. Meanwhile, the combination of decentralized ledgers and blockchain sharding provides a reliable basis for subsequent information uploads and privacy protection.
[0240] (2) Formation stage
[0241] When the formation mission M(1) is initiated, vehicle v i The anonymized historical evaluation data b1, b2, ..., b n The data is reported to the nearest Road Test Unit (RSU). The RSU then uses the vehicle-uploaded data to enter the Road Filtering MAB layer, which selects a set of candidate vehicles from a global perspective. The size of the candidate pool was determined by referring to the "multiple-choice secretary problem". Subsequently candidate set Upon entering the vehicle matching MAB layer, a "vehicle-role" UCB matrix A is constructed, and the Hungarian algorithm is used to solve the optimal matching problem between the lead vehicle and the following vehicles. The selected vehicle immediately receives the RSU allocation result x. * In the V2V network, identity verification, communication channel establishment and formation positioning are completed, thereby forming a high-yield and energy-efficient formation structure.
[0242] (3) Payment stage
[0243] The lead vehicle (PL, Platoon Leader) is responsible for navigation and decision-making, while the follower vehicles (PF, Platoon Follower) maintain formation according to a collaborative control strategy. A blockchain-driven multi-party consensus mechanism continuously collects driving data and violation event hashes (VR). If an anomaly is detected, the information spreads rapidly through the V2V network, creating an "immune" warning. At the end of the mission, the system calculates the marginal contribution and compensation for each vehicle based on the social welfare increment formula, especially to compensate the lead vehicle for additional energy consumption. The payment results, along with new environment-benefit observations, are written back to their respective ledgers. RSUs synchronize the records to on-chain shards, enabling data traceability and continuous model iteration.
[0244] 2.2 Social Welfare
[0245] The MAB algorithm maximizes cumulative reward across K arms with an unknown reward distribution by balancing "exploration-exploitation" within a finite number of rounds. In vehicle platooning, each arm is mapped to a different mission road and platoon position. Vehicles use historical exploration data to determine the probabilistic relationship between their own consumption and mission rewards.
[0246] Definition 1 (Decision MAB Model). Based on the above description, this study constructs the vehicle platooning decision MAB model as a quintuple. The specific meanings are as follows:
[0247] (1) It is the set of all candidate vehicles participating in formation task M(n), v n This represents the nth vehicle in the candidate set. The candidate vehicle set is selected by the Road Filtering Action (MAB) from the set of candidate vehicles. The initial screening results, and meet the requirements It is a constant; For the final selected vehicle set, existing research experiments show that the optimal number of vehicles in a single formation should be 3 ≤ |x * |≤6.
[0248] (2) It is a set of available mission paths. Represents the first in the set of task paths A road, where (v:r) is defined below to represent vehicle-road matching.
[0249] (3) U represents vehicle v i Expected in different locations The reporting efficiency of U i,p The set (t) contains vehicles v, where PL represents the lead vehicle and PF represents the following vehicle. i The expected value of expenditure and revenue for a single task M(n).
[0250] (4) P represents vehicle v i Complete road r n The current average reward R generated during formation missions r (t) is a historical set. In order to reflect the impact of historical factors on future returns, it is beneficial to calculate the historical return function μ(t) by P. The generalization of μ(t) can more comprehensively describe the actual return of the vehicle.
[0251] (5) This is a violation constraint factor. Among them, H(ID) v V is the hash value that uniquely identifies the vehicle. t For vehicle violation types, T sL is the timestamp, L is the geolocation hash value, and Sig is the multi-signature verification information. When vehicle v j When violations occur in task M(n), directly recording the information of the violating vehicles may pose privacy and security risks. Therefore, it is crucial to protect user information security using anonymization techniques. Thus, the following approach is introduced... Constraint factors and constraint information recording formats protect vehicle privacy and security.
[0252] Definition 2 (Marginal Contribution). Let vehicle v i Participation in formation strategy x * On the road r n Task M(n), the total reward of the formation is W when the formation task is completed. * Among them, vehicle v i The incremental utility contribution to the entire formation is described as follows: That is, vehicle v i The change in total utility brought about by task M(n) is called the marginal contribution. The marginal contribution is used to determine the payoff rule π for each participant. i This ensures the fair allocation of resources.
[0253] Definition 3 (Social Welfare). Given an optimal allocation strategy x... * Below, the total utility can be calculated based on the vehicle set. Different strategic arrangements within a vehicle platoon can yield varying sums of utility. Therefore, given a vehicle resource x... * In this case, the strategy satisfies the requirement of maximizing utility and The strategy is said to maximize social welfare.
[0254] Definition 4 (Individual Rationality). For the optimal allocation strategy x... * Vehicles v1, v2, ..., v n Based on its own "exploration" strategy, the path r under the current task M(n) is... n Different positions Perform a true valuation U i,p (t), the actual utility gained after the task is completed. Should satisfy For vehicles not participating in the strategy x * The utility of strategy x at that time is called strategy x. * To satisfy individual rationality.
[0255] Definition 5 (Excitation Compatibility). Assume vehicles v1, v2, ..., v n In the optimal allocation strategy x * The vehicle will be offered at its own price. i Choosing to honestly report expected utility U i,p(t), vehicle v j Selecting malicious price gouging to report expected utility In the final profit scenario, vehicle v i Actual utility satisfaction Vehicle v j Actual utility satisfaction Then we call strategy x * Incentive compatibility is satisfied.
[0256] 3. MAB-based vehicle platooning method
[0257] This section describes a vehicle platooning scheme based on a multi-armed slot machine algorithm, comprising four stages: road selection, vehicle matching, and penalty and payment. The UCB strategy and Hungarian algorithm are employed to optimize resource allocation, and a compensation mechanism is designed to balance differences in vehicle energy consumption. Blockchain technology is used to ensure data security, transparency, and trust.
[0258] 3.1 Vehicle Utility Model
[0259] Vehicles waiting to be platooned at any time The set of selectable mission paths is At time t, vehicle v i Take on a role The utility of time is:
[0260] U i,p (t)=v i,p (t)-c i,p (t)+r i,p (t)+s i,p (t)
[0261] U i,p (t) is v i The revenue and expenditure metrics during task M(n) execution assist in MAB parameter updates and decision-making processes.
[0262] v i,p (t) is the base valuation, which is the vehicle v i The expected value of the current task is related to contextual factors such as the current environment:
[0263]
[0264] Where, ξ p >0 represents the benefit coefficient corresponding to character p, d i,p (t) represents vehicle v i The distance to the mission destination, τ > 0, is a distance-sensitive parameter. When the mission distance d i,p (t) The further away the basic valuation v i,pThe lower the value of (t), the greater the distance, the fewer tasks the user can complete per unit time, and the greater the energy consumption and wear on the vehicle. This also leads to higher road uncertainty and greater risk. Therefore, the task distance d... i,p (t) and the underlying valuation v i,p (t) shows a negative correlation.
[0265] c i,p (t) represents energy consumption expenditure, which is the vehicle's energy consumption expenditure. i Energy consumption when playing role p in a formation mission, and vehicle speed v i (t) and the wind resistance experienced by different roles p can be expressed as:
[0266] c i,p (t)=α p ·f(v i (t))
[0267] Where, α p >0 represents a specific energy consumption coefficient for the character, f(v) i (t) is the energy consumption function. PL needs to withstand greater wind resistance, therefore α p Greater than PF, and with vehicle speed v i (t) increases vehicle energy consumption expenditure c i,p (t) will also increase accordingly. As the navigator vehicle is the control hub of the formation, it can control the overall speed of the formation while ensuring the quality of mission completion, and further achieve energy conservation.
[0268] r i,p (t) represents the compensation revenue, and to incentivize vehicles to take on high-energy-consuming positions, compensation revenue is provided to PL:
[0269]
[0270] χ p >0 represents the corresponding compensation coefficient. This is an indicator function that takes the value 1 only when p = PL and 0 when p = PF.
[0271] s i,p (t) represents the synergistic effect, reflecting the positive benefits that vehicles gain from cooperation with each other. For example, following vehicles may benefit from aerodynamic wakes or safety is improved when multiple vehicles cooperate. It can be expressed as:
[0272]
[0273] Where, N i For vehicle v i The set of neighboring vehicles, κ ij This represents the coordination strength coefficient. Numerous experiments have shown that the vehicle spacing remains at a fixed, relatively small distance d. ij(t) yields the best energy-saving effect, therefore this study uses h(d) ij The optimal cooperative distance is described by (t)). The closer the distance between vehicles is to the optimal cooperative distance, the stronger the cooperative effect; the farther the distance, the weaker the cooperative effect.
[0274] 3.2 MAB Dynamic Filtering
[0275] 3.2.1 Road Filtering MAB
[0276] This plan assumes that vehicles agree to make reasonable use of their idle time and calculates the period T. (0) Under the premise that the vehicle v i When operating under low load, the onboard computing unit can utilize idle computing power to execute the "exploration" strategy of the MAB algorithm, performing on-site assessments of the surrounding road environment and updating the data on the blockchain. This process does not affect the vehicle's normal driving functions and can quickly access existing experience data when a task is triggered, enabling rapid decision-making without significantly increasing vehicle hardware costs. The road selection MAB model designed in this section, such as... Figure 4 As shown.
[0277] For the task path set When driving alone or in convoy, the vehicle selects the appropriate road. n After passage is completed, the system calculates the current average actual benefit R based on the comprehensive utility feedback from the vehicle under different conditions. r (t), while updating the vehicle's local ledger, recording road conditions under different environments. n Historical mission overall benefits μ r (t), the revenue update expression:
[0278] μ i,r (t)=φμ i,r (t-1)+(1-φ)R r (t)
[0279] Where φ∈(0,1) is the time decay factor, and μ is the vehicle's comprehensive gain at time t. i,r (t) The update is not based solely on the current average real return R. r (t) is also related to the gains at time t-1, which means that it takes into account two important influencing factors: historical data and the current performance of the vehicle. This ensures that the data information updates are highly timely and also improves the availability of information when using the MAB strategy for road screening.
[0280] If a novice vehicle with no "exploration" experience applies to join a platoon, it can be initialized using vehicles with similar driving experience, reducing the cold start for novice vehicles in the first layer of screening. For vehicle v i For road rn The initial estimate can be matched with v from historical data. i Similar features inherit Right now:
[0281]
[0282] It can effectively increase the probability of novice vehicles being assigned to the best mission routes and enhance their attractiveness to novice vehicles, avoiding the impact of invalid assignments due to the lack of historical data on vehicles on the stability of the fleet.
[0283] To address the issue of vehicles being unable to fully explore all roads, a strategy based on the Upper Confidence Bound (UCB) is adopted for road exploration.
[0284]
[0285] Where, n i,r (t) represents vehicle v i Driving on the road n Historical record count; γ > 0, β > 0 are the exploration adjustment coefficients; UCB i,r (t) can reflect the vehicle's exploration level on the corresponding road; N r (t) represents road r n Historical selection count, via r * (t) can reflect the road r n Based on the historical best exploration value of UCB, through the vehicle's own UCB i,r (t) and the corresponding road requirements r * (t) Compare and filter to identify those that do not match the road r n Vehicles that meet the requirements will be included in the candidate vehicle set. among.
[0286] In a decision cycle After the internal screening is completed, when the candidate vehicle set is... The number of vehicles satisfies the constraints of the "multiple secretary selection problem" or is not less than the minimum task requirement |x * At that time, it will no longer wait for the next cycle. The process proceeds directly to the next level of decision-making, and the relatively optimal formation combination x at time t is determined. * It will definitely be in the candidate vehicle collection This process generates the following: In multi-road task scenarios, by combining the vehicle's historical "exploration" with the initial screening based on the current environment (v:r), the optimal multi-road matching problem at time t is solved.
[0287] 3.2.2 Vehicle matching MAB
[0288] After applying the MAB (Mobility Allocation) strategy for road selection in the previous layer, a sufficient set of candidate vehicles is obtained. The purpose of the second-level vehicle matching MAB strategy, which performs optimal (v:r) filtering, is to determine the candidate vehicle set. The optimal vehicle formation is matched to form the best formation for mission driving. The matching model in this section is as follows: Figure 5 As shown.
[0289] Use μ i,p (t) represents vehicle v i In the position of role The average utility estimate at time μ i,r (t) differs from μ i,p (t) More attention is paid to vehicles on the road. n Specific utility estimation when serving as PL or PF, μ i,r (t) More attention is paid to vehicles on the road. n The average utility estimate on, therefore μ i,p (t) can be greater than μ i,r (t) A more detailed description of the vehicle's position on the road. n Utility estimation when playing role p. Using U i,p (t) represents vehicle v i If the actual utility observed at time t while playing role p is given, then its historical average utility update can be expressed as:
[0290] μ i,p (t)=φμ i,p (t-1)+(1-φ)U i,p (t)
[0291] To measure uncertainty, the UCB value of the vehicle matching MAB layer is defined as follows:
[0292]
[0293] Where, n i,p (t) represents vehicle v i The number of times a vehicle has been selected in its role (p) in historical records. The number of times a vehicle has played different roles in history reflects its experience to some extent; vehicles with more experience are better able to handle various unexpected situations that may arise during formation.
[0294] After obtaining the candidate vehicle set After that, it is necessary to One PL and one |x were selected. * -1 PF, and maximize the profit W *At this point, the problem is transformed into an "Assignment Problem," and the Hungarian algorithm is used to find the optimal matching strategy x. * For ease of description, assume that in the current formation task, a total of N vehicles need to be selected (i.e., 1 PL and N-1 PF).
[0295] Before proceeding with the formal screening, a "vehicle-role" UCB value matrix A needs to be constructed. The size of the matrix depends on the set of candidate vehicles. Given the number of vehicles N required for the task, define a matrix row to store the corresponding... All vehicles, totaling The rows and columns correspond to the N role positions to be assigned in the formation. These N positions include: 1 PL and N-1 PF.
[0296] like If the value is greater than N, then there will be more vehicles than needed. However, the Hungarian algorithm requires that the number of rows and columns be equal, so matrix A needs to be expanded: if You can add at the end of the column. A virtual slot, making its UCB i,p (t) = 0, ensuring the number of columns equals the number of rows. After confirming the matrix size, corresponding values need to be filled in the corresponding positions for easy filtering. Therefore, in vehicle v... i At the intersection of the corresponding row and the navigator position column, enter the UCB of the vehicle when it served as the historical navigator. i,PL (t)(i=1,2,3···); At the intersection of the following position columns, fill in the UCB when the vehicle served as the historical following vehicle. i,PF (t). If there is a virtual slot, fill in 0 in the corresponding position. This will give you a A square matrix, the values in which represent the maximum profit W. * .
[0297] Assume the set of candidate vehicles in a task is The task requires 3 vehicles N, meaning 1 PL and 2 PF vehicles need to be selected. The column slots are represented as {PL, PF1, PF2}, therefore two virtual slots need to be added, resulting in a 5×5 matrix A, represented as:
[0298]
[0299] Take a from the matrix ij (representing the element in the i-th row and j-th column of the "vehicle-role" UCB value matrix A) is transformed into max(a ij )-a ij , where max(a ijThe maximum value in the entire matrix is 0, thus transforming the maximization problem into a minimization problem. For each row, find the minimum value of that row and subtract it from all elements in that row; for each column, find the minimum value of that column and subtract it from all elements in that column. This ensures that each row and each column has at least one 0 element.
[0300] In the processed matrix, attempt to mark some 0 elements so that each row and each column can be marked with at most one 0; try to cover all 0 elements with the minimum number of column lines. If the number of lines equals the matrix dimension, a feasible solution can be found; otherwise, further adjustments need to be made to the elements not covered by lines until all 0s can be covered with a number of lines equal to the dimension.
[0301] A perfect match exists when all zeros can be covered by lines of the same order as the matrix. Based on the marked zero elements, a column index is assigned to each row sequentially, ultimately yielding the optimal match x. * .
[0302] From the matching result x * It can be known that: vehicle v i Which position is it assigned to? If this column indicates a navigator role, it means vehicle v i Select the lead vehicle; if the column belongs to the follower role, i.e., vehicle v i Selected vehicle to follow; if matched with a "virtual slot", it means that the vehicle has not been selected into the formation and will continue to wait for the next round of selection.
[0303] 3.3 Reward and Punishment Mechanism and Payment Rules
[0304] After MAB dynamic filtering is completed, the optimal vehicle platooning matching result is obtained. * However, in actual implementation, the issue of potential vehicle violations still needs to be addressed. To ensure vehicles participate honestly and comply with platooning rules, this section describes a blockchain-based reward and punishment mechanism. This mechanism is closely integrated with the MAB allocation scheme, influencing the UCB value of vehicles in subsequent MAB decisions by monitoring vehicle behavior in real time and recording violation information, thus forming a closed-loop feedback system.
[0305] 3.3.1 Reward and Punishment Mechanism
[0306] The behavior of a vehicle during mission M(n) is monitored through collective verification by other vehicles in the formation. When a violation is detected, the other vehicles in the formation verify the behavior through multi-party consensus:
[0307]
[0308] Among them, O i(j) represents the detection value of the violation by vehicle i against vehicle j (1 indicates a confirmed violation, 0 indicates no violation was observed), and n represents the number of vehicles participating in the consensus. This is the consensus threshold. When C v If ≥Γ, then vehicle v is confirmed. j There were violations.
[0309] Violations are categorized into three types based on severity: T = {V} L V M V H}. Among them, V L V M V H These represent different degrees of violation, corresponding to different levels of harm to the formation. When a violation occurs and causes economic losses to the formation mission, the violating vehicle v j The fleet vehicles and the mission initiator should be compensated according to the revenue generated when the mission is completed normally.
[0310]
[0311] Among them, P i The outstanding amount for vehicles involved in violations. The amount paid for marginal contribution loss, u i This represents the amount of loss incurred by the task initiator. This further increases the penalties for violating regulations; when the cost of violation outweighs the benefits, it can effectively deter such behavior.
[0312] Once the violation is confirmed through the above checks, the system records the violation information via the blockchain. However, to protect user privacy, the violation record uses blind signature and zero-knowledge proof technology, and the recording format is as follows:
[0313] R j (t)=(H(ID v ),V t ,T s ,L,Sig)
[0314] Among them, H(ID) v V is the hash value that uniquely identifies the vehicle. t For the type of violation, T s `L` is the timestamp, `L` is the geolocation hash, and `Sig` is the multi-signature verification information. By recording the behavioral outcome without specific details, user privacy and security issues are prevented.
[0315] To prevent vehicles that have already caused harm from joining other platoons and causing even greater economic losses, this solution adopts a mechanism of spreading violations, helping more platoons to receive early warnings. The propagation of these violations through blockchain records follows the mathematical model below:
[0316]
[0317] Among them, R j (t) represents vehicle v j The set of violation records possessed at time t, N i (t) represents vehicle v i The set of other vehicles encountered at time t. At time t, vehicle v i The existing early warning information is R i (t). Vehicle v i By using set N i The vehicles in (t) communicate with each other to obtain their early warning information R. j (t). Vehicle v i At the next time point t+1, it will exchange its existing records with those of neighboring vehicles. In this way, violation records can be naturally spread among different vehicles, forming an "immune system"-like early warning mechanism.
[0318] Based on the violation history recorded on the blockchain, the system dynamically adjusts the vehicle's future platooning permissions. (Vehicle v is defined.) j The historical impact factors of violations are:
[0319]
[0320] Where, n L n M n H w represents the number of violations at different levels. L w M w H Let w be the weighting coefficient for each level of violation, and w L <w M <w H . Let t be the time decay function. r This is the timestamp of the violation record, where t is the current time. The impact of the violation record will gradually decrease over time.
[0321] During the vehicle matching MAB phase, the UCB value of non-compliant vehicles is adjusted:
[0322] UCB′ i,p (t)=UCB i,p (t)·P(H v )
[0323] Among them, UCB′ i,p (t) represents the adjusted UCB value. The UCB value is a key factor affecting whether a vehicle can be selected into the formation. Therefore, a UCB value that is smaller may result in the vehicle being rejected from joining the formation.
[0324] P(H v ) is the penalty function for violations:
[0325]
[0326] When H v If a certain threshold is exceeded, the vehicle will be unable to participate in certain types of formation missions for a certain period of time.
[0327]
[0328] Where σ = 3 is the critical value for serious violations; T(v) is the prohibition periodic function, which is proportional to the violation history factor. During the T(v) period, vehicles are prohibited from gaining benefits by joining vehicle platoons. This restricts violating vehicles while ensuring that they have the opportunity to regain platoon participation rights after a period of time.
[0329] In addition, the system allows vehicles to restore trust through behavioral proofs:
[0330]
[0331] When a vehicle successfully completes a platooning maneuver without violations m times, its historical violation impact factor K will decrease proportionally.
[0332] 3.3.2 Payment Rules
[0333] Based on the actual service of all vehicles An optimal vehicle formation N can be calculated. * The total revenue can be expressed as:
[0334]
[0335] If vehicle v is removed from any vehicle platoon i The optimal profit that the system can achieve on the remaining vehicles is Then, the vehicle v is defined according to its marginal contribution. i The payment is defined as:
[0336]
[0337] In order to consider the contribution of vehicles to the overall formation system, it is necessary to calculate the true economic benefit of each vehicle, v. i The final net income is:
[0338]
[0339] Where π iThe amount a vehicle needs to pay based on its marginal contribution can actually reduce its net revenue if it misreports, thus encouraging honest reporting. The rational choice for a vehicle is to report truthfully.
[0340] Marginal contribution payments ensure that honest reporting remains the dominant strategy in cases of vehicle violations. By leveraging distributed trust records on the blockchain, misconduct is constrained without the need for third-party escrow, reducing system complexity and trust requirements.
[0341] 4. Theoretical Analysis
[0342] The design of this scheme needs to satisfy incentive compatibility with users while also meeting the latency requirements of vehicle platooning in vehicle-road cooperative scenarios. Therefore, this section will verify the incentive compatibility with users, and then use time complexity analysis to verify that the scheme meets the latency requirements of platooning.
[0343] 4.1 Incentive Compatibility
[0344] Lemma 4-1. This scheme satisfies individual rationality.
[0345] Proof: Individual rationality requires vehicle v i Expected choice to improve one's utility U i,p The strategy of maximizing (t) is to verify that vehicle v is optimal in this scheme. i If the user is honest, their utility will not be lower than that in any other situation, meaning that participation in the mechanism is voluntary and beneficial for each vehicle.
[0346] The efficiency of a convoy is expressed as follows:
[0347] U i,p (t)=v i,p (t)-c i,p (t)+r i,p (t)+s i,p (t)
[0348] Assume vehicle v i The utility when driving alone is Since vehicles do not offer compensatory benefits or synergistic effects, their utility can be expressed as:
[0349]
[0350] in, and These are the basic estimate and energy consumption expenditure when driving alone, and the utility U. i,p (t) and To meet the physical requirements, that is Compensation income r i,p (t) and synergistic effect si,p (t) While ensuring a reasonable compensation mechanism for the platooning mechanism, more vehicles are encouraged to participate in the platooning.
[0351] For PL, its energy consumption expenditure Because it needs to withstand greater wind resistance and additional communication overhead, a compensation function is set. To compensate for additional costs, synergy i,p (t) While improving the driving experience of the fleet on the road, it indirectly saves energy and achieves the effect of improving efficiency; PF, due to the aerodynamic wake effect, makes And there is inter-vehicle cooperative effect s i,p (t)>0, even if the compensation benefit r i,p When (t) = 0, utility can be maximized.
[0352] Considering the specific scenario, let's assume... but:
[0353]
[0354] For PL Because r exists i,p (t) compensates for the lead vehicle, and s i,p If (t)≥0, then
[0355] For PF, s i,p (t)>0, even if r i,p Since (t) = 0, we have:
[0356]
[0357] Therefore
[0358] Therefore, this solution satisfies That is, the utility of all vehicles when participating in platooning is no less than when not participating, satisfying the individual rational requirements for vehicles in different positions.
[0359] Lemma 4-2. This scheme satisfies excitation compatibility for PL.
[0360] Proof: Incentive compatibility requires that PL maximizes its net utility when honestly reporting its true utility. Let the lead vehicle v be... 1,pl The real utility is But it may report a false utility. Therefore, it needs to be demonstrated that any deviation from the truthful reporting does not increase its net utility.
[0361] Vehicle v i Payment P iDefined as its marginal contribution, that is, its impact on the sum of the utilities of other vehicles in the system, vehicle v i The net utility is:
[0362]
[0363] Because of P i It is based on the utility calculations of other vehicles and does not directly depend on... Leading vehicles cannot pass false reporting Directly affects P i However, if false reporting results in it not being selected as a PL, then its utility becomes... As known from Lemma 4-1 Therefore, leaving the formation will not increase its effectiveness.
[0364] Assume the lead vehicle overstates its utility to obtain a higher P. i The system chooses the distribution that maximizes social welfare, that is:
[0365]
[0366] like The system may consider vehicle v i The vehicle's contribution to the formation was low, leading to the selection of other vehicles, resulting in v i If not selected, the utility is reduced to Conversely, if As long as it remains selected, P i The net utility remains unchanged, but if the candidate is not selected as a result, the utility also decreases. Therefore, honest reporting is essential. It is a dominant strategy that ensures that it is correctly selected and obtains the maximum net utility.
[0367] Therefore, this solution satisfies the excitation compatibility requirement for PL.
[0368] Lemma 4-3. This scheme satisfies excitation compatibility for PPF.
[0369] Proof: The incentive compatibility of PF also requires that it maximizes net utility when honestly reporting true utility. The utility of PF is... Typically P i =0 , s i,p (t)>0 and If the payment is non-zero, then P i The net utility is still determined by the marginal contribution:
[0370]
[0371] Similar to PL, P i Not dependent If PF falsely reports This may result in the PF not being selected, reducing its utility to [missing value]. and because If false reporting As long as it remains selected, P i If the net utility remains unchanged, it remains unchanged; if it is not selected, the utility decreases.
[0372] In MAB matching, the role assignment of PF is based on historical utility estimation. And the UCB value. If misreporting affects the estimate, it may lead to suboptimal allocation. The payment rules guarantee that social welfare is maximized under honest reporting; deviation will only reduce the probability of selection or net utility. Therefore, honest reporting is the optimal strategy for PF.
[0373] Therefore, this scheme satisfies the excitation compatibility for PF.
[0374] Theorem 4-1. This scheme is excitation compatible.
[0375] Proof: Lemma 4-1 shows that this scheme satisfies individual rationality, and the utility of all vehicles participating in the formation is no less than the utility of not participating. Lemma 4-2 proves that PL maximizes net utility when honestly reporting true utility, and deviation leads to a decrease in utility. Lemma 4-3 proves that PF maximizes net utility when honestly reporting true utility, and false reporting is of no benefit. Incentive compatibility requires that the optimal strategy for all participants in the system is to honestly report their private information. This scheme, through payoff rules and reasonable utility design, ensures that neither PL nor PF can benefit from false reporting, and that participation itself is superior to non-participation. Combining Lemmas 4-1, 4-2, and 4-3, the entire formation scheme satisfies incentive compatibility among all vehicles.
[0376] Therefore, this scheme is incentive compatible.
[0377] 4.2 Time Complexity
[0378] The Road Selection Tool (MAB) uses vehicle history to "explore" multiple roads and vehicles, selects the optimal road-vehicle combination, and uses this to determine the candidate vehicle set. The solution uses the UCB strategy for decision-making, which requires calculating the upper confidence bound for each road. The calculation of the upper confidence bound is based on the historical average return and the number of explorations.
[0379] Let the size of the task road set be... Calculating the UCB value for each road requires basing it on historical average returns. and the number of times the road is selected, N r (t), the time complexity of calculating a single UCB value is O(1). Each decision requires calculating all The UCB value is calculated for each road, and then the maximum value is selected. Therefore, the time complexity of a single decision is O(n log n). The Road Selection MAB utilizes vehicle idle computing cycles for exploration and leverages existing data at the start of the task. Therefore, it can be assumed that the UCB value for each road is pre-calculated using historical data at task execution, and the decision-making process requires only one selection step, resulting in a complexity of O(n log n). Furthermore, by incorporating the concept of the "multiple-choice secretary problem," the set of candidate vehicles to proceed to the next stage was determined. The quantity, since this part mainly involves threshold comparison and the maximum number of candidate vehicles to be screened is 15, has a complexity of O(1). Therefore, the time complexity of MAB screening on roads is O(1). It refers to the number of roads.
[0380] The goal of vehicle matching MAB is to select from the candidate vehicle set Select the best formation match x * Determine the role assignments for PL and PF. Assume the size of the candidate vehicle set is [size missing]. The task requirement is to select N cars, where The Hungarian algorithm is used to solve the (v:r) assignment problem, constructing a... The matrix. The matrix elements are the historical average utility estimates (UCB) of the vehicles in their respective roles. i,p (t), the filling complexity is O(t). The time complexity of the algorithm is
[0381] The calculation of UCB values is also used at this stage to evaluate the utility of vehicles in different roles. The UCB value is calculated for each role corresponding to each vehicle, with a complexity of O(1) per calculation. Assuming calculations are needed for |V| vehicles and 2 roles, the total complexity is O(|V|). However, these UCB values are usually updated in historical data and only need to be queried during the task, so they can be considered as preprocessing costs and are not included in the main complexity. For a task, the Hungarian algorithm only needs to run once, therefore the complexity of matching vehicles to the MAB is O(|V|). 3 Therefore, the time complexity of matching the vehicle to the MAB is O(|V|). 3 ).
[0382] The penalty warning data for each vehicle participating in the platoon is updated on the blockchain. Each update takes O(1) time, and the total complexity for N vehicles is O(N).
[0383] Calculating the marginal contribution of each vehicle requires calculating the total revenue W. * And the optimal benefit after removing the car W * The sum of the utilities of K vehicles has a complexity of O(N). The time complexity of calculating the sum of the utilities of the remaining N-1 vehicles is O(N). Calculating the marginal contribution for each vehicle has a total complexity of O(N×N) = O(N). 2 ).
[0384] The total time complexity is in Let V be the number of roads involved in the task, V be the number of candidate vehicles, and N be the number of vehicles required for the task. Since V and N are constant terms in the second-stage MAB, the total time complexity can be simplified to... In summary, the optimal time complexity is... The worst-case time complexity is
[0385] 4.3 Security Analysis
[0386] Lemma 4-4. This solution can effectively defend against vehicle utility misrepresentation.
[0387] Proof: Let vehicle v i The real utility is However, for their own benefit, they falsified the report. According to the marginal contribution payment rule, a vehicle's payment depends on its actual contribution to the system, rather than its reported utility. When vehicle v i Exaggerated efficacy When the allocation changes, according to Lemmas 4-2 and 4-3, two situations may occur: one is that the person who should have been selected is not selected due to false reporting, and their utility decreases to... Secondly, obtaining an unsuitable role through false reporting reduces its actual effectiveness. This is lower than the case of honest reporting. As Lemma 4-1 shows, false reporting will inevitably lead to a suboptimal match, ultimately reducing the vehicle's net utility.
[0388] In addition, this scheme updates historical return records μ i,r (t) allows the discrepancy between the reported utility and actual performance to gradually become apparent over time. Combined with the UCB exploration strategy, even undervalued vehicles have the opportunity to be re-evaluated.
[0389] Therefore, this solution can effectively prevent vehicles from falsely reporting their utility.
[0390] Lemma 4-5. This solution can effectively defend against strategic breaches of contract by vehicles.
[0391] Proof: Strategic breach refers to the act of a vehicle intentionally violating the agreement after joining a platoon. This scheme uses a multi-party consensus detection method to detect C. v To ensure the objectivity of violation determination, once a vehicle violation is confirmed, it will be determined according to the violation category {V}. L V M VH Implement tiered penalties (P) i To ensure that the cost of violation outweighs the benefits, thus deterring violations, violation records are maintained using blind signatures and zero-knowledge proofs (R). j (t)=(H(ID v ),V t ,T s (,L,Sig) to protect user privacy. From the diffusion formula R i (t+1), violation records are rapidly propagated through the blockchain network, forming an "immune" early warning system. Based on the violation history, the system dynamically adjusts the vehicle's future platooning permissions using UCB′. i,p (t) Adjustment.
[0392] Therefore, this solution can effectively defend against strategic breaches of contract by vehicles.
[0393] Lemma 4-6. This solution can effectively defend against collusive attacks by vehicles.
[0394] Proof: Collusion attacks refer to multiple vehicles conspiring to cheat in order to manipulate system decisions. This solution utilizes blockchain technology to build a multi-layered defense mechanism. This is achieved through anonymous hash records of R... j (t)=(H(ID v ),V t ,T s (L, Sig) ensures that violation records cannot be tampered with; even a colluding group cannot modify historical data already on the blockchain. H(ID) v Γ is a cryptographic hash function, possessing tamper-resistant and collision-resistant properties. Secondly, the multi-signature verification mechanism requires multiple parties to jointly confirm violations; given a consensus threshold Γ, a violation is only possible if and only if C is exceeded. v Violations are only recorded when at least one vehicle is confirmed to have committed a violation. This significantly increases the difficulty of forging violation records, as attackers would need to control a large number of vehicles. Roadside units, as relatively neutral infrastructure, can identify anomalous utility reporting patterns from a global perspective. Given a set of vehicles... The utility report allows the roadside unit to detect statistical anomalies.
[0395] This scheme separates time and space. The vehicle's historical records include multiple time points and road segments, making it difficult for conspirators to maintain consistent false information in all scenarios. Historical records accumulate over time, and the inconsistency of information is amplified as the number of interactions increases.
[0396] Therefore, this solution can effectively defend against collusive attacks by vehicles.
[0397] 5 Experiments
[0398] 5.1 Experimental Setup
[0399] The experimental hardware environment consisted of an AMD Ryzen 9 7940H processor, 16.0GB DDR5 4800MT / s memory, an AMD Radeon 780M graphics card, and an NVIDIA GeForce RTX 4060 graphics card. The experimental software environment consisted of a Windows 11 64-bit operating system and a Python 3.11 programming language. The simulation parameters are shown in Table 1.
[0400] Table 1 Simulation Parameters
[0401]
[0402] This experiment comprehensively evaluates the performance of a vehicle platooning resource allocation scheme that integrates a Multi-Armed Slots (MAB) framework. The comparative test design focuses on the following metrics: social welfare, role utility, system latency, and system robustness against malicious users.
[0403] 5.2 Experimental Results and Analysis
[0404] (1) Social Welfare Analysis
[0405] Figure 6 The study demonstrates a comparison of the social welfare of this approach with that of autonomous vehicles, assuming the optimal number of vehicles in a platoon.
[0406] The results show that the social welfare of this scheme is higher than that of autonomous driving under different vehicle number configurations. When the number of vehicles increases from 3 to 6, the average social welfare of the platooning scheme shows an increasing trend, while the social welfare of the autonomous driving scheme remains relatively stable within a certain range. With 6 vehicles, the social welfare of this scheme is approximately 2.5 times that of the autonomous driving scheme, because the aerodynamic wake effect in this scheme reduces PF energy consumption (c). i,p (t) and synergistic effect s i,p (t) This results in overall energy savings. Theoretically, the optimal platoon size is between 3 and 6 vehicles. Through social welfare comparison, this scheme is superior to autonomous driving.
[0407] (2) Role utility analysis
[0408] Figure 7 The average utility of PL, PF and autonomous vehicles was compared.
[0409] The results showed that the average utility of the PL was the highest at 0.238, and the PF was 0.155, while the utility of the autonomous vehicle was the lowest at 0.121. Among these, the PL bore a greater share of wind resistance and communication overhead, resulting in higher energy consumption (c). i,p (t) is relatively high, but through the compensation mechanism Significant compensation was achieved, resulting in optimal overall utility; PF benefits from the aerodynamic wake effect, reducing energy expenditure c. i,p(t) decreased significantly, while enjoying the synergistic effect s i,p The additional benefits brought by (t) are therefore higher than those of autonomous driving; autonomous vehicles do not enjoy any benefits from platooning, neither energy consumption reduction nor synergy and compensation benefits, therefore their utility is the lowest. Energy consumption model c i,p (t) reflects the energy consumption coefficient α under different roles. p The differences.
[0410] The results demonstrate that this scheme successfully balances the overall utility of vehicles in each role, ensuring that all participants benefit while maintaining the incentive for PL participation.
[0411] (3) Comprehensive utility analysis
[0412] Figure 8 The differences in the average utility composition of vehicles in different positions within a formation were analyzed in detail.
[0413] The basic valuations of vehicles at different locations are relatively balanced, reflecting... Characteristics in. Distance factor d i,p (t) caused a slight difference, but the role coefficient ξ p The impact is relatively small; energy consumption expenditure c i,p (t) shows a clear trend, with PL consuming the most energy and PF decreasing significantly. This reflects c i,p (t) Characteristics of the model, where c PL i,p (t)>c PF i,p (t) reflects the fact that the lead vehicle bears greater wind resistance. The energy consumption curve clearly shows the physical characteristic that the aerodynamic wake effect weakens with increasing distance; PL has a significant compensation benefit, reflecting the compensation benefit r. i,p (t)=χ p The definition of I{p=PL}. The compensation benefit value is positively correlated with the number of following vehicles and utility, and the value is approximately 60-70% of the energy consumption cost of the lead vehicle, effectively balancing the additional burden on PL; the synergy effect at position 2 is the strongest, and then slightly weakens, reflecting the synergy effect. Module. Location 2: The closest optimal collaborative distance between the vehicle and PL is h(d). ij (t)), thus obtaining the maximum benefit; while PL has no car in front, and the synergistic effect is relatively weak.
[0414] In summary, the total utility of PL mainly comes from the basic valuation and compensation benefits, which can offset the high energy consumption; while the utility of PF mainly comes from low energy consumption and high synergy. This utility allocation mechanism ensures that vehicles in different positions can obtain reasonable benefits, which is in line with the compensation mechanism proposed in this scheme that takes into account both fairness and incentives. It can effectively balance the differences in energy consumption burden of vehicles in different formation positions and ensure formation stability and efficiency.
[0415] (4) Group delay experiment
[0416] Figure 9 The detailed decomposition of decision allocation delay under optimal formation conditions is presented. Experimental results show that the total system delay increases non-linearly with the number of vehicles, but the growth rate varies significantly across different stages.
[0417] The latency during the road selection phase remains relatively stable, fluctuating within the range of 0.101-0.106 ms, indicating that this phase has low complexity and is significantly unaffected by the number of vehicles. The latency during the location allocation phase increases significantly with the number of vehicles, rising from 0.207 ms with 3 vehicles to 0.301 ms with 5 vehicles, consistent with its theoretical complexity O(V0). 3 The characteristics are defined by V, where V represents the number of candidate vehicles. The payment calculation phase saw the most significant increase in latency, surging from 0.201ms with 3 vehicles to 2.024ms with 6 vehicles, an increase of approximately 10 times, validating its... The expected complexity Represents the number of roads.
[0418] Even with a 6-vehicle configuration, the total system latency remains below 2.5ms, far below the maximum tolerable latency threshold of 300-500ms for vehicle-to-everything (V2X) communication systems. This result validates the time complexity analysis proposed in the paper: with a small number of vehicles, the overall decision-making latency is proportional to the cube of the number of vehicles, but the absolute value is small, fully meeting real-time requirements. Furthermore, the strategy of utilizing vehicle idle time computing resources to rationally allocate the exploration phase to historical trajectories significantly reduces real-time computing overhead.
[0419] (5) Large-scale model delay experiment
[0420] Figure 10 The estimated allocation latency of a single RSU is shown in different scale scenarios, including: 100 roads / 600 vehicles, 500 roads / 3000 vehicles, and 1000 roads / 6000 vehicles.
[0421] Experimental results show that when the system is scaled up to a large-scale scenario, the proportions of each component of latency change significantly: location allocation time dominates, reaching 216.502 seconds in a configuration of 1000 roads / 6000 vehicles, far exceeding other stages. Payment calculation time in this configuration is only 0.579 seconds, approximately 0.27% of the location allocation time, a stark contrast to small-scale scenarios. The latency of the road selection stage remains relatively stable with scale, increasing from 0.331 seconds in 100 roads / 600 vehicles to 3.301 seconds in 1000 roads / 6000 vehicles. Location allocation time surges from 0.268 seconds in 100 roads / 600 vehicles to 216.502 seconds in 1000 roads / 6000 vehicles, an increase of over 800 times, while the number of vehicles and roads only increases tenfold during the same period, fully demonstrating the impact of cubic complexity.
[0422] The above experiments analyzed the processing capacity of a single RSU under different scales (v:r). In urban-level vehicle-road cooperative systems, RSUs are densely deployed devices, and tasks are naturally spatially fragmented. Parallel computing by multiple RSUs significantly reduces the impact of time complexity in the location allocation stage, causing the overall end-to-end latency to fluctuate slightly around the best-case scenario. Combined with mechanisms such as batch scheduling, release of unsold vehicles, and "multiple-selection secretary" to limit the number of candidate vehicles, even in the continuous operation of tens of millions of vehicles, the system can still meet the real-time platooning scheduling requirements, far below the upper bound given by the extreme assumption of a single RSU.
[0423] (6) Elasticity Index Experiment
[0424] Figure 11 The results demonstrate the feasibility and robustness of this solution under different malicious vehicle percentages in formation missions: 0%, 10%, 20%, 30%, 40%, and 50%, verifying the feasibility and robustness of this solution under different attack types: utility misrepresentation, strategic breach of contract, and collusion attack.
[0425] The system resilience index reflects the system's ability to maintain basic functionality, performance, and stability under varying proportions of malicious vehicles, by considering factors such as the reputation of honest vehicles, the detection rate of malicious behavior, and social welfare. Results show that when the malicious vehicle proportion is between 0% and 20%, the resilience index exceeds 0.9, demonstrating extremely high stability. At 30% and below, the index remains above 0.807, indicating excellent robustness in low to moderate malicious environments. Above 30%, the optimal fleet size is 3-6 vehicles; a malicious vehicle proportion exceeding 40% significantly impacts the resilience index. When the malicious vehicle proportion reaches 50%, the resilience index is 0.664-0.675, still exhibiting a certain degree of stability. This solution performs exceptionally well with a malicious vehicle proportion below 30% and remains feasible in high-malicious scenarios.
[0426] This invention demonstrates through simulation and experimentation that the vehicle platooning scheme based on the Multi-Arm Slot Machine (MAB) is effective. This invention also relates to specific application areas or related products.
[0427] This invention belongs to the field of vehicle platooning technology, and particularly relates to a vehicle platooning method and system based on a multi-armed vehicle platooning mechanism. Against the backdrop of rapid development in wireless communication technology and intelligent transportation systems, this technical solution demonstrates significant commercial potential, addressing the core needs of improving road utilization, reducing fuel consumption, and enhancing safety. In intelligent logistics and sharing platforms, its dynamic and efficient platooning capabilities can support small and medium-sized logistics enterprises and individual truck drivers to share economies of scale through dynamic platooning. In the field of autonomous driving fleet operation, this technology can provide operators with better fleet scheduling and management solutions, significantly improving vehicle utilization and transportation service competitiveness. Simultaneously, the trusted data recording system based on blockchain technology can generate value-added services for the Internet of Vehicles, including vehicle credit assessment, customized insurance pricing, and serving as a qualification certificate for participating in advanced collaborative tasks such as distributed energy trading. It is expected to drive the rapid development of multiple sub-markets such as intelligent logistics, autonomous driving transportation services, and value-added services for the Internet of Vehicles, forming new economic growth points and industrial ecosystems.
[0428] II. Evidence related to the technical effects obtained by the embodiments of the present invention.
[0429] This aims to improve social welfare (up to 2.5 times that of autonomous driving) and ensure fairness in role utility through a compensation mechanism. Both PL and PF have better utility than autonomous driving, as supported by theoretical analysis and experimental data (e.g., Figure 6 , Figure 7 This jointly verified the rationality of resource allocation and the effectiveness of incentives. Experiments show that the system's total decision-making latency is less than 2.5ms in small-scale scenarios. Figure 9 This meets the real-time needs of vehicle networking; in large-scale scenarios, it utilizes multi-RSU parallel computing and sharding technology. Figure 10 Maintaining scalability. Theoretical analysis further demonstrates that the distributed architecture and blockchain technology optimize time complexity, supporting efficient operation in dynamic traffic environments. The solution maintains a high elasticity index (≥0.807) even when the proportion of malicious vehicles is ≤30%. Figure 11 Security is enhanced by combining blockchain with a penalty mechanism. In terms of energy consumption, the aerodynamic wake effect and synergistic effect significantly reduce PF energy consumption. Figure 8 The theoretical model and experimental data are consistent, verifying the energy-saving advantages of the formation scheme.
[0430] It should be noted that embodiments of the present invention can be implemented in hardware, software, or a combination of both. The hardware portion can be implemented using dedicated logic; the software portion can be stored in memory and executed by a suitable instruction execution system, such as a microprocessor or dedicated-design hardware. Those skilled in the art will understand that the above-described devices and methods can be implemented using computer-executable instructions and / or included in processor control code, for example, such code provided on a carrier medium such as a disk, CD, or DVD-ROM, a programmable memory such as read-only memory (firmware), or a data carrier such as an optical or electronic signal carrier. The devices and modules of the present invention can be implemented by hardware circuitry such as very large-scale integrated circuits or gate arrays, semiconductors such as logic chips, transistors, or programmable hardware devices such as field-programmable gate arrays, programmable logic devices, etc., or by software executed by various types of processors, or by a combination of the above-described hardware circuitry and software, such as firmware.
[0431] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any modifications, equivalent substitutions, and improvements made by those skilled in the art within the scope of the technology disclosed in the present invention, and within the spirit and principles of the present invention, should be covered within the scope of protection of the present invention.
Claims
1. A vehicle platooning method based on a multi-armed slot machine, characterized in that, The vehicle platooning method based on multi-armed slot machines includes the following steps: Step 1, Vehicle Utility Model; Vehicles waiting to be platooned at any time The set of selectable mission paths is At time t, vehicle v i Take on a role The utility of time is: U i,p (t)=v i,p (t)-c i,p (t)+r i,p (t)+s i,p (t), U i,p (t) is v i The revenue and expenditure indicators when performing task M(n) assist in the MAB parameter update and decision-making process; v i,p (t) is the base valuation, which is the vehicle v i The expected value of the current task is related to the current environmental context factors: Where, ξ p >0 represents the benefit coefficient corresponding to character p, d i,p (t) represents vehicle v i The distance to the mission destination, τ > 0 is a distance-sensitive parameter; c i,p (t) represents energy consumption expenditure, which is the vehicle's energy consumption expenditure. i Energy consumption when playing role p in a formation mission, and vehicle speed v i (t) and the wind resistance experienced by different roles p can be expressed as: c i,p (t)=α p ·f(v i (t)), Where, α p >0 represents a specific energy consumption coefficient for the character, f(v) i (t) is the energy consumption function; r i,p (t) As compensation revenue, to incentivize vehicles to assume high-energy-consuming positions, compensation revenue is provided to the lead vehicle PL: χ p >0 represents the corresponding compensation coefficient. This is an indicator function that takes the value 1 only when p = PL and 0 when p = PF, where PL is the lead vehicle and PF is the follow vehicle. s i,p (t) represents the synergistic effect, reflecting the positive benefits that vehicles gain from cooperating with each other. For example, following vehicles may benefit from aerodynamic wakes or safety may be improved when multiple vehicles cooperate. It can be represented as: Where, N i For vehicle v i The set of neighboring vehicles, κ ij As the coordination strength coefficient, the vehicle spacing is maintained at a fixed, small distance d. ij (t) achieves the best energy saving effect, using the function h(d) ij The optimal cooperative distance is described by (t). The closer the distance between vehicles is to the optimal cooperative distance, the stronger the cooperative effect; the farther the distance, the weaker the cooperative effect. Step 2, MAB dynamic filtering; Step 3: Reward and punishment mechanism and payment rules.
2. The vehicle platooning method based on a multi-armed slot machine as described in claim 1, characterized in that, The MAB dynamic filtering: 1) Road Filtering MAB For the task path set When driving alone or in convoy, the vehicle selects the appropriate road. n After passage is completed, the system calculates the current average actual benefit R based on the comprehensive utility feedback from the vehicle under different conditions. r (t), while updating the vehicle's local ledger, recording road conditions under different environments. n Historical mission overall benefits μ r (t), the revenue update expression: m i,r (t)=φμ i,r (t-1)+(1-φ)R r (t), Where φ∈(0,1) is the time decay factor, and μ is the vehicle's comprehensive gain at time t. i,r (t) The update is not based solely on the current average real return R. r (t) is also related to the payoff at time t-1; Assuming novice vehicles with no "exploration" experience apply for platooning, they can be initialized using vehicles with similar driving experience, reducing the cold start for novice vehicles in the first layer of screening; For vehicle v i For road r n The initial estimate can be matched with v from historical data. i Similar features The estimated historical revenue from inheriting this similar vehicle is denoted as Right now: To address the issue of vehicles being unable to fully explore all roads, a strategy criterion based on upper confidence bounds (UCB) is adopted for road exploration: Where, n i,r (t) represents vehicle v i Driving on the road n The number of historical records, γ>0, β>0 are the exploration adjustment coefficients; μ i,r (t) represents the vehicle's comprehensive revenue at time t; UCB i,r (t) can reflect the vehicle's exploration level on the corresponding road; N r (t) represents road r n Historical selection count, via r * (t) can reflect the road r n Based on the historical best exploration value of UCB, through the vehicle's own UCB i,r (t) and the corresponding road requirements r * (t) Compare and filter to identify those that do not match the road r n Vehicles that meet the requirements will be included in the candidate vehicle set. among; 2) Vehicle matching MAB.
3. The vehicle platooning method based on a multi-armed slot machine as described in claim 2, characterized in that, The vehicle is equipped with MAB: Use μ i,p (t) represents vehicle v i In the position of role The average utility estimate at time μ i,r (t) differs from μ i,p (t) More attention is paid to vehicles on the road. n Specific utility estimation when serving as PL or PF, μ i,r (t) More attention is paid to vehicles on the road. n The average utility estimate on, therefore μ i,p (t) can be greater than μ i,r (t) A more detailed description of the vehicle's position on the road. n Utility estimation when playing role p; using U i,p (t) represents vehicle v i If the actual utility observed at time t while playing role p is given, then its historical average utility update can be expressed as: m i,p (t)=φμ i,p (t-1)+(1-φ)U i,p (t), To measure uncertainty, the UCB value of the vehicle matching MAB layer is defined as follows: Where, n i,p (t) represents vehicle v i The number of times a vehicle has been selected in the history of playing role p; the number of times a vehicle has played different roles reflects the vehicle's experience to some extent. Vehicles with more experience are better able to handle various unexpected situations that occur during formation. After obtaining the candidate vehicle set After that, it is necessary to One PL and one |x were selected. * -1 PF, and maximize the profit W * At this point, the problem is transformed into an "assignment problem," and the Hungarian algorithm is used to solve for the optimal matching strategy x. * For ease of description, assume that in the current formation task, a total of N vehicles need to be selected, namely 1 PL and N-1 PF vehicles. Before proceeding with the formal screening, a "vehicle-role" UCB value matrix A needs to be constructed. The size of the matrix depends on the set of candidate vehicles. Given the number of vehicles N required for the task, define a matrix row to store the corresponding... All vehicles, totaling The rows and columns correspond to the N vehicle positions to be assigned in the formation, where the N vehicle positions include: 1 PL and N-1 PF. like If the value is greater than N, then there will be more vehicles than needed. However, the Hungarian algorithm requires that the number of rows and columns be equal, so matrix A needs to be expanded: if You can add at the end of the column. A virtual slot, making its UCB i,p (t) = 0, ensuring the number of columns equals the number of rows; after confirming the matrix size, corresponding values need to be filled in the corresponding positions for easy filtering, therefore, in vehicle v i At the intersection of the corresponding row and the navigator position column, enter the UCB of the vehicle when it served as the historical navigator. i,PL (t)(i=1,2,3···); At the intersection of the following position columns, fill in the UCB when the vehicle served as the historical following vehicle. i,PF (t); If there is a virtual slot, fill in 0 in the corresponding position; this will give you a A square matrix, the values in which represent the maximum profit W. * ; Take a from the matrix ij Transform into max(a) ij )-a ij , where a ij This represents the element in the i-th row and j-th column of the "vehicle-role" UCB value matrix A, where max(a ij ) is the maximum value among all elements of the matrix, thus transforming the maximization problem into a minimization problem; for each row, find the minimum value of that row and subtract it from all elements of that row; for each column, find the minimum value of that column and subtract it from all elements of that column; this ensures that each row and each column has at least one 0 element; In the processed matrix, try to mark some 0 elements so that each row and each column can be marked with at most one 0; try to cover all 0 elements with the minimum number of column lines; if the number of lines is equal to the matrix dimension, a feasible solution can be found; otherwise, it is necessary to continue to make further adjustments to the elements not covered by lines until all 0s can be covered with the same number of lines as the dimension. A perfect match exists when all zeros can be covered by lines of the same order as the matrix. Based on the marked zero elements, a column index is assigned to each row sequentially, ultimately yielding the optimal matching strategy x. * ; By the optimal matching strategy x * It can be known that: vehicle v i Which position is it assigned to; if the column belongs to the navigator role, it indicates that vehicle v i Select the lead vehicle; if the column belongs to the follower role, i.e., vehicle v i Selected vehicle to follow; if matched with a "virtual slot", it means that the vehicle has not been selected into the formation and will continue to wait for the next round of selection.
4. The vehicle platooning method based on a multi-armed slot machine as described in claim 1, characterized in that, The aforementioned reward and punishment mechanism: The behavior of a vehicle during mission M(n) is monitored through collective verification by other vehicles in the formation; when a violation is detected, the other vehicles in the formation verify the behavior through multi-party consensus. Among them, O i (j) represents the detection value of the violation by vehicle i against vehicle j, where 1 indicates a confirmed violation and 0 indicates no violation was observed, and n is the number of vehicles participating in the consensus. This is the consensus threshold; when C v If ≥Γ, then confirm vehicle v j There were violations; Violations are categorized into three types based on severity: T = {V} L V M V H }; where V L V M V H These represent different degrees of violation, corresponding to different levels of harm to the formation; when a violation occurs and causes economic losses to the formation mission, the violating vehicle v j The vehicles in the formation and the person who initiated the mission should be compensated according to the revenue generated when the mission is completed normally. Among them, P i The outstanding amount for vehicles involved in violations. The amount paid for marginal contribution loss, u i The amount lost by the task initiator; Once the violation is confirmed through the above checks, the system records the violation information via the blockchain; however, to protect user privacy, the violation record uses blind signature and zero-knowledge proof technology, and the recording format is as follows: R j (t)=(H(ID v ),V t ,T s ,L,Sig), Among them, H(ID) v V is the hash value that uniquely identifies the vehicle. t For the type of violation, T s L is the timestamp, L is the geolocation hash, and Sig is the multi-signature verification information; By leveraging the spread of violations to help more vehicles in a convoy receive early warnings, the propagation of blockchain records follows the mathematical model below: Among them, R j (t) represents vehicle v j The set of violation records possessed at time t, N i (t) represents vehicle v i The set of other vehicles encountered at time t; vehicle v at time t. i The existing early warning information is R i (t); vehicle v i By using set N i The vehicles in (t) communicate with each other to obtain their early warning information R. j (t); vehicle v i At the next time point t+1, it will interact with its original records and the records of neighboring vehicles; Based on the violation history recorded by the blockchain, the system dynamically adjusts the vehicle's future platooning permissions; defining vehicle v j The historical impact factors of violations are: Where, n L n M n H w represents the number of violations at different levels. L w M w H Let w be the weighting coefficient for each level of violation, and w L <w M <w H ; Let t be the time decay function. r Here, t represents the current time, θ > 0 is the time decay coefficient, and r represents R. j (t) A single violation record in the set; the impact of the violation record will gradually weaken over time; During the vehicle matching MAB phase, the UCB value of non-compliant vehicles is adjusted: UCB′ i,p (t)=UCB i,p (t)·P(H v ), Among them, UCB′ i,p (t) represents the adjusted UCB value. The UCB value is a key factor affecting whether a vehicle can be selected into the formation. Therefore, a UCB value that is smaller may be directly rejected from joining the formation. P(H v Let ω be the penalty function for violations, and ω > 0 be the penalty intensity coefficient. When H v If a certain threshold is exceeded, the vehicle will be unable to participate in certain types of formation missions for a certain period of time. Where σ = 3 is the critical value for serious violations; T(v) is the prohibition periodic function, which is proportional to the violation history factor; within the T(v) period, vehicles are prohibited from gaining benefits by joining vehicle platoons. This restricts violating vehicles while ensuring that vehicles have the opportunity to regain platoon participation rights after a period of time. In addition, the system allows vehicles to restore trust through behavioral proofs: After a vehicle successfully completes a platooning maneuver without violations m times, its historical violation impact factor K will decrease proportionally.
5. The vehicle platooning method based on a multi-armed slot machine as described in claim 1, characterized in that, The payment rules are as follows: Based on the actual service of all vehicles An optimal vehicle formation N can be calculated. * The total revenue can be expressed as: If vehicle v is removed from any vehicle platoon i The optimal profit that the system can achieve on the remaining vehicles is Then, the vehicle v is defined according to its marginal contribution. i The payment is defined as: In order to consider the contribution of vehicles to the overall formation system, it is necessary to calculate the true economic benefit of each vehicle, v. i The final net income is: Where π i The amount a vehicle needs to pay based on its marginal contribution can actually reduce its net revenue if it misreports, thus encouraging honest reporting. The rational choice for a vehicle is to report truthfully.
6. A vehicle platooning system based on a multi-armed machine, implementing the vehicle platooning method based on any one of claims 1-5, characterized in that, The vehicle platooning system based on multi-armed slot machines includes: Vehicle utility module, used for vehicle utility model; Dynamic filtering module, used for dynamic filtering of MABs; The reward and punishment module is used for reward and punishment mechanisms and payment rules.
7. A computer device, characterized in that, The computer device includes a memory and a processor, the memory storing a computer program that, when executed by the processor, causes the processor to perform the steps of the vehicle platooning method based on any one of claims 1-5.
8. A computer-readable storage medium storing a computer program that, when executed by a processor, causes the processor to perform the steps of the vehicle platooning method based on a multi-armed slot machine as described in any one of claims 1-5.
9. An information data processing terminal, characterized in that, The information data processing terminal is used to implement the vehicle platooning system based on the multi-armed slot machine as described in claim 6.