Internet of vehicles cooperative sensing resource allocation method based on region and reputation driving

By dividing perception areas in the Internet of Vehicles environment and selecting reputable RSU nodes, combining blockchain and multi-agent optimization algorithms, the problem of collaborative perception in crossroads scenarios is solved, and data redundancy reduction, security improvement and energy efficiency optimization are achieved.

CN120282177APending Publication Date: 2025-07-08CHONGQING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510595889.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-09
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

In the Internet of Vehicles environment, in the crossroads scenario, collaborative perception has perceived data redundancy, data consistency maintenance difficulties and data security problems. Traditional collaborative perception strategies are difficult to adapt to highly dynamic vehicle networks, and mobile edge computing brings security risks.

Method used

The Internet of Vehicles collaborative perceptual resource allocation method based on region and reputation-driven is adopted. By dividing the global perceptual region into overlapping and non-overlapping regions, the RSU with high reputation is dynamically selected as the consensus node, and combined with blockchain technology and multi-agent near-end strategy optimization algorithm, resource allocation is optimized to reduce redundancy and improve security.

Benefits of technology

Effectively reduce perceived data redundancy, ensure data security and privacy protection, reduce system delay and energy consumption, and meet real-time and energy efficiency optimization needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120282177A_ABST
    Figure CN120282177A_ABST
Patent Text Reader

Abstract

The invention relates to an Internet of Vehicles cooperative sensing resource allocation method based on region and reputation driving, and belongs to the technical field of automatic driving and mobile communication. The method comprises the following steps: S1, dividing a global sensing area into an overlapping area and a non-overlapping area in an Internet of Vehicles scene, and adopting a differentiated fusion method according to data characteristics of different areas; s2, providing a global map dynamic updating mechanism based on gridding; s3, by comprehensively analyzing the current state and historical behavior data of the RSU, constructing a reputation model of the RSU for calculating the reputation value of the RSU; s4, providing a resource allocation target combining energy consumption and delay optimization, and comprehensively considering optimization requirements of energy consumption and delay; and S5, solving an optimization problem through an A-MAPPO algorithm. According to the method, the system delay and the energy consumption weighted sum can be effectively reduced, and the real-time performance and energy efficiency optimization requirements are met.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical fields of autonomous driving and mobile communication, and relates to a method for allocating cooperative perception resources in a vehicle-to-everything (V2X) network driven by region and reputation. Background Art

[0002] Autonomous driving technology is profoundly transforming the future transportation system, and has achieved remarkable results in reducing traffic accidents, improving safety and energy efficiency. As key applications in this field, autonomous vehicles and connected automated vehicles (CAVs) are increasingly widely deployed, which not only bring convenience to personal travel, but also lay an important foundation for the construction of smart cities and the digital transformation of transportation systems. In the V2X network system, the integration of communication, sensing, and computing is extremely crucial. Vehicles collect a large amount of sensing data through sensors such as LIDAR and cameras. These data are the core basis for constructing an accurate global map and assisting driving decisions, and are directly related to the safety and reliability of autonomous driving. However, the sensing range and capabilities of a single vehicle are limited, and it is difficult to comprehensively and timely obtain complex and changing environmental information. Therefore, cooperative perception technology has emerged. It effectively makes up for the deficiencies of single-vehicle perception through data sharing and collaborative processing among multiple vehicles and devices, and has become the key to improving the sensing ability of the V2X network. However, the transmission and processing of a large amount of sensing data pose extremely high requirements on communication resources and computing capabilities. At the same time, the issues of security and privacy protection in the data processing process have also attracted much attention, making it urgent to develop efficient cooperative perception strategies and resource allocation schemes.

[0003] The V2X network environment is complex, and different scenarios face different challenges. In the crossroads scenario, the cooperative perception problem is particularly prominent. On the one hand, there is serious redundancy in sensing data. Repeated sensing of the same area by multiple vehicles and roadside units (RSUs) generates a large amount of overlapping data, wasting communication bandwidth, increasing the data processing burden, and seriously affecting the real-time performance of the system. On the other hand, it is difficult to maintain data consistency. The accuracy and reliability of sensing data from different devices vary greatly, and local errors are easily amplified during the fusion process, affecting the accuracy of the global map and threatening the safety of autonomous driving. In addition, traditional cooperative perception mostly adopts a single-layer perception model, and a single fusion method is difficult to adapt to a highly dynamic vehicle network. Moreover, although mobile edge computing can make up for the insufficient computing power of intelligent vehicles, the open environment brings new risks, such as false interactions caused by malicious behaviors of intelligent vehicles, and problems such as the difficulty in guaranteeing the authenticity and integrity of data, which not only threaten user privacy but also may disrupt the normal operation of the V2X network system. Therefore, breaking through the limitations of the single-layer perception model, developing an efficient multi-layer fusion strategy, and solving the problems of data security and privacy protection are urgent issues to be solved in the current field of V2X network cooperative perception. Summary of the Invention

[0004] In view of this, the purpose of the present invention is to provide a method for allocating cooperative perception resources in a vehicle networking based on region and reputation driving, so as to reduce the redundancy of perception data, ensure the security of perception data, reduce the weighted sum of system latency and energy consumption in cooperative perception, and meet the requirements of real-time and energy efficiency optimization.

[0005] To achieve the above object, the present invention provides the following technical solutions:

[0006] A method for allocating cooperative perception resources in a vehicle networking based on region and reputation driving, the method comprising:

[0007] S1. In a vehicle networking scenario including CAV (Connected Automatic Vehicles), RSU, MES (Mobile Edge Server), TA (Trusted Authority) and a central server, divide the global perception regions of CAV and RSU into overlapping regions and non-overlapping regions; the CAV collects surrounding point cloud data and converts it to the global coordinate system, and uploads data according to the region characteristics. For the overlapping regions, upload the processed feature data, and for the non-overlapping regions, upload the original point cloud data to the MES for fusion; the RSU is used to collect surrounding environmental information, expand the perception range of CAV, and is selected as a blockchain consensus node under specific conditions to help process perception data and participate in blockchain verification; the MES is connected to the RSU serving as a consensus node one-to-one to provide computing services; the TA is responsible for entity identity authentication and registration, issues certificates using encryption technology, and collaborates with the central server to update certificate information to ensure the security and stability of the vehicle networking; the central server dynamically updates the global map according to the uploaded data, broadcasts the latest perception results through grid management, and stores the information related to RSU and CAV. In addition, the central server participates in the computing process of the blockchain as the main blockchain consensus node;

[0008] S2. Comprehensively analyze the current status and historical behavior data of the RSU, and at the same time construct a reputation model of the RSU. The reputation model dynamically evaluates the reputation value of the RSU through direct reputation value and indirect reputation value, and screens out the RSU with a reputation value greater than the threshold as the blockchain consensus node;

[0009] S3. Based on the blockchain, propose a dynamic update mechanism for the global map based on grid. The central server generates a global high-precision map according to the perception data uploaded by the RSU, and then distributes it to the RSU and CAV through hierarchical broadcasting;

[0010] S4. Build the energy consumption and latency calculation models for the blockchain, the energy consumption and latency calculation models for processing sensed data on the CAV side, and the energy consumption and latency calculation models for processing sensed data on the RSU side; based on the above energy consumption and latency calculation models, build a joint optimization objective function for energy consumption and latency, and comprehensively balance the total energy consumption and total latency of the vehicle network sensing tasks to simultaneously meet the real-time requirements of the sensing tasks and the energy efficiency optimization needs;

[0011] S5. Use the multi-agent proximal policy optimization algorithm based on the attention mechanism to learn the resource allocation scheme of the weighted sum of latency and energy consumption.

[0012] Furthermore, in step S2, the reputation model is expressed as:

[0013]

[0014] In the formula, C r (t) represents the reputation value of RSU r, ω2 represents the reputation value adjustment weight, represents the direct reputation value of RSU r, represents the indirect reputation value of RSU r.

[0015] Among them, the direct reputation value is reflected in the consensus success rate and reputation growth rate demonstrated by the RSU during the past consensus process;

[0016] The consensus success rate is an indicator to measure the completion of the RSU in the consensus link. The factors affecting the consensus success rate of the RSU include the number of times that the consensus fails due to equipment failures or attacks; the consensus success rate is expressed as:

[0017]

[0018] In the formula, su r (t) represents the cumulative number of times that RSU r has successfully completed the consensus, and fa r (t) represents the cumulative number of times that RSU r has not successfully completed the consensus;

[0019] The reputation growth rate is used to measure the stability of the RSU node and reflects the reputation change trend of the RSU; the reputation growth rate is expressed as:

[0020]

[0021] In the formula, SU r (t - 1) represents the reputation value of RSU r in the (t - 1)-th round. If it is the first round currently, then SU r (t - 1) = SU r (t);

[0022] The direct reputation value is expressed as:

[0023]

[0024] Wherein, ω1 represents the direct trust adjustment weight, which can adjust the ratio of the consensus success rate and the reputation growth rate;

[0025] The indirect reputation value is the evaluation of RSU r by other RSU in the blockchain system; to prevent the abnormal evaluation of malicious nodes, the evaluation of other RSU on RSU r needs to refer to the direct reputation value of the RSU making the evaluation at the same time; let R r′ (t) represents the credibility of the RSU r' making the evaluation, and R r′ (t) is expressed as:

[0026]

[0027] In the formula, represents the direct reputation value of RSU r' in the (t-1)th round; D max and D min represent the thresholds of credibility, exceeding D max means that RSU r' is completely credible, lower than D min means that RSU r' is completely untrustworthy;

[0028] Then the indirect reputation value of RSU r is expressed as:

[0029]

[0030] In the formula, represents the evaluation of RSU r by RSU r', and R represents the total number of RSU.

[0031] Furthermore, in step S3, the grid-based global map dynamic update mechanism includes:

[0032] S31. Each CAV entering the cooperative perception system needs to submit account information to the TA for identity verification and registration; after the TA verification is passed, an identity authentication certificate is distributed to the CAV; after the initialization is completed, the CAV starts to collect point cloud data;

[0033] S32. The blockchain selects the RSU with a reputation value exceeding the threshold as the consensus node according to the reputation values of each RSU calculated by the reputation model;

[0034] S33. The CAV and RSU convert the coordinates of the collected point cloud data into global coordinates, and upload the coordinate data to the central server through the RSU. The central server determines the grid to which the point cloud data of the CAV and RSU belongs according to the coordinate information, and marks the overlapping and non-overlapping areas of each data to form a regional division result;

[0035] S34. After completing the regional division, the central server returns the regional division result to the relevant RSU and CAV;

[0036] S35. The CAV further processes the perception data: for the data in the non-overlapping area, directly upload the original point cloud data to the RSU for generating the global map; for the data in the overlapping area, the CAV preferentially selects local processing. If the computing resources of the CAV are insufficient, the CAV chooses to offload the data to the MES equipped with the RSU for processing, and then upload it by the RSU;

[0037] S36. After the CAV uploads the perception data to the RSU, the blockchain node verifies the perception data according to the PBFT consensus mechanism;

[0038] S37. The RSU processes the perception data from the CAV and simultaneously processes the perception data collected by itself;

[0039] S38. The RSU uploads the processed perception data to the central server;

[0040] S39. The central server receives the processed data from all RSUs, fuses the original perception data in the non-overlapping area and the feature data in the overlapping area to generate a global high-precision map, and then distributes it to the CAV and RSU through hierarchical broadcasting.

[0041] Among them, the central processor distributes the global high-precision map to the CAV and RSU through hierarchical broadcasting. The first layer is distributed from the central server to the RSU to ensure the update of the local map in the scene; the second layer is that the RSU distributes the received global map to the CAVs within its coverage area to ensure that each CAV obtains the updated global map in real time.

[0042] Further, in step S4, the constructed energy consumption and delay calculation model of the blockchain is expressed as:

[0043] T b (t) = t b1 (t) + t b2 (t) + t b3 (t) + t b4 (t) + t b5 (t)

[0044]

[0045] where, T b (t) represents the delay in completing the consensus process in the blockchain, and t b1 (t)~t b5 (t) respectively represent the delays in the request stage, pre-prepare stage, prepare stage, confirm stage, and reply stage of the PBFT consensus algorithm when the consensus process in the blockchain is completed based on the PBFT consensus algorithm; E b (t) represents the energy consumption in completing the consensus process in the blockchain, and k c represents the effective energy coefficient, and f c,b (t) represents the computing resources provided by the central server for the blockchain, and c b1 (t)~c b5 (t) respectively represent the energy consumptions in the request stage, pre-prepare stage, prepare stage, confirm stage, and reply stage of the PBFT consensus algorithm;

[0046] The constructed energy consumption and delay calculation model for the CAV side to process perception data is expressed as:

[0047] T CAV (t)=max{T1(t), T2(t),..., T I (t)}

[0048] T i (t)=β i (t)T i mes-ol (t)+(1-β i (t))T i V-ol (t)+T i nol (t)

[0049]

[0050] where, T CAV (t) represents the delay of the CAV side in overall processing perception data, and T i (t) represents the delay of CAV i in processing perception data, i = 1, 2,..., I, and I is the total number of CAVs; β i (t) represents the decision of CAV i whether to offload perception data to the RSU; T i V -ol (t) represents the delay of CAV i in locally processing overlapping area data, and T i mes-ol (t) represents the delay of CAV i in selecting to offload overlapping area data to the RSU for processing, and T i nol (t) represents the delay of CAV i in processing non-overlapping area data; Ej,i (t) represents the energy consumption of MES j for processing sensed data, represents the energy consumption of CAV i for selecting to offload the overlapping area data to RSU j for processing, represents the energy consumption of CAV i for processing non - overlapping area data;

[0051] The constructed energy consumption and delay calculation model for processing sensed data on the RSU side is expressed as:

[0052] T RSU (t) = max{T1(t), T2(t),..., T J (t)}

[0053]

[0054] In the formula, T RSU (t) represents the overall delay of processing sensed data on the RSU side, T j (t) represents the processing of sensed data by RSU j, and respectively represent the delays of RSU j for processing overlapping area data and non - overlapping area data; E j (t) represents the energy consumption of RSU j for processing sensed data, represents the energy consumption of RSU j for processing sensed data, represents the energy consumption of the central server for processing the sensed data of RSU j;

[0055] Then the total energy consumption and total delay of the vehicle - to - everything (V2X) sensing task are respectively expressed as:

[0056]

[0057] T total (t) = max{T CAV (t), T RSU (t)} + T b (t)

[0058] In the formula, J represents the total number of RSUs acting as consensus nodes, that is, the total number of MESs.

[0059] Further, in step S4, based on the total energy consumption and total delay of the V2X sensing task, a joint optimization objective function of energy consumption and delay is constructed:

[0060]

[0061]

[0062] Where, F1(t) and F2(t) respectively represent the computational resource allocation strategies of the MES for the consensus process and the perception data processing process; F3(t) represents the computational resource allocation strategy of the CAV for the perception data processing process; β(t) represents the CAV offloading decision; α∈[0,1] represents the adjustment factor of the collaborative perception system cost, which is used to adjust the system delay cost and the system energy consumption cost;

[0063] The constraint condition C1 means that the delay of generating the overall global map cannot exceed the maximum tolerable delay T of the system max ; The constraint condition C2 means that the energy consumption of generating the overall global map cannot exceed the maximum tolerable energy consumption E of the system max ; The constraint condition C3 means that the bandwidth B i (t) allocated to the CAV i cannot exceed the total system bandwidth B max ; The constraint condition C4 means that the computational resource f i,i (t) used by the CAV i to process the perception data cannot exceed its maximum available computational resource f i (t); The constraint condition C5 means that when the MES equipped with the RSU j processes the perception data of the RSU j itself and the perception data offloaded by the CAV i, the computational resources used cannot exceed the maximum computational resource f of the MES j (t), f j,j (t) represents the computational resource used by the RSU j to process its own perception data, and f j,i (t) is the computational resource allocated by the MES j for processing the perception data from the CAV i; The constraint condition C6 means that the computational resource f j,b (t) allocated by the MES j to the blockchain cannot exceed the maximum computational resource f of the MES j j (t); The constraint condition C7 means that the computational resource of the central server for processing the perception data uploaded by the RSU cannot exceed the maximum computational resource f of the central server c (t), f c,j (t) represents the computational resource allocated by the central server for processing the perception data from the RSU j.

[0064] Furthermore, step S5 includes modeling the optimization problem constructed in step S4 as a Markov decision process, and the set of agents of this Markov decision process is the perception device Q; The Markov decision process includes:

[0065] First, the agent q obtains the current observation value from the environment and executes the action given by the actor network, synchronizes the observation and action of the agent q to the central server, and after the central server performs the reward evaluation, the agent q receives the returned evaluation reward; Store the relevant information in the experience area;

[0066] Secondly, the agent q samples some experience batches from its own buffer pr q (t) represents the sampled action a q (t)'s log probability, and s(t) represents the state;

[0067] Then, in each update, the actor and the critic use the policy loss and the global state value loss to update the actor network parameters and the critic network parameters respectively; among them, the loss of the actor q is expressed as:

[0068]

[0069] In the formula, and represent the old policy and the current policy respectively; represents the advantage function A q (t) = Q q (s(t), a q (t)) - V q (s(t))'s estimated value, Q q (s(t), a q (t)) represents the value function of performing the action a q (t) in the state s(t); ∈ represents a hyperparameter;

[0070] The state value function estimated by the critic of the agent q is Then the loss of the critic q is expressed as:

[0071]

[0072] In the formula, ξ q represents the parameters of the q-th critic network, and V q (s(t)) represents the cumulative discounted reward;

[0073] After introducing the attention mechanism, the observation vectors of each agent first pass through their respective MLPs for feature extraction to obtain the eigenvalue e q , and the eigenvalues of all agents are sent to the attention head, and the attention value x is obtained through the following formula q :

[0074]

[0075] In the formula, e q′ represents the eigenvalue of the agent q'; d key represents 's variance; the matrix W key converts e q′ into a key, and the matrix W v converts e q into a query; Wq′ represents the weight matrix; send x q and o q (t) to the MLP to obtain the estimated state value

[0076] Finally, download the updated parameters to the agent q, and determine whether the algorithm iteration is completed. If the iteration is not completed, iterate again; otherwise, end the algorithm.

[0077] The Markov decision process described includes two types of agents: CAV and MES.

[0078] Among them, the elements of the Markov decision process of agent CAV i include the observation o i (t), the action a i (t), and the reward r i (t);

[0079] The observation o i (t) is expressed as L i (t) represents the amount of data sensed by CAV i itself, represents the amount of data in the overlapping area, represents the amount of data in the non - overlapping area, f i (t) represents the computing resources of CAV i, Y i,j (t - 1) represents the computing load for MES j to process the offloaded data of CAV i;

[0080] The action a i (t) is expressed as a i (t) = {β i (t), f i,i (t)}, where β i (t) represents the offloading decision of CAV i, and f i,i (t) represents the computing resources of CAV i itself for sensing data processing;

[0081] The reward r i (t) is expressed as r i (t) = -αmax{T CAV (t), T RSU (t)}-(1 - α)(E i (t)+E j,i (t))}, where E i (t) represents the energy consumption of CAV i for processing its own sensed data, and E j,i (t) represents the energy consumption of MES j for processing the sensed data of CAV i;

[0082] The elements of the Markov decision process of agent MES j include the observation o j (t), the action aj (t) and reward r j (t);

[0083] where the observation o j (t) is expressed as L j (t) represents the amount of sensed data of RSU j itself, represents the amount of data in the overlapping area, represents the amount of data in the non - overlapping area, f j (t) represents the computing resource of MES j itself;

[0084] action a j (t) is expressed as a j (t) = {f j,i (t), f j,j (t)}, f j,i (t) and f j,j (t) respectively represent the computing resources allocated by MES j for processing the sensing tasks of RSU j and CAV i;

[0085] reward r j (t) is expressed as E j,j (t) represents the energy consumption of MES j for processing its own sensed data, α represents the adjustment factor of the system cost, f j,b (t) represents the computing resources allocated by MES j to the blockchain, c b2 (t) ~ c b5 (t) respectively represent the energy consumption in the pre - prepare stage, prepare stage, confirm stage and reply stage of the PBFT consensus algorithm, T total (t) represents the total delay of the vehicle - to - everything sensing task.

[0086] The beneficial effects of the present invention are as follows: Aiming at the challenges faced by IoVs collaborative sensing in the cross - intersection street scenario, including sensing data redundancy, difficulty in maintaining data consistency, and data security, etc., the present invention proposes a method for allocating IoVs collaborative sensing resources based on region and reputation driving.

[0087] To reduce the redundancy of perception data, the present invention proposes a dynamic region classification and differential fusion strategy, which divides the global perception region into overlapping regions and non-overlapping regions, and adopts differential processing methods according to the data characteristics of different regions. That is, in the overlapping regions, the medium-term fusion strategy is preferentially adopted to upload the processed feature data to reduce the communication volume; in the non-overlapping regions, the early fusion strategy is adopted to upload the original point cloud data to ensure the accuracy of global perception. At the same time, to ensure the security of perception data, the present invention proposes a reputation-driven consensus mechanism based on the Practical Byzantine Fault Tolerance (PBFT) consensus algorithm, constructs a roadside unit reputation model to screen consensus nodes, fully considers the reputation of consensus nodes, improves the security and efficiency of the data verification process, and effectively guarantees the credibility and privacy protection of perception data in the vehicle-to-everything (V2X) environment. In addition, the present invention proposes to construct an objective function for joint optimization of energy consumption and delay, and uses the A-MAPPO algorithm to solve this optimization problem, which can effectively reduce the sum of system delay and energy consumption weighted in cooperative perception and meet the requirements of real-time and energy efficiency optimization.

[0088] Other advantages, objects, and features of the present invention will be described to some extent in the following specification, and to some extent, will be apparent to those skilled in the art based on the study of the following text, or can be taught from the practice of the present invention. The objects and other advantages of the present invention can be realized and obtained through the following specification. BRIEF DESCRIPTION OF THE DRAWINGS

[0089] In order to make the objects, technical solutions, and advantages of the present invention clearer, the present invention will be described in detail preferably with reference to the accompanying drawings, where:

[0090] Figure 1 is a schematic diagram of a cooperative perception scenario;

[0091] Figure 2 is a schematic diagram of the key process of the cooperative perception scheme;

[0092] Figure 3 is a network structure diagram of the A-MAPPO algorithm. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0093] The following specific examples illustrate the implementation manners of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific implementation manners. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the drawings provided in the following embodiments only schematically illustrate the basic concept of the present invention. Without conflict, the following embodiments and the features in the embodiments can be combined with each other.

[0094] This embodiment provides a vehicle networking collaborative perception resource allocation method based on region and reputation drive, which can optimize the collaborative perception resource allocation strategy of the vehicle networking and reduce the execution delay of the collaborative perception.

[0095] The steps of this method are as follows:

[0096] S1: In the vehicle networking scenario, a multi-layer collaborative perception concern region scheme based on relative position is proposed. The global perception region is divided into overlapping regions and non-overlapping regions, and different fusion methods are adopted according to the data characteristics of different regions. Among them, the collaborative perception scenario is as Figure 1 shown.

[0097] The vehicle networking scenario consists of CAV, RSU, MES, a central server, and TA. The CAV is equipped with sensors such as LiDAR, collects surrounding point cloud data and converts it to the global coordinate system, uploads data according to the region characteristics (uploads the processed feature data in the overlapping region and uploads the original point cloud data to MES for fusion in the non-overlapping region), and at the same time determines the grid to which the data belongs and uploads relevant information. The vehicle set is denoted as I = {1, 2,..., i,..., I}, and the computing resource of CAV i is denoted as f i (t). The RSU is deployed on fixed facilities at intersections, collects surrounding environmental information, expands the perception range of the CAV, can be used as a blockchain consensus node, is closely connected to the MES, helps process perception data and participate in blockchain verification. The set of all RSUs is R = {1, 2,..., r,..., R}, and the set of RSUs elected as consensus nodes is J = {1, 2,..., j,..., J}, and stores relevant information of the CAV when interacting with the CAV. The MES can provide computing services. The MES corresponds to the RSU one by one and is connected by a wired link. The computing resource of MES j is f j(t). The central server dynamically updates the global map based on the uploaded data, manages through grid division and broadcasts the latest sensing results. At the same time, it stores the information related to RSU and CAV and participates as the main node of blockchain consensus. TA is responsible for entity identity authentication and registration, ensures the security and legality of data interaction through digital certificate management, issues certificates using encryption technology, and collaborates with the central server to update certificate information to ensure the security and stability of the vehicle network. The entire vehicle network collaborative sensing system operates within discrete time slots T = {1, 2,... t,... T}. The consensus node selection, offloading decision, and resource allocation strategies remain unchanged within the time slots. The set of sensing devices is Q = {1, 2,…, q,…, I + J}, where the first I are CAVs and the last J are RSUs. Each device collaborates to complete sensing, data transmission, and fusion to accurately generate the global map.

[0098] By dividing the global sensing area into overlapping and non - overlapping areas and adopting different fusion methods according to the data characteristics of different areas. On the one hand, uploading the processed feature data in the overlapping area can reduce the communication volume. On the other hand, uploading the original point cloud data in the non - overlapping area can ensure the accuracy of global sensing.

[0099] S2: Propose a reputation - driven consensus mechanism based on the PBFT consensus algorithm. Specifically, comprehensively analyze the current state and historical behavior data of RSU, and at the same time construct a reputation model of RSU. This reputation model dynamically evaluates the reputation value of RSU through direct reputation value and indirect reputation value. According to the reputation value of RSU dynamically evaluated by the reputation model, select the RSUs with reputation values greater than the threshold as consensus nodes. These high - reputation RSUs have significant advantages in the reliability and stability of data processing. When CAV uploads data, the system uses the selected high - reputation RSU consensus nodes to verify the consistency of the data through the PBFT algorithm. Compared with the traditional PBFT algorithm, the present invention fully considers the reputation of consensus nodes, improves the security and efficiency of the data verification process, and effectively guarantees the credibility and privacy protection of the sensing data in the vehicle network environment.

[0100] 1) Direct reputation value: Define the direct reputation value as the consensus success rate and reputation growth rate demonstrated by RSU in the past consensus process.

[0101] Among them, the consensus success rate is a key indicator to measure the completion of RSU in the consensus link. The core factor affecting the consensus success rate of RSU is the number of times it fails to complete the consensus due to equipment failure or being attacked. With this indicator, it is possible to preferentially select the RSUs that perform better in completing the consensus. The consensus success rate SU r (t) is defined as the ratio of the number of successful consensus completions to the total number of consensus participations and can be expressed as:

[0102]

[0103] In the formula, su r (t) represents the cumulative number of successful consensus completions of RSU r; fa r (t) represents the cumulative number of failures of RSU r.

[0104] The reputation growth rate is used to measure the stability of the RSU node. By evaluating it, the changing trend of the RSU's reputation can be effectively understood. Since the reputation of the RSU is in dynamic change, its reputation growth rate will also be adjusted in real time accordingly. The reputation growth rate can be expressed as:

[0105]

[0106] In the formula, SU r (t - 1) represents the reputation value in the (t - 1)-th round. If the current is the first round, then SU r (t - 1) = SU r (t). Maintaining a reasonable reputation growth rate helps ensure the stable and reasonable growth of the RSU's reputation and prevent malicious nodes from interfering with the normal operation of the system by intermittently misbehaving. Therefore, when G r (t) ≥ 0.8, set G r (t) in the consensus selection of this round to 0. The complete direct reputation value of RSU r can be defined as:

[0107]

[0108] In the formula, ω1 is the direct trust adjustment weight, which can adjust the ratio of the consensus success rate and the reputation growth rate. Before the end of each period, the current primary node will initiate a request to update the consensus success rate and the reputation growth rate. After receiving this request, each node will calculate the reputation value and the reputation growth rate of each node based on the previously accumulated data, realizing the dynamic update of the reputation status of each node.

[0109] 2) Indirect reputation value: Define the indirect reputation value as the evaluation of RSU r by other RSU in the blockchain system. In the described reputation model, the reputation evaluation of RSU r by other RSU nodes takes into account the reputation evaluation of RSU r by other RSU in the region. It should be noted that to prevent malicious nodes from maliciously evaluating normal nodes and giving high scores to other malicious nodes, the evaluation of RSU r by other RSU should also refer to the direct reputation value of the RSU making the evaluation. For example, when a certain RSU r' evaluates RSU r, to prevent malicious nodes from maliciously evaluating normal nodes, it is necessary to refer to the direct reputation value of RSU r'.

[0110] Let R r′(t) represents the credibility of RSU r′. To prevent malicious nodes from damaging the fairness of the reputation value, R r′ (t) is expressed as:

[0111]

[0112] In the formula, represents the direct reputation value of RSU r′ in the (t - 1)-th round; D max and D min represent the thresholds of credibility. Exceeding D max means completely credible, and being lower than D min means completely non-credible.

[0113] Therefore, the indirect reputation value of RSU r can be expressed as follows:

[0114]

[0115] In the formula, represents the evaluation of RSU r by RSU r′. By combining the direct reputation value and the indirect reputation value, the reputation value of RSU r is expressed as:

[0116]

[0117] In the formula, ω2 is the reputation value adjustment weight, which can adjust the ratio of the direct reputation value and the indirect reputation value. Finally, the nodes with a reputation value greater than the threshold C min are selected as consensus nodes and added to the consensus node set J of the RSU.

[0118] S3: Propose a grid-based global map dynamic update mechanism, as Figure 2 shown, including:

[0119] ① Initialization: Each CAV entering the cooperative perception system needs to submit account information to the TA. This information includes the account, identity details, public key, and private key signature for identity verification and registration. After being verified by the TA, the TA will distribute an identity authentication certificate to the CAV for subsequent communication encryption and identity verification. After initialization is completed, the CAV can start collecting point cloud data as the basic data source for generating the global map in the future;

[0120] ② Consensus node selection: The blockchain system will calculate the reputation values of each RSU according to the RSU reputation model and select the RSUs with a reputation exceeding the threshold as consensus nodes.

[0121] ③ CAV and RSU upload pre - processed information: The CAV and RSU convert the coordinates of the collected point cloud data into global coordinates, and upload the coordinate data to the central server through the RSU. The central server determines the grid to which the point cloud data of the CAV and RSU belongs according to the coordinate information, and marks the overlapping and non - overlapping areas of each data, forming the area division result.

[0122] ④ Return area division information: After completing the area division, the central server will promptly return the area division result to the relevant RSU and CAV. These division information clearly mark whether the point cloud data of each device is in the overlapping area or the non - overlapping area, providing clear guidance for subsequent data processing and transmission.

[0123] ⑤ CAV uploads data to RSU: The CAV further processes the perception data. For the data in the non - overlapping area, the raw point cloud data is directly uploaded to the RSU for generating the global map. For the data in the overlapping area, the CAV preferentially selects local processing. If the computing resources are insufficient, that is, the time for the CAV to process the perception data itself is significantly longer than the time for offloading and processing the perception data, the CAV can also offload the data to the MES equipped on the RSU for processing, and then the RSU uploads it.

[0124] ⑥ Data verification: After the CAV uploads the perception data to the RSU, the blockchain nodes (the central server and the RSU selected as the consensus nodes) verify the perception data according to the PBFT consensus mechanism.

[0125] ⑦ RSU processes perception data and receives data uploaded by CAV: On the one hand, the RSU has to receive the perception data from the CAV, which may include the raw data and the feature data preliminarily processed by the CAV. On the other hand, since the RSU is equipped with MES and has strong computing power, the RSU will choose to process the raw perception data uploaded by the CAV and its own collected perception data by itself.

[0126] ⑧ RSU uploads the processed feature data: The RSU uploads the processed perception data (including its own data and CAV data) to the central server.

[0127] ⑨ Return the global map: The central server receives the processed data from all RSUs, and fuses the raw perception data in the non - overlapping area and the feature data in the overlapping area to generate a global high - precision map. After the global map is generated, it will be distributed through hierarchical broadcasting. The first layer is distributed by the central server to the RSU to ensure the update of the local map in the scene; the second layer, the RSU distributes the received global map to the CAVs within its coverage area to ensure that each CAV obtains the updated global map in real time.

[0128] S4: Propose a resource allocation objective that jointly optimizes energy consumption and latency, and comprehensively considers the optimization requirements of energy consumption and latency.

[0129] (1) Blockchain computing model: Assume that all nodes require α1 and α2 CPU cycles respectively to generate or verify a signature and a MAC. Since the CAV will communicate with the RSU regardless of whether it chooses to offload the processing of sensed data, in the process of generating a global map once, each CAV is regarded as a transaction, and a total of I transactions need to be verified. Denote the computing resource allocation strategy of the MES equipped with the RSU for the blockchain system as F1(t) = {f j,b (t)|j ∈ J}. Since the consensus process occurs before the sensed data processing model, it can be considered that the computing resources required for this process do not conflict with the computing resources required for sensed data processing. f j,b (t) represents the computing resources allocated to the blockchain by RSU j.

[0130] The PBFT consensus algorithm includes 5 stages, which are respectively:

[0131] ① Request stage: Before the CAV transmits the data to the blockchain node for verification, the data needs to be encrypted. Then the CAV sends the request to the primary node (central server), indicating the start of the request and entering the consensus process. When the primary node receives the message from the CAV, the primary node needs to verify the signature and MAC when receiving the transaction, and then generate 1 signature and J - 1 MACs and broadcast them to other nodes. The computing cost and latency of the primary node can be expressed as:

[0132] c b1 (t) = I(α1 + α2) + α1 + α2(J - 1)

[0133]

[0134] In the formula, c b1 (t) is the computing cost of the primary node, t b1 (t) is the latency of the primary node, f c,b (t) represents the computing resources provided by the central server for the blockchain

[0135] ② Pre - prepare stage: When the replica node receives the pre - prepare message from the primary node, it needs to continue to verify the signature and MAC of the transaction. The computing cost and latency of the replica node can be expressed as:

[0136] c b2 (t) = (I + 1)(α1 + α2)

[0137]

[0138] In the formula, c b2(t) is the computing cost of the replica node in the pre-preparation stage, t b2 (t) is the latency of the replica node in the pre-preparation stage.

[0139] ③ Preparation stage: After the replica node receives the pre-preparation message broadcast by the primary node, the replica node needs to verify the signature and MAC of the transaction. It should be noted that in the preparation stage, if more than a certain number 2f (f is the number of Byzantine nodes that can be tolerated, f = (J - 1) / 3) of the same requests are received, the confirmation stage is entered. Therefore, in the preparation stage, the replica node only needs to verify the 2f signatures and MACs from other nodes, and then generate 1 signature and J - 1 MACs to broadcast to other nodes. In this process, the computing cost and latency of each replica node can be expressed as:

[0140] c b3 (t) = 2f(α1 + α2) + α1 + α2(J - 1)

[0141]

[0142] In the formula, c b3 (t) is the computing cost of the replica node in the pre-preparation stage, t b3 (t) is the latency of the replica node in the pre-preparation stage.

[0143] ④ Confirmation stage: In this stage, messages need to be exchanged between all nodes. Similarly, each node verifies the 2f signatures and MACs of other nodes, and then generates 1 signature and J - 1 MACs to broadcast to other nodes. In this process, the computing cost c of a single node b4 = c b3 , and the latency cost can be expressed as:

[0144]

[0145] ⑤ Reply stage: When each node collects 2f verification success messages, the new block becomes a valid block and is then broadcast to the blockchain system. In this process, the computing cost and latency cost can be expressed as:

[0146] c b5 (t) = I(α1 + α2)

[0147]

[0148] In the formula, c b5 (t) is the computing cost of a single node in the reply stage, t b5 (t) is the maximum latency among all nodes in the reply stage.

[0149] Based on the above analysis, the total latency cost and total energy consumption cost for completing tasks in the blockchain system can be expressed as:

[0150] T b (t)=t b1 (t)+t b2 (t)+t b3 (t)+t b4 (t)+t b5 (t)

[0151] Wherein, t b1 (t), t b2 (t), t b3 (t), t b4 (t), t b5 (t) respectively represent the delays in the five stages of the PBFT consensus algorithm.

[0152]

[0153] Wherein, k c is the effective energy coefficient, f c,b (t) represents the computing resources provided by the central server for the blockchain, c b1 (t), c b4 (t), c b5 (t) respectively represent the computing costs in the corresponding stages of the PBFT consensus algorithm, f j,b (t) represents the computing resources provided by MES j for the blockchain.

[0154] (2) CAV-side perception data processing model: When processing the data in the overlapping area, if the CAV processes it locally, the considered delay can be expressed as:

[0155]

[0156] Wherein, d represents the number of CPU cycles required to process one bit of data, represents the data volume in the overlapping area, f i,i (t) represents the computing resources used by the CAV itself for perception data processing.

[0157] During the process of executing the computing task in the CAV, the energy consumption cost of the CAV can be expressed as:

[0158]

[0159] If the CAV chooses to offload the data in the overlapping area to the RSU for processing, the considered delay can be expressed as:

[0160]

[0161] r i,j (t)=B i (t)log2(1 + SINRi,j (t))

[0162]

[0163] L(D i,j (t)) = 28.0 + 22log 10 (D i,j (t)) + 20log 10 (f i (t))

[0164] D i,j (t) = [(d i,j (t)) 2 +(h r -h v ) 2 0.5

[0165] In the formula, r i,j (t) represents the transmission rate from CAV to RSU, B i (t) represents the bandwidth allocated by the vehicle networking cooperative perception system to CAVi, SINR i,j (t) is the signal-to-noise ratio, p i (t) represents the transmission power of CAV i, h i,j (t) represents the channel gain from CAV i to RSU j, N0 represents the noise power spectral density, L(D i,j (t)) represents the path loss between CAV i and RSU j, D i,j (t) is the Euclidean distance, f i (t) represents the subcarrier frequency. f j,i (t) represents the computing resources allocated by RSUj to CAVi, d i,j (t) represents the horizontal distance from CAVi to RSUj, h r 、h v respectively represent the antenna heights of MES and CAV;

[0166] The CAV selects to offload the data in the overlapping area to the RSU for processing. The considered energy consumption is the energy consumption of the RSU connected to the MES and can be expressed as:

[0167]

[0168] When the CAV processes the data in the non-overlapping area, the considered delay can be calculated as:

[0169]

[0170] In the formula, represents the data volume in the non-overlapping area. ​

[0171] The considered energy consumption can be expressed as:

[0172]

[0173] Let β(t) = {β1(t), β2(t),..., β i (t),..., β I (t)} represent the offloading decision of the CAV, where β i (t) = 1 indicates offloading, and β i (t) = 0 indicates local processing.

[0174] Therefore, the total delay for CAV i to process (including CAV i and MES j) can be expressed as:

[0175] T i (t) = β i (t)T i mes-ol (t) + (1 - β i (t))T i V-ol (t) + T i nol (t)

[0176] The energy consumption required for CAV i to process is expressed as:

[0177]

[0178] The energy consumption of MES j is expressed as:

[0179]

[0180] The delay for the CAV side to process the sensed task data as a whole is expressed as:

[0181] T CAV (t) = max{T1(t), T2(t),..., T I (t)}

[0182] In addition, the number of CPU cycles required for the data of CAV i offloaded to MES j is expressed as The number of CPU cycles required for the data volume of all CAVs offloaded to MES j is expressed as:

[0183]

[0184] (3) RSU-side sensed data processing model: For the RSU, it is equipped with a MES. Therefore, the data processing delay of the RSU for the overlapping area can be expressed as:

[0185]

[0186] In the formula, represents the amount of data in the overlapping area, and f j,j (t) represents the computing resources used by the RSU itself for sensing data processing.

[0187] During the process of the MES system executing a computing task, the energy consumption cost is considered as the energy consumption of the MES for processing the data collected by the RSU itself, and can be expressed as:

[0188]

[0189] The data processing delay of the RSU for the non-overlapping area includes data upload and data processing time, and can be expressed as:

[0190]

[0191] In the formula, r j,c (t) represents the transmission rate from the RSU to the central server; f c,j (t) represents the computing resources provided by the central server for RSU j.

[0192] In summary, when RSU j processes sensing data, the total delay and total energy consumption can be expressed as:

[0193]

[0194] Among them, represents the energy consumption of the central server for processing the sensing data of RSU j;

[0195] The overall delay of the RSU side for processing sensing data is expressed as:

[0196] T RSU (t) = max{T1(t), T2(t),..., T J (t)}

[0197] Then the total delay of the vehicle-to-everything (V2X) collaborative sensing system can be expressed as:

[0198] T total (t) = max{T CAV (t), T RSU (t)} + T b (t)

[0199] The total energy consumption consumed by the V2X collaborative sensing system is expressed as:

[0200]

[0201] (4) Construct a joint optimization objective function to comprehensively balance the computing energy consumption, communication energy consumption, blockchain consensus energy consumption, and total latency of the sensing task, so as to simultaneously meet the real-time requirements of the sensing task and the energy efficiency optimization needs. The optimization objective function is expressed as follows:

[0202]

[0203] In the formula, F1(t) and F2(t) respectively represent the computing resource allocation strategies of the MES for the consensus process and the sensing data processing process; F3(t) represents the computing resource allocation strategy of the CAV for the sensing data processing process; β(t) represents the CAV offloading decision; α∈[0,1] represents the adjustment factor of the system cost, which can adjust the system latency cost and the system energy consumption cost; f c (t) represents the computing resources of the central server.

[0204] The constraint condition C1 means that the total latency of generating the global map cannot exceed the maximum tolerable latency of the system; the constraint condition C2 means that the total energy consumption of generating the global map cannot exceed the maximum tolerable energy consumption of the system; the constraint condition C3 means that the bandwidth allocated to the CAV cannot exceed the total system bandwidth; the constraint condition C4 means that the computing resources used by the CAV to process the sensing data cannot exceed its maximum available computing resources; the constraint condition C5 means that when the MES equipped with the RSU j processes the sensing data of the RSU itself and the sensing data offloaded by the CAV, the computing resources used cannot exceed the maximum computing resources of the MES; the constraint condition C6 means that the computing resources used for the blockchain cannot exceed the maximum computing resources of the MES. Since the consensus process occurs before the sensing data processing, the computing resources used for the blockchain do not conflict with the sensing data processing; the constraint condition C7 means that the computing resources of the central server used to process the sensing data uploaded by the RSU cannot exceed the maximum computing resources of the central server.

[0205] S5: To solve the optimization problem constructed in step S4, an Attention mechanism-Multi agent proximal strategy optimization (A-MAPPO) algorithm is proposed to learn the resource allocation scheme of the weighted sum of system latency and energy consumption; among them, the network structure of the A-MADDPG algorithm is as Figure 3 shown.

[0206] Specifically, the optimization problem can be modeled as a Markov decision process, where the agent set is the sensing device Q. The state space is S, and the action space is A={a1×a2…a q ×…a I+J}。The agent can improve its policy by interacting with the environment. At each step t, each agent q obtains the current observation o from the global environmental state q (t), takes an action a q (t), then obtains a reward r q (t), and the environment transfers to a new state s(t + 1).

[0207] This Markov decision process involves two types of agents, namely the MES equipped with CAV and RSU, and their sets are I and J. At the beginning of each step t, the CAV requests offloading according to the perception data situation. After that, the MES obtains the action and serves to allocate computing resources for the CAV and the connected RSU.

[0208] The elements of the Markov decision process of CAV include the observation o i (t), the action a i (t) and the reward r i (t), which are specifically as follows:

[0209] Observation: The CAV can obtain its own perceived data volume L i (t), the data volume in the overlapping area the data volume in the non - overlapping area and the computing resource f i (t). In addition, since the perception data processing task load of the MES is highly correlated with the offloading decision of the CAV, the computing load Y i,j (t - 1) of the MES is also taken into consideration. The observation of CAV i can be expressed as:

[0210]

[0211] Action: The action of the CAV involves the offloading decision β i (t) and the computing resource f i,i (t) for perception data processing. Therefore, for CAV i, the action is decomposed as:

[0212] a i (t) = {β i (t), f i,i (t)}

[0213] Reward: The CAV needs to consider its own impact on the total weighted delay and energy consumption and the load of the MES to which it offloads. Therefore, the reward of each CAV i should include the energy consumption and delay of CAV i itself and the energy consumption of the MES to which it offloads. Therefore, the reward r i (t) of each CAV i is expressed as r i (t) = -αmax{T CAV (t), TRSU (t)}-(1 - α)(E i (t)+E j,i (t))}, E i (t) represents the energy consumption of CAVi processing its own perceived data, and E j,i (t) represents the energy consumption of MESj processing the perceived data of CAVi;

[0214] The elements of the Markov decision process of MES include the observation o j (t), the action a j (t) and the reward r j (t), which are specifically as follows:

[0215] Observation: MES can obtain the amount of perceived data L j (t) of the RSU, the amount of data in the overlapping area, the amount of data in the non - overlapping area and its own computing resources f j (t). In addition, MES can also obtain the perceived data task information from CAVs, and its observation can be expressed as:

[0216]

[0217] Action: MES first needs to complete the consensus process. Then, after receiving a CAV request, MES needs to adjust the computing resources allocated to the RSU and CAV according to the perceived tasks of the connected RSU, the perceived tasks offloaded by the CAV, and its own computing resources. Therefore, the action of MES is as follows:

[0218] a j (t) = {f j,i (t), f j,j (t)}

[0219] Reward: MES needs to consider the consensus time, the completion time of the RSU perceived task, and its own consumption, that is

[0220]

[0221] In the formula, E j,j (t) represents the energy consumption of MESj processing its own perceived data, α represents the adjustment factor of the system cost, f j,b (t) represents the computing resources allocated by MES j to the blockchain, c b2 (t)~c b5 (t) respectively represent the energy consumption in the pre - prepare stage, prepare stage, confirm stage, and reply stage of the PBFT consensus algorithm, and T total (t) represents the total delay of the vehicle - to - everything perception task.

[0222] The agent repeatedly inputs the observation o q (t) into the actor network to obtain a q (t) and r q (t), and stores the experience in the buffer. Then, at the end of an episode, the agent updates their policy. First, they sample some experience batches from their respective buffers where pr q (t) represents the log probability of the sampled action a q (t). Then, in each update, the actor and critic update their parameters using the policy loss and the global state value loss respectively. The loss of actor q is calculated as:

[0223]

[0224] In the formula, and represent the old policy and the current policy respectively; is the advantage function A q (t) = Q q (s(t), a q (t)) - V q (s(t)) is the estimated value, Q q (s(t), a q (t)) represents the value function of performing the action a q (t) in the state s(t); ∈ is a hyperparameter. In A-MAPPO, Generalized Advantage Estimation (GAE) is adopted to improve the performance, and its definition is as follows:

[0225]

[0226] In the formula, γ is the discount factor; λ is the GAE parameter used to balance bias and variance in the estimation, and l represents the offset of the time step. V q (s(t)) is the cumulative discounted reward and also represents the state value function, represents the state value function of agent q in the state s(t + l).

[0227] The state value function estimated by the critic of agent q is The loss of critic q is:

[0228]

[0229] In the formula, ξ q is the parameter of the q-th critic network.

[0230] After introducing the attention mechanism, the observation vectors of each agent are first subjected to feature extraction through their respective 3-layer MLP to obtain the eigenvalue e q . Then, the eigenvalues of all agents are sent to the attention head, and the attention value x is obtained through the following formula q :

[0231]

[0232] where e q′ is the eigenvalue of agent q'; d key is variance. The matrix W key converts e q′ to a key, and the matrix W v converts e q to a query. W q′ represents the weight matrix. Finally, x q and o q (t) are sent to the MLP to obtain the estimated state value

[0233] The A-MAPPO algorithm described above specifically includes the following steps:

[0234] 1. Initialize the actor network parameters, critic network parameters, and experience pool for each intelligent agent CAV and MES;

[0235] 2. Start iteration. In each iteration, the CAV and MES obtain the current observation from the environment and execute the actions given by the actor network;

[0236] 3. Calculate the log probability of each CAV and MES executing actions at the current time step;

[0237] 4. Synchronize the observations and actions of each CAV and MES to the central server;

[0238] 5. The central server conducts reward evaluation, and each CAV and MES receives the returned evaluation rewards;

[0239] 6. Store the relevant information in the experience area;

[0240] 7. Update the actor network parameters and critic network parameters according to the information in the experience area and the corresponding formula;

[0241] 8. Download the updated actor network parameters to the CAV and MES;

[0242] 9. If the iteration is not completed, return to step 2, otherwise end.

[0243] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the purpose and scope of the present technical solution, and they should all be covered within the scope of the claims of the present invention.

Claims

1. A vehicle networking collaborative perception resource allocation method based on region and reputation driving, characterized in that The method includes: S1. In a vehicle networking scenario including CAV (Connected Automatic Vehicles), RSU, MES (Mobile Edge Server), TA (Trusted Authority), and a central server, divide the global perception areas of CAV and RSU into overlapping areas and non-overlapping areas; the CAV collects surrounding point cloud data and converts it to the global coordinate system, and uploads data according to the area characteristics. For the overlapping areas, upload the processed feature data, and for the non-overlapping areas, upload the original point cloud data to the MES for fusion; the RSU is used to collect surrounding environment information, expand the CAV perception range, and is selected as a blockchain consensus node under specific conditions to help process perception data and participate in blockchain verification; the MES is connected to the RSU acting as a consensus node one by one to provide computing services; the TA is responsible for entity identity authentication and registration, issues certificates using encryption technology, and collaborates with the central server to update certificate information to ensure the security and stability of the vehicle networking; the central server dynamically updates the global map according to the uploaded data, broadcasts the latest perception results through grid management, and stores the information related to RSU and CAV. In addition, the central server participates in the calculation process of the blockchain as the main blockchain consensus node; S2. Comprehensively analyze the current status and historical behavior data of the RSU, and at the same time construct a reputation model of the RSU. This reputation model dynamically evaluates the reputation value of the RSU through direct reputation value and indirect reputation value, and filters out the RSU with a reputation value greater than the threshold as the blockchain consensus node; S3. Based on the blockchain, propose a dynamic update mechanism for the global map based on grid. The central server generates a global high-precision map according to the perception data uploaded by the RSU, and then distributes it to the RSU and CAV through hierarchical broadcasting; S4. Construct an energy consumption and delay calculation model for the blockchain, construct an energy consumption and delay calculation model for processing perception data on the CAV side, and construct an energy consumption and delay calculation model for processing perception data on the RSU side; comprehensively consider the above energy consumption and delay calculation models, construct a joint optimization objective function for energy consumption and delay, and comprehensively balance the total energy consumption and total delay of the vehicle networking perception task to simultaneously meet the real-time requirements of the perception task and the energy efficiency optimization requirements; S5. Through the multi-agent proximal policy optimization algorithm based on the attention mechanism, learn to obtain a resource allocation scheme for the weighted sum of delay and energy consumption.

2. The method according to claim 1, wherein In step S2, the reputation model is expressed as: Where C r (t) represents the reputation value of RSU r, and ω2 represents the reputation value adjustment weight. represents the direct reputation value of RSU r. represents the indirect reputation value of RSU r.

3. The method according to claim 2, characterized in that, The direct reputation value is reflected in the consensus success rate and reputation growth rate demonstrated by the RSU in the past consensus process; The consensus success rate is an indicator to measure the completion of the RSU in the consensus link. The factors affecting the consensus success rate of the RSU include the number of times of failure to complete the consensus due to equipment failure or being attacked; the consensus success rate is expressed as: where, su r (t) represents the number of times that RSU r has successfully completed consensus cumulatively, and fa r (t) represents the number of times that RSU r has unsuccessfully completed consensus cumulatively; The reputation growth rate is used to measure the stability of the RSU node and reflects the reputation change trend of the RSU; the reputation growth rate is expressed as: Wherein, SU r (t - 1) represents the reputation value of RSU r in the (t - 1)-th round. If the current is the 1st round, then SU r (t - 1) = SU r (t); The direct reputation value is expressed as: In the formula, ω1 represents the direct trust adjustment weight, which can adjust the ratio of the consensus success rate and the reputation growth rate; The indirect reputation value is the evaluation of RSU r by other RSUs in the blockchain system. To prevent malicious nodes from making abnormal evaluations, the evaluation of RSU r by other RSUs needs to refer to the direct reputation value of the RSU making the evaluation at the same time. Let R r′ (t) represent the credibility of the RSU r' making the evaluation, and R r′ (t) is expressed as: In the formula, represents the direct credibility value of the RSU r′ in the (t - 1)-th round; D max and D min represent the threshold of credibility. Exceeding D max means that the RSU r′ is completely credible. Being lower than D min means that the RSU r′ is completely non-credible. Then the indirect reputation value of RSU r is expressed as: In the formula, represents the evaluation of RSU r′ on RSU r, and R represents the total number of RSUs.

4. The method according to claim 1, wherein In step S3, the grid-based global map dynamic update mechanism includes: S31. Each CAV entering the cooperative perception system needs to submit account information to the TA for identity verification and registration; after the TA verifies and passes, it distributes an identity authentication certificate to the CAV; after the initialization is completed, the CAV starts to collect point cloud data; S32. The blockchain selects the RSU with a reputation value exceeding the threshold as a consensus node according to the reputation values of each RSU calculated by the reputation model; S33. The CAV and the RSU convert the collected point cloud data coordinates into global coordinates, and upload the coordinate data to the central server through the RSU. The central server determines the grid to which the point cloud data of the CAV and the RSU belongs according to the coordinate information, and marks the overlapping area and non-overlapping area of each data to form a regional division result; S34. After completing the regional division, the central server returns the regional division result to the relevant RSU and CAV; S35. The CAV further processes the perception data: for the data in the non-overlapping area, directly upload the original point cloud data to the RSU for generating the global map; for the data in the overlapping area, the CAV preferentially selects local processing. If the CAV's computing resources are insufficient, the CAV selects to offload the data to the MES equipped by the RSU for processing, and then the RSU uploads it; S36. After the CAV uploads the perception data to the RSU, the blockchain node verifies the perception data according to the PBFT consensus mechanism; S37. The RSU processes the perception data from the CAV and simultaneously processes the perception data collected by itself; S38. The RSU uploads the processed perception data to the central server; S39. The central server receives the processed data from all RSUs, fuses the original perception data in the non-overlapping area and the feature data in the overlapping area to generate a global high-precision map, and then distributes it to the CAV and the RSU through hierarchical broadcasting.

5. The method according to claim 1 or 4, characterized in that, The central processor distributes the global high-precision map to the CAV and the RSU through hierarchical broadcasting. The first layer is distributed from the central server to the RSU to ensure the update of the local map in the scene; the second layer is that the RSU distributes the received global map to the CAVs within its coverage area to ensure that each CAV obtains the updated global map in real time.

6. The method according to claim 1, wherein In step S4, the constructed energy consumption and delay calculation model of the blockchain is expressed as: T b (t) = t b1 (t) + t b2 (t) + t b3 (t) + t b4 (t) + t b5 (t) where, T b (t) represents the delay in completing the consensus process in the blockchain, and t b1 (t) ~ t b5 (t) respectively represent the delays in the request phase, pre-prepare phase, prepare phase, confirm phase, and reply phase of the PBFT consensus algorithm when completing the consensus process in the blockchain based on the PBFT consensus algorithm; E b (t) represents the energy consumption in completing the consensus process in the blockchain, and k c represents the effective energy coefficient, and f c,b (t) represents the computing resources provided by the central server for the blockchain, and c b1 (t) ~ c b5 (t) respectively represent the energy consumptions in the request phase, pre-prepare phase, prepare phase, confirm phase, and reply phase of the PBFT consensus algorithm; The constructed energy consumption and delay calculation model for the CAV side to process perception data is expressed as: T CAV (t) = max{T1(t), T2(t),..., T I (t)} where, T CAV (t) represents the delay of the overall processing of the perception data on the CAV side, and T i (t) represents the delay of the perception data processed by CAV i, where i = 1, 2, …, I and I is the total number of CAVs; β i (t) represents the decision of whether CAV i unloads the perception data to the RSU; T i V-ol (t) represents the latency of the CAV i for processing the data in the local processing overlapping region, T i mes-ol (t) represents the latency of the CAV i for selecting to offload the overlapping region data to the RSU for processing, T i nol (t) represents the latency of the CAV i for processing the non - overlapping region data; E j,i (t) represents the energy consumption of the MES j for processing the perception data, represents the energy consumption of the CAV i for selecting to offload the overlapping region data to the RSU j for processing, represents the energy consumption of the CAV i for processing the non - overlapping region data; The constructed energy consumption and delay calculation model for the RSU side to process perception data is expressed as: T RSU (t) = max{T1(t), T2(t),..., T J (t)} Where, T RSU (t) represents the latency of the RSU side to process the perceived data, T j (t) represents the RSU j to process the perceived data, and respectively represent the latency of the RSU j to process the data in the overlapping area and the non - overlapping area; E j (t) represents the energy consumption of the RSU j to process the perceived data, represents the energy consumption of the RSU j to process the perceived data, represents the energy consumption of the central server to process the perceived data of the RSU j; Then the total energy consumption and total delay of the vehicle networking perception task are respectively expressed as: T total (t) = max{T CAV (t), T RSU (t)} + T b (t) In the formula, J represents the total number of RSUs serving as consensus nodes, that is, the total number of MESs.

7. The method according to claim 6, wherein In step S4, based on the total energy consumption and total delay of the vehicle networking perception task, a joint optimization objective function of energy consumption and delay is constructed: s.t.C1:T total (t)≤T max C2:E total (t) ≤ E max C4:f i,i (t) ≤ f i (t) C6:f j,b (t) ≤ f j (t) Where, F1(t) and F2(t) respectively represent the computational resource allocation strategies of MES for the consensus process and the perceived data processing process; F3(t) represents the computational resource allocation strategy of CAV for the perceived data processing process; β(t) represents the CAV offloading decision; α∈[0,1] represents the adjustment factor of the collaborative perception system cost, which is used to adjust the system delay cost and the system energy consumption cost; The constraint condition C1 indicates that the delay in generating the overall global map cannot exceed the system's maximum tolerable delay T max ; The constraint condition C2 indicates that the energy consumption in generating the overall global map cannot exceed the system's maximum tolerable energy consumption E max ; The constraint condition C3 indicates the bandwidth B i (t) allocated to CAV i cannot exceed the total system bandwidth B max ; The constraint condition C4 indicates that the computing resource f i,i (t) used by CAV i to process the sensed data cannot exceed its maximum available computing resource f i (t); The constraint condition C5 indicates that when the MES equipped in RSU j is used to process the sensed data of RSU j itself and the sensed data offloaded by CAV i, the computing resource used cannot exceed the maximum computing resource f j (t), f j,j (t) represents the computing resource used by RSU j to process its own sensed data, and f j,i (t) is the computing resource allocated by MES j for processing the sensed data from CAV i; The constraint condition C6 indicates that the computing resource f j,b (t) allocated by MES j to the blockchain cannot exceed the maximum computing resource f j (t) of MES j; The constraint condition C7 indicates that the computing resource of the central server used to process the sensed data uploaded by the RSU cannot exceed the maximum computing resource f c (t), f c,j (t) represents the computing resource allocated by the central server for processing the sensed data from RSU j.

8. The method according to claim 1, characterized in that, Step S5 includes modeling the optimization problem constructed in step S4 as a Markov decision process, and the agent set of this Markov decision process is the sensing device Q; The Markov decision process includes: First, the agent q obtains the current observation from the environment and executes the action given by the actor network, synchronizes the observation and action of the agent q to the central server, and after the central server performs the reward evaluation, the agent q receives the returned evaluation reward; stores the relevant information in the experience area; Secondly, the agent q samples some batches of experience from its own buffer pr q (t) represents the sampled action a q (t)'s log probability, s(t) represents the state; Then, in each update, the actor and the critic respectively use the policy loss and the global state value loss to update the actor network parameters and the critic network parameters; among them, the loss of the actor q is expressed as: In the formula, and represent the old policy and the current policy respectively; represents the advantage function A q (t) = Q q (s(t), a q (t)) - V q (s(t)) is the estimated value, Q q (s(t), a q (t)) represents the value function of performing the action a q (t) in the state s(t); ∈ represents a hyperparameter; The state value function estimated by the critic of agent q is Then the loss of critic q is expressed as: where ξ q represents the parameters of the q-th critic network, and V q (s(t)) represents the cumulative discounted reward; After introducing the attention mechanism, the observation vectors of each agent first pass through their respective MLPs for feature extraction to obtain the feature value e q , and the feature values of all agents are sent to the attention head to obtain the attention value x through the following formula q : where, e q′ represents the eigenvalue of the agent q′; d key represents the variance of; the matrix W key converts e q′ into a key, and the matrix W v converts e q into a query; W q′ represents the weight matrix; x q and o q (t) are sent to the MLP to obtain the estimated state value Finally, download the updated parameters to the agent q, and judge whether the algorithm iteration is completed. If the iteration is not completed, iterate again, otherwise end the algorithm.

9. The method according to claim 8, wherein This Markov decision process includes two types of agents, CAV and MES; Among them, the elements of the Markov decision process of the intelligent agent CAV i include the observation o i (t), the action a i (t), and the reward r i (t); Observation o i is expressed as L i (t) represents the amount of data sensed by CAV i itself, represents the amount of data in the overlapping area, represents the amount of data in the non - overlapping area, f i (t) represents the computing resources of CAV i, Y i,j (t - 1) represents the computing load for MES j to process the offloaded data of CAV i; Action a i (t) is represented as a i (t) = {β i (t), f i,i (t)}, β i (t) represents the offloading decision of CAV i, f i,i (t) represents the computing resources used by CAVi itself for sensing data processing; Reward r i (t) is expressed as r i (t) = -α max{T CAV (t), T RSU (t)}-(1 - α)(E i (t)+E j,i (t))}, E i (t) represents the energy consumption of CAVi processing its own perception data, E j,i (t) represents the energy consumption of MESj processing CAVi perception data; The elements of the Markov decision process for MES j include the observation o j (t), the action a j (t), and the reward r j (t); Among them, the observation o j (t) is expressed as L j (t) represents the amount of sensed data of RSU j itself, represents the amount of data in the overlapping area, represents the amount of data in the non-overlapping area, f j (t) represents the computing resources of MES j itself; Action a j (t) is represented as a j (t) = {f j,i (t), f j,j (t)}, f j,i (t) and f j,j (t) respectively represent the computing resources allocated by MES j for processing the perception tasks of RSU j and the perception tasks of CAV i; Reward r j (t) is expressed as E j,j (t) represents the energy consumption of MESj to process its own sensed data, α represents the adjustment factor of the system cost, f j,b (t) represents the computing resources allocated by MES j to the blockchain, c b2 (t) ~ c b5 (t) respectively represent the energy consumption in the pre-prepare phase, prepare phase, confirm phase, and reply phase of the PBFT consensus algorithm, T total (t) represents the total delay of the vehicle network sensing task.