Trusted and efficient Internet of Vehicles data sharing method, system, equipment and medium

By designing a trust evaluation model and deep reinforcement learning algorithm in the Internet of Vehicles system, the problem of malicious vehicle attacks and insufficient utilization of spectrum resources is solved, and an efficient and trustworthy Internet of Vehicles data sharing environment is achieved.

CN120075759APending Publication Date: 2025-05-30XIDIAN UNIV

Patent Information

Application Number
CN202510234034.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-28
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The existing Internet of Vehicle data sharing solutions have problems of malicious vehicle attacks and insufficient utilization of spectrum resources, making it difficult to effectively identify malicious vehicles and efficiently utilize spectrum resources.

Method used

By designing a trust evaluation model based on the sharing behavior of vehicle data and vehicle data, and combining deep reinforcement learning algorithms to allocate communication resources, multi-dimensional evaluation of vehicle trust values ​​and optimized allocation of spectrum resources are realized.

Benefits of technology

Effectively identify and prevent attacks from malicious vehicles, improve the trust and communication rate of vehicle data sharing, and optimize the utilization efficiency of spectrum resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120075759A_ABST
    Figure CN120075759A_ABST
Patent Text Reader

Abstract

The invention discloses a credible and efficient Internet of Vehicles data sharing method, system and device and a medium. The method comprises the following steps: requesting for data sharing; data sharing authorization; performing communication resource allocation based on a D3QN algorithm; data sharing is carried out, and trust value updating is carried out based on a multi-dimensional trust value model; the system, the device and the medium are used for bearing and implementing the method. The method comprises the following steps: uploading an updated trust value of a vehicle and a current data sharing record to a block chain account book; by considering the multi-dimensional trust value of the vehicle and allocating the communication resources of the Internet of Vehicles based on the deep reinforcement learning algorithm, the trust degree and the communication rate of vehicle data sharing can be improved, and the method has the advantages of reliability and high efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of wireless communication technology, and particularly relates to a trustworthy and efficient vehicle networking data sharing method, system, device and medium. Background Art

[0002] With the rapid development of intelligent transportation systems, Vehicular Ad-hoc Networks (VANETs) play an increasingly important role in modern traffic management and autonomous driving. However, the VANET environment is highly dynamic and complex, posing great challenges to VANET data sharing. In a VANET system, vehicles need to frequently interact with Road Side Units (RSUs) and other vehicles to achieve functions such as cooperative perception, path optimization, and traffic management. Most traditional VANET data sharing schemes rely on centralized servers or cloud storage systems, which not only have the risk of single point of failure but also may lead to problems such as data privacy leakage, low management transparency, and insufficient security.

[0003] In recent years, the integration of blockchain technology and VANETs has become a research hotspot. The distributed ledger of blockchain can ensure the immutability and traceability of data, making the data sharing process more transparent and trustworthy. For example, existing research has explored decentralized data sharing architectures based on blockchain, which improve the robustness of the network while ensuring data security.

[0004] However, although blockchain technology can provide a decentralized and highly trustworthy solution for VANET data sharing, vehicles still face many challenges in data sharing. First, since vehicles participating in data sharing in VANETs have a certain degree of self-interest, some malicious vehicles may attack the data sharing system by sending false data, tampering with information, or refusing to provide data, thus endangering the security and trust of the entire system. Although existing work has designed the trust value of vehicles in combination with vehicle behavior to ensure the trust environment of VANETs, the consideration of vehicle communication behavior and multi-dimensional trust values is still lacking, and malicious vehicles cannot be detected more effectively.

[0005] Second, due to the dynamic nature of the VANET communication environment and the limited spectrum resources, existing data sharing schemes often cannot efficiently utilize spectrum resources, resulting in problems such as excessive communication overhead and resource waste. Although some schemes apply deep reinforcement learning algorithms to the communication resource allocation of VANETs, they still lack in complex scenarios with obstacles and relay assistance in VANETs.

[0006] In addition, in existing research (W.Sun, D.Yuan, E.G. and F. "Cluster-Based Radio Resource Management for D2D-Supported Safety-Critical V2X Communications," in IEEE Transactions on Wireless Communications, vol. 15, no. 4, pp. 2756-2769, April 2016, doi: 10.1109 / TWC.2015.2509978.) mostly considers the vehicle direct transmission communication scenario to solve the problems of limited spectrum resources and identifying malicious vehicles in vehicle-to-everything (V2X) networks and dealing with malicious attacks.

[0007] In summary, (1) The existing technologies only consider the impact of some vehicle behaviors on the vehicle trust value, while parameters such as vehicle communication behaviors and data quality are also important references for judging the vehicle trust value in data sharing. It is difficult to effectively identify malicious vehicles by only considering some social behaviors of vehicles. (2) The existing technologies only consider the vehicle spectrum sharing problem in the direct transmission scenario and cannot ensure the efficient communication of vehicles in more complex scenarios. Summary of the Invention

[0008] To overcome the above deficiencies of the existing technologies, the purpose of the present invention is to provide a trustworthy and efficient vehicle-to-everything (V2X) data sharing method, system, device and medium. By considering the multi-dimensional trust value of vehicles and allocating communication resources in the V2X network based on the deep reinforcement learning algorithm, the trust degree and communication rate of vehicle data sharing can be improved, and it has the advantages of reliability and efficiency.

[0009] To achieve the above purpose, the technical solution adopted by the present invention is:

[0010] A trustworthy and efficient vehicle-to-everything (V2X) data sharing method includes the following steps:

[0011] Step 1: The data requester V j finds the data provider V with the optimal trust value i and sends a data sharing request.

[0012] Step 2: The data provider V i accepts the sharing request and verifies the identity of the data requester V j for authorization.

[0013] Step 3: Use the D3QN algorithm for communication resource allocation: After the data provider V i confirms that the identity of the data requester V j meets the requirements, the data provider V iSend the relevant information of this data sharing to the nearby Road Side Unit (RSU), and request communication resources for data sharing; after receiving the request, the Road Side Unit (RSU) performs algorithm analysis of deep reinforcement learning based on the data sharing vehicle pairs and communication resources in the vehicle network, obtains the optimal allocation strategy that conforms to the communication resources in the vehicle network, and sends it to the corresponding data provider V i ;

[0014] Step 4: Data sharing and trust value update. The data provider V i After receiving the corresponding allocation strategy, it will start data sharing with the data requester V j After the data sharing is completed, the data requester V j The vehicle evaluates the trust of the data provider V i based on the data provided by the vehicle and the behavior of the data provider V i in the data sharing i ;

[0015] Step 5: Share the transaction record. After completing the data sharing transaction, the data requester V j broadcasts the transaction record and the trust value scoring result to the vehicle chain, and other vehicle consensus nodes are responsible for verifying and auditing the broadcast content. After blockchain consensus, the transaction record is permanently stored in the blockchain, and at the same time the blockchain ledger is synchronized to the Road Side Unit (RSU); the Road Side Unit (RSU) updates the reputation value of the corresponding vehicle based on the archived scoring information to improve the trust management and data sharing of the entire vehicle network system.

[0016] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0017] 1. By combining the data provided by the vehicle and the vehicle's behavior in data sharing, the present invention comprehensively calculates the direct trust value, indirect trust value, local trust value, and comprehensive trust value of the vehicle, and introduces a blacklist mechanism and a time forgetting factor, which can effectively identify malicious vehicles and prevent malicious vehicles from continuously doing evil.

[0018] 2. By constructing the spectrum sharing problem of vehicle network relay-assisted and direct connection transmission, based on the multi-agent D3QN deep reinforcement learning algorithm, the present invention obtains the optimal spectrum allocation strategy, which can effectively improve the transmission rate of vehicle data sharing in complex scenarios where relay-assisted and direct connection transmission coexist.

[0019] 3. In the existing research on vehicle network communication resource allocation, the complex scenario with relays is not considered. The present invention considers the influence of obstacles and distance in vehicle communication and introduces relays for cooperative transmission, optimizes the communication rate of the overall vehicle network system, and meets the communication needs of different vehicles.

[0020] 4. Existing research on vehicle trust values only considers the social behavior of vehicles and does not take into account various factors during vehicle data transmission. The present invention simultaneously considers parameters such as the social behavior and communication behavior of vehicles, establishes a multi-dimensional trust value for vehicles, further refines the trust level of vehicles, and effectively identifies malicious vehicles.

[0021] In summary, the present invention provides a trusted and efficient data sharing environment for vehicle data sharing by designing a trust evaluation model based on vehicle data and vehicle data sharing behavior and allocating communication resources in vehicle data sharing based on a deep reinforcement learning algorithm, enhancing the reliability and efficiency of vehicle data sharing. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] Figure 1 is a flowchart of the method of the present invention.

[0023] Figure 2 is a schematic diagram of a vehicle-to-vehicle (V2V) data sharing scenario provided by an embodiment of the present invention.

[0024] Figure 3 is a schematic diagram of a multi-agent model based on D3QN in an embodiment of the present invention.

[0025] Figure 4 is a schematic diagram of the convergence of the multi-agent model based on D3QN in an embodiment of the present invention.

[0026] Figure 5 is a schematic diagram of the performance analysis of the total channel capacity and data packet changes in a vehicle-to-everything (V2X) network in an embodiment of the present invention.

[0027] Figure 6 is a schematic diagram of the performance analysis of the data transmission success rate and load changes of the V2V link in a V2X network in an embodiment of the present invention.

[0028] Figure 7 is a schematic diagram of the analysis of the malicious vehicle detection rate under different trust thresholds in an embodiment of the present invention.

[0029] Figure 8 is a schematic diagram of the analysis of the malicious vehicle detection rate and the probability of malicious vehicles sending error messages in an embodiment of the present invention.

[0030] Figure 9 is a schematic diagram of the analysis of the vehicle trust value after being attacked by a switch in an embodiment of the present invention.

[0031] Figure 10 is a schematic diagram of the analysis of the vehicle trust value and the change of the vehicle channel capacity in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0032] The present invention will be further described below in conjunction with the accompanying drawings and specific embodiments.

[0033] In view of the problems of vulnerability to malicious vehicle attacks and insufficient utilization of spectrum resources in vehicle data sharing, this invention explores the design and optimization of a trusted data sharing scheme for the Internet of Vehicles based on blockchain.

[0034] Firstly, due to the diversity of vehicle users in the Internet of Vehicles, some malicious vehicles may launch attacks on the system by uploading false data, tampering with information, or refusing to participate in sharing, thus endangering the security and trust of the entire data sharing network. This invention designs multi-dimensional trust values for vehicles in view of their social and communication behaviors in vehicle data sharing, to more effectively identify malicious and passive vehicles and impose penalties.

[0035] Secondly, there are also problems with the utilization efficiency of spectrum resources in vehicle-to-vehicle (V2V) data sharing. Due to the highly dynamic communication environment of the Internet of Vehicles and the instability of communication links between vehicles, spectrum resources may be unreasonably allocated or inefficiently used during data sharing, resulting in increased communication overhead, data transmission delay, and a decline in the overall system performance. In view of the communication resource allocation problem with relay assistance, this invention designs a communication resource allocation scheme for the Internet of Vehicles based on the deep reinforcement learning D3QN algorithm, regarding the vehicle V2V link as an Agent, to optimize the communication rate of vehicles, meet the communication requirements of different types of vehicle links, and maximize the utilization of communication resources in the Internet of Vehicles. Specifically, when there are vehicles with data requirements in the Internet of Vehicles, the data requester vehicle will send a specific data request to the roadside unit (RSU). The RSU will find the corresponding data provider according to the blockchain ledger of the vehicle chain and arrange for the vehicles to share data.

[0036] As Figure 1 shown, a trusted and efficient vehicle-to-vehicle data sharing method includes the following steps:

[0037] Step 1: The data requester vehicle V j searches for the data provider vehicle V i with the optimal trust value and sends a data sharing request;

[0038] The specific method of Step 1 is as follows:

[0039] When searching for the data provider vehicle V i , the data requester vehicle V j first checks the latest block synchronized by the nearby roadside unit (RSU) and searches for the data of interest in the data index. When the data requester vehicle V j finds the data of interest, the data requester vehicle V j sends a request to the roadside unit (RSU) to query the trust value of the data provider vehicle V i ;

[0040] The data requester vehicle Vj Select a data provider V with the optimal trust value based on the trust value and data description of the queried data provider V i ; the data requester V i sends a data sharing request to the data provider V j ; i Send a data sharing request;

[0041] Step 2: The data provider V i accepts the sharing request and verifies the identity of the data requester V j for authorization;

[0042] The specific method of the said Step 2 is as follows:

[0043] Step 2: Data sharing authorization. When the data provider V i receives the sharing request, first evaluate the comprehensive trust value of the data requester V j . Based on this, the data provider V i decides whether to allow the data requester V j to obtain the original data. At the same time, to prevent malicious vehicles from frequently accessing the shared data of the same data provider V i and trying to reveal the actual identity of the data provider V i , the data provider V i first checks whether the data requester V j is on the blacklist. If the data provider V i has the data requester V j on its blacklist, the request will be rejected; in addition, if the comprehensive trust value of the data requester V j is lower than the sharing threshold φ, the data provider V i will also reject the request. The sharing threshold φ can be dynamically adjusted according to specific application scenarios or vehicle safety requirements to optimize the performance and security of the system.

[0044] After completing the above checks, the data provider V i will further verify whether the data requester V j has the behavior of frequent access. For this purpose, each vehicle data provider V i maintains an access list PIDvisited, which records the vehicle ID PID of the vehicles that have accessed the data provider V i in the past 10 minutes and the corresponding access timestamp Timestampvisited. If the ID of the data requester V j vehicle appears in this access list PIDvisited, it means that the vehicle has repeated access in a short period of time, and the data provider V i rejects the sharing request of this vehicle. If all verifications pass, the data provider Vi Generate a transaction and record the transaction in the relevant information of this data sharing;

[0045] Step 3: Use the D3QN algorithm for communication resource allocation: Among the data providers V i After confirming that the identity of the data requester V j meets the requirements, the data provider V i sends the relevant information of this data sharing to the nearby Road Side Unit (RSU) and requests the communication resources for data sharing;

[0046] After receiving the request, the Road Side Unit (RSU) performs algorithm analysis of deep reinforcement learning based on the current vehicle pairs for data sharing and communication resources in the vehicle network, obtains the optimal allocation strategy that conforms to the current communication resources in the vehicle network, and sends it to the corresponding data provider V i to achieve the optimal utilization of communication resources;

[0047] Furthermore, the process of using the D3QN algorithm for communication resource allocation in Step 3 is as follows:

[0048] Construct a dynamic vehicle network communication network model, which includes M available idle spectrum resources and K vehicle-to-vehicle (V2V) communication links. Vehicles travel in different directions on multiple lanes. In each service cycle, vehicles can be divided into three categories, namely data provider V i vehicles, data requester V j vehicles, and idle vehicles (i.e., selectable relay vehicles); Assume that all transceivers use single antennas, and use the sets M = {1,…, M} and K = {1,…, K} to represent the idle spectrum resources and K vehicle-to-vehicle (V2V) communication links in the dynamic vehicle network communication network respectively;

[0049] Step 3.1, construct a vehicle-to-vehicle (V2V) communication model, including a path loss model and a shadow fading model, and calculate the channel gain:

[0050] (1) Construct a path loss model: In the vehicle-to-vehicle (V2V) communication model, calculate the path loss through the line-of-sight path loss PL Los and the non-line-of-sight path loss PL NLos The specific formula calculation is as follows:

[0051] When the vehicle-to-vehicle (V2V) is line-of-sight, the line-of-sight path loss PL Los The formula is:

[0052]

[0053] Among them, in the path loss model, A 1 represents when the distance d ≤ d 0The path loss coefficient at a certain time, which reflects the propagation loss speed of millimeter-wave signals over short distances; A 2 is the path loss reference value at this time, which is related to the fixed factors of the propagation environment, B 1 is a coefficient related to frequency dependence, which reflects the influence of frequency on path loss, especially when the carrier frequency is f c When the unit is GHz, log10(fc / F 0 ) represents the contribution of path loss, where F 0 is the standardized frequency reference value. Usually, 5 GHz is taken as the reference frequency to standardize the influence of frequency; when the distance d > d0, A 3 As the path loss coefficient, A 3 is greater than A 1 , because the signal attenuation is more serious during long-distance propagation; A 4 is a constant used to compensate for the influence of other environmental factors on path loss, such as the attenuation of signals by buildings, terrain, etc. A 5 and A 6 are the attenuation coefficients of the transmitting vehicle antenna height h 0 and the receiving vehicle antenna height h 1 respectively, which represent the influence of antenna height on path loss. The higher the antenna height, the smaller the signal attenuation; B 2 is similar to B 1 , B 2 is used for frequency dependence compensation at long distances;

[0054] When the line-of-sight (LOS) between vehicle-to-vehicle (V2V) is blocked, the non-line-of-sight path loss PL NLos The formula is:

[0055]

[0056] In the non-line-of-sight (NLoS) path loss model, PL NLoS (d a , d b ) represents the path loss in the presence of obstacles. Its calculation formula is based on the line-of-sight (LoS) path loss PL NLoS (d a ) and additionally considers the influence of obstacles on signal propagation. Specifically, C 1 represents the reference additional loss of non-line-of-sight propagation, which reflects the additional attenuation of signals by obstacles; C 2 is multiplied by the path loss adjustment factor n j to quantify the influence of obstacle complexity on path loss; C 3 represents the additional influence of the propagation distance d b of the signal after passing through the obstacle on path loss. As the distance increases, the path loss increases significantly; C 4is the additional path loss coefficient related to frequency, which reflects the greater attenuation degree of high-frequency signals during propagation. In addition, F 0 is used as the reference frequency to standardize the dependence of path loss on frequency, and usually the reference frequency band commonly used in the communication field is selected.

[0057] (2) Construct the shadow fading model, which is expressed as: The shadow fading takes into account the large-scale occlusion effect in the environment.

[0058] Among them, Δd represents the change in the distance between vehicles, reflecting the change in the relative position of vehicles during movement; d dec is the decorrelation distance between vehicles, which is used to describe the significant reduction in signal correlation within this distance and is usually related to the complexity of the environment; is a standard normal distribution noise with a mean of 0 and a standard deviation of 3, which is used to simulate the randomness of shadow fading; represents the attenuation part of shadow fading, which gradually weakens as the distance between vehicles changes; represents the random change part of shadow fading, reflecting the impact of environmental noise on signal propagation;

[0059] (3) Calculate the channel gain: The orthogonal frequency division multiplexing (OFDM) technology is used to convert the frequency-selective wireless channel into multiple parallel flat channels. Several consecutive subcarriers are combined into a spectral subband. It is assumed that the channel fading within a subband is approximately the same, and the channel fading between different subbands is independent;

[0060] Within a coherence time period, the channel power gain g k [m] of the kth V2V link on the mth subband (occupied by the mth V2V link) follows the following formula:

[0061] g k [m] = α k h k [m],

[0062] Among them, h k [m] is the small-scale fading power component dependent on frequency, which is assumed to follow an exponential distribution with a mean of 1; α k captures the large-scale fading effect, including path loss and shadow fading, and is assumed to be independent of frequency;

[0063] Step 3.2, select the set of relay vehicles from the idle vehicles through the physical transmission distance and the comprehensive trust value:

[0064] A relay vehicle is a vehicle that is idle in the vehicle network communication network at a certain time. Since the selected relay vehicle needs to provide relay assistance services, multiple aspects should be considered when selecting a relay vehicle to minimize the interruption probability during the relay transmission process. In the present invention, when selecting the set of idle relay vehicles, both the trust value and the physical distance are considered respectively.

[0065] In terms of physical distance, in the actual vehicle networking scenario, when a vehicle pair selects a distant idle node as a relay, the probability of encountering obstacles during the signal transmission process will increase significantly, resulting in an increased risk of communication link interruption. At the same time, considering distant idle nodes will also significantly increase the detection times and the complexity of selection, which is not conducive to rapid decision-making. Therefore, it is necessary to screen the idle nodes in the vehicle network communication network according to the physical transmission distance at the initial stage. The screening range is centered on the data provider vehicle V i and the data requester vehicle V j to draw a circle with the straight-line distance between them as the diameter. The relay nodes within the area covered by this circle are the candidate nodes that meet the physical transmission distance constraint. For a vehicle link pair, the set of relay nodes that meet its physical transmission distance constraint is expressed as:

[0066]

[0067] where, represents the distance between the data provider V i and the data requester V j in the vehicle link pair; and respectively represent the coordinates of the data provider V i and the data requester V j nodes in the vehicle link pair; (x j , y j ) represents the position coordinates of any idle vehicle node n j in the vehicle network:

[0068] In terms of the comprehensive trust value, since the relay vehicle is also part of the vehicle link and there will also be a data transmission process, the update of its comprehensive trust value is the same as that of other directly connected data provider vehicles. On this basis, the present invention uses the comprehensive trust value threshold R TH to further screen the relay nodes in the relay node set , and the specific expression is as follows:

[0069]

[0070] where, n represents the relay node set The number of idle vehicles after the first round of screening in China, R i,j Is the trust value of the idle vehicle. If the trust value of the idle vehicle is greater than or equal to R TH , then the idle vehicle is retained in the relay node set ; Otherwise, delete the node;

[0071] After screening by physical transmission distance and comprehensive trust value, the final set of candidate relay nodes suitable is N r ;

[0072] Step 3.3, calculate the service priority of the vehicle link:

[0073] For different data requesters V j 's comprehensive trust value and the communication requirements of the vehicle link pair, the present invention designs a vehicle service based on priority, where Y i (t) represents the priority value of the i-th vehicle at time t;

[0074] On the one hand, according to the comprehensive trust value of the data provider V i in the vehicle link pair, determine the priority related to the comprehensive trust value, specifically:

[0075]

[0076] On the other hand, according to the communication requirements of the vehicle link pair, it is divided into a high-speed rate type for real-time transmission and a high-reliability type with a rate meeting the requirements and being stable;

[0077] The service priority of the vehicle link at high speed is higher than that of the vehicle link at high reliability, specifically:

[0078]

[0079] At the same time, if a vehicle has a record of providing relay cooperation transmission services for other vehicles recently, its service priority will be increased by 1 additionally as a relay incentive;

[0080] Finally, the service priority of each vehicle link is expressed as:

[0081] Y i (t) = Y(r) + Y(re) + ρ 0

[0082] Among them, ρ 0 represents whether there is a record of providing relay cooperation transmission services as a relay within the specified time interval of the vehicle. If so, it is 1, otherwise it is 0;

[0083] Step 3.4, based on the vehicle-to-vehicle (V2V) communication model constructed in Step 3.1, construct a channel transmission model:

[0084] As Figure 2 shown, in the vehicle - to - everything (V2X) vehicle data sharing scenario of the present invention, the roadside unit (RSU) first selects a dedicated or shared transmission link for the vehicle according to the service priority of the vehicle link calculated in step 3.3, and then selects whether relay assistance is required for the vehicle according to the vehicle's communication environment; therefore, the vehicle set in the V2X communication network is divided into four types, namely the vehicle set for direct dedicated transmission the vehicle set for direct multiplexed transmission the vehicle set for relay dedicated transmission the vehicle set for relay multiplexed transmission Assume that the total vehicle set is N V , then:

[0085]

[0086] In the relay - assisted transmission process, the present invention mainly adopts the amplify - and - forward (AF) relay strategy, aiming to reduce the computational complexity of relay vehicle decoding and reduce the delay in the relay cooperative transmission process. Based on this strategy, that is, the nth relay vehicle amplifies the signal it receives by β n,k times and forwards the received signal to the kth target vehicle. The entire relay - assisted transmission process can be divided into two stages: the downlink transmission stage and the relay forwarding stage; in the multi - vehicle V2X scenario, the receiving vehicle will be interfered by the synchronous transmission signals of other transmitting vehicles. Therefore, the present invention considers the interference effects of these two stages in the direct transmission process and the relay - assisted transmission process respectively. The specific calculation processes of the signal - to - noise ratio and the channel capacity are as follows:

[0087] (1) Direct dedicated transmission: For the direct dedicated transmission process of the kth receiving vehicle, it is only interfered by Gaussian white noise and not interfered by any other vehicle during the entire transmission process. Then the signal - to - noise ratio of its transmission is expressed as:

[0088]

[0089] The channel capacity is:

[0090]

[0091] Among them, is the transmission power of the k z th target vehicle on the mth sub - band, is the channel gain of the k z th target vehicle on the mth sub - band, σ 2 is the noise power, and W is the bandwidth;

[0092] (2) Relay - specific transmission: For the relay - specific transmission process of the \(k\) - th receiving vehicle, the transmission process includes a downlink transmission stage and a relay forwarding process. During the entire transmission process, it is only affected by additive white Gaussian noise and not interfered by any other vehicles. Then, the signal - to - noise ratio of its transmission is expressed as:

[0093] The downlink transmission stage is expressed as:

[0094]

[0095] The relay forwarding process is expressed as:

[0096]

[0097] Among them, are the transmit powers of the \(i\) - th target vehicle in the \(m\) - th sub - band during the downlink transmission stage and the \(j\) - th relay vehicle in the \(m\) - th sub - band during the relay forwarding stage respectively, are the channel gains of the two stages respectively, and \(\sigma^{2}\) is the noise power; Since there is relay - assisted transmission and the vehicle - to - vehicle link communication is divided into two stages, when calculating the channel capacity of the entire communication process, the 2 factor is introduced. The specific calculation is as follows:

[0098]

[0099]

[0100] (3) Direct - link multiplexing transmission: Due to the existence of relay vehicles, for the direct - link multiplexing transmission process of the \(k\) - th receiving vehicle, there are two stages and the two stages are interfered by the downlink transmission processes and relay forwarding processes from other vehicles. Specifically, for the first stage, the direct - link multiplexing vehicle - to - vehicle link is interfered by the downlink transmission processes of other direct - link transmission links and the downlink transmission process of the selected relay, which is expressed as:

[0101]

[0102] Among them, are the transmit powers of the downlink transmission stages of other direct - link vehicles and relay - assisted vehicles in the \(m\) - th sub - band respectively, are the corresponding channel gains; and are binary spectrum allocation indicator variables; when the value of the binary spectrum allocation indicator variable is 1, it means that the \(k\) - th V2V link uses the \(m\) - th sub - band, and when it is 0, it means not occupying. At the same time, it is assumed that each V2V link only accesses one sub - band, that is:

[0103] ​​

[0104] Then the signal-to-noise ratio SINR in the first stage is:

[0105]

[0106] For the second stage, the direct-reuse vehicle link is interfered by the downlink transmission process of other direct transmission links and the relay forwarding process of the relay forwarding link, which is specifically expressed as:

[0107]

[0108] The signal-to-noise ratio SINR in the second stage is:

[0109]

[0110] Based on the signal-to-noise ratios of the above two stages, when calculating the channel capacity, the factor is also introduced for processing:

[0111]

[0112] (4) Relay-reuse transmission: For the vehicle link with relay-reuse, the transmission process is also divided into two stages, namely the downlink transmission stage and the relay forwarding process, and the interferences received in the two stages are also different. For the downlink transmission stage, the vehicle link with relay-reuse is interfered by the downlink transmission process of other direct transmission links and the downlink transmission process of the selected relay, which is specifically expressed as:

[0113]

[0114] Among them, are the transmission powers of the downlink transmission stages of other direct vehicle and relay-assisted vehicle on the m-th sub-band respectively, are the corresponding channel gains respectively; then the signal-to-noise ratio SINR in the downlink transmission stage is:

[0115]

[0116] For the relay forwarding process, the vehicle link with relay-reuse is interfered by the downlink transmission process of other direct transmission links and the relay forwarding process of the relay forwarding link, which is specifically expressed as:

[0117]

[0118] Among them, are the transmission powers of the downlink transmission stages of other direct vehicle and relay-assisted vehicle on the m-th sub-band respectively; are the corresponding channel gains respectively; then the signal-to-noise ratio SINR of the relay forwarding process is:

[0119]

[0120] Then the final channel capacity is calculated as:

[0121]

[0122] 3.5, construct a joint optimization problem:

[0123] The optimization objective of the present invention is to meet the communication requirements of different vehicle link pairs in the vehicle-to-everything (V2X) network by optimizing the joint decision-making of spectrum resource allocation and relay assistance.

[0124] On the one hand, for link pairs with high data rate requirements, such as real-time video streaming, the optimization objective is to maximize the total channel capacity of high-data-rate links:

[0125]

[0126] On the other hand, for link pairs with high reliability requirements, such as periodic safety messages, the optimization objective is to reliably and stably transmit safety-critical messages within the time budget T; model the high-reliability link pairs as transmitting data packets of size B within each time budget, that is:

[0127]

[0128] where ΔT is the channel coherence time, and t represents different coherence time slots;

[0129] To solve the optimization problem in step three, a multi-agent collaborative communication system is built, and a multi-agent collaborative communication strategy based on D3QN is designed. Existing research on intelligent algorithms for V2X mainly focuses on the design of single-agent direct transmission. In the single-agent scenario, the agent does not need to model or predict the behavior of other agents in the environment. However, in a multi-agent environment, the agent needs to simultaneously learn the strategies of other agents, that is, the decisions and rewards obtained among multiple agents affect each other. Secondly, in the V2X scenario, the communication transmission between vehicles is affected by obstacles and distance, and relays need to be introduced to assist in transmission to ensure the stability of vehicle transmission. Therefore, the present invention builds a multi-agent-based collaborative communication system to improve the robustness and stability of vehicle communication transmission.

[0130] As Figure 3 shown, the present invention first Figure 2The vehicle - to - everything (V2X) vehicle data sharing scenario constructed is modeled as an environment. Secondly, the vehicle nodes in the V2X communication network are modeled as agents, which obtain experience by interacting with the environment. Multiple vehicle agents cooperate with each other to jointly optimize the communication rate of all vehicles in the V2X, meet the communication needs of different vehicles, and make decisions on communication resource allocation, so as to maximize the transmission rate of all legitimate vehicles. Specifically, at each time slot t, each agent first obtains the current state information s t ∈S, and selects the action a t ∈A according to the state information, where S and A represent the state space and the action space respectively. The actions of multiple agents act on the environment together and cause changes in the current environment, that is, enter the next environment state s t+1 ; At the same time, the environment evaluates the actions of each agent and feedbacks the reward r t back to the agent through the reward module; So far, each agent obtains the training sample (s t , a t , r t , s t+1 ), and stores it in the experience replay pool for the optimization training of the agent model; For each agent, the expression of its Markov chain is: {s t , a t , s t+1 , a t+1 ,...};

[0131] In an actual multi - agent scenario, if each agent does not know the information of other agents, large fluctuations may occur during the training process, affecting the convergence and stability of the model. Therefore, in order to enable multiple agents to cooperate with each other to jointly improve the overall performance of the system, the agent not only needs to consider local observation information, but also needs to consider the information of other agents globally. Multiple agents need to exchange information with the environment regularly to help each other make more reasonable decisions and jointly improve the overall performance of the system.

[0132] In a multi - agent cooperative communication system, the decision variables of legitimate vehicles are the decisions of spectrum allocation and relay selection, which are discrete decision variables; The deep reinforcement learning algorithm based on DQN uses a deep neural network to learn the mapping relationship between high - dimensional state and action spaces, and can effectively handle decision - making problems in discrete action spaces. However, the max operator in DQN uses the same value for action selection and action value estimation, which makes the DQN model more likely to select overestimated values, that is, the "over - estimation" problem. The present invention constructs a multi - agent model based on the D3QN algorithm as Figure 3 shown. The D3QN algorithm uses two different networks, a training network and a target network, for action selection and action evaluation, that is, uses the training network to obtain st+1 The optimal action in the state is determined, and the action value of this action is calculated using the target network. Through the interaction between the training network and the target network, the "overestimation" problem in the DQN algorithm is effectively avoided. The detailed multi-agent design is as follows.

[0133] 3.6, Construct the state and observation space:

[0134] In the resource sharing problem under the multi-agent reinforcement learning (multi-agent RL) framework, each V2V link k acts as an agent and explores the unknown environment simultaneously. Mathematically, the resource sharing problem is modeled as a Markov decision process (MDP). As Figure 3 shown, at each channel coherence time step t, given the current environmental state S t , each V2V agent k receives an observation of the environment This observation is determined by the observation function O, that is Then it takes an action to form a joint action A t ; After that, the agent receives a reward R t+1 , and the environment transfers to the next state S with probability p(s',r|s,a) t+1 ; Each agent receives a new observation All V2V agents share the same reward in the multi-agent cooperative communication system to encourage cooperative behavior among them.

[0135] The true environmental state S t , including the global channel condition and the behavior of all agents, is unknown to each individual V2V agent. Each V2V agent can only obtain knowledge of the environment through the perspective of the observation function; the observation space of V2V agent k contains local channel information, and the local channel information includes: for all m ∈ M, the channel gain g k [m] of V2V agent k itself; for all k′ = k and m ∈ M, the interference channel gain g k ' ,k [m] from other V2V transmitters; the channel information in the observation space is accurately estimated by the receiver of the k-th V2V link at the beginning of each time slot t, and it is assumed that the channel information in the observation space can also be obtained instantaneously at the transmitter through delay-free feedback. For all m ∈ M, the received interference power I k [m] on all frequency bands can be measured at the V2V receiver and is also introduced into the local observation; in addition, the local observation space includes the remaining V2V load B k and the remaining time budget T k to better capture the state of each V2V link. Therefore, the observation function of V2V agent k is expressed as:

[0136] O(S t ,k) = {B k ,T k ,{I k [m]} m∈M ,{G k [m]} m∈M},

[0137] Among them, G k [m] = {g k [m], g k',k [m]}.

[0138] 3.7, Construct the action space:

[0139] The resource sharing design of the vehicle link is reduced to the spectrum sub - band selection and transmission power control of the V2V link. Although the spectrum is naturally divided into M non - overlapping sub - bands, each sub - band is occupied by a V2V link, the V2V transmission power usually takes continuous values in most of the existing power control literature. In the present invention, the power control options are limited to four levels, namely [23, 10, 5, - 100] dBm; note that selecting - 100 dBm actually means zero V2V transmission power. Therefore, the dimension of the action space is 4×M, and each action corresponds to a combination of a specific spectrum sub - band and power selection;

[0140] 3.8, Design the reward:

[0141] In the resource sharing problem constructed in step 3.6 of the present invention, there are two objectives: maximizing the total capacity of the high - rate V2V links and simultaneously increasing the probability of successful transmission of the high - reliability V2V link load within a certain time constraint T;

[0142] The first objective: Take the instantaneous total capacity of all V2V links at each time step t:

[0143]

[0144] as the reward for each step;

[0145] The second objective: For each V2V agent k, set the reward L k to the effective V2V transmission rate until the load transmission is completed; afterwards, the reward is set to a constant β, and the constant β is greater than the maximum V2V transmission rate; therefore, the V2V - related reward at each time step t is set as:

[0146]

[0147] 3.9, Construct the learning objective:

[0148] The goal of learning is to find an optimal policy π* (a probability mapping from the state set S to the action set A) to maximize the expected return starting from any initial state s. The cumulative discounted reward is defined as the return Gt with a discount rate γ, that is:

[0149]

[0150] If the discount rate γ is set to 1, then a larger cumulative reward will translate into a high-reliability V2V link transmitting more data before the load transmission is completed. Therefore, maximizing the expected cumulative reward encourages transmitting more data for the V2V link when the remaining load is still non-zero (i.e., B k ≥ 0). In addition, the learning process will obtain as much reward of β as possible, which will lead to a higher probability of successful V2V load transmission.

[0151] 3.10, Adjust the hyperparameter β:

[0152] In practice, β is a hyperparameter that needs to be adjusted empirically. During training, β is adjusted to be greater than the maximum V2V transmission rate obtained by running a few steps of random resource allocation, but should not be "too large", ideally less than twice the maximum value in the adjustment experience; this design reflects the consideration of the trade-off between the final goal and learning efficiency in the reward design of the scheme.

[0153] Therefore, the following formula reward is designed to balance the reward design; the reward is set at each time step t as:

[0154]

[0155] where λ c and λ d are positive weights used to balance the goals of high-rate V2V links and high-reliability V2V links;

[0156] 3.11, Design the D3QN algorithm:

[0157] The agent based on the D3QN algorithm searches for an optimal policy to maximize the long-term cumulative reward, that is where γ ∈ [0, 1) is the discount rate of the reward;

[0158] The optimal policy can be solved through the Bellman equation, and its elements include: a discrete state set S (V) , a discrete action set A (V) and the state transition probability Therefore, for time slot t, the state-action function (Q function) of the VU agent is:

[0159]

[0160] Among them, π is the policy of the agent in the t time slot; correspondingly, the update expression of the Q function is:

[0161]

[0162] The D3QN algorithm uses a neural network model to approximate the Q function. The neural network model consists of a training network and a target network; the training network is used to select actions for the current neural network model, and the network parameters are Ω t ; the target network is used to evaluate actions, and its network parameters are The goal of the training network is:

[0163]

[0164] Therefore, the basis for updating the model parameters is to minimize the loss function for each time slot: The target value and the error δ between the estimated value Q(s t ,a t |Ω t ) is called the temporal difference error and is expressed as: Therefore, the D3QN algorithm updates the parameters of its training network according to ; where φ l represents the learning rate of the training network Ω, represents the first-order partial derivative; in the practice of the D3QN algorithm, random batch sample data is used for model training and parameter update, and its loss function expression is:

[0165]

[0166] Step 4: Data sharing and trust value update. After the data provider V i receives the corresponding allocation policy, it will start data sharing with the data requester V j . After the data sharing is completed, the data requester V j vehicle evaluates the trust of the data provider V i vehicle based on the data provided by the data provider V i vehicle and the behavior of the data provider V i in the data sharing;

[0167] Furthermore, the trust value update process in the above Step 4 is as follows:

[0168] Construct the multi-dimensional trust value of the vehicle, including: direct trust value, indirect trust value, and comprehensive trust value;

[0169] Step 4.1, calculate the direct trust value:

[0170] The direct trust value is based on the data requester V j and the data provider V i obtained from the direct data sharing interaction between them; First, when the vehicle conducts data interaction, the data requester V j The vehicle is based on the data provided by the data provider V i and the parameters in the behavioral interaction during data sharing by the data provider V i to calculate the data credibility f ji for this interaction. The calculation formula is:

[0171]

[0172] where d ji is the distance between the data provider V i and the data occurrence location when collecting this data; Q ji represents the quality of the shared data. If it is good, then Q ji is between 0.5 - 1, and if it is bad, then Q ji is between 0 - 0.5; C ji and t ji respectively represent the communication channel capacity and communication delay of this interaction; t 0 represents the standard communication delay; R i represents the trust value of the data provider; a, b, c, k, k1 are preset parameters used to adjust f ji value;

[0173]

[0174] By calculating the value of f ji , the data credibility of this interaction is obtained; Then, based on the value of f ji , it is judged whether this interaction type is a positive interaction or a negative interaction; After multiple data interactions, the number of positive interactions α and the number of negative interactions β are obtained; Finally, the direct trust value is obtained according to the ternary logic model, and the specific calculation is as follows:

[0175]

[0176] u j→i =(1 - q j→i ).

[0177] where b j→i represents the degree of trust of the data requester V j in the data provider V i , and b j→i reflects the trust level of the data requester V j in the data provider V i providing reliable data based on historical interactions and current information; dj→i Represents the data requester V j Degree of distrust towards the data provider V i , relative to the degree of trust, d j→i Measures the data requester V j Degree of suspicion of the data provider V i Providing unreliable or false data; u j→i Represents the data requester V j Degree of uncertainty about the data provider V i , u j→i The higher it is, the more data interaction information the data requester V j Needs to make a trust assessment; q j→i Represents the probability of successful data transmission, q j→i Reflects that under the current network conditions, data is transferred from the data provider V i The possibility of successful transmission to the requester. A higher q j→i Value means higher reliability during data transmission. The data requester V j Is more likely to receive complete and accurate data; σ and λ are preset parameters, and their values are different for different data sharing scenarios, indicating the influence degree of positive and negative interactions on the trust value;

[0178] The direct trust value calculated by the ternary logic model is:

[0179]

[0180] η = e -γ .

[0181] Among them, α j→i Represents the influence degree of uncertainty on the direct trust value, F j→i Represents the data provider V i And the data requester V j Interaction frequency between, M j→i Represents the data provider V i And the current data requester V j Interaction times between, while M j→X Represents the data provider V i Total interaction times with all vehicles. η is a penalty coefficient, indicating the influence of malicious messages on the direct trust value; when a vehicle sends a malicious message, the value of the malicious factor γ increases, then the value of η decreases, and the direct trust value also decreases accordingly;

[0182] Step 4.2, calculate the indirect trust value:

[0183] The calculation of direct trust values plays a crucial role in the trust evaluation model in the vehicle networking environment. This value mainly depends on the direct interaction records between vehicles. Only when there is direct interaction can the direct trust value be accurately calculated. However, in the actual vehicle networking scenario, vehicles frequently join and leave the network, and the probability of new vehicles emerging is relatively high. Therefore, relying solely on direct trust values is difficult to meet the requirements of practical applications, especially when there is a lack of direct interaction records between vehicles. To make up for this deficiency, data requesters need to collect the recommended trust values of other vehicles and then calculate the indirect trust values. This process can not only expand the scope of trust evaluation but also improve the adaptability and flexibility of the model.

[0184] 1) Screening of recommended vehicles: To obtain accurate indirect trust values while ensuring computational efficiency, the present invention proposes a recommendation screening mechanism. Specifically, vehicles that have direct interaction records with both the data provider V i and the data requester V j are selected as the recommended vehicle V k . This strategy effectively avoids invalid recommendations and improves the credibility of the recommended trust values.

[0185] 2) Recommendation confidence: Generally, the recommended trust values provided by vehicles with higher trust values are more credible. To calculate the indirect trust values more accurately, the trust evaluation model introduces recommendation confidence to evaluate the credibility of the recommended vehicles. Specifically, based on the comprehensive trust value k of the recommended vehicle V and the direct trust value R j between the data requester V k and the recommended vehicle V (j,k)(dr) , the recommendation confidence of the recommended vehicle V k is calculated; assuming that there are M recommended vehicles after screening, the calculation formula for their confidence values is as follows:

[0186]

[0187] 3) Selection of pre-trusted vehicles: However, in the vehicle networking environment, there is a potential threat that multiple malicious vehicles may form a malicious group to improve their own trust values and reduce the trust values of other vehicles through mutual cover and malicious recommendations, disrupting the normal operation of the vehicle networking network. This attack method is called a collusion attack. To effectively resist such attacks, this trust evaluation model introduces pre-trusted vehicles among the recommended vehicles.

[0188] Pre-trusted vehicles refer to those vehicles that are pre-recognized by the system as more trustworthy and communication-stable. They have special identifiers and are considered more credible than ordinary vehicles in the trust evaluation model. The mechanism of introducing pre-trusted vehicles can significantly enhance the model's resistance to collusion attacks.

[0189] Based on this, the trust evaluation model selects vehicles with higher rankings as pre-trusted vehicles according to the comprehensive trust value ranking, such as fire trucks, ambulances, police cars, and buses. These vehicles generally have a high trust value due to their critical roles and high reliability in social services.

[0190] 4) Calculate the indirect trust value: The trust evaluation model classifies the screened recommended vehicles into two categories: pre-trusted vehicles and ordinary vehicles; assume that among the recommended vehicles, the numbers of pre-trusted vehicles and ordinary vehicles are n and m respectively, and n + m = M. On this basis, the indirect trust value calculation formula for vehicle V i is as follows:

[0191]

[0192] where represents the direct trust values of the corresponding pre-trusted vehicles and ordinary vehicles with the data provider V i , and ξ represents the weight of the pre-trusted vehicles, which is used to adjust the influence of the pre-trusted vehicles in the calculation of the indirect trust value;

[0193] Step 4.3, calculate the local trust value:

[0194] After obtaining the direct trust value and indirect trust value of the vehicle, the data requester V j vehicle can further calculate the local trust value of the data provider V i vehicle

[0195]

[0196] where ρ is the recommendation degree, indicating the influence degree of the recommendation trust value on the local trust value;

[0197] Step 4.4, calculate the comprehensive trust value:

[0198] In the vehicle networking system, vehicles are encouraged to actively participate in data sharing, and vehicles are punished for malicious attacks and passive inaction. Therefore, the trust value of vehicles should also change over time. To solve this problem, the trust evaluation model introduces a forgetting factor to calculate the comprehensive trust value of vehicles to prevent the passive inaction of vehicles. The specific calculation process is as follows:

[0199]

[0200] where t c is the current time, t lis the time when the vehicle participated in the tasks in the vehicle network last time. This time interval represents the vehicle's negative time. The larger the interval, the longer the vehicle's negative time, and the greater the decrease in its trust value; δ is the forgetting factor, indicating the influence degree of the negative time interval on the comprehensive trust value;

[0201] At the same time, due to the existence of malicious vehicles, they may continue to act maliciously in the vehicle network. Therefore, the trust evaluation model introduces a blacklist mechanism to prevent continuous attacks by malicious vehicles. Each data requester V j The vehicle maintains a local blacklist, which contains all malicious vehicles identified by the data requester V j The data requester V j will not communicate with any vehicle in the blacklist; the definition of the blacklist is as follows:

[0202]

[0203] Generally speaking, in the vehicle network environment, the dynamic update of trust values is crucial for maintaining the security and reliability of the network. To achieve this goal, the roadside unit (RSU) updates the comprehensive trust value of the vehicle at the end of each time interval.

[0204] Step Five: Share transaction records. After completing the data sharing transaction, the data requester V j will broadcast the transaction record and the trust value scoring result to the vehicle chain. Other vehicle consensus nodes are responsible for verifying and auditing the broadcast content to ensure the authenticity and integrity of the data. After being consensus by the blockchain, the transaction record will be permanently stored in the blockchain, and at the same time, the blockchain ledger will also be synchronized to the roadside unit (RSU); the roadside unit (RSU) updates the reputation value of the corresponding vehicle based on the scored information after deposit, improving the trust management and data sharing of the entire vehicle network system.

[0205] Experimental Analysis

[0206] Figure 4 shows the process of the model gradually converging during training. From Figure 4 it can be seen that: the reward steadily increases as the training progresses and finally reaches a relatively stable state, indicating that the trained model can perform effective decision-making strategies.

[0207] Figure 5 is the schematic diagram of the performance analysis of the total channel capacity and data packet changes in the vehicle network of the embodiment of the present invention. From Figure 5 it can be seen that as the data packet size increases, the performance of all schemes decreases. The increased data packet size will result in a longer V2V transmission delay, and may require higher transmission power to improve the success probability of high-reliability V2V link data packet transmission. This will inevitably increase the interference to the high-rate V2V link and thus affect its capacity performance. At the same time, fromFigure 5 It can be seen that the D3QN-MARL method proposed by the present invention (the red curve in the figure) performs better than other MARL schemes (DQN-MARL and DDQN-MARL) under different data packet sizes. In addition, from Figure 5 it can also be seen that the multi-agent-based scheme is superior to the single-agent and random schemes, further proving the effectiveness of the MARL method.

[0208] Figure 6 It is a schematic diagram of the analysis of the data transmission success rate and load change performance of the V2V link in the vehicle-to-everything (V2V) network according to the embodiment of the present invention. From Figure 6 it can be seen that as the V2V data packet size increases, the V2V data packet transmission success probability of all distributed algorithms decreases. However, from Figure 5 it can be seen that the D3QN-MARL method proposed by the present invention (the black curve in the figure) performs better than other MARL schemes (DQN-MARL and DDQN-MARL) under different data packet sizes. In addition, from Figure 6 it can also be seen that the multi-agent-based scheme is superior to the single-agent and random schemes, further proving the effectiveness of the MARL method.

[0209] Figure 7 It is a schematic diagram of the analysis of the malicious vehicle detection rate under different trust thresholds according to the embodiment of the present invention. The present invention is respectively compared and analyzed with the trust value evaluation method based on the triple subjective logic model in Document [1] (J. Kang et al., "Blockchain for Secure and Efficient Data Sharing in Vehicular Edge Computing and Networks," in IEEE Internet of Things Journal, vol. 6, no. 3, pp. 4660-4670, June 2019, doi: 10.1109 / JIOT.2018.2875542.) and the trust management model based on recommendation in Document [2] (A. M. Shabut, K. P. Dahal, S. K. Bista and I. U. Awan, "Recommendation Based Trust Model with an Effective Defence Scheme for MANETs," in IEEE Transactions on Mobile Computing, vol. 14, no. 10, pp. 2101-2115, 1 Oct. 2015, doi: 10.1109 / TMC.2014.2374154.). From Figure 7As can be seen, the trust evaluation model proposed by the present invention has obvious advantages in malicious vehicle identification and can achieve a higher detection rate at a lower trust threshold.

[0210] Figure 8 It is a schematic diagram for analyzing the detection rate of malicious vehicles and the probability of malicious vehicles sending error messages in the embodiments of the present invention. The present invention is respectively compared and analyzed with the trust value evaluation method based on the triple subjective logic model in Document [1] (J. Kang et al., "Blockchain for Secure and Efficient Data Sharing in Vehicular Edge Computing and Networks," in IEEE Internet of Things Journal, vol. 6, no. 3, pp. 4660 - 4670, June 2019, doi: 10.1109 / JIOT.2018.2875542.) and the trust management model based on recommendation in Document [2] (A. M. Shabut, K. P. Dahal, S. K. Bista and I. U. Awan, "Recommendation Based Trust Model with an Effective Defence Scheme for MANETs," in IEEE Transactions on Mobile Computing, vol. 14, no. 10, pp. 2101 - 2115, 1 Oct. 2015, doi: 10.1109 / TMC.2014.2374154.). From Figure 8 As can be obtained, when dealing with the problem of malicious vehicle detection, the trust evaluation model proposed by the present invention not only has high accuracy but also can effectively cope with a higher probability of sending error messages, which makes the trust evaluation model have strong robustness in complex and unstable network environments.

[0211] Figure 9 It is a schematic diagram for analyzing the vehicle trust value after being attacked by a switch in the embodiments of the present invention. The present invention is compared with the anti - attack - based trust evaluation method AATMS in Document [3] (J. Zhang, K. Zheng, D. Zhang and B. Yan, "AATMS: An Anti - Attack Trust Management Scheme in VANET," in IEEE Access, vol. 8, pp. 21077 - 21090, 2020, doi: 10.1109 / ACCESS.2020.2966747.). From Figure 9It can be seen that the change of vehicle trust value shows an obvious fluctuating trend. Especially during the attack event (marked as "Attack zone"), the method proposed in the present invention (red curve) shows strong sensitivity in the attack interval. The vehicle trust value drops rapidly and is lower than the trust threshold at the end of the attack. After being defined as a malicious vehicle, the vehicle trust value does not rise but drops slowly because of the existence of the blacklist mechanism and the time penalty factor. After a malicious vehicle is pulled into the blacklist, it cannot share data with others, so its trust value will drop slowly. While for AATMS (blue curve), although the trust value drops after the vehicle launches an attack, it can still continue to share data and accumulate trust value for subsequent attacks after the attack. This shows that the proposed scheme can effectively prevent switch attacks and prevent malicious vehicles from accumulating trust value to continue attacks after the attack. At the same time, as can be seen from Figure 9 the green curve in, when the vehicle maintains normal behavior, the trust value will rise slowly.

[0212] Figure 10 This is a schematic diagram for the analysis of the change of vehicle trust value and vehicle channel capacity under the embodiment of the present invention. From Figure 10 it can be seen that when the vehicle channel capacity rises, its trust value also rises. This is because a higher channel capacity means higher data credibility and also means the quality of this data sharing in vehicle data sharing, so the vehicle will obtain a higher trust value. This shows that the communication channel capacity of the vehicle during data sharing and its trust value interact with each other.

[0213] Aiming at the problems of vulnerability to malicious vehicle attacks and insufficient utilization of spectrum resources when sharing data for vehicles, the present invention studies a trusted and efficient data sharing scheme for the Internet of Vehicles based on blockchain, and proposes corresponding trust evaluation mechanisms and spectrum resource allocation strategies, solving the problems of vulnerability to malicious attacks and limited spectrum resources in vehicle data sharing. In the present invention, first, considering problems such as the self-interest and malicious behaviors of vehicles, a corresponding trust value update strategy is designed, combining vehicle data sharing behaviors, communication channel capacities, data qualities, etc., to design a more complete trust value update strategy, effectively detecting malicious and negative behaviors of vehicles and removing them in a timely manner. Then, aiming at problems such as limited spectrum resources, high dynamicity, and obstacles in the Internet of Vehicles, relay-assisted communication and the D3QN algorithm based on deep reinforcement learning are proposed to build a multi-agent collaborative communication system for the Internet of Vehicles, improving the utilization rate of spectrum resources in the Internet of Vehicles and optimizing the overall communication rate of the entire Internet of Vehicles system. Finally, based on the decentralization and immutability of blockchain technology, the data sharing behaviors and trust values of vehicles are saved to the blockchain ledger to ensure the authenticity and reliability of vehicle trust values. The simulation results of the model show that the method proposed in the present invention can, with the support of blockchain technology, effectively identify malicious vehicles, improve the trust environment in the Internet of Vehicles, optimize the utilization efficiency of spectrum resources, and improve the communication rate of the overall network, thereby enhancing the security and overall performance of the vehicle data sharing system.

[0214] The key points and protected points of the present invention include but are not limited to:

[0215] 1. Based on the D3QN algorithm of deep reinforcement learning, considering the complex scenarios of relay assistance and limited spectrum resources in the Internet of Vehicles, optimizing the vehicle communication rate to meet the communication needs of different vehicles, that is, the content of step three.

[0216] 2. Based on the social and communication behaviors of vehicle data sharing, conducting multi-dimensional design of the trust values of vehicles to more effectively identify malicious vehicles in the Internet of Vehicles and carry out effective prevention. That is, the content of step four.

[0217] 3. Linking the trust values of vehicles with the allocation of communication resources, interacting with each other, better regulating vehicle behaviors, and serving the Internet of Vehicles system, that is, part of the content of steps three and four.

[0218] In the existing research on communication resource allocation based on deep reinforcement learning algorithms, only the communication scenarios of direct V2I and V2V links are considered, and the complex scenarios of relay-assisted communication are not considered. Moreover, the selection of relays is also crucial in the complex scenarios of relay assistance. At the same time, in the selection of relays and the allocation of communication resources, the trust value of vehicles should be considered, and the trust value of vehicles should be combined with the priority of vehicle services to ensure the credibility and efficiency of the entire vehicle network communication, which is an issue ignored by many studies. By using the D3QN deep reinforcement learning algorithm, based on the trust value of vehicles and the state of the environment, jointly designing relay-assisted communication and communication resource allocation in the vehicle network can meet the communication needs of different vehicles and maximize the communication rate of vehicles in the vehicle network.

[0219] Therefore, there is no other alternative that can fully achieve the purpose of the present invention.

[0220] The present invention also provides a trustworthy and efficient vehicle network data sharing system, including:

[0221] A data sharing request module, used to implement the data requester V in step one j Find the data provider V with the optimal trust value i , and send a data sharing request;

[0222] A data sharing authorization module, used to implement the data provider V in step two i Accept the sharing request and verify the identity of the data requester V j for authorization;

[0223] A communication resource allocation module, used to implement the data provider V in step three i After confirming that the identity of the data requester V j meets the requirements, the data provider V i Sends the relevant information of this data sharing to the nearby Road Side Unit (RSU) and requests the communication resources for data sharing;

[0224] After receiving the request, the Road Side Unit (RSU) performs an algorithm analysis of deep reinforcement learning based on the data sharing vehicle pairs and communication resources in the vehicle network, obtains the optimal allocation strategy that conforms to the communication resources in the vehicle network, and sends it to the corresponding data provider V i ;

[0225] A data sharing and trust value update module, used to implement the data provider V in step four i After receiving the corresponding allocation strategy, it will start data sharing with the data requester V j . After the data sharing is completed, the data requester V j The vehicle, based on the data provided by the data provider V i Vehicle-provided data and data provider Vi Behavior in data sharing for data provider V i Conduct trust assessment;

[0226] Trust management and data sharing module, used to implement the shared transaction record in step five. After completing the data sharing transaction, data requester V j Broadcasts the transaction record, trust value scoring result, and other relevant information to the vehicle chain. Other vehicle consensus nodes are responsible for verifying and auditing the broadcast content. After blockchain consensus, the transaction record is permanently stored in the blockchain, and at the same time, the blockchain ledger is synchronized to the roadside unit (RSU); the roadside unit (RSU) updates the reputation value of the corresponding vehicle based on the archived scoring information to improve the trust management and data sharing of the entire vehicle networking system.

[0227] The present invention also provides a trustworthy and efficient vehicle networking data sharing device, including:

[0228] Memory: Stores the computer program of the above-mentioned trustworthy and efficient vehicle networking data sharing method, and is a computer-readable device;

[0229] Processor: Used to implement the above-mentioned trustworthy and efficient vehicle networking data sharing method when executing the computer program.

[0230] The present invention also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it can implement the above-mentioned trustworthy and efficient vehicle networking data sharing method.

Claims

1. A reliable and efficient vehicle network data sharing method, characterized in that: The following steps are involved: Step 1: Data Requester V j Find the data provider V with the best trust value i , send a data sharing request; Step 2: Data Provider V i Accept the sharing request and verify the data requester V j Identity authorization; Step 3: Use the D3QN algorithm to allocate communication resources: i Confirm data requester V j After the identity meets the requirements, the data provider V i Send the relevant information of this data sharing to the nearby roadside unit (RSU) and request the communication resources for data sharing; after receiving the request, the roadside unit (RSU) performs deep reinforcement learning algorithm analysis based on the data sharing vehicle pairs and communication resources in the Internet of Vehicles, obtains the best allocation strategy for communication resources in the Internet of Vehicles, and sends it to the corresponding data provider V i ; Step 4: Data sharing and trust value update, data provider V i After receiving the corresponding allocation strategy, it will communicate with the data requester V j Start data sharing. After data sharing is completed, the data requester V j Vehicle based on data provider V i Data provided by the vehicle and data provider V i Behavior in data sharing for data providers V i Conduct trust assessments; Step 5: Share transaction records. After completing the data sharing transaction, the data requester V j The transaction records and trust value scoring results are broadcast to the vehicle chain, and other vehicle consensus nodes are responsible for verifying and auditing the broadcast content. After blockchain consensus, the transaction records are permanently stored in the blockchain, and the blockchain ledger is synchronized to the roadside unit (RSU); the roadside unit (RSU) updates the reputation value of the corresponding vehicle based on the stored scoring information, improving the trust management and data sharing of the entire Internet of Vehicles system.

2. A reliable and efficient vehicle networking data sharing method according to claim 1, characterized in that: The specific method of step one is: Find data provider V i , data requester V j First, check the latest block synchronized by the nearby roadside unit (RSU) and search for the data of interest in the data index. j When the data of interest is found, the data requester V j Send a request to the roadside unit (RSU) to query the data provider V i Trust value; Data Requester V j According to the queried data provider V i Choose a data provider V with the best trust value based on the trust value and data description i , data requester V j To the data provider V i Send a data sharing request.

3. A reliable and efficient vehicle networking data sharing method according to claim 1, characterized in that: The specific method of step 2 is: Data sharing authorization, when the data provider V i When receiving a sharing request, the data requester V j The comprehensive trust value of the data provider V i Decide whether to allow data requester V j Get the original data, data provider V i First check the data requester V j Is there a blacklist? If the data provider V i The blacklist contains data requester V j , the request will be rejected; if the data requester V j The comprehensive trust value of the data provider V is lower than the sharing threshold φ. i The request will also be rejected; Data Provider V i Verify data requester V j Whether there is frequent access behavior, each vehicle data provider V i Maintain an access list PIDvisited, which records the data provider V that has been visited in the past 10 minutes i The vehicle number PID of the vehicle and the corresponding access timestamp Timestampvisited. If the data requester V j The vehicle ID has appeared in the access list PIDvisited, indicating that the vehicle has visited repeatedly in a short period of time. The data provider V i Reject the sharing request of the vehicle. If all verifications pass, the data provider V i Generate a transaction and record the transaction in the relevant information of this data sharing.

4. A reliable and efficient vehicle networking data sharing method according to claim 1, characterized in that: In step 3, the D3QN algorithm is used to allocate communication resources to obtain the best allocation strategy for communication resources in the Internet of Vehicles as follows: A dynamic IoV communication network model is constructed, including M available idle spectrum resources and K vehicle-to-vehicle (V2V) communication links. Vehicles travel in different directions on multiple lanes. In each service cycle, vehicles are divided into three categories, namely, data providers V i Vehicle, data requester V j Vehicles and idle vehicles; all transceivers use a single antenna, and the sets M = {1, ..., M} and K = {1, ..., K} represent the idle spectrum resources and K vehicle-to-vehicle (V2V) communication links in the dynamic vehicle network communication network, respectively; Step 3.1: construct a vehicle-to-vehicle (V2V) communication model, including a path loss model and a shadow fading model, and calculate the channel gain: (1) Constructing a path loss model: In the vehicle-to-vehicle (V2V) communication model, the line-of-sight path loss PL Los and non-line-of-sight path loss PLN Los Calculate the path loss. The specific formula is as follows: When there is line of sight between vehicles (V2V), the line of sight path loss RL Los The formula is: Among them, in the path loss model, A1 represents the path loss coefficient when the distance d≤d0, reflecting the propagation loss speed of the millimeter wave signal in a short distance; A2 is the path loss reference value at this time, which is related to the fixed factors of the propagation environment; B1 is a coefficient related to frequency dependence, reflecting the influence of frequency on path loss, especially at the carrier frequency f c When the unit is GHz, log10(fc / F0) represents the contribution of path loss, where F0 is the standardized frequency reference value, which is used to standardize the influence of frequency; when the distance d>d0, A3 is used as the path loss coefficient, A3 is greater than A1, A4 is a constant, which is used to compensate for the influence of other environmental factors on path loss, A5 and A6 are the attenuation coefficients of the antenna height h0 of the transmitting vehicle and the antenna height h1 of the receiving vehicle, respectively, which represent the influence of antenna height on path loss, and B2 is used for frequency dependence compensation for long distances; When there is non-line-of-sight between vehicles (V2V), the non-line-of-sight path loss PLN Los The formula is: In the non-direct (NLoS) path loss model, PL NLoS (d a ,d b ) represents the path loss when there are obstacles, C1 represents the reference additional loss of non-direct propagation, reflecting the additional attenuation of the signal by obstacles; C2 is related to the path loss adjustment factor n j Multiplication is used to quantify the impact of obstacle complexity on path loss; C3 represents the propagation distance d of the signal behind the obstacle. b The additional impact on the path loss, C4 is the frequency-dependent additional path loss coefficient, and F0 is used as the reference frequency to normalize the frequency dependence of the path loss; (2) Construct a shadow fading model, expressed as: Among them, Δd represents the change in the distance between vehicles, reflecting the change in the relative position of the vehicles during movement; d dec is the decorrelation distance between vehicles, which is used to describe the significant reduction in signal correlation within this distance. is a standard normal distribution noise with a mean of 0 and a standard deviation of 3, which is used to simulate the randomness of shadow fading; represents the attenuation part of the shadow fading; It represents the random variation part of shadow fading; (3) Calculation of channel gain: Orthogonal frequency division multiplexing (OFDM) technology is used to convert the frequency selective wireless channel into multiple parallel flat channels. Several consecutive subcarriers are combined into a spectrum subband. It is assumed that the channel fading within a subband is approximately the same and the channel fading between different subbands is independent. In a coherent time period, the channel power gain g of the kth V2V link on the mth subband is k [m] follows the following formula: g k [m]=α k h k [m], Among them, h k [m] is the frequency-dependent small-scale fading power component, α k Capture large-scale fading effects, including path loss and shadow fading; Step 3.2, select a set of relay vehicles from idle vehicles based on the physical transmission distance and comprehensive trust value: In the initial stage, the idle nodes in the vehicle network communication network are screened according to the physical transmission distance. The screening range is based on the data provider V i Vehicle and Data Requester V j The center of the vehicle is the center of the circle, and the straight-line distance between the two is the diameter of the circle. The relay nodes in the area covered by the circle are the candidate nodes that meet the physical transmission distance constraint. For a vehicle link pair, the set of relay nodes that meet its physical transmission distance constraint is expressed as: in, Represents the vehicle link alignment data provider V i and data requester V j The distance between and Respectively represent the vehicle link alignment data provider V i and data requester V j The coordinates of the node; (x j ,y j ) represents any idle vehicle node n in the Internet of Vehicles j The location coordinates are: The comprehensive trust value threshold R is used TH Relay node set The relay nodes in continue to screen, the specific expression is as follows: Where n represents the set of relay nodes The number of idle vehicles after the first round of screening, R i,j is the trust value of the idle vehicle. If the trust value of the idle vehicle is greater than or equal to R TH , then in the relay node set Keep the idle vehicle in the node; otherwise, delete the node; After screening by physical transmission distance and comprehensive trust value, the final set of candidate relay nodes is N r ; Step 3.3, calculate the service priority of the vehicle link: For different data requesters V j The comprehensive trust value of the vehicle link pair and the communication requirements of the vehicle link pair are used to design a priority-based vehicle service, where Y i (t) represents the priority value of the i-th vehicle at time t; According to the vehicle link alignment data provider V i The comprehensive trust value of the system determines the priority related to the comprehensive trust value, specifically: According to the communication requirements of vehicle links, they are divided into high-speed category with real-time transmission and high-reliability category with stable and high-speed that meets the requirements; The service priority of the vehicle link at high speed is higher than that of the vehicle link at high reliability, specifically: At the same time, if a vehicle has recently provided relay cooperative transmission services to other vehicles, its service priority will be increased by 1 as a relay incentive; Finally, the service priority of each vehicle link is expressed as: Y i (t)=Y(r)+Y(re)+ρ0 Among them, ρ0 indicates whether the vehicle has a record of providing relay cooperative transmission service as a relay within the specified time interval. If yes, it is 1, otherwise it is 0; Step 3.4, based on the service priority of the vehicle link calculated in step 3.3, construct a channel transmission model: The roadside unit (RSU) first selects a dedicated or shared transmission link for the vehicle according to the service priority of the vehicle link calculated in step 3.3, and then selects whether the vehicle needs relay assistance according to the communication environment of the vehicle; the vehicle collection in the Internet of Vehicles communication network is divided into four types: direct connection dedicated transmission vehicle collection Vehicle collection for direct multiplexing transmission A collection of vehicles relaying dedicated transmissions Relay multiplexing transmission vehicle collection Assume that the total number of vehicles is N V ,but: In the relay-assisted transmission process, the amplify-and-forward (AF) relay strategy is adopted, that is, the nth relay vehicle amplifies the signal it receives by β n,k times, and forwards the received signal to the kth target vehicle. The specific calculation process of the signal-to-noise ratio and channel capacity is as follows: (1) Direct connection dedicated transmission: For the direct connection dedicated transmission process of the kth receiving vehicle, the signal-to-noise ratio of the transmission is expressed as: The channel capacity is: in, is the kth z The transmission power of a target vehicle on the mth subband, is the kth z The channel gain of the target vehicle in the mth subband, σ 2 is the noise power, W is the bandwidth; (2) Relay-dedicated transmission: For the relay-dedicated transmission process of the kth receiving vehicle, the transmission process includes the downlink transmission stage and the relay forwarding process. The signal-to-noise ratio of the transmission is expressed as: The downlink transmission phase is expressed as: The relay forwarding process is expressed as: in, They are respectively the downlink transmission phase The transmission power of the target vehicle on the mth subband is proportional to the transmission power of the relay forwarding phase. The transmission power of the relay vehicle on the mth subband, are the channel gains of the two stages, σ 2 is the noise power; When calculating the channel capacity of the entire communication process, we introduce The specific calculation of factors is as follows: (3) Direct multiplexing transmission: For the first stage, the direct multiplexing vehicle link is interfered by the downlink transmission process from other direct transmission links and the downlink transmission process of the selected relay, which is expressed as: in, are the transmission powers of other directly connected vehicles and relay auxiliary vehicles in the downlink transmission phase on the mth subband, are the corresponding channel gains respectively; and is a binary spectrum allocation indicator variable; when the value of the binary spectrum allocation indicator variable is 1, it means that the kth V2V link uses the mth subband, and when it is 0, it means that it is not occupied; at the same time, it is assumed that each V2V link only accesses one subband, that is: Then the signal-to-noise ratio SINR of the first stage is: For the second stage, the direct multiplexing vehicle link is interfered by the downlink transmission process of other direct transmission links and the relay forwarding process of the relay forwarding link, which is specifically expressed as: The signal-to-noise ratio (SINR) of the second stage is: Based on the signal-to-noise ratio of the above two stages, when calculating the channel capacity, we also introduce Factor processing: (4) Relay multiplexing transmission: For the vehicle link with relay multiplexing, the transmission process is also divided into two stages, namely the downlink transmission stage and the relay forwarding process. In the downlink transmission stage, the vehicle link with relay multiplexing is interfered by the downlink transmission process of other direct transmission links and the downlink transmission process of the selected relay, which can be specifically expressed as: in, are the transmission power of the other directly connected vehicles in the downlink transmission phase and the relay-assisted vehicle in the downlink transmission phase on the mth subband, are the corresponding channel gains respectively; then the signal-to-noise ratio SINR in the downlink transmission stage is: For the relay forwarding process, the relay multiplexing vehicle link is interfered by the downlink transmission process of other direct transmission links and the relay forwarding process of the relay forwarding link, which is specifically expressed as: in, are the transmission power of the other directly connected vehicles in the downlink transmission phase and the relay-assisted vehicle in the downlink transmission phase on the mth subband, respectively; are the corresponding channel gains respectively; then the signal-to-noise ratio SINR of the relay forwarding process is: The final channel capacity is calculated as: 3.5, construct joint optimization problem: For link pairs that require high rates, the optimization goal is to maximize the total channel capacity of the high-rate link: For link pairs that require high reliability, the optimization goal is to stably transmit security-critical messages within the time budget T. The high-reliability link pair is modeled as transmitting a data packet of size B within each time budget, that is: Where ΔT is the channel coherence time, and t represents different coherence time slots; Build a multi-agent cooperative communication system and design a multi-agent cooperative communication strategy based on D3QN. First, the constructed Internet of Vehicles communication network model is modeled as the environment. Secondly, the vehicle nodes in the Internet of Vehicles communication network are modeled as agents. At each time slot t, each agent first obtains the current state information s from the environment. t ∈S, and select the action a to be performed according to the state information t ∈A, where S and A represent the state space and action space respectively. The actions of multiple agents act together in the environment and cause changes in the current environment, that is, entering the next environmental state s t+1 At the same time, the environment evaluates the actions of each agent and sends the reward r through the reward module. t Feedback to the agent; at this point, each agent obtains training samples (s t ,a t ,r t ,s t+1 ) and stored in the experience replay pool for optimization training of the agent model; for each agent, the expression of its Markov chain is: {s t ,a t ,s t+1 ,a t+1 ,...}; In the multi-agent cooperative communication system, the decision variables of the legal vehicle are the spectrum allocation and relay selection decisions, which are discrete decision variables. The multi-agent model is constructed based on the D3QN algorithm. The D3QN algorithm uses the training network and the target network for action selection and action evaluation, that is, the training network is used to obtain s t+1 The best action in the state and use the target network to calculate the action value of the action; 3.6, construct state and observation space: In the resource sharing problem under the multi-agent reinforcement learning (multi-agent RL) framework, each V2V link k acts as an agent, and the resource sharing problem is modeled as a Markov decision process (MDP). At each channel coherence time step t, given the current environment state S t , each V2V agent k receives an observation of the environment The observation is determined by the observation function O, namely Then take an action Forming joint action A t ; Afterwards, the agent receives the reward R t+1 , the environment transfers to the next state S with probability p(s',r|s,a) t+1 ; Each agent receives new observations All V2V agents share the same reward in a multi-agent cooperative communication system; The actual environment state S t , including the global channel conditions and the behaviors of all agents. Each V2V agent acquires knowledge of the environment from the perspective of the observation function. The observation space of V2V agent k contains local channel information, which includes: for all m∈M, the channel gain g of V2V agent k itself k [m]; for all k′=k and m∈M, the interference channel gain g k',k [m] comes from other V2V transmitters; for all m∈M, the received interference power I on all frequency bands is measured at the V2V receiver k [m], and is also introduced into the local observation; in addition, the local observation space includes the remaining V2V load B k and the remaining time budget T k , the observation function of V2V agent k is expressed as: O(S t ,k)={B k ,T k ,{I k [m]} m∈M ,{G k [m]} m∈M }, Among them, G k [m] = {g k [m],g k',k [m]}; 3.7, construct action space: By limiting the power control options to four levels, the dimension of the action space is 4 × M, and each action corresponds to a specific combination of spectrum subband and power selection; 3.8, Design Rewards: In the resource sharing problem constructed in step 3.6, there are two objectives: maximizing the total capacity of the high-rate V2V link and increasing the probability of successful load transmission of the high-reliability V2V link within a certain time constraint T; The first goal is to calculate the instantaneous total capacity of all V2V links at each time step t: As a reward for each step; Second goal: For each V2V agent k, a reward L k is set to the effective V2V transmission rate until the load transfer is completed; the reward is set to a constant β, which is greater than the maximum V2V transmission rate; the V2V-related reward at each time step t is set to: 3.9, Construct learning objectives: The cumulative discounted reward is defined as the return Gt, with a discount rate of γ, that is: 3.10, adjust the hyperparameter β: During training, β is adjusted to be larger than the maximum V2V transmission rate obtained by running several steps of random resource allocation and smaller than twice the maximum value obtained in the adjustment experience; The following formula reward is designed to balance the reward design; at each time step t, the reward is set to: Among them, λ c and λ d is a positive weight used to balance the goals of high-rate V2V links and high-reliability V2V links; 3.11, Design D3QN algorithm: The D3QN algorithm-based agent seeks the best allocation strategy for communication resources to maximize the long-term cumulative reward, namely Where γ∈[0,1) is the discount rate of the reward; The optimal allocation strategy of communication resources is solved by the Bellman equation, whose elements include: discrete state set S (V) , a discrete action set A (V) and state transition probability Therefore, for time slot t, the state-action function (Q function) of the VU agent is: Where π is the strategy of the agent in time slot t; accordingly, the update expression of the Q function is: The D3QN algorithm uses a neural network model to fit the Q function. The neural network model consists of a training network and a target network. The training network is used to select the action of the current neural network model. The network parameter is Ω t ; The target network is used to evaluate the action, and its network parameters are The goals of training the network are: Therefore, the model parameter update is based on minimizing the loss function at each time slot: Target value With the estimated value Q(s t ,a t |Ω t ) is called the time difference error, which is expressed as: Therefore, the D3QN algorithm follows Update the parameters of its training network; where φ l represents the learning rate of training network Ω, Represents the first-order partial derivative; in the practice of the D3QN algorithm, random batch sample data is used To train the model and update the parameters, the loss function is expressed as:

5. A reliable and efficient vehicle networking data sharing method according to claim 1, characterized in that: The trust value update process in step 4 is as follows: Construct multi-dimensional trust value of vehicles, including direct trust value, indirect trust value and comprehensive trust value; Step 4.1, calculate the direct trust value: The direct trust value is based on the data requester V j With data provider V i The trust value obtained by direct data sharing interaction between vehicles; First, when vehicles interact with each other, the data requester V j Vehicle based on data provider V i Data provided and data provider V i Parameters in the behavioral interaction in data sharing, the credibility of the data in this interaction f ji Calculate, the calculation formula is: Among them, d ji For data provider V i When the data was collected, the distance from where the data occurred; Q ji Indicates the quality of shared data. If it is good, Q ji Between 0.5-1, bad is Q ji Between 0-0.5; C ji With t ji They represent the communication channel capacity and communication delay of this interaction respectively; t0 represents the standard communication delay; R i represents the trust value of the data provider; a, b, c, k, k1 are preset parameters used to adjust f ji The value of By calculating f ji The value of f is used to obtain the data credibility of this interaction; then according to f ji The value of is used to judge whether the interaction type is positive or negative; after multiple data interactions, the number of positive interactions α and the number of negative interactions β are obtained; finally, the direct trust value is obtained based on the ternary logic model, which is calculated as follows: u j→i =(1-q j→i ). Among them, b j→i Indicates data requester V j For data providers V i The degree of trust, b j→i Reflect data requester V j Based on historical interactions and current information, data provider V i Provides a level of confidence in reliable data; d j→i Indicates data requester V j For data providers V i The degree of distrust; j→i Indicates data requester V j For data providers V i The degree of uncertainty; q j→i represents the probability of successful data transmission, q j→i Reflecting the current network conditions, data is from data provider V i The probability of successful transmission to the requester; σ and λ are preset parameters, indicating the influence of positive and negative interactions on the trust value; The direct trust value calculated by the ternary logic model is: the=e -γ . Among them, a j→i Indicates the degree of influence of uncertainty on the direct trust value, F j→i Represents the data provider V i With data requester V j The interaction frequency between j→i Represents the data provider V i With the current data requester V j The number of interactions between j→X Represents the data provider V i The total number of interactions with all vehicles, η is a penalty coefficient, which indicates the impact of malicious messages on the direct trust value; when a vehicle sends a malicious message, the value of the malicious factor γ increases, then the value of η decreases, and the direct trust value also decreases accordingly; Step 4.2, calculate the indirect trust value: 1) Selection of recommended vehicles: Select and data provider V i and data requester V j The vehicles with direct interaction records are recommended as vehicles V k ; 2) Recommendation confidence: Based on recommended vehicle V k The comprehensive trust value And the data requester V j With recommended vehicle V k The direct trust value R (j,k)(dr) , calculate the recommended vehicle V k The confidence of recommendation; assuming that there are M recommended vehicles after screening, the calculation formula of its confidence value is as follows: 3) Select a pre-trusted vehicle: The trust assessment model selects vehicles with higher rankings as pre-trusted vehicles based on the comprehensive trust value ranking; 4) Calculate indirect trust value: The trust evaluation model divides the screened recommended vehicles into two categories: pre-trusted vehicles and ordinary vehicles. Assume that the number of pre-trusted vehicles and ordinary vehicles in the recommended vehicles is n and m respectively, and n+m=M. On this basis, vehicle V i The indirect trust value calculation formula is as follows: in, Represents the corresponding pre-trusted vehicles and ordinary vehicles and data providers V i The direct trust value of the vehicle, ξ represents the weight of the pre-trusted vehicle, which is used to adjust the influence of the pre-trusted vehicle in the calculation of the indirect trust value; Step 4.3, calculate the local trust value: Data Requester V j Vehicle calculation data provider V i The local trust value of the vehicle Among them, ρ is the recommendation degree, which indicates the influence of the recommended trust value on the local trust value; Step 4.4, calculate the comprehensive trust value: The trust evaluation model introduces the forgetting factor to calculate the comprehensive trust value of the vehicle. The specific calculation process is as follows: Among them, t c is the current time, t l is the time when the vehicle last participated in a task in the Internet of Vehicles. This time interval represents the vehicle's passive time. The larger the interval, the longer the vehicle's passive time, and the greater the decrease in its trust value. δ is the forgetting factor, which indicates the degree of influence of the passive time interval on the comprehensive trust value. The trust evaluation model introduces a blacklist mechanism to prevent continuous attacks from malicious vehicles. Each data requester V j The vehicle maintains a local blacklist containing data requesters V j All malicious vehicles identified, data requesters V j No communication will be made with any vehicle in the blacklist; the blacklist is defined as follows: The roadside unit (RSU) updates the comprehensive trust value of the vehicle at the end of each time interval.

6. A reliable and efficient vehicle networking data sharing system based on the method according to any one of claims 1 to 5, characterized in that: include: Data sharing request module, used to implement data requester V j Find the data provider V with the best trust value i , send a data sharing request; Data sharing authorization module, used to implement data provider V i Accept the sharing request and verify the data requester V j Identity authorization; Communication resource allocation module, used to implement the data provider V i Confirm data requester V j After the identity meets the requirements, the data provider V i Send relevant information about the data sharing to nearby roadside units (RSUs) and request communication resources for data sharing; After receiving the request, the roadside unit (RSU) performs deep reinforcement learning algorithm analysis based on the data sharing vehicle pairs and communication resources in the Internet of Vehicles, obtains the best allocation strategy for communication resources in the Internet of Vehicles, and sends it to the corresponding data provider V i ; Data sharing and trust value update module, used to implement data provider V i After receiving the corresponding allocation strategy, it will communicate with the data requester V j Start data sharing. After data sharing is completed, the data requester V j Vehicle based on data provider V i Data provided by the vehicle and data provider V i Behavior in data sharing for data providers V i Conduct trust assessments; Trust management and data sharing module, used to realize shared transaction records. After completing the data sharing transaction, the data requester V j The transaction records, trust value scoring results and other related information are broadcast to the vehicle chain, and other vehicle consensus nodes are responsible for verifying and auditing the broadcast content. After blockchain consensus, the transaction records are permanently stored in the blockchain, and the blockchain ledger is synchronized to the roadside unit (RSU); the roadside unit (RSU) updates the reputation value of the corresponding vehicle based on the stored scoring information, improving the trust management and data sharing of the entire Internet of Vehicles system.

7. A reliable and efficient Internet of Vehicles data sharing device, characterized in that: include: Memory: a computer program storing a reliable and efficient vehicle networking data sharing method as described in any one of claims 1 to 5, which is a computer-readable device; Processor: used to implement a reliable and efficient vehicle network data sharing method as described in any one of claims 1-5 when executing the computer program.

8. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it can implement a reliable and efficient vehicle network data sharing method as described in any one of claims 1-5.

Citation Information

Patent Citations

  • Block chain technology-based secure Internet of Vehicles social network construction method

    CN111064800A

  • Inter-vehicle data safety sharing method and system based on block chain

    CN111967051A

  • Internet of vehicles edge computing sharing method based on block chain

    CN114945022A

  • Internet of vehicles data sharing method based on cross-chain technology

    CN114980023A

  • Internet of vehicles trusted data sharing method and system based on deep reinforcement learning

    CN116684442A

Cited By

  • Public facility abnormal data processing and analysis method based on space-time diagram neural network

    CN121580247A

  • Privacy-enhanced Internet of Vehicles data sharing and analysis system

    CN121598427A