Internet of vehicles communication security situation assessment and decision optimization method based on reinforcement learning

By employing a multi-source data fusion and federated learning approach based on reinforcement learning, the limitations of static strategies and the problem of data silos in vehicle-to-everything (V2X) information security are addressed. This enables accurate perception and dynamic defense against complex threats, thereby improving the security and availability of V2X.

CN122027353APending Publication Date: 2026-05-12SOUTHEAST UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SOUTHEAST UNIV
Filing Date
2026-04-09
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

The Internet of Vehicles (IoV) faces information security challenges, including the limitations of static strategies, insufficient single-dimensional data perception, crude decision-making and resource imbalance, lack of online adaptive capabilities, insufficient distributed collaboration and privacy issues. Traditional security defenses are unable to cope with rapidly changing attack methods and complex threats.

Method used

We employ a reinforcement learning-based approach, collecting multi-source data, preprocessing and extracting features, using a deep Q-network enhanced with an attention mechanism for situation assessment and defense action selection, and combining federated learning to achieve dynamic policy optimization. We integrate multi-dimensional data for global threat perception and collaborative defense.

Benefits of technology

It enables multi-dimensional situational awareness of the vehicle-to-everything (V2X) environment, dynamically optimizes defense strategies, improves the comprehensiveness and accuracy of threat detection, ensures the security and availability of vehicle communication, responds promptly to complex threats, and enhances the overall security posture while protecting user privacy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122027353A_ABST
    Figure CN122027353A_ABST
Patent Text Reader

Abstract

The invention discloses an Internet of Vehicles information security situation assessment and decision optimization method based on reinforcement learning, which is applied to vehicle-vehicle and vehicle-cloud communication scenes. The method comprises the following steps: step 1, collecting multi-source Internet of Vehicles security data; 2, preprocessing the multi-source data, and extracting a feature vector reflecting the security situation of the Internet of Vehicles; 3, inputting the feature vector into a reinforcement learning model, and outputting a defense action by the model according to the current security situation; 4, executing a defense action and obtaining a reward signal, wherein the reward signal is generated based on the change of the security situation of the Internet of Vehicles; and 5, updating parameters of the reinforcement learning model according to the reward signal, and optimizing the security situation assessment and defense decision of the Internet of Vehicles. The system adopts a multi-agent cooperation mechanism of federal learning, and the defense knowledge is shared between each vehicle node and the road infrastructure, so that the cooperative defense capability is improved. The method has the advantages of adaptability and distributed cooperation, and can effectively deal with diversified security threats in the Internet of Vehicles.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of vehicle network information security technology, specifically to a method for vehicle network communication security situation assessment and decision optimization based on reinforcement learning. Background Technology

[0002] With the development of intelligent connected vehicles and vehicle-to-everything (V2X) technology, communication between vehicles and between vehicles and the cloud is becoming increasingly frequent. However, open wireless communication and vehicle internal buses pose serious information security challenges to V2X. Attackers can exploit vulnerabilities in V2X to carry out various attacks, such as man-in-the-middle attacks that intercept and tamper with vehicle communication data, message spoofing attacks that forge emergency messages or identity information, GPS signal spoofing that interferes with vehicle positioning, and attacks that inject malicious commands into the vehicle's CAN bus. These attacks can directly threaten driving safety and network stability, and traditional security measures are insufficient to respond promptly and effectively. Current V2X security defenses mainly face the following core challenges:

[0003] Limitations of static strategies: Traditional defense systems for connected vehicles often rely on preset static security strategies (such as fixed communication authentication mechanisms and predefined filtering rules). In the rapidly changing vehicular network environment, such manually rule-based defenses are often slow to respond to new and unknown attacks and are unable to cope with dynamically evolving attack methods.

[0004] Insufficient single-dimensional data perception: Traditional vehicle security monitoring may only focus on one type of data (such as monitoring only V2X network traffic or only detecting vehicle logs), failing to integrate multi-dimensional information such as vehicle internal bus data, vehicle location information, and external threat intelligence. This single perspective leads to a one-sided understanding of the security situation, making it difficult to detect combined attacks or complex threats in a timely manner, increasing the risk of misjudgment and missed detection.

[0005] Lax decision-making and resource imbalance: Existing security decisions mostly rely on fixed rules or simple threshold triggers, lacking intelligent optimization and making it difficult to achieve a balance between security protection and vehicle performance. In the Internet of Vehicles, if the defense strategy is too conservative, it may cause delays in critical traffic information or even failure to deliver vehicle control commands in a timely manner due to excessive blocking of normal communication; conversely, if the strategy is too lenient, it may allow attacks to succeed.

[0006] Lack of online adaptive capabilities: Traditional machine learning models and rule-based systems cannot learn and update online according to environmental changes, and lack timely response mechanisms to emerging attack methods or system vulnerabilities in connected vehicles. The highly dynamic vehicle environment means that security strategies cannot effectively contain threats if they are not optimized in real time to adapt to new attack scenarios.

[0007] Insufficient distributed collaboration and privacy issues: Vehicle-to-everything (V2X) nodes are scattered across different vehicles and roadside infrastructure. Isolated security defense by individual vehicles creates a "data silo" phenomenon, where each vehicle makes decisions based on its limited perception, failing to grasp the overall threat situation. Furthermore, directly centralizing all vehicle data in the cloud for unified analysis can lead to privacy leaks and bandwidth burdens.

[0008] Therefore, to address the above problems, it is necessary to provide an intelligent security defense method that can adapt to the low-latency, high-dynamic, and heterogeneous distributed environment of the Internet of Vehicles (IoV) to enhance the ability to assess and optimize the information security situation of the IoV. This invention proposes a reinforcement learning-based method for assessing and optimizing the security situation of IoV communication to solve the aforementioned technical problems. Summary of the Invention

[0009] This invention addresses the aforementioned problems in vehicle network security protection through the following technical solution: This invention includes the following steps:

[0010] Step 1: Collect multi-source vehicle-to-everything (V2X) security data. This multi-source data includes vehicle-to-vehicle / vehicle-to-cloud communication network traffic data, in-vehicle system log information, vehicle device status information, etc.

[0011] Step 2: Preprocess the collected multi-source data and extract feature vectors containing vehicle network security situation characteristics;

[0012] Step 3: Input the feature vector into the reinforcement learning model, which outputs defensive actions based on the current security situation;

[0013] Step 4: Execute defensive actions and obtain a reward signal, which is generated based on changes in the security situation;

[0014] Step 5: Update the reinforcement learning model parameters based on the reward signal to optimize the vehicle-to-everything (V2X) security situation assessment and defense decision-making strategies.

[0015] Furthermore, the multi-source data in step 1 also includes in-vehicle network data, security incident alarm information, historical security policy data, and external threat intelligence data. By integrating multi-dimensional data from the Internet of Vehicles, a comprehensive security situation awareness system can be constructed. Among them, in-vehicle network data can reflect the interaction and abnormal situations of various control units within the vehicle. By monitoring the frequency and patterns of key control messages, abnormal command injection or unauthorized operations can be identified, thereby significantly improving the system's ability to perceive and respond to in-vehicle attacks (such as CAN bus injection attacks). Security event alarm information comes from in-vehicle intrusion detection systems, in-vehicle firewalls, roadside safety monitoring equipment, or cloud-based security management platforms, and includes fields such as alarm type, severity level, trigger rule number, affected vehicles / components, and confidence level. Historical security policy data records the configuration, trigger frequency, interception effect, and impact on in-vehicle communication and vehicle functions of various defense policies at different times (including in-vehicle firewall rules, IDS policy configuration, policy execution logs, and policy rollback records). External threat intelligence data is obtained in real time from the vehicle network security intelligence platform through standardized interfaces, introducing security intelligence from third parties or industry-shared sources, including suspicious vehicle IDs or device identifiers, malicious communication node IP addresses, forged signal source characteristics, known attack tools and their signatures, and vulnerability announcement information. After the system performs confidence assessment and labeling on the acquired intelligence data, it correlates and matches it with internally collected data to achieve rapid prediction and proactive defense against known threats.

[0016] Furthermore, the preprocessing in step 2 includes data cleaning, normalization, feature dimensionality reduction, and temporal feature extraction. Specifically, it involves: removing outliers from the collected data using an outlier detection algorithm, filling in missing data using a missing value interpolation method to ensure the integrity and reliability of the input data; and using a normalization method to unify the scale of numerical data to eliminate the impact of differences in feature dimensions from different sources on model training.

[0017]

[0018] in This represents the feature value of the i-th sample in the collected original data sequence. Let represent the feature value of the i-th sample after normalization, and x represent the original data set formed by this feature in the collected samples.

[0019] Based on a set time sliding window, extract d-dimensional normalized feature data from N consecutive time steps within the window and construct a feature matrix X.

[0020]

[0021]

[0022] Where N represents the total number of sample sequences contained within the time sliding window; m represents the index of the feature dimension. ; This represents the m-th normalized eigenvalue of the i-th sample in the feature matrix X; This represents the empirical mean of the m-th feature dimension within the time window; This represents the eigenvalues ​​after centering, which are determined by all... Construct a centered feature matrix .

[0023] Based on the centralized feature matrix Calculate the covariance between each feature dimension, construct the covariance matrix C, perform eigenvalue decomposition on C, and obtain the eigenvalues ​​and eigenvectors of the principal components.

[0024]

[0025]

[0026] in This represents the j-th eigenvalue obtained by solving the problem. Representation and eigenvalues The corresponding j-th eigenvector, and satisfying Furthermore, the d eigenvalues ​​obtained are arranged in descending order from largest to smallest, i.e., satisfying... .

[0027] Principal component analysis (PCA) combined with a time sliding window is used for feature dimensionality reduction and temporal feature extraction, retaining k such that the cumulative variance contribution rate is... 0.95.

[0028]

[0029] Furthermore, the reinforcement learning model in step 3 is an attention-enhanced deep Q-network (A-DQN). This model achieves vehicle-to-everything (V2X) security situation assessment and defense action selection in the following ways: First, it uses a convolutional neural network to extract spatial local features from the input security feature data; then, it calculates the weight distribution of different feature dimensions through a self-attention mechanism, dynamically highlighting the features most critical to security situation determination; finally, based on the weighted feature vectors, it outputs the value scores of each candidate defense action through a fully connected network and selects the optimal defense action with the highest value. The defense actions may include, but are not limited to, specific measures such as updating communication firewall rules (e.g., temporarily blocking suspicious vehicle nodes or communication channels), adjusting intrusion detection strategy parameters (e.g., tightening anomaly detection thresholds, enabling advanced detection rules), and prioritizing vulnerability remediation (e.g., prioritizing updates to in-vehicle software with high-risk vulnerabilities).

[0030] Furthermore, the reward signal in step 4 is designed as a multi-dimensional weighted function, specifically including:

[0031] Positive rewards: percentage reduction in attack traffic, index reduction in system vulnerability risk, and retention rate of normal business traffic;

[0032] Negative penalties: resource consumption cost of defensive actions, and duration of normal communication interruption caused by misjudgment; the reward function expression is:

[0033]

[0034]

[0035]

[0036]

[0037]

[0038]

[0039] Among them, α, β, γ, δ, These are the weighting coefficients. This indicates the size of the attack traffic detected by the system before executing defensive actions. This indicates the size of the attack traffic detected by the system after the defensive action was performed. This represents the overall system vulnerability risk score obtained through situational assessment before implementing defensive actions. This indicates the residual vulnerability risk score of the system after defensive actions have been performed. It is the ratio of the average CPU utilization during the execution of the defense strategy to the system's maximum CPU capacity. It is the ratio of network bandwidth consumption to total bandwidth capacity when implementing a defense strategy. It is the ratio of memory usage when executing a defense strategy to the system's maximum available memory. It is the first The actual downtime caused by the execution of defensive actions by the first business node is the same as that of the second. The ratio of the maximum tolerable downtime threshold allowed by the service level agreement for each business node.

[0040] Furthermore, the model parameter update in step 5 employs an improved reinforcement learning training strategy, combining an experience replay buffer and a target network technique to enhance training stability and efficiency. Specifically, this includes: continuously storing historical data of vehicle-threat interactions (including states, actions, rewards, and next states) in the experience replay buffer; uniformly and randomly sampling a batch of interaction data from the buffer for training, breaking the temporal correlation of data to avoid bias caused by high correlation between adjacent samples and improving sample utilization efficiency; introducing a target network mechanism to periodically update the parameters of the main reinforcement learning network to the target network, using the target network to calculate a stable Q-value reference, avoiding estimation oscillations caused by frequent updates of the main network parameters during training, thereby stabilizing the training process of the reinforcement learning model and accelerating convergence. The loss function and target network parameter updates during training are as follows:

[0041]

[0042]

[0043] in This represents the loss function of the main reinforcement learning network at the current training iteration step; This represents the model parameters in the current main reinforcement learning network; denoted by ; D represents the experience replay buffer region; s represents the current state in the sampled historical interaction sequence; a represents the defensive action taken by the system in state s; r represents the immediate reward value obtained by the system after performing action a in state s. This indicates the next state that the system transitions to after executing action 'a'; Indicates the next state To take one action from all possible actions; This represents the reward discount factor, used to weigh the relative importance of immediate rewards versus long-term future rewards; This represents the estimated value of taking action a in state s, obtained from the evaluation of the main network. This indicates the result obtained by the target network evaluation in the next state. Take all possible actions The maximum estimated value of action that can be obtained in the middle; This represents the soft update coefficient of the target network.

[0044] Furthermore, the method also includes a dynamic defense strategy adjustment mechanism. Specifically, after executing defense actions, a real-time monitoring module continuously assesses the security posture of the vehicle network. This monitoring module integrates an anomaly detection algorithm based on isolated forests and a time series prediction model to perform real-time analysis of vehicle communication and status data. When signs of new attack types are detected or the system security risk index shows an upward trend for three consecutive monitoring cycles, steps 3 to 5 are re-executed to adaptively adjust the defense strategy. By promptly feeding the detected anomalies into the reinforcement learning decision loop, dynamic re-optimization of the strategy is achieved.

[0045]

[0046]

[0047] in As weight, This indicates the percentage of attack traffic. This indicates the system vulnerability risk index. This indicates the percentage of normal business traffic affected. This indicates the frequency of security incidents. The threshold for "a continuous increase in the security risk index for three consecutive monitoring periods" is set based on the reasonableness of changes in the cybersecurity situation. for .

[0048] Furthermore, the reinforcement learning model of this invention is deployed using a multi-agent collaborative mechanism within a federated learning framework, enabling multiple vehicle nodes and roadside infrastructure to collaboratively optimize security strategies. Specifically, different vehicles, roadside units, and other network nodes are defined as independent agents. Each agent runs a reinforcement learning model locally and trains it based on local observation data. Each agent periodically shares encrypted model parameter gradients through a secure communication channel, rather than directly exchanging raw sensitive data, ensuring that each vehicle's private data does not leave its local machine. When a central coordination node (such as a cloud security server or a roadside aggregation node) aggregates global model parameters, differential privacy technology is introduced to perturb and protect the gradient information uploaded by each node, protecting the data privacy of each node at the protocol and algorithm levels. Simultaneously, a reward function for federated collaboration is designed, incorporating global defense effectiveness indicators such as cross-vehicle attack blocking rates into the reward calculation to encourage collaborative optimization of strategies among agents towards overall security optimization.

[0049] The present invention has the following advantages over the prior art:

[0050] Multi-source data fusion for comprehensive situational awareness: This invention integrates multi-source information from the vehicle-to-everything (V2X) environment, including vehicle-to-vehicle communication, vehicle logs, internal bus data, historical policies, and external intelligence. It overcomes the bottleneck of traditional security monitoring being limited to a single data source, achieving comprehensive awareness of abnormal behavior within vehicles, known attack characteristics, and unknown risk signs. Through multi-dimensional data correlation analysis, the system can accurately identify complex threats such as forged messages and abnormal control commands, improving the comprehensiveness and accuracy of threat detection.

[0051] Intelligent decision-making optimization balances security and availability: Employing reinforcement learning combined with attention mechanisms and multi-dimensional reward design, the system dynamically weighs the effectiveness of security protection against the availability of vehicle communication services. The model can autonomously learn the optimal defense measures at different times and in different scenarios, avoiding delays in important messages or waste of resources due to over-defense. By assigning weights to key features and imposing penalties on resource consumption, an optimized balance between security and real-time performance is achieved, ensuring that the vehicle-to-everything (V2X) system can maintain normal operation of core functions even when attacked.

[0052] Real-time Adaptation and Rapid Response: This invention constructs a real-time monitoring and policy adaptive adjustment mechanism, which can promptly capture new abnormal patterns or attack attempts in the rapidly changing environment of the Internet of Vehicles (IoV) and quickly retrain or adjust defense strategies. Compared with traditional static rules, this invention significantly shortens the time from threat detection to response. When sudden security events such as GPS signal spoofing or new types of message forgery occur, the system can automatically adjust its strategy to eliminate the threat with extremely low latency, improving the timeliness of driving safety protection.

[0053] Federated collaborative defense enhances overall security: A multi-vehicle collaborative defense framework based on federated learning enables vehicles and infrastructure nodes to share security insights and form collaborative combat capabilities. Without exposing their raw data, multiple nodes jointly train more comprehensive security policies, effectively blocking the propagation of attacks across regions and vehicles. This collaborative mechanism overcomes the bottleneck of unifying security policies among heterogeneous vehicle systems, achieving globally consistent threat defense while protecting user privacy, and significantly improving the overall security posture of the connected vehicle network. Attached Figure Description

[0054] Figure 1 This is a flowchart of the present invention. Detailed Implementation

[0055] The embodiments of the present invention are described in detail below. These embodiments are implemented based on the technical solution of the present invention, and provide detailed implementation methods and specific operation processes. However, the scope of protection of the present invention is not limited to the following embodiments.

[0056] Example: Figure 1As shown, this embodiment provides a technical solution: a method for assessing and optimizing the security situation of vehicle-to-everything (V2X) communication based on reinforcement learning, comprising the following steps:

[0057] Step 1: Collect multi-source vehicle-to-everything (V2X) security data. This multi-source data includes vehicle-to-vehicle / vehicle-to-cloud communication network traffic data, in-vehicle system log information, vehicle device status information, etc.

[0058] Step 2: Preprocess the collected multi-source data and extract feature vectors containing vehicle network security situation characteristics;

[0059] Step 3: Input the feature vector into the reinforcement learning model, which outputs defensive actions based on the current security situation;

[0060] Step 4: Execute defensive actions and obtain a reward signal, which is generated based on changes in the security situation;

[0061] Step 5: Update the reinforcement learning model parameters based on the reward signal to optimize the vehicle-to-everything (V2X) security situation assessment and defense decision-making strategies.

[0062] The multi-source data in step 1 includes network traffic data (such as V2X message messages, vehicle location and status broadcast information, etc.) generated during vehicle-to-vehicle (V2V) and vehicle-to-cloud (V2C) communication, vehicle system log information (such as vehicle operating system logs, application logs, and security event logs), and basic data such as vehicle equipment status information (such as the operating status of each electronic control unit (ECU), sensor readings, and communication module status). In addition, to enrich the dimensions of security situation awareness, the following information is preferably also collected: vehicle internal network data, namely message data and diagnostic information on the vehicle internal communication bus (such as CAN bus), covering the communication status between various modules inside the vehicle; security event alarm information, from real-time alarms from the vehicle intrusion detection system (IDS), firewall, vehicle operating system security monitoring, and roadside / cloud security management platform, recording the occurrence of potential attacks or abnormal events; historical security policy data, including the security policy configurations, policy execution results and logs previously adopted by the vehicle and related systems (such as previously implemented filtering rules, authentication policy adjustments and their effectiveness), and records of policy changes or rollbacks; external threat intelligence data, obtained through vehicle-to-everything (V2X) security intelligence sharing channels, including information on recently known attack methods against V2X, lists of suspicious node identifiers (such as malicious vehicle IDs, spoofed base station information), disclosed vulnerability information and patches, and intelligence on hacker group activities. The aforementioned multi-source data is collected collaboratively through vehicle-mounted terminals, roadside units, and cloud servers: vehicles collect local data and received communication data through vehicle-mounted terminal units, roadside units monitor communications within the area, and the cloud security platform integrates information from various sources and distributes the latest threat intelligence. Through multi-source data fusion, the system can comprehensively perceive the security situation of the Internet of Vehicles.

[0063] The preprocessing in step 2 includes: First, using an outlier detection algorithm to identify and remove obvious outliers in sensor data and logs, and then filling in missing data parts using interpolation or default values. If the data is numerical (such as network traffic rate, system log frequency), mean or median interpolation is preferred; if the data has temporal correlation (such as sequence data of device status changes over time), linear interpolation is used to preserve the temporal trend of the data. Regarding normalization methods, normalization transformations are applied to numerical features from different sources to scale them to similar scales. Min-Max normalization is used to map the data to the [0,1] interval, effectively preserving the relative relationship of data distribution and avoiding model bias due to differences in units.

[0064]

[0065] in This represents the feature value of the i-th sample in the collected original data sequence. Let represent the feature value of the i-th sample after normalization, and x represent the original data set formed by this feature in the collected samples.

[0066] Based on a set time sliding window, extract d-dimensional normalized feature data from N consecutive time steps within the window and construct a feature matrix X.

[0067]

[0068]

[0069] Where N represents the total number of sample sequences contained within the time sliding window; m represents the index of the feature dimension. ; This represents the m-th normalized eigenvalue of the i-th sample in the feature matrix X; This represents the empirical mean of the m-th feature dimension within the time window; This represents the eigenvalues ​​after centering, which are determined by all... Construct a centered feature matrix .

[0070] Based on the centralized feature matrix Calculate the covariance between each feature dimension, construct the covariance matrix C, perform eigenvalue decomposition on C, and obtain the eigenvalues ​​and eigenvectors of the principal components.

[0071]

[0072]

[0073] in This represents the j-th eigenvalue obtained by solving the problem. Representation and eigenvalues The corresponding j-th eigenvector, and satisfying Furthermore, the d eigenvalues ​​obtained are arranged in descending order from largest to smallest, i.e., satisfying... .

[0074] Principal component analysis (PCA) combined with a time sliding window is used for feature dimensionality reduction and temporal feature extraction, retaining k such that the cumulative variance contribution rate is... 0.95.

[0075]

[0076] This setting achieves both dimensionality compression and maximum preservation of key information regarding cybersecurity characteristics, balancing dimensionality reduction efficiency with information integrity. The time sliding window size is set to 5-10 minutes because changes in the cybersecurity situation have short-term continuity (e.g., attack traffic bursts typically last for several minutes). This window length can effectively capture temporal characteristics (e.g., traffic fluctuation trends, changes in the frequency of security events) while avoiding information delays caused by excessively large windows.

[0077] The reinforcement learning model in step 3 is an attention-enhanced deep Q-network (A-DQN), which achieves situation assessment and action selection in the following way:

[0078] The model structure consists of three parts: a convolutional neural network, a self-attention layer, and a Q-value estimation layer. First, the feature vectors extracted in step 2 are further refined using a convolutional neural network. Two convolutional layers are set: the first layer is equipped with 32 filters, and the second layer is equipped with 64 filters. The kernel size of both convolutional layers is 3×3, which is suitable for the feature extraction requirements of high-dimensional feature matrices. The activation function is ReLU to handle the non-linear features of the data and alleviate the gradient vanishing problem. After the convolutional layers, a 2×2 max pooling layer is set to compress the feature dimension and improve computational efficiency. Then, the self-attention mechanism layer is entered. The self-attention mechanism adopts scaled dot product attention, setting 2 to 4 attention heads. This number can balance computational efficiency and feature weight discrimination, and effectively identify the difference between high-risk features (such as attack type correlation and vulnerability impact range index) and low-risk features (such as normal business traffic frequency). The fully connected layer is used to approximate the Q-value function, setting 2 hidden layers: the first layer contains 128 neurons and the second layer contains 64 neurons, avoiding overfitting due to excessive model complexity. The number of neurons in the output layer is consistent with the number of defense actions. The activation function of the hidden layer is ReLU, and the output layer uses a linear activation function. The system sets a set of security defense actions that can be taken. The model will calculate the corresponding value score and select the action with the largest Q value to execute.

[0079] Examples of defensive actions include: adjusting vehicle communication firewall policies (such as isolating suspicious node communications, switching secure communication channels or protocols), modifying intrusion detection system parameter configurations (such as lowering the anomaly detection sensitivity threshold to reduce false alarms, or increasing sensitivity to promptly intercept suspicious behavior, adaptively adjusting according to the current situation), and managing in-vehicle software and firmware updates (such as dynamically adjusting the priority and time window of patch installation for detected vulnerability threats). Through the above structure, convolutional networks effectively capture key anomaly patterns in multi-source data, the self-attention mechanism highlights the role of high-risk features in decision-making, and the fully connected layer performs fine-grained evaluation of each action, ultimately achieving accurate decision-making for defensive actions in complex vehicular network environments.

[0080] The reward signal in step 4 is designed as a multi-dimensional weighted function, specifically including:

[0081] The main reward components include: (1) Attack suppression effect: If the defense action effectively reduces malicious communication traffic (such as intercepting data injection of man-in-the-middle attacks and preventing the spread of forged messages), a positive reward will be given, and the greater the reduction, the higher the reward; if the security risk of the system is reduced (such as patching vulnerabilities to reduce the overall risk index of the system), a positive reward will also be given; (2) Communication service guarantee: If the defense action successfully maintains the continuity of normal vehicle communication and functions while blocking the attack (such as most legitimate V2V messages are still delivered smoothly without triggering false alarms or false actions), a positive reward will be given according to the proportion of normal business traffic maintained, so as to encourage the strategy to take into account service availability; (3) Resource and side effect cost: The defense action will consume certain computing, communication and other resources. If a certain action occupies a large amount of vehicle computing resources or bandwidth, a negative penalty will be given according to the degree of resource consumption; at the same time, if the normal operation is intercepted or interfered with due to misjudgment (such as mistakenly blocking the vehicle's emergency braking signal or causing the vehicle communication to be suspended), a negative penalty will be given according to the duration and scope of the business interruption.

[0082]

[0083]

[0084]

[0085]

[0086]

[0087]

[0088] Among them, α, β, γ, δ, These are the weighting coefficients. This indicates the size of the attack traffic detected by the system before executing defensive actions. This indicates the size of the attack traffic detected by the system after the defensive action was performed. This represents the overall system vulnerability risk score obtained through situational assessment before implementing defensive actions. This indicates the residual vulnerability risk score of the system after defensive actions have been performed. It is the ratio of the average CPU utilization during the execution of the defense strategy to the system's maximum CPU capacity. It is the ratio of network bandwidth consumption to total bandwidth capacity when implementing a defense strategy. It is the ratio of memory usage when executing a defense strategy to the system's maximum available memory. It is the first The actual downtime caused by the execution of defensive actions by the first business node is the same as that of the second. The ratio of the maximum tolerable downtime threshold allowed by the service level agreement for each business node.

[0089] Step 5, parameter updating, employs an improved policy gradient algorithm combined with an experience replay buffer and a target network technique. Specifically, historical data generated from the interaction between the model and the network environment (including the current situation feature vector, executed defensive actions, obtained reward signals, and the next situation feature vector) is first stored in the experience replay buffer. This breaks the temporal correlation of data and avoids model training fluctuations caused by data correlation in traditional policy gradient algorithms. During training, batch training is performed using historical data uniformly sampled from the buffer, increasing the diversity of training samples, improving sample utilization efficiency, and alleviating the problem of sample inefficiency. Simultaneously, a target network is introduced. The target network has the same structure as the main network, but its parameter update frequency is lower. The parameters of the main network are periodically synchronized to the target network, and the target network is used to calculate the target Q-value, replacing the main network's own calculation of the target Q-value. This avoids Q-value function oscillations caused by frequent updates of the main network parameters. The loss function and target network parameter updates during training are as follows:

[0090]

[0091]

[0092] in This represents the loss function of the main reinforcement learning network at the current training iteration step; This represents the model parameters in the current main reinforcement learning network; denoted by ; D represents the experience replay buffer region; s represents the current state in the sampled historical interaction sequence; a represents the defensive action taken by the system in state s; r represents the immediate reward value obtained by the system after performing action a in state s. This indicates the next state that the system transitions to after executing action 'a'; Indicates the next state To take one action from all possible actions; This represents the reward discount factor, used to weigh the relative importance of immediate rewards versus long-term future rewards; This represents the estimated value of taking action a in state s, as determined by the main network. This indicates the result obtained by the target network evaluation in the next state. Take all possible actions The maximum estimated value of action that can be obtained in the middle; This represents the soft update coefficient of the target network.

[0093] The method also includes a dynamic defense strategy adjustment mechanism, specifically: after the defense action is performed, the network security situation is continuously evaluated through a real-time monitoring module. The monitoring module integrates an isolated forest anomaly detection algorithm and a time series prediction model. When a new attack type is detected (such as the deviation of the traffic pattern of an unknown threat exceeding the threshold) or the security risk index rises for three consecutive monitoring cycles, steps 3 to 5 are re-executed to achieve adaptive adjustment of the defense strategy.

[0094]

[0095]

[0096] in As weight, This indicates the percentage of attack traffic. This indicates the system vulnerability risk index. This indicates the percentage of normal business traffic affected. This indicates the frequency of security incidents. The threshold for "a continuous increase in the security risk index for three consecutive monitoring periods" is set based on the reasonableness of changes in the cybersecurity situation. for .

[0097] Steps 1 to 5 above describe the reinforcement learning security decision-making process on any single node (e.g., a vehicle or a roadside unit). To fully leverage the collaborative defense potential of distributed nodes in the vehicle-to-everything (V2X) network, this invention further utilizes a federated learning framework for collaborative training and decision optimization among multiple agents. Specifically, each connected vehicle and its associated roadside infrastructure equipment are considered as independent agents. Each agent executes the aforementioned reinforcement learning algorithm to evaluate its own environment and output a local defense strategy. Each agent accumulates a certain amount of training updates locally (e.g., after experiencing several state-action interactions), and then uploads its model gradients or updated policy parameters, after encryption, to a central coordination server for federated learning in the cloud. After collecting the encrypted parameters uploaded by multiple agents, the central server uses a federated average aggregation algorithm to merge the updates from each local model, obtaining globally shared model parameters. Differential privacy and other technologies are used to protect the aggregation process to prevent the inference of the original data features of individual nodes. The server then distributes the updated global model back to each vehicle node. Local agents update their own models accordingly, thereby enabling each node to share the defense experience learned from each other. This process is iterated periodically, allowing the model to gradually converge in federated collaboration. Simultaneously, to ensure policy consistency among agents within the federated learning framework, this invention incorporates collaborative metrics into the reward design of reinforcement learning. For example, a global attack blocking rate metric is set: when multiple vehicles cooperate to successfully prevent a cross-vehicle attack propagation event, each participating agent receives an additional bonus in the reward; if a vehicle's defense decision contributes to improving the overall security posture of the vehicle network, this positive effect is also reflected to that vehicle through the reward function. Through this collaborative mechanism, the system encourages nodes to take actions that benefit global security rather than merely optimizing local performance. Ultimately, federated learning multi-agent collaboration achieves globally linked vehicle network security defense: while ensuring that sensitive data of each vehicle is not leaked, nodes share knowledge and collaboratively evolve stronger security strategies capable of coping with large-scale distributed attacks.

[0098] Although embodiments of the present invention have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those skilled in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of the present invention.

Claims

1. A method for security situation assessment and decision optimization in vehicle-to-everything (V2X) communication based on reinforcement learning, characterized in that, Includes the following steps: Step 1: Collect multi-source vehicle network security data, including network traffic data of vehicle-to-vehicle (V2V) and vehicle-to-cloud (V2C) communication, vehicle system log information, and vehicle equipment status information; Step 2: Preprocess the collected multi-source data and extract feature vectors containing vehicle network security situation characteristics; Step 3: Input the feature vector into the reinforcement learning model, and the model outputs defensive actions based on the current security situation; Step 4: Execute defensive actions and obtain reward signals, which are generated based on changes in the vehicle network security situation; Step 5: Update the reinforcement learning model parameters based on the reward signal to optimize the vehicle-to-everything (V2X) security situation assessment and defense decision-making strategies.

2. The method for vehicle-to-everything (V2X) security situation assessment and decision optimization based on reinforcement learning according to claim 1, characterized in that: The multi-source data in step 1 also includes vehicle internal network data, security event alarm information, historical security policy data, and external threat intelligence data. Among them, vehicle internal network data obtains communication information of various control units in the vehicle through the vehicle bus; security event alarm information is generated by security devices such as intrusion detection and firewalls in the vehicle or infrastructure; historical security policy data records the security policies implemented at different time periods and their effects; and external threat intelligence data is obtained in real time from security intelligence sources through standard interfaces.

3. The method for vehicle-to-everything (V2X) security situation assessment and decision optimization based on reinforcement learning according to claim 1, characterized in that: Step 2 is as follows: outlier points in the collected data are removed by outlier detection algorithm, missing data is filled by missing value interpolation method to ensure the integrity and reliability of input data; normalization method is used to unify the scale of numerical data to eliminate the impact of differences in feature units from different sources on model training. in This represents the feature value of the i-th sample in the collected original data sequence. Let represent the feature value of the i-th sample after normalization, and x represent the original data set formed by this feature in the collected samples. Based on a defined time sliding window, d-dimensional normalized feature data from N consecutive time steps within the window are extracted to construct a feature matrix X. Where N represents the total number of sample sequences contained within the time sliding window; m represents the index of the feature dimension. ; This represents the m-th normalized eigenvalue of the i-th sample in the feature matrix X; This represents the empirical mean of the m-th feature dimension within the time window; This represents the eigenvalues ​​after centering, which are determined by all... Construct a centered feature matrix , Based on the centralized feature matrix Calculate the covariance between each feature dimension, construct the covariance matrix C, perform eigenvalue decomposition on C, and obtain the eigenvalues ​​and eigenvectors of the principal components. in This represents the j-th eigenvalue obtained by solving the problem. Representation and eigenvalues The corresponding j-th eigenvector, and satisfying Furthermore, the d eigenvalues ​​obtained are arranged in descending order from largest to smallest, i.e., satisfying... , Principal component analysis (PCA) combined with a time sliding window is used for feature dimensionality reduction and temporal feature extraction, retaining k such that the cumulative variance contribution rate is... 0.95, 。 4. The method for vehicle-to-everything (V2X) security situation assessment and decision optimization based on reinforcement learning according to claim 1, characterized in that: The reinforcement learning model in step 3 is an attention-enhanced deep Q-network (A-DQN). The model extracts the local spatial features of the feature vector through a convolutional neural network, calculates the weight distribution of different feature dimensions through a self-attention mechanism, and outputs a value score of the defense action and selects the optimal action based on the weighted feature vector through a fully connected network. The defense actions include updating communication firewall rules, adjusting intrusion detection strategies, and prioritizing vulnerability repairs.

5. The method for vehicle-to-everything (V2X) security situation assessment and decision optimization based on reinforcement learning according to claim 1, characterized in that: The reward signal in step 4 is designed as a multi-dimensional weighted function, specifically including: Positive rewards: percentage reduction in attack traffic, index reduction in system vulnerability risk, and retention rate of normal business traffic; Negative penalties: resource consumption cost of defensive actions, and service interruption duration caused by misjudgment; the reward function expression is: Where α, β, γ, δ and Here are the weighting coefficients: ΔA represents the attack traffic reduction rate, and ΔV represents the vulnerability risk reduction rate. C represents availability retention rate, D represents resource consumption index, and D represents service interruption duration.

6. The method for vehicle-to-everything (V2X) security situation assessment and decision optimization based on reinforcement learning according to claim 1, characterized in that: Step 5 parameter updates employ an improved policy gradient algorithm, combined with an empirical replay buffer and target network techniques: Store historical interaction data in the experience replay buffer; Batch training is performed using data from a uniformly sampled buffer to reduce data correlation; Regularly synchronize the main network parameters to the target network to stabilize the Q-value function training process and improve the model convergence speed.

7. The method for vehicle-to-everything (V2X) security situation assessment and decision optimization based on reinforcement learning according to claim 1, characterized in that: The method also includes a dynamic defense strategy adjustment mechanism, specifically: after the defense action is performed, the security situation of the Internet of Vehicles is continuously evaluated through a real-time monitoring module. The monitoring module integrates the isolated forest anomaly detection algorithm and the time series prediction model. When a new attack type is detected or the security risk index rises for three consecutive monitoring cycles, steps 3 to 5 are re-executed to achieve adaptive adjustment of the defense strategy.

8. The method for vehicle-to-everything (V2X) security situation assessment and decision optimization based on reinforcement learning according to claim 1, characterized in that: The multi-agent cooperation mechanism using the federated learning framework is as follows: Different vehicle nodes or roadside infrastructure are defined as independent intelligent agents, and each intelligent agent maintains a local reinforcement learning model. The agents periodically share encrypted model parameter gradients through a secure channel, rather than the original data; when the central coordinator aggregates global model parameters, it uses differential privacy technology to protect the privacy data of each node. Design a collaborative reward function that incorporates the cross-regional attack blocking rate into the reward calculation to promote collaborative optimization of strategies among agents.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements a reinforcement learning-based method for assessing the security situation and optimizing decisions in the Internet of Vehicles, as described in any one of claims 1 to 8.

10. A computer-readable storage medium storing computer instructions thereon, characterized in that, When executed by the processor, the computer instructions implement a reinforcement learning-based method for assessing the security situation and optimizing decisions in the Internet of Vehicles (IoV) as described in any one of claims 1-8.