Training method and allocation method of Internet of Vehicles resource allocation model

By introducing a RIS-assisted resource allocation model in the Internet of Vehicles, utilizing self-attention and cross-attention mechanisms, and combining diffusion strategies with the SAC algorithm, the problems of insufficient signal coverage and high power consumption of 5G millimeter wave technology in the Internet of Vehicles are solved, and the signal coverage is expanded and the communication quality is improved.

CN120632783APending Publication Date: 2025-09-12NANCHANG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510760510.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-09
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

In the Internet of Vehicles environment, 5G millimeter wave technology has insufficient signal coverage and high system power consumption under non-line-of-sight conditions. Traditional solutions increase deployment costs and energy consumption, making it difficult to apply them on a large scale in the Internet of Vehicles field.

Method used

A vehicle network resource allocation model assisted by reconfigurable intelligent reflecting surface (RIS) is adopted. Self-attention and cross-attention mechanisms are used to enhance multi-agent collaborative decision-making. Diffusion strategy and SAC algorithm are combined to optimize signal coverage and communication performance.

Benefits of technology

It has achieved the expansion of signal coverage and improvement of communication quality under non-line-of-sight conditions, reduced system power consumption, improved the information sharing and collaborative optimization capabilities of multi-agent systems, and solved the limitations of traditional technologies in coverage and efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120632783A_ABST
    Figure CN120632783A_ABST
Patent Text Reader

Abstract

The invention discloses a training method and a distribution method of an Internet of Vehicles resource distribution model. The training method specifically comprises the following steps: acquiring a state vector of an intelligent agent; inputting the state vector of the agent into an actor network to obtain a first strategy action vector; capturing key information of the state vector of each agent by adopting a self-attention mechanism so as to generate self-attention output; realizing information sharing among the intelligent agents by adopting a cross attention mechanism so as to generate cross attention output; fusing the self-attention output and the cross-attention output to output a fused feature representation; inputting the obtained fusion feature representation into a first MLP network for dimension reduction processing, adding noise to the feature representation after dimension reduction, and inputting the feature representation into a diffusion model to output predicted noise; obtaining a second strategy action vector based on the predicted noise; according to the invention, a multi-agent soft actor-commentator algorithm with an enhanced attention mechanism is provided, and global resource optimization and communication performance improvement of the RIS-assisted Internet of Vehicles are realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of vehicle networking, and in particular to a training method and an allocation method for a vehicle networking resource allocation model. Background Art

[0002] With the rapid development of wireless communication technology and artificial intelligence, intelligent transportation systems are becoming highly interconnected and intelligent in the era of the Internet of Things (IoT), leading to increasing global attention for the Internet of Vehicles (IoV). IoV is a quintessential application of the deep integration of AI and transportation. Its core is the comprehensive collaboration of "people, vehicles, roads, and the cloud." This not only improves road traffic efficiency but also effectively reduces traffic accidents by over 30%. IoV enables information exchange and interoperability between vehicles and various entities, including various interaction methods such as vehicle-to-vehicle (V2V), vehicle-to-infrastructure (V2I), and vehicle-to-pedestrian (V2P), collectively referred to as vehicle-to-everything (V2X) communication.

[0003] The application scenarios of the Internet of Vehicles (IoV) are extremely diverse, encompassing autonomous driving, real-time traffic information sharing, immersive in-vehicle entertainment experiences, and safety warning systems. However, these applications pose significant challenges to communication systems, including requirements for data transmission speed, real-time performance, network coverage, user fairness, deployment costs, and security. In IoV environments, particularly due to the high speeds of vehicles and the complex and ever-changing urban environment, direct line of sight (LoS) links are often blocked by various moving obstacles. Furthermore, the dynamically changing communication conditions between vehicles lead to significant differences in communication quality between users in different vehicles. Therefore, how to improve the fairness and security of IoV communication systems within the context of limited communication resources has become a key issue that needs to be addressed.

[0004] The development of fifth-generation mobile networks (5G) has brought new opportunities to the connected vehicle (IoV). Compared to previous generations of wireless communication systems, 5G offers higher data rates, lower latency, and more reliable communication performance. 5G millimeter-wave technology, in particular, with its abundant spectrum resources and concentrated energy, is considered a key technology for achieving high-capacity data transmission in the IoV. Theoretically, it can support transmission rates of up to gigabits per second, significantly improving the IoV's data processing capabilities.

[0005] However, 5G millimeter wave technology also faces significant challenges. Due to its high frequency, millimeter wave signals have weak penetration capabilities and significant path loss, making their communication performance heavily dependent on the presence of line-of-sight links. In the connected vehicle (IoV) scenario, millimeter wave signals are easily obstructed by obstacles on the road, which limits their effective coverage. Although traditional solutions include deploying wireless relay nodes, using large-scale antenna arrays, and applying beamforming technology, these methods generally increase system deployment costs and energy consumption, limiting the large-scale application of 5G millimeter wave technology in the IoV field. Therefore, how to enhance signal coverage and reduce system power consumption under non-line-of-sight conditions has become a key technical challenge in the application of 5G millimeter wave technology in the IoV field.

[0006] Recently, reconfigurable intelligent reflective surface (RIS) technology has been introduced as a promising innovative solution for connected vehicle communication systems. RIS consists of numerous low-cost, energy-efficient passive reflective units, each capable of precisely controlling the amplitude and phase of the received signal. This radically overturns the passive, untamperable nature of traditional wireless channels. Unlike traditional wireless relays, which require additional signal processing and transmission modules, RIS enhances communication link quality solely through passive reflection, eliminating the need for additional energy for signal amplification and processing, resulting in significant energy efficiency advantages.

[0007] In connected vehicle applications, RIS-based communication systems can effectively expand signal coverage, optimize non-line-of-sight (NLOS) performance, and enhance communication quality, even without the addition of roadside units (ROS). Compared to traditional multi-antenna systems, RIS is not only more cost-effective and requires less hardware investment, but also flexible enough to be installed on various curved surfaces, meeting the needs of diverse connected vehicle environments. The integration of RIS technology into connected vehicle communication systems represents an innovative breakthrough, potentially overcoming the coverage and performance limitations of traditional technologies and unlocking the potential value of connected vehicles in diverse future application scenarios.

[0008] RIS-assisted wireless communication systems have been extensively researched in various fields, including physical layer security, multi-cell cooperative communications, drone networks, cognitive radio, and non-orthogonal multiple access. Despite its numerous advantages, RIS technology still faces challenges in complex and volatile network environments in connected vehicles. Traditional optimization methods often struggle to adapt to the real-time decision-making requirements of highly dynamic scenarios. Summary of the Invention

[0009] The purpose of the present invention is to improve and innovate the shortcomings and problems existing in the background technology, and to provide a training method and allocation method for a vehicle network resource allocation model.

[0010] According to a first aspect of the present invention, a method for training a vehicle network resource allocation model is provided, which specifically includes the following steps: Get the state vector of the agent; Input the agent’s state vector into the SAC actor network to obtain the first policy action vector; A self-attention mechanism is used to capture the key information of each agent's state vector to generate self-attention outputs; a cross-attention mechanism is used to achieve efficient information sharing between agents to generate cross-attention outputs; the self-attention outputs and cross-attention outputs are fused to output a fused feature representation; The obtained fusion feature representation is input into the first MLP network for dimensionality reduction. Noise is added to the reduced feature representation and then input into the diffusion model to output predicted noise. The second strategy action vector is obtained based on the predicted noise. fusing the first policy action vector and the second policy action vector to generate a target policy action vector; The parameters of the diffusion model, cross-attention mechanism and self-attention mechanism are updated according to the predicted noise output by the diffusion model and the actual noise added; the target strategy action vector is used to extract the target Q values ​​corresponding to the target dual-Q networks corresponding to the SAC target critic network; the loss value of the dual-Q network corresponding to the critic network is calculated according to the smaller value of the target Q value, and the parameters of the dual-Q network are updated based on the loss value of the dual-Q network; and the loss value of the SAC actor network is calculated according to the smaller value of the loss value of the dual-Q network, and the parameters of the actor network are updated based on the loss value of the SAC actor network.

[0011] A further solution is that the self-attention mechanism is used to capture the key information of each agent's state vector to generate the self-attention output, specifically including: Through the state encoding network Extract features: ; in, Representing an agent The state encoding feature, is the state encoding network, It is an intelligent agent The state vector of Divide the state encoding features of the agent into several sub-vectors; Input the sub-vector of the agent's state encoding feature into the self-attention mechanism to map the sub-vector of the state encoding feature to different subspaces; The feature vectors corresponding to the different subspaces obtained by mapping are concatenated to generate the self-attention output.

[0012] A further solution is to use the cross-attention mechanism to achieve efficient information sharing between agents to generate cross-attention output, specifically including: Generate state encoding features of each agent through a state encoding network with shared parameters; Agent Generate feature representations through an encoder with shared parameters: ; in, Representing an agent The state encoding feature, is the state encoding network, It is an intelligent agent The original state vector of The state encoding features of each agent are composed into a sequence and input into the cross attention mechanism, and the state encoding features of each agent are mapped to different subspaces to generate the state encoding features of each agent. Cross attention output.

[0013] A further solution is that the fusion of the self-attention output and the cross-attention output to output the fusion feature representation specifically includes: The self-attention output and cross-attention output of each agent are fused, and the calculation formula is as follows: ; ; in, represents the self-attention output; represents the cross attention output; represents the fusion of self-attention output and state encoding features, Represents the fusion of cross-attention output and fused features, LayerNorm represents layer normalization, and Dropout represents dropout regularization; Further feature extraction through feedforward network: ; in are the learnable weights and bias parameters of the feedforward network.

[0014] Use residual connection to output fusion feature representation: ; in is the output fusion feature representation.

[0015] A further solution is that fusing the first strategy action vector and the second strategy action vector to generate a target strategy action vector specifically includes: The first strategy action is fused with the second strategy action through linear combination to generate the target strategy action vector: ; ; in, Representing an agent The first policy action vector of Representing an agent In the time slot Channel selection, Representing an agent In the time slot Power control, Representing an agent In the time slot The RIS phase is recommended, the mixing coefficient Dynamic adjustment as training progresses: , t is the number of iterations, initial value , attenuation rate .

[0016] According to a second aspect of the present invention, a method for allocating a vehicle network resource allocation model is provided, which specifically includes the following steps: Get the state vector of the agent; Input the agent's state vector into the trained SAC actor network to obtain the first policy action vector; A self-attention mechanism is used to capture the key information of each agent's state vector to generate self-attention outputs; a cross-attention mechanism is used to achieve efficient information sharing between agents to generate cross-attention outputs; the self-attention outputs and cross-attention outputs are fused to output a fused feature representation; The obtained fusion feature representation is input into the first MLP network for dimensionality reduction. The reduced feature representation is directly input into the trained diffusion model to generate the second strategy action vector. The first strategy action vector and the second strategy action vector are fused to generate a target strategy action vector, and guidance is provided to the agent's channel selection, power control and RIS phase recommendation according to the target strategy action vector.

[0017] A further solution is that the self-attention mechanism is used to capture the key information of each agent's state vector to generate the self-attention output, specifically including: Through the state encoding network Extract features: ; in, Representing an agent The state encoding feature, is the state encoding network, It is an intelligent agent The state vector of Divide the state encoding features of the agent into several sub-vectors; Input the sub-vector of the agent's state encoding feature into the self-attention mechanism to map the sub-vector of the state encoding feature to different subspaces; The feature vectors corresponding to the different subspaces obtained by mapping are concatenated to generate the self-attention output.

[0018] A further solution is to use the cross-attention mechanism to achieve efficient information sharing between agents to generate cross-attention output, specifically including: Generate state encoding features of each agent through a state encoding network with shared parameters; Agent Generate feature representations through an encoder with shared parameters: ; in, Representing an agent The state encoding feature, is the state encoding network, It is an intelligent agent The original state vector of The state encoding features of each agent are composed into a sequence and input into the cross attention mechanism, and the state encoding features of each agent are mapped to different subspaces to generate the state encoding features of each agent. Cross attention output.

[0019] A further solution is that the fusion of the self-attention output and the cross-attention output to output the fusion feature representation specifically includes: The self-attention output and cross-attention output of each agent are fused, and the calculation formula is as follows: ; ; in, represents the self-attention output; represents the cross attention output; represents the fusion of self-attention output and state encoding features, Represents the fusion of cross-attention output and fused features, LayerNorm represents layer normalization, and Dropout represents dropout regularization; Further feature extraction through feedforward network: ; in are the learnable weights and bias parameters of the feedforward network.

[0020] Use residual connection to output fusion feature representation: ; in is the output fusion feature representation.

[0021] A further solution is that fusing the first strategy action vector and the second strategy action vector to generate a target strategy action vector specifically includes: The first strategy action is fused with the second strategy action through linear combination to generate the target strategy action vector: ; ; in, Representing an agent The first policy action vector of Representing an agent In the time slot Channel selection, Representing an agent In the time slot Power control, Representing an agent In the time slot RIS phase recommendations for hybrid systems is a constant.

[0022] Compared with existing technologies, the present invention offers the following advantages: It provides a training and allocation method for an IoV resource allocation model and proposes an Enhanced Multi-Agent Soft Actor-Critic-Diffusion (EMASAC-D) algorithm with enhanced attention mechanisms, enabling global resource optimization and communication performance improvement for IoV systems assisted by RIS. Compared with existing technologies, the present invention offers the following technical advantages: 1. Build a multi-agent collaborative decision-making architecture: Design a distributed perception and parallel decision-making mechanism to support information sharing and collaborative optimization among multiple vehicles, breaking through the local observation limitations of a single agent.

[0023] 2. Design a hierarchical reward mechanism: Through a multi-level reward function, balance individual goals with global performance, and coordinate the multi-objective optimization problem of minimizing AoI and maximizing V2V task completion rate.

[0024] 3. Introducing an attention enhancement mechanism: Through self-attention and cross-attention mechanisms, the agent's ability to perceive key environmental information is enhanced, improving decision-making quality. The self-attention mechanism enables each agent to capture key information from its own observation state, thereby highlighting important features and suppressing irrelevant information, thereby improving decision accuracy. The cross-attention mechanism enables efficient information sharing between agents, overcoming some observability issues in multi-agent systems. The cross-attention mechanism allows each agent to pay attention to the state of other agents, thereby obtaining global information and improving collaboration capabilities. Its design principles are as follows: 4. Integrated diffusion strategy model: Integrate the diffusion strategy with the SAC algorithm to enhance the action exploration capability and improve the strategy diversity and robustness. BRIEF DESCRIPTION OF THE DRAWINGS In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0025] Figure 1 This is a flow chart of a method for training a vehicle network resource allocation model provided by an embodiment of the present invention; Figure 2 It is a flowchart of a method for allocating resources in an Internet of Vehicles provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0026] In order to make the objects, features and advantages of the present invention more obvious and easy to understand, the specific embodiments of the present invention are described in detail below with reference to the accompanying drawings.

[0027] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this invention pertains. The terms used in this specification of the present invention are for the purpose of describing specific embodiments only and are not intended to limit the present invention. The term "and / or" as used herein includes any and all combinations of one or more of the associated listed items.

[0028] See also Figure 1 The present invention provides a method for training a vehicle network resource allocation model, which specifically includes the following steps: Step 101: Obtain the state vector of the agent; In this embodiment, the present invention considers each vehicle as an independent intelligent entity. In the connected vehicle network, vehicle-to-vehicle (V2V) and vehicle-to-infrastructure (V2I) communications, based on C-V2X technology, are widely used to ensure traffic efficiency, road safety, and real-time information services. More specifically, V2V communication refers to the exchange of information (such as traffic conditions and vehicle location) between vehicles and is commonly used in vehicle collision avoidance safety systems. V2I communication, on the other hand, refers to communication between vehicles and road infrastructure to obtain road management information, such as the timing of traffic lights, and is used for vehicle monitoring and management. By sharing time-critical information in real time, V2V and V2I communications can effectively avoid or minimize traffic accidents, ensuring road safety.

[0029] In order to fully characterize the dynamic Internet of Vehicles environment, the present invention constructs a multi-dimensional state space, which enables the intelligent agent to accurately perceive the changes in the environment and respond accordingly. In the time slot The state vector It includes the following key components: (a).RIS phase state The phase state of RIS (Reconfigurable Smart Surface) reflects the control effect of the current RIS on the wireless channel and is defined as: ; in Representing an agent For the first The phase configuration of each RIS reflection unit is in the range of , considering the computational complexity and practical application requirements, the phase value of RIS is usually discretized as quantization level, in this model , that is, each reflection unit supports 8 discrete phase configurations, namely ; Among them, for discretization of multiple quantization levels, those skilled in the art can make a selection according to actual conditions, and the present invention does not make any specific limitation.

[0030] To facilitate learning representation, it is normalized to For each element in the interval, divide by ,get To achieve multi-agent collaborative control of RIS, the system aggregates the phase suggestions of each agent through weighted average to form the final RIS configuration information: .

[0031] (b) Channel state information Channel state information includes the fast fading and path loss information of the V2I link and the V2V link, which are defined as follows: V2I link fast fading and path loss information: ; in Representing an agent Use subcarriers The instantaneous channel gain of the V2I link (dB), Indicates the corresponding long-term path loss. V2V link fast fading and path loss information: ; in Representing an agent With the agent Subcarriers The instantaneous channel gain of the V2V link, represents the corresponding long-term path loss.

[0032] Absolute channel gain: The system also records the absolute channel gain and , used to accurately calculate the signal-to-noise ratio and transmission rate.

[0033] Taking into account the time-varying characteristics of the channel caused by the mobility of the intelligent agent, the system updates the fast fading parameters every 10ms and the path loss parameters every 100ms, balancing the channel tracking accuracy and system overhead.

[0034] (c) Task and timeliness parameters The task and timeliness parameters reflect the current communication needs and service quality status of the agent. ; in Representing an agent In the time slot The remaining tasks, The preset maximum transmission task volume. This parameter reflects the task completion progress of the V2V link and is normalized to the range [0, 1].

[0035] Normalized AoI (Age of Information) level: ; in Representing an agent In the time slot AoI, is the maximum allowed information age threshold. This parameter measures the timeliness of V2I link information and is normalized to the range [0, 1].

[0036] 3. Interference Heatmap Interference heatmaps provide an overview of resource usage across the entire network, helping agents avoid high-interference areas.

[0037] Channel occupancy: ; in Representing an agent In subcarrier The interference level perceived by the agent. By calculating the average interference of the entire network, the occupancy status of each subcarrier is reflected. Normalized interference intensity: Normalize the interference level to the range [0, 1] to facilitate intelligent learning and decision-making: ; in The default setting is -85dBm.

[0038] Combining the above components, the state vector of agent k in time slot n is The complete representation of is: .

[0039] Step 102: Input the state vector of the agent into the SAC actor network to obtain a first strategic action vector, wherein the first strategic action vector includes channel selection, power control, and RIS phase suggestion components; In the Soft Actor-Critic (SAC) algorithm, an actor network generates action policies, while a dual-critic network evaluates action values. Through collaborative optimization of the actor and dual-critic networks, the policy is gradually improved. The actor network outputs a probability distribution from which specific actions can be sampled. Because SAC employs a maximum entropy reinforcement learning framework, the actor network's goal is not only to maximize expected reward but also to maximize the entropy of the policy, thereby encouraging exploration.

[0040] Specifically, the first policy action vector is directly generated through the SAC actor network:

[0041] in, Representing an agent In the time slot The first policy action vector, Indicated by the parameter The defined SAC actor network, Representing an agent In the time slot The state vector of .

[0042] In this embodiment, the first strategy action vector is a three-dimensional vector, the elements of which correspond to the channel selection, power control, and RIS phase recommendation components, respectively. Its mathematical expression is as follows: .

[0043] Step 103: Use a self-attention mechanism to capture key information of each agent's state vector to generate a self-attention output; use a cross-attention mechanism to achieve efficient information sharing between agents to generate a cross-attention output; fuse the self-attention output and the cross-attention output to output a fused feature representation; Specifically, the self-attention mechanism enables each agent to capture key information in its own state vector, making it particularly suitable for high-dimensional, heterogeneous state spaces. It weights each dimension of its own state, highlighting important features and suppressing irrelevant information, thereby improving decision-making accuracy. The specific process is as follows: (a) State feature extraction: First, through the state encoding network Extract features: ; in, Representing an agent The state encoding feature, is the state encoding network, It is an intelligent agent The state encoding network consists of three fully connected layers (256-128-128) and uses the GELU activation function to map the original state to a high-dimensional feature space, providing a basis for subsequent attention calculations.

[0044] (b) Divide the state encoding features of the agent into several sub-vectors: For example, since the state coding network consists of three fully connected layers, and the dimension of the last fully connected layer is 128, through the state coding network After extracting the features, we will get a state encoding feature with a dimension of 128 ; Then the state encoding feature It is equally divided into four sub-vectors, and the dimension of each sub-vector is reduced to 31.

[0045] It should be noted that, regarding the equal division into multiple sub-vectors, those skilled in the art may make a selection according to actual conditions, and the present invention does not impose any specific limitation thereto.

[0046] (b) Input the sub-vector of the agent's state encoding feature into the self-attention mechanism, mapping the sub-vector of the state encoding feature to different subspaces: The calculation formula of the self-attention mechanism is as follows: ; Where, , , ;in, 、 and Represents the intelligent agent The query, key, and value matrices of 、 and is a learnable weight matrix; and Each row in the state encoding feature Through linear transformation, the sub-vectors of the state encoding features are mapped to different subspaces, thereby capturing the key information in the state vector of the agent.

[0047] (c) Concatenate the feature vectors corresponding to the different subspaces obtained by mapping to generate the self-attention output: Since step (b) maps the subvectors of the state encoding features to different subspaces, multiple feature vectors can be obtained. These feature vectors are concatenated, that is, the multiple feature vectors obtained are connected in series in the dimension, and finally the self-attention output is obtained. .in, Representing an agent The self-attention output.

[0048] In some preferred implementations, to further improve feature extraction capabilities, the self-attention mechanism can adopt a multi-head design (4 heads by default), where each attention head focuses on a different feature subspace, further improving feature extraction capabilities:

[0049] in, Indicates the The output of an attention head, is the output projection matrix, Represents the dimension of the output of each attention head. Through multi-head design, it can capture information from different subspaces and improve the expressiveness of the model.

[0050] It should be noted that the main advantages of the self-attention mechanism are: (1) capturing long-range dependencies between different state dimensions; (2) adaptively assigning attention weights to highlight key information; and (3) processing variable-length sequence inputs and adapting to dynamic environmental changes. For RIS-assisted IoV scenarios, this mechanism particularly strengthens the understanding of the association between channel states, interference heat maps, and task parameters.

[0051] Specifically, the cross-attention mechanism enables efficient information sharing between agents, overcoming some observability issues in multi-agent systems. It allows each agent to pay attention to the status of other agents, thereby obtaining global information and improving collaboration capabilities. The specific process is as follows: (a) Generate the state encoding features of each agent through the state encoding network with shared parameters: For example, the agent Generate feature representations through an encoder with shared parameters: ; in, Representing an agent The state encoding feature, is the state encoding network, It is an intelligent agent The original state vector of . Among them, the shared parameter encoder can reduce model parameters and improve generalization ability.

[0052] (b) The state encoding features of each agent are sequenced and input into the cross attention mechanism to map the state encoding features of each agent to different subspaces to generate the state encoding features of each agent. The cross attention output of: Among them, the calculation formula of the cross attention mechanism is as follows: ; Where, , , ;in, 、 and denote the query, key, and value matrices corresponding to the agent sequence, respectively. 、 and is a learnable weight matrix; and Each row in corresponds to the state encoding feature of an agent. Through linear transformation, the state encoding features of each agent are mapped to different subspaces, so that each agent can pay attention to the state of other agents, thereby obtaining global information and improving cooperation ability.

[0053] By calculating the dot product between the query and the key, scaling and soft max normalization, we get the attention weights. After multiplying the attention weights with the value matrix corresponding to the agent sequence, we can generate the cross-attention output matrix corresponding to the agent sequence. Each row of the cross-attention output matrix represents the cross-attention output of each agent. Representing an agent Cross attention output.

[0054] The cross-attention mechanism enables dynamic information selection between agents, allowing each agent to adaptively determine the importance of other agents. In highly dynamic connected vehicle environments, this mechanism significantly improves the system's global state perception and the efficiency of multi-agent collaboration.

[0055] Specifically, the fusion of the self-attention output and the cross-attention output to output the fusion feature representation specifically includes: (1) The self-attention output and cross-attention output of each agent are fused, and the calculation formula is as follows: ; ; in, represents the fusion of self-attention output and state encoding features, LayerNorm represents the fusion of the cross-attention output and the fused features, Dropout represents dropout regularization, and LayerNorm and Dropout regularization improve the stability and generalization of the model.

[0056] The cross-attention mechanism enables dynamic information selection among agents through attention weights, enabling each agent to adaptively determine the importance of other agents. In highly dynamic connected vehicle environments, this mechanism significantly improves the system's global state perception and the efficiency of multi-agent collaboration.

[0057] (2) Feedforward network processing: The feedforward network uses a two-layer structure to further extract features: ; in are the learnable weights and bias parameters of the feedforward network.

[0058] (3) Residual connection and final output The final output uses residual connection: ; in The output fusion feature representation will be sent to the subsequent network for subsequent processing. In order to facilitate the representation, this paper uses .

[0059] It should be noted that the self-attention mechanism, the cross-attention mechanism, and the feedforward network form a hierarchical attention network. The core advantages of the hierarchical attention network of the present invention are: (1) effectively balancing local and global information; (2) enabling controllable information flow between agents; and (3) extracting high-order features while preserving original state information. Each component in the network is optimized for the characteristics of the Internet of Vehicles environment. For example, the encoder pays special attention to channel status and interference patterns, while the cross-attention layer focuses on processing the spatial relationship between vehicles and task priorities.

[0060] The application of the attention mechanism not only improves algorithm performance but also enhances model interpretability. By analyzing attention weights, key state dimensions and important agents in the connected vehicle environment can be identified, providing intuitive guidance for system optimization. For example, in scenarios with severe channel contention, the attention mechanism automatically focuses on features related to the interference heat map; in mission-critical situations, it prioritizes AoI and remaining load information.

[0061] Step 104: Input the obtained fusion feature representation into the first MLP network for dimensionality reduction processing, add noise to the reduced feature representation and input it into the diffusion model to output predicted noise; obtain the second strategy action vector based on the predicted noise; wherein the network structure of the diffusion model is the second MLP network; Specifically, an MLP network (Multi-layer Perceptron) is a feedforward artificial neural network model. The MLP consists of an input layer, hidden layers, and an output layer. The neurons in each layer are fully connected to the neurons in the next layer, that is, the neurons in each layer are connected to the neurons in the next layer through fully connected layers. In this embodiment, the first MLP network consists of five fully connected layers, with the dimensions of each layer being 512-256-128-64-3, respectively. Since the dimension of the last layer of the first MLP network is set to 3, when the output fusion feature is input into the network, the network will output a three-dimensional vector.

[0062] It's important to note that diffusion models are a type of generative model whose core idea is to generate actions by gradually adding noise to the data (the forward process) and then learning how to gradually remove the noise (the backward process). In the context of reinforcement learning, diffusion models are able to capture the complex distribution of action spaces and are particularly well-suited for processing multimodal action spaces.

[0063] In this embodiment, the core of the diffusion model is a noise prediction network based on the second MLP network, whose training goal is to accurately predict the noise added to the action.

[0064] Among them, the second MLP network consists of four fully connected layers, and the dimensions of each fully connected layer are 256-128-64-3 respectively. Of course, those skilled in the art can also choose according to actual conditions, and those skilled in the art will not make specific limitations.

[0065] It should be noted that the diffusion model consists of three stages: the noise addition stage, the noise prediction network training stage, and the noise prediction network execution stage. The noise addition stage of the diffusion model generates action samples by gradually adding noise, aiming to gradually transform the original action into pure noise, providing a foundation for the subsequent denoising process. The noise prediction network training stage of the diffusion model serves as the core engine of the diffusion strategy, performing the reverse process of action sample generation. The noise training stage of the diffusion model outputs predicted noise and trains the noise prediction network by minimizing a loss function, enabling it to accurately identify the noise components in the action. The noise prediction network execution stage of the diffusion model utilizes the trained noise prediction network to generate high-quality actions from random noise through iterative denoising.

[0066] It should be further explained that the diffusion model noise prediction network training phase trains the diffusion model by using the difference between the predicted noise and the actual noise, so that the weights and biases of the diffusion model are updated by using the difference between the predicted noise and the actual noise; however, the diffusion model training phase can also generate action vectors based on the predicted noise.

[0067] Its mathematical expression in the noise prediction network training stage is as follows: ; in, Indicates the The agent parameters are The noise prediction network, the input is the fusion feature representation , noise action and time steps , the output is the predicted noise.

[0068] The training objective of the noise prediction network is to minimize the mean square error between the predicted noise and the actual noise: ; in, Indicates the The training loss of the agent noise prediction network, Expressing hope, represents the standard normal distribution, Representing an agent The fusion feature representation of Indicates noise action, represents the time step, Indicates the The output of the noise prediction network of each agent, represents the actual noise, Represents the weight coefficient, which is used to balance the losses at different time steps. for: ; is the noise scheduling coefficient, using cosine scheduling: ; Cosine scheduling can make the noise addition process smoother, thereby improving the quality of generated samples.

[0069] During the noise prediction network execution phase, the noise prediction network gradually generates structured actions from random noise through multiple iterations.

[0070] (a) Initialization: Sample initial noise from a standard normal distribution: ; As the starting point of the denoising process, the quality of the initial noise directly affects the quality of the final generated action.

[0071] (b) Iterative denoising: From time step arrive Reverse iteration: ; in, Indicates passing The action after denoising, Representing an agent The fusion feature representation of Indicates passing The action after adding step noise, represents the current time step, It is the additional noise introduced in the denoising process to increase the diversity of actions. Represents the noise coefficient, which is used to control the degree of exploration: ; Exploration coefficient It decays linearly from 0.8 to 0.2 as the training progresses.

[0072] (c) Action generation: After completing the denoising process, we get the structured action , where structured actions Including channel selection component , power control component and RIS phase suggestion component , where the final selection is determined using the argmax operation for discrete actions (channel selection): arg max( ); For continuous actions (power control and RIS phase), a tanh function is used to map to the valid range: tanh( )· ; ; Finally, these components are combined to form the complete second policy action vector As mentioned above, the diffusion model noise prediction network training stage trains the diffusion model by using the difference between the predicted noise and the actual noise, so that the weights and biases of the diffusion model are updated by using the predicted noise and the actual noise; however, the diffusion model training stage can also generate an action vector based on the predicted noise, that is, the diffusion model training stage can also generate a second strategy action vector.

[0073] Step 105: Fusing the first strategy action vector and the second strategy action vector to generate a target strategy action vector; It should be noted that the fusion of the diffusion strategy and the SAC strategy can fully utilize the complementary advantages of the two strategies: the SAC strategy provides deterministic and efficient actions, while the diffusion strategy provides diverse exploration capabilities.

[0074] (a) SAC strategic advantages: By maximizing expected rewards and entropy learning strategies, it provides stable gradient updates. It directly optimizes based on the current value function, is computationally efficient, and performs well in areas with known environmental dynamics.

[0075] (b) Diffusion strategy advantages: The ability to model complex multimodal action distributions can also provide richer and deeper action space exploration, while showing stronger adaptability to environmental changes.

[0076] By combining these two strategies, the present invention can significantly improve the expressiveness and exploration efficiency of the strategy while maintaining learning stability.

[0077] Specifically, the present invention fuses the SAC strategy action with the diffusion strategy action through linear combination to generate the target strategy action vector: ; ; in, Representing an agent In the time slot Channel selection, Representing an agent In the time slot Power control, Representing an agent In the time slot The RIS phase is recommended, the mixing coefficient Dynamic adjustment as training progresses: , t is the number of iterations, initial value , attenuation rate This allows the system to focus on exploration (where the diffusion strategy dominates) in the early stages of training and on exploitation (where the SAC strategy dominates) in the later stages. This dynamic balancing mechanism enables the present invention to adaptively adjust the strategy combination to maximize benefits at different stages of training.

[0078] Step 106: Update the parameters of the diffusion model, cross-attention mechanism, and self-attention mechanism based on the predicted noise output by the diffusion model and the added actual noise; extract the target Q values ​​corresponding to the target dual-Q networks corresponding to the SAC target critic network using the target strategy action vector; calculate the loss value of the dual-Q network corresponding to the critic network based on the smaller value of the target Q value, and update the parameters of the dual-Q network based on the loss value of the dual-Q network; and calculate the loss value of the SAC actor network based on the smaller value of the loss value of the dual-Q network, and update the parameters of the actor network based on the loss value of the SAC actor network; Among them, the diffusion model is updated: The diffusion model learns to generate actions by predicting noise and minimizing the mean square error between the predicted noise and the actual noise. The specific update steps are as follows: (a) Sampling training data: Sample a batch of training data from the experience replay buffer in, Indicates status, Indicates action, Indicates reward, Indicates the next state, Indicates whether it is a terminal state. These data are used to train the diffusion strategy network so that it can accurately predict noise.

[0079] (b) Calculate the gradient: Calculate the gradient based on the noise prediction loss:

[0080] in, represents the loss function of the noise prediction network, Expressing hope, represents the experience replay buffer, Indicates status, Indicates action, represents the time step, represents the output of the noise prediction network, represents the actual noise, Represents the weight coefficient.

[0081] (c) Parameter update: Update the parameters using the Adam optimizer:

[0082] in, represents the learning rate, Represents the Adam optimizer, which is used to update the parameters of the noise prediction network .

[0083] To ensure training stability and convergence, this paper uses gradient clipping (maximum norm 10.0) to avoid gradient explosion. Gradient clipping can limit the range of gradients, preventing excessive gradients from causing training instability.

[0084] In addition, when iteratively updating the parameters of the second MLP network in the diffusion model, the parameters of the first MLP network, the cross attention mechanism, and the self-attention mechanism are iteratively updated synchronously based on the loss function of the noise prediction network. 2. Optimizing the Critic Network The critic network uses a dual-Q network architecture to improve learning stability and sample efficiency: The target Q value is calculated using Clipped Double-Q technology, which selects the smaller value from the estimates of the two target double-Q networks to effectively suppress overestimation: ; in Representing an agent In the time slot Immediate rewards obtained after performing an action; is a discount factor used to measure the importance of future rewards; Indicates the The target Q network evaluates the value of the next state-action pair; The current policy network is based on the next state Generated action, entropy term Encourage exploration. By taking the minimum of the two target network estimates, this method effectively alleviates the problem of overestimation of Q values.

[0085] (b) Loss function: The two Q networks are updated completely independently, each optimizing for TD error. ; ; in Represents the experience replay buffer Find the expectation of the sampled data, It is the importance sampling weight to correct the distribution bias introduced by priority sampling.

[0086] (c) Parameter update: The two Q networks are updated independently using the Adam optimizer. The Adam optimizer with momentum is used for parameter update to enhance training stability. ; ; The learning rate , which is used to control the amplitude of parameter updates. In order to ensure the stability and convergence of training, the present invention adopts a linear annealing scheduling strategy to gradually reduce the learning rate.

[0087] (d) Shared feature extraction: To balance computational efficiency and representation diversity, the two Q networks share a state encoding layer but maintain independent evaluation heads. The state encoding layer extracts feature representations of the state, while the independent evaluation head assesses the value of state-action pairs. This design reduces the number of parameters while ensuring a certain degree of independence in the Q network's value assessment, thereby improving learning efficiency and performance.

[0088] 3. Actor Network Gradient Calculation The actor network generates the optimal action by maximizing the expected return and introducing policy entropy regularization. To achieve this goal, the policy network needs to learn a policy function that outputs the optimal action based on the current state. The specific update steps are as follows: (a) Gradient calculation:

[0089] The first is the entropy gradient, which encourages strategy diversity, and the second is the performance gradient, which improves expected returns.

[0090] (b) Action Reparameterization: Use the reparameterization trick to stabilize gradient computation:

[0091] By transferring randomness from the policy distribution to external noise, we achieve stable gradient propagation. The reparameterization technique can make the policy gradient updates smoother, thereby improving the stability of training.

[0092] (c) Loss function: defined based on soft Q value and policy entropy:

[0093] (d) Parameter update: Update the policy network using the Adam optimizer: ; Learning rate .

[0094] Through the above steps, the actor network is able to learn a policy function that maximizes the expected return and maintains policy diversity, thereby achieving the goal of multi-agent collaborative control.

[0095] 4. Adaptive adjustment of temperature parameters Temperature parameters Control the balance between exploration and exploitation and dynamically adjust it by maximizing the policy entropy. The specific update steps are as follows: (a) Target entropy: Based on the action space dimension setting: ; For this model The target entropy is -6.

[0096] (b) Temperature loss function: ; (c) Parameter update: Update the temperature parameter using gradient descent: ; Learning rate .

[0097] 5. Target network soft synchronization To improve the stability and performance of the algorithm, this paper uses a soft update mechanism to synchronize the target network parameters. The target network provides a stable value estimate, which can reduce fluctuations during training. The specific steps are as follows: (a) Soft update formula: ; in is the target Q network parameter, is the current Q network parameter.

[0098] (b) Time-varying update coefficient: ; The initial value , gradually increases with the progress of training, the maximum value , making the system more stable in the early stages of training and faster in the later stages of updating.

[0099] Through the soft update mechanism, the target network can slowly track the parameters of the main network, thereby providing a more stable and accurate Q value estimation, thereby improving the overall performance of the algorithm.

[0100] It should be noted that effective reward function design is key to the success of multi-agent reinforcement learning. This paper constructs a hierarchical multi-objective reward function to balance the real-time information of the V2I link with the transmission reliability of the V2V link, while coordinating individual goals with group benefits.

[0101] intelligent In the time slot The reward function A hierarchical structure is adopted, including two levels: individual rewards and team rewards: ; in is the weight coefficient, which is used to balance individual and group goals.

[0102] Individual rewards reflect the efficiency of task completion and information timeliness of a single vehicle, and include two key components: ; Task completion incentive (first item): ; in Indicates vehicle In the time slot The remaining tasks, The maximum number of tasks to be transferred is preset. This reward incentivizes the agent to complete data transfers efficiently by penalizing incomplete tasks with a negative value. The coefficient -10 is used to map the task penalty term to the interval [-10, 0], ensuring a reasonable magnitude match with the AoI penalty term.

[0103] When the transmission rate of the V2V link Does not exceed the minimum threshold When the transmission rate of the V2V link is Exceeding the minimum threshold When , the remaining task amount is updated according to the following rules: ; = ; in is the time slot length, is the signal-to-noise ratio of the V2V link. When , the task is completed and the reward reaches the maximum value of 0.

[0104] Information timeliness constraint (second item): ; in For vehicles In the time slot AoI, is the maximum allowable AoI threshold. The exponential function design ensures that the AoI penalty increases non-linearly with the age of the information, ensuring that the system imposes severe penalties on information that has not been updated for a long time and prevents it from becoming too outdated.

[0105] The AoI value is dynamically updated based on the V2I link transmission rate: ; in It is the minimum rate threshold of the V2I link to ensure the quality of data updates.

[0106] (2) Team reward design Team awards reflect the overall performance of the network's communication system and include two key indicators: ; Average V2I rate (first item): ; Calculates the average transmission rate of the entire network's V2I links in Mbps. This encourages the improvement of overall network capacity to ensure that all vehicles can receive satisfactory service quality.

[0107] V2I link rate The calculation formula is: ; in is the subcarrier bandwidth, is the signal-to-noise ratio of the V2I link.

[0108] V2V comprehensive success rate (second item): ; Indicates the proportion of vehicles that successfully complete the task in the current time slot. This indicator directly reflects the overall success rate of V2V task transmission and encourages the system to improve the overall task completion efficiency. If and only if the vehicle The value is 1 when the task is completed (the remaining task amount is zero), otherwise it is 0.

[0109] (3) Dynamic weight adjustment mechanism In order to adapt to the needs of different scenarios, this framework designs a dynamic weight adjustment mechanism, which is adaptively adjusted according to the network status and task characteristics. value: ; in is the base weight (default 0.7), is the adjustment coefficient (default 0.5). This mechanism enables the system to automatically reduce individual weights in high AoI situations, increase attention to global performance, and ensure information timeliness.

[0110] The design innovations of the multi-objective reward function are: first, balancing individual and team goals through a hierarchical reward architecture; second, introducing exponential penalties to ensure the timeliness of AoI constraints; and finally, adapting to changes in network status through dynamic weight adjustment.

[0111] It's worth noting that to improve sample utilization, this paper employs a prioritized experience replay mechanism based on temporal difference (TD). Prioritized experience replay assigns different sampling probabilities to different experiences, allowing the model to focus on important experiences and thus accelerating the learning process.

[0112] (a) Experience storage format: , including state, action, reward, next state, termination flag and priority. The termination flag is used to indicate whether the episode is over in order to distinguish the experience of different episodes.

[0113] (b) Priority calculation: The importance of a sample is determined based on the TD error. The larger the TD error, the greater the prediction error of the sample, and the more worthy the model is to learn: ; in, is a small constant used to avoid the situation where the priority is 0. The TD error is: ; in, Rewards, represents the discount factor, represents the smaller value of the target Q value in the target dual-Q network, Indicates the smaller value of Q value in the double Q network, which is used to reduce the over-estimation bias.

[0114] (c) Sampling probability: The sampling probability is determined based on the priority. The higher the priority, the greater the probability of being sampled: ; in, is the priority index, which controls the degree of sampling preference. The larger it is, the more inclined it is to select samples with high priority.

[0115] Importance sampling weight: Since priority sampling introduces distribution bias, importance sampling weight is needed to correct this bias to ensure the correctness of learning: ; in, is the size of the playback buffer, For samples The sampling probability of Used to adjust the weight of importance sampling. It increases linearly from 0.4 to 1.0 and adjusts as the training progresses. , we can pay more attention to high-priority samples in the early stage of training, and pay more attention to all samples in the later stage of training, thereby balancing the stability and convergence speed of learning.

[0116] Figure 2 This is a flow chart of a method for allocating resources in an Internet of Vehicles according to an embodiment of the present invention. Figure 2 As shown, the vehicle network resource allocation method according to an embodiment of the present invention specifically includes the following steps: Step 201: Obtain the state vector of the agent; Among them, the state vector of agent k in time slot n is obtained The complete representation of is: Among them, multi-intelligent RIS configuration information, V2I link fast fading and path loss information, V2V link fast fading and path loss information, V2I link instantaneous channel gain, V2V link instantaneous channel gain, intelligent agent In the time slot The amount of remaining tasks, agent In the time slot The specific processing process of the AoI and the average interference information of the entire network on the subcarrier is similar to step 101 and will not be repeated here.

[0117] Step 202: Input the state vector of the agent into the trained SAC actor network to obtain a first strategic action vector, wherein the first strategic action vector includes channel selection, power control, and RIS phase suggestion components; In the Soft Actor-Critic (SAC) algorithm, an actor network generates action policies, while a dual-critic network evaluates action values. Through the collaborative optimization of the actor and dual-critic networks, policies are gradually improved.

[0118] Specifically, the first policy action vector is directly generated through the SAC actor network:

[0119] in, Representing an agent In the time slot The first policy action vector, Indicated by the parameter The defined SAC actor network, Representing an agent In the time slot The state vector of .

[0120] The first strategy action vector is a three-dimensional vector, corresponding to the channel selection, power control and RIS phase recommendation components respectively. Its mathematical expression is as follows: .

[0121] Step 203: Use a self-attention mechanism to capture key information of each agent's state vector to generate a self-attention output; use a cross-attention mechanism to achieve efficient information sharing between agents to generate a cross-attention output; fuse the self-attention output and the cross-attention output to output a fused feature representation; Among them, the self-attention mechanism enables each agent to capture key information in its own state vector, making it particularly suitable for high-dimensional, heterogeneous state spaces. By weighting each dimension of its own state, it highlights important features and suppresses irrelevant information, thereby improving decision accuracy. The cross-attention mechanism enables efficient information sharing between agents, overcoming the partial observability problem in multi-agent systems. By having each agent pay attention to the states of other agents, it obtains global information and improves collaboration. Fusion of the outputs of the self-attention mechanism and the cross-attention mechanism achieves an effective balance between local and global information, enabling controlled information exchange between agents. This process preserves original state information while extracting high-level features.

[0122] It should be noted that the specific processing of the self-attention mechanism, the cross-attention mechanism, and the fusion of the self-attention output and the cross-attention is similar to step 103 and will not be repeated here.

[0123] Step 204: Input the obtained fusion feature representation into the first MLP network for dimensionality reduction. The reduced feature representation is directly input into the trained diffusion model to generate a second strategy action vector. The diffusion model has a network structure of the second MLP network. It should be noted that the noise prediction network execution phase of the diffusion model uses the trained noise prediction network to generate high-quality actions through iterative denoising from the fusion feature representation after dimensionality reduction. In the execution phase of the diffusion model, it is not necessary to add noise to the input; instead, the feature representation after dimensionality reduction is directly input into the trained diffusion model. The diffusion model gradually generates structured actions from random noise through multiple iterations, where structured actions Including channel selection component , power control component and RIS phase components are recommended , where the final selection is determined using the argmax operation for discrete actions (channel selection): arg max( ); For continuous actions (power control and RIS phase), a tanh function is used to map to the valid range: tanh( )· ; ; Finally, these components are combined to form the complete second policy action vector .

[0124] Step 205: Fusing the first strategy action vector and the second strategy action vector to generate a target strategy action vector, and providing guidance for channel selection, power control, and RIS phase recommendation of the agent according to the target strategy action vector; Specifically, the present invention fuses the SAC strategy action with the diffusion strategy action through linear combination to generate the target strategy action vector: ; .

[0125] in, Representing an agent In the time slot Channel selection, Representing an agent In the time slot Power control, Representing an agent In the time slot RIS phase recommendations for hybrid systems Of course, those skilled in the art can make a selection based on actual conditions, and the present invention does not make any specific limitation.

[0126] Because the intelligent In the time slot The target strategy action vector consists of channel selection, power control and RIS phase recommendation. Representing an agent In the time slot Channel selection, Representing an agent In the time slot Power control, Representing an agent In the time slot Therefore, the channel selection, power control and RIS phase recommendation of the agent can be guided according to the target strategy action vector.

[0127] For discrete channel selection, each vehicle agent selects an optimal channel from the available subcarriers for communication. The discrete channel selection is represented by one-hot coding; one-hot coding is based on the agent. In the time slot Channel selection The discrete channel selection is converted into a vector with only one position being 1 and all other positions being 0, where the dimension of the vector is L; if agent k selects one of the subcarriers, the discrete channel selection of agent k is 1 at the position corresponding to the subcarrier and all other positions are 0.

[0128] For continuous power control, each agent also needs to determine the actual transmission power to balance communication quality and energy consumption, where the actual transmission power is determined by the agent. In the time slot Power control It is obtained through function mapping, where the actual transmission power of each agent is ;in The maximum allowed transmit power is set to 23dBm by default.

[0129] For RIS phase recommendations, each vehicle agent is based on the agent In the time slot RIS phase recommendations Update of RIS phase configuration recommendations: ; in Indicates vehicle For the first Considering the quantization accuracy and computational complexity, this system adopts 8-level quantization ( ),Right now: ; To facilitate learning representation, it is normalized to For each element in the interval, divide by ,get .

[0130] In summary, this invention provides a training and allocation method for an IoV resource allocation model and proposes an Enhanced Multi-Agent Soft Actor-Critic-Diffusion (EMASAC-D) algorithm with enhanced attention mechanisms to achieve global resource optimization and communication performance improvement for IoV systems using RIS. Compared to existing technologies, this invention has the following technical benefits: 1. Build a multi-agent collaborative decision-making architecture: Design a distributed perception and parallel decision-making mechanism to support information sharing and collaborative optimization among multiple vehicles, breaking through the local observation limitations of a single agent.

[0131] 2. Design a hierarchical reward mechanism: Through a multi-level reward function, balance individual goals with global performance, and coordinate the multi-objective optimization problem of minimizing AoI and maximizing V2V task completion rate.

[0132] 3. Introducing an attention enhancement mechanism: Through self-attention and cross-attention mechanisms, the agent's ability to perceive key environmental information is enhanced, improving decision-making quality. The self-attention mechanism enables each agent to capture key information from its own observation state, thereby highlighting important features and suppressing irrelevant information, thereby improving decision accuracy. The cross-attention mechanism enables efficient information sharing between agents, overcoming some observability issues in multi-agent systems. The cross-attention mechanism allows each agent to pay attention to the state of other agents, thereby obtaining global information and improving collaboration capabilities. Its design principles are as follows: 4. Integrated diffusion strategy model: Integrate the diffusion strategy with the SAC algorithm to enhance the action exploration capability and improve the strategy diversity and robustness.

[0133] Although the embodiments of the present invention have been disclosed above, they are not limited to the applications listed in the description and implementation methods. They can be fully applied to various fields suitable for the present invention. For those familiar with the art, additional modifications can be easily implemented. Therefore, without departing from the general concept defined by the claims and the scope of equivalents, the present invention is not limited to the specific details and embodiments shown and described herein.

Claims

1. A training method for a vehicle network resource allocation model, characterized in that: The specific steps include: Get the state vector of the agent; Input the agent's state vector into the SAC actor network to obtain the first policy action vector; A self-attention mechanism is used to capture the key information of each agent’s state vector to generate self-attention output; Adopting cross-attention mechanism to achieve efficient information sharing between agents to generate cross-attention output; Fuse the self-attention output and the cross-attention output to output the fused feature representation; The obtained fusion feature representation is input into the first MLP network for dimensionality reduction. Noise is added to the reduced feature representation and then input into the diffusion model to output predicted noise. Obtaining a second strategy action vector based on the predicted noise; fusing the first policy action vector and the second policy action vector to generate a target policy action vector; The parameters of the diffusion model, cross-attention mechanism, and self-attention mechanism are updated based on the predicted noise output by the diffusion model and the actual noise added; the target strategy action vector is used to extract the target Q values ​​corresponding to the target dual-Q networks corresponding to the SAC target critic network; the loss value of the dual-Q network corresponding to the critic network is calculated based on the smaller value of the target Q value, and the parameters of the dual-Q network are updated based on the loss value of the dual-Q network; The loss value of the SAC actor network is calculated based on the smaller value of the loss value of the dual Q network, and the parameters of the actor network are updated based on the loss value of the SAC actor network.

2. The method for training a vehicle network resource allocation model according to claim 1, characterized in that: The self-attention mechanism is used to capture the key information of each agent's state vector to generate the self-attention output, specifically including: Through the state encoding network Extract features: ; in, Representing an agent The state encoding feature, is the state encoding network, It is an intelligent agent The state vector of Divide the state encoding features of the agent into several sub-vectors; Input the sub-vector of the agent's state encoding feature into the self-attention mechanism to map the sub-vector of the state encoding feature to different subspaces; The feature vectors corresponding to the different subspaces obtained by mapping are concatenated to generate the self-attention output.

3. The method for training a vehicle network resource allocation model according to claim 2, characterized in that: The cross-attention mechanism is used to achieve efficient information sharing between agents to generate cross-attention outputs, specifically including: Generate state encoding features of each agent through a state encoding network with shared parameters; Agent Generate feature representations through an encoder with shared parameters: ; in, Representing an agent The state encoding feature, is the state encoding network, It is an intelligent agent The original state vector of The state encoding features of each agent are composed into a sequence and input into the cross attention mechanism, and the state encoding features of each agent are mapped to different subspaces to generate the state encoding features of each agent. Cross attention output.

4. The method for training a vehicle network resource allocation model according to claim 3, characterized in that: The fusion of the self-attention output and the cross-attention output to output the fusion feature representation specifically includes: The self-attention output and cross-attention output of each agent are fused, and the calculation formula is as follows: ; ; in, represents the self-attention output; represents the cross attention output; represents the fusion of self-attention output and state encoding features, Represents the fusion of cross-attention output and fused features, LayerNorm represents layer normalization, and Dropout represents dropout regularization; Further feature extraction through feedforward network: ; in are the learnable weights and bias parameters of the feedforward network; Use residual connection to output fusion feature representation: ; in is the output fusion feature representation.

5. The method for training a vehicle network resource allocation model according to claim 1, characterized in that: The fusing of the first strategy action vector and the second strategy action vector to generate a target strategy action vector specifically includes: The first strategy action is fused with the second strategy action through linear combination to generate the target strategy action vector: ; ; in, Representing an agent The first policy action vector of Representing an agent In the time slot Channel selection, Representing an agent In the time slot Power control, Representing an agent In the time slot The RIS phase is recommended, the mixing coefficient Dynamic adjustment as training progresses: , t is the number of iterations, initial value , attenuation rate .

6. A method for allocating resources in a vehicle network allocation model, characterized in that: The specific steps include: Get the state vector of the agent; Input the agent's state vector into the trained SAC actor network to obtain the first policy action vector; A self-attention mechanism is used to capture the key information of each agent’s state vector to generate self-attention output; Adopting cross-attention mechanism to achieve efficient information sharing between agents to generate cross-attention output; Fuse the self-attention output and the cross-attention output to output the fused feature representation; The obtained fusion feature representation is input into the first MLP network for dimensionality reduction. The reduced feature representation is directly input into the trained diffusion model to generate the second strategy action vector. The first strategy action vector and the second strategy action vector are fused to generate a target strategy action vector, and guidance is provided to the agent's channel selection, power control and RIS phase recommendation according to the target strategy action vector.

7. The method for allocating Internet of Vehicles resources according to claim 6, characterized in that: The self-attention mechanism is used to capture the key information of each agent's state vector to generate the self-attention output, specifically including: Through the state encoding network Extract features: ; in, Representing an agent The state encoding feature, is the state encoding network, It is an intelligent agent The state vector of Divide the state encoding features of the agent into several sub-vectors; Input the sub-vector of the agent's state encoding feature into the self-attention mechanism to map the sub-vector of the state encoding feature to different subspaces; The feature vectors corresponding to the different subspaces obtained by mapping are concatenated to generate the self-attention output.

8. The method for allocating Internet of Vehicles resources according to claim 7, characterized in that: The cross-attention mechanism is used to achieve efficient information sharing between agents to generate cross-attention outputs, specifically including: Generate state encoding features of each agent through a state encoding network with shared parameters; Agent Generate feature representations through an encoder with shared parameters: ; in, Representing an agent The state encoding feature, is the state encoding network, It is an intelligent agent The original state vector of The state encoding features of each agent are composed into a sequence and input into the cross attention mechanism, and the state encoding features of each agent are mapped to different subspaces to generate the state encoding features of each agent. Cross attention output.

9. The method for allocating Internet of Vehicles resources according to claim 8, characterized in that: The fusion of the self-attention output and the cross-attention output to output the fusion feature representation specifically includes: The self-attention output and cross-attention output of each agent are fused, and the calculation formula is as follows: ; ; in, represents the self-attention output; represents the cross attention output; represents the fusion of self-attention output and state encoding features, Represents the fusion of cross-attention output and fused features, LayerNorm represents layer normalization, and Dropout represents dropout regularization; Further feature extraction through feedforward network: ; in are the learnable weights and bias parameters of the feedforward network; Use residual connection to output fusion feature representation: ; in is the output fusion feature representation.

10. The method for allocating Internet of Vehicles resources according to claim 6, characterized in that: The fusing of the first strategy action vector and the second strategy action vector to generate a target strategy action vector specifically includes: The first strategy action is fused with the second strategy action through linear combination to generate the target strategy action vector: ; ; in, Representing an agent The first policy action vector of Representing an agent In the time slot Channel selection, Representing an agent In the time slot Power control, Representing an agent In the time slot RIS phase recommendations for hybrid systems is a constant.