Vehicle communication network resource allocation method and device, and storage medium

Through real-time channel state data and predicted channel state data, combined with the deep reinforcement learning model, the resource allocation strategy is dynamically adjusted, and the problem that the static preconfiguration mechanism cannot adapt to wireless channel changes is solved, and the efficient utilization and reliability of communication resources of autonomous driving vehicles is achieved.

CN120343732APending Publication Date: 2025-07-18BEIJING TRUNK TECHNOLOGY CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510497402.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-18
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The existing static preconfiguration mechanism cannot adapt to the dynamic changes in wireless channel state and complex communication environments, resulting in lack of flexibility in resource allocation and cannot meet the needs of autonomous driving vehicles for high throughput, low latency and high reliability.

Method used

By acquiring the real-time channel state data of the vehicle, using the channel state prediction model to predict future channel states, and dynamically adjusting resource allocation strategies, including spectrum, transmission power and modulation coding strategies, to adapt to the needs of different service data types.

Benefits of technology

It realizes dynamic and flexible resource allocation, improves communication resource utilization, reduces latency, enhances communication reliability, meets the real-time and reliability requirements in autonomous driving scenarios, and adapts to dynamic changes in wireless channel state.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120343732A_ABST
    Figure CN120343732A_ABST
Patent Text Reader

Abstract

The invention provides a vehicle communication network resource allocation method and device and a storage medium, and relates to the technical field of automatic driving wireless communication. The method comprises the following steps: inputting acquired real-time channel state data into a pre-trained channel state prediction model to obtain predicted channel state data corresponding to the real-time channel state data; determining a resource allocation result corresponding to the target business data type of the vehicle according to the real-time channel state data and the predicted channel state data; and dynamically allocating communication resources corresponding to the resource allocation result to the vehicle, wherein the vehicle is used for communicating according to the communication resources. According to the method, the real-time channel state data and the predicted channel state data are combined, resource allocation is dynamically and flexibly adjusted, different service requirements are met, the communication resource utilization rate is increased, delay is reduced, and the requirements of automatic driving for real-time performance and reliability are met.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of autonomous driving wireless communication technology, and in particular, to a method, device, and storage medium for allocating vehicle communication network resources. Background Art

[0002] With the rapid development of autonomous driving technology, vehicle systems need to process and transmit a large amount of data in high-speed mobile scenarios, including but not limited to sensor data, map navigation information, high-precision map update data, and vehicle networking control instructions. Such data has characteristics such as high throughput, low latency tolerance, and strong transmission reliability, and has relatively high requirements for the real-time performance, reliability, and resource allocation efficiency of the communication network. Among them, solving the problem of efficient resource allocation in a dynamic wireless channel environment has become a key challenge in the field of autonomous driving wireless communication.

[0003] In related technologies, wireless communication resource allocation strategies are based on static pre-configuration mechanisms. For example, a fixed channel is allocated to each networked vehicle, and resources such as fixed bandwidth, frequency, and transmit power are configured in this channel.

[0004] However, static allocation lacks flexibility and cannot adapt to the dynamic changes of wireless channel states and complex communication environments. Summary of the Invention

[0005] This application provides a method, device, and storage medium for allocating vehicle communication network resources to solve the problem that static resource allocation lacks flexibility and cannot adapt to the dynamic changes of wireless channel states and complex communication environments.

[0006] In a first aspect, this application provides a method for allocating vehicle communication network resources, including:

[0007] Obtain real-time channel state data of the vehicle;

[0008] Input the real-time channel state data into a pre-trained channel state prediction model to obtain predicted channel state data corresponding to the real-time channel state data;

[0009] Determine a resource allocation result corresponding to the target service data type for the vehicle according to the real-time channel state data and the predicted channel state data;

[0010] Dynamically allocate communication resources corresponding to the resource allocation result for the vehicle, and the vehicle communicates based on the communication resources.

[0011] In a possible implementation manner, determining a resource allocation result corresponding to the target service data type for the vehicle according to the real-time channel state data and the predicted channel state data includes:

[0012] Input the real-time channel state data and the predicted channel state data into a pre-constructed deep reinforcement learning model to obtain a resource allocation result corresponding to the target service data type for the vehicle. The state space of the deep reinforcement learning model includes real-time channel state data, predicted channel state data, and service data type, and the action space of the deep reinforcement learning model includes at least one of a spectrum allocation policy, a transmit power allocation policy, and a modulation and coding allocation policy.

[0013] In a possible implementation manner, the reward function of the deep reinforcement learning model is determined based on multi-objective rewards of spectrum utilization rate, transmission delay, and communication loss, and the reward function is used to evaluate the influence degree of the resource allocation result on vehicle communication;

[0014] And / or, the resource allocation method is applied to a cloud server, the deep reinforcement learning model is deployed on an edge computing node, and the cloud server and the edge computing node work in cooperation.

[0015] In a possible implementation manner, the deep reinforcement learning model is optimized in the following way:

[0016] Store the reward value corresponding to the resource allocation result, as well as the state space data and action space data corresponding to the reward value, into a buffer;

[0017] Sample from the buffer to obtain sampled data;

[0018] Optimize the model parameters of the deep reinforcement learning model according to the sampled data.

[0019] In a possible implementation manner, the deep reinforcement learning model is obtained through model compression, and the model compression includes pruning and / or quantization.

[0020] In a possible implementation manner, dynamically allocate communication resources corresponding to the resource allocation result for the vehicle, including:

[0021] Based on the pre-set correspondence between channels and resource allocation, determine the target channel for the resource allocation result corresponding to the target service data type;

[0022] Dynamically allocate the target channel for the vehicle, and the target channel is used to transmit service data corresponding to the target service data type.

[0023] In a possible implementation manner, input the real-time channel state data into a pre-trained channel state prediction model, including:

[0024] Perform preprocessing on the real-time channel state data to obtain processed channel state data, and the preprocessing includes data cleaning and normalization processing;

[0025] Input the processed channel state data into the pre-trained channel state prediction model.

[0026] In a second aspect, the present application provides a vehicle communication network resource allocation device, including:

[0027] An acquisition module, configured to acquire the real-time channel state data of the vehicle;

[0028] An input module, configured to input the real-time channel state data into the pre-trained channel state prediction model to obtain predicted channel state data corresponding to the real-time channel state data;

[0029] A determination module, configured to determine a resource allocation result corresponding to the target service data type for the vehicle according to the real-time channel state data and the predicted channel state data;

[0030] An allocation module, configured to dynamically allocate communication resources corresponding to the resource allocation result for the vehicle, and the vehicle communicates based on the communication resources.

[0031] In a possible implementation manner, the determination module is specifically configured to: input the real-time channel state data and the predicted channel state data into a pre-constructed deep reinforcement learning model to obtain a resource allocation result corresponding to the target service data type for the vehicle. The state space of the deep reinforcement learning model includes the real-time channel state data, the predicted channel state data, and the service data type, and the action space of the deep reinforcement learning model includes at least one of a spectrum allocation strategy, a transmit power allocation strategy, and a modulation and coding allocation strategy.

[0032] In a possible implementation manner, the reward function of the deep reinforcement learning model is determined based on multi-objective rewards of spectrum utilization rate, transmission delay, and communication loss. The reward function is used to evaluate the influence degree of the resource allocation result on vehicle communication; and / or, the resource allocation method is applied to a cloud server, and the deep reinforcement learning model is deployed on an edge computing node, and the cloud server and the edge computing node cooperate.

[0033] In a possible implementation manner, the deep reinforcement learning model is optimized in the following manner:

[0034] Store the reward value corresponding to the resource allocation result, as well as the state space data and action space data corresponding to the reward value, into a buffer; sample from the buffer to obtain sampled data; optimize the model parameters of the deep reinforcement learning model according to the sampled data.

[0035] In a possible implementation manner, the deep reinforcement learning model is obtained through model compression, and the model compression includes pruning and / or quantization.

[0036] In a possible implementation, the allocation module is specifically configured to: determine a target channel for a resource allocation result corresponding to a target service data type based on a preset correspondence between channels and resource allocations; dynamically allocate the target channel to a vehicle, where the target channel is used to transmit service data corresponding to the target service data type.

[0037] In a possible implementation, the input module is specifically configured to: preprocess real-time channel state data to obtain processed channel state data, where the preprocessing includes data cleaning and normalization processing; input the processed channel state data into a pre-trained channel state prediction model.

[0038] In a third aspect, the present application provides a device, including: a memory, a processor;

[0039] The memory stores computer-executable instructions;

[0040] The processor executes the computer-executable instructions stored in the memory, so that the processor executes the above first aspect and / or various possible implementations of the first aspect.

[0041] In a fourth aspect, the present application provides a computer-readable storage medium, where computer-executable instructions are stored in the computer-readable storage medium, and when the computer-executable instructions are executed, they are used to implement the above first aspect and / or various possible implementations of the first aspect.

[0042] In a fifth aspect, the present application provides a computer program product, including a computer program, and when the computer program is executed, it is used to implement the above first aspect and / or various possible implementations of the first aspect.

[0043] The vehicle communication network resource allocation method, device, and storage medium provided by the present application ensure timely response to the current channel condition by obtaining the real-time channel state data of the vehicle. The real-time channel state data is input into a pre-trained channel state prediction model to obtain predicted channel state data corresponding to the real-time channel state data. By predicting the change trend of the future channel state, it provides a basis for making the optimal resource allocation result. Combining the real-time channel state data and the predicted channel state data, the resource allocation result corresponding to the target service data type of the vehicle is determined, realizing dynamic and flexible adjustment of resource allocation to adapt to different service requirements. The communication resources corresponding to the resource allocation result are dynamically allocated to the vehicle, and the vehicle communicates based on the communication resources, thereby improving the utilization rate of communication resources, reducing latency, enhancing communication reliability, and meeting the requirements for real-time performance and reliability in the autonomous driving scenario; in addition, by dynamically allocating resources, it can better adapt to the dynamic changes of the wireless channel state and the complex communication environment. Description of the Drawings

[0044] The accompanying drawings here are incorporated into and form a part of this specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application.

[0045] Figure 1 It is a schematic diagram of the scenario of the vehicle communication network resource allocation method provided for the embodiments of the present application;

[0046] Figure 2 It is a schematic flowchart of the vehicle communication network resource allocation method provided for the embodiments of the present application;

[0047] Figure 3 It is a schematic structural diagram of the vehicle communication network resource allocation device provided for the embodiments of the application;

[0048] Figure 4 It is a schematic structural diagram of the device provided for the embodiments of the present application.

[0049] Through the above accompanying drawings, specific embodiments of the present application have been shown, and there will be more detailed descriptions hereinafter. These drawings and written descriptions are not intended to limit the scope of the concept of the present application in any way, but to illustrate the concept of the present application to those skilled in the art by referring to specific embodiments. Specific Embodiments

[0050] Exemplary embodiments will be described in detail here, and examples are shown in the accompanying drawings. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present application. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.

[0051] In the related art, the static pre-configuration mechanism allocates a fixed channel for each accessing vehicle, and resources such as fixed bandwidth, frequency, and transmission power are configured in this channel. Affected by factors such as environmental interference and vehicle mobility, the state of the wireless channel changes frequently and is difficult to predict. The static pre-configuration mechanism does not consider the changing characteristics of wireless communication, such as channel time-variation and spatial variation, and cannot adapt to the dynamic changes of the channel environment, nor can it adjust according to real-time requirements, resulting in non-optimal resource allocation. For example, in a certain area, the signal attenuation is severe, but the static allocation cannot increase the power or switch the frequency, resulting in a decline in communication quality; some vehicles may only need a small amount of bandwidth, but are allocated a fixed large amount of resources, causing resource waste; the sudden change of the channel state (such as the sudden drop in channel capacity caused by Doppler frequency shift) leads to transmission delay of high-priority data (such as emergency braking instructions) due to resource preemption failure. In addition, different types of service data (such as emergency braking instructions, navigation updates, entertainment traffic, etc.) have significant differences in delay and bandwidth requirements. Static allocation cannot adapt the resource allocation according to the type of service data, and may cause high-priority services (such as safety instructions) to be delayed due to insufficient fixed resources.

[0052] Therefore, the main problem of the static pre-configuration mechanism lies in the lack of flexibility and adaptability, unable to adjust according to the real-time channel conditions and the type of service data, resulting in data transmission timeouts, low resource utilization efficiency, and possible situations of resource waste or shortage, while affecting network performance and user experience.

[0053] To solve the above technical problems, the present application provides a method for allocating resources in a vehicle communication network. Based on the real-time channel state data of the vehicle, an intelligent model is used to predict the change trend of the future channel state, providing a basis for making the optimal resource allocation result. Combining the real-time channel state data and the predicted channel state data, the resource allocation is dynamically and flexibly adjusted to adapt to different service requirements in the vehicle, and the communication resources corresponding to the resource allocation result are dynamically allocated for the vehicle, thereby improving the utilization rate of communication resources, reducing delays, enhancing communication reliability, and better adapting to the dynamic changes of the wireless channel state and the complex communication environment through dynamic resource allocation.

[0054] The following uses specific embodiments to elaborate in detail on the technical solution of the present application and how the technical solution of the present application solves the above technical problems. These several specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of the present application will be described below with reference to the accompanying drawings.

[0055] Figure 1 It is a schematic diagram of the scenario of the method for allocating resources in a vehicle communication network provided by the embodiments of the present application. As Figure 1As shown in the figure, this scenario includes vehicle 11 and cloud server 12. Cloud server 12 is communicatively connected to vehicle 11. During the automatic driving process of vehicle 11, for example, the vehicle uses in-vehicle communication equipment to collect channel state data. Vehicle 11 actively sends the real-time channel state data to cloud server 12, or cloud server 12 sends a request to vehicle 11 to obtain the real-time channel state data of vehicle 11. A channel state prediction model is deployed on cloud server 12. Through this intelligent model, the predicted channel state data corresponding to the real-time channel state data is obtained. Then, based on the real-time channel state data and the predicted channel state data, the resource allocation result corresponding to the target service data type for vehicle 11 is determined, and the communication resources corresponding to the resource allocation result are dynamically allocated to vehicle 11. After receiving the communication resources allocated by cloud server 12, vehicle 11 transmits the service data corresponding to the target service data type to cloud server 12, or other surrounding vehicles, or third-party platforms, etc., based on this communication resource.

[0056] It should be noted that cloud server 12 can be a server cluster, a virtualized server, an edge cloud server, etc. The application scenario provided in this embodiment of the present application places no restrictions on the number of cloud servers 12 and vehicles 11. In addition, vehicle 11 and cloud server 12 are not limited to communicating through data relay with gateways, cellular networks, Wi-Fi, or satellite communications, etc.

[0057] Next, in combination with Figure 1 the application scenario of Figure 2 this application, with reference to Figure 1 this figure, the vehicle communication network resource allocation method provided in this embodiment of the present application will be described. It should be noted that this resource allocation method can be executed by Figure 1 the cloud server 12 in

[0058] Figure 2 This is a schematic flowchart of the vehicle communication network resource allocation method provided in this embodiment of the present application. As Figure 2 shown, the vehicle communication network resource allocation method provided in this embodiment of the present application includes:

[0059] S201. Obtain the real-time channel state data of the vehicle.

[0060] Among them, the channel state reflects the physical characteristics of the channel at a certain moment, and the channel state data includes signal-to-noise ratio, bandwidth, interference level, noise attenuation, Doppler frequency shift, time delay, and multipath effect, etc.

[0061] Exemplarily, real-time channel state data of a vehicle during driving is collected by in-vehicle devices (such as in-vehicle communication modules, sensors, etc.), roadside units, or base stations.

[0062] It should be noted that in the embodiments of the present application, there is no limitation on the execution subject of the vehicle communication network resource allocation method. Taking the execution subject as the cloud server as an example, the cloud server obtains the real-time channel state data of the vehicle from in-vehicle devices, roadside units, or base stations.

[0063] S202. Input the real-time channel state data into a pre-trained channel state prediction model to obtain predicted channel state data corresponding to the real-time channel state data.

[0064] The channel state prediction model is a deep learning time series model, which can capture time correlation from historical channel parameters to infer future channel states. The channel state prediction model is constructed based on, for example, Long Short-Term Memory (LSTM), Gated Recurrent Unit (GRU), or Transformer.

[0065] In this embodiment, the channel state prediction model is used to dynamically predict the channel state data of the next time slot or multiple future time slots according to the real-time channel state data to obtain predicted channel state data.

[0066] Exemplarily, a trained channel state prediction model is deployed on the cloud server. After obtaining the real-time channel state data, the real-time channel state data is used as the input of the channel state prediction model. After the channel state prediction model performs intelligent calculations, it outputs the predicted channel state data.

[0067] S203. Determine the resource allocation result corresponding to the target service data type for the vehicle according to the real-time channel state data and the predicted channel state data.

[0068] Among them, the service data type can be regarded as the data corresponding to the service requirements, such as emergency braking instructions, collision warnings, map update data, entertainment information, power instructions, remote OTA updates, vehicle location sharing, model training, etc. Different service data types have different requirements for communication. For example, emergency braking instructions require low latency and high reliability, while entertainment information may require high bandwidth but less strict requirements for latency.

[0069] During the vehicle's driving process, the service requirements are dynamically changing, and the service data types can be dynamically updated based on the predicted channel state data. For example, if the current vehicle speed is too fast, resulting in a deteriorated channel state of the vehicle, the service requirement for the next moment may be an emergency braking requirement; if it is predicted that the vehicle is about to enter a signal blind area, the service requirement may be high-precision map downloading.

[0070] It can be understood that when the channel state is poor (caused by, for example, bad weather (heavy rain), too fast vehicle speed, etc.), the data transmission delay becomes higher, and the vehicle's response speed may slow down. In the autonomous driving scenario, real-time data transmission is required. If the channel state is poor, it will seriously affect safety. Therefore, the service requirements need to be adjusted. It is necessary to prioritize ensuring the transmission of critical data, such as emergency braking instructions, and reduce the transmission of non-critical data, such as entertainment system data.

[0071] Based on the current real-time channel state data and the predicted channel state data, resources are dynamically allocated to meet the service requirements of different service data types. Among them, resources include bandwidth, frequency band, transmit power, modulation and coding scheme, etc. For example, when the channel state is good, resources are reserved for high-priority service data types, and low-priority service data types are allowed to use the remaining resources (for example, the high-frequency band is allocated to high-priority service data types (such as emergency control instructions), and the low-frequency band is allocated to low-priority service data types (such as map update data)). When the channel state suddenly deteriorates, the resource allocation for low-priority service data types is suspended to ensure the transmission of high-priority services. When the channel state is poor, the transmit power of high-priority service data types (such as emergency control instructions, vehicle shared location) is increased to ensure data transmission reliability, and the transmit power of low-priority service data types is reduced to reduce energy consumption. When the channel state is good, high-order modulation (such as 64-QAM) is allocated to improve the transmission efficiency, and when the channel state is poor, low-order modulation (such as QPSK) is allocated to improve the transmission reliability.

[0072] When determining the specific implementation of the resource allocation result, for example, heuristic algorithms, optimization models, reinforcement learning, etc. are adopted.

[0073] S204. Dynamically allocate the communication resources corresponding to the resource allocation result for the vehicle, and the vehicle communicates based on the communication resources.

[0074] Among them, the communication resources include channels, and different channels are configured with different resources such as bandwidth, frequency, and transmit power. For example Figure 1As shown, it can be understood that the cloud server 12 dynamically allocates communication resources adapted to the target service data type for the vehicle 11, and sends the communication resources to the vehicle 11. The vehicle terminal 11 extracts the configuration parameters of the communication resources, adjusts its own communication parameters, and uses the communication resources to interact and communicate with the cloud server 12, or other surrounding vehicles, or third-party platforms, etc. For example, the vehicle 11 extracts the configuration parameters of the received communication resources (such as the transmission power of 50W), adjusts its own transmission power to the allocated value of 50W, and sends collision warning information to the surrounding vehicles through the allocated channel.

[0075] In the embodiments of the present application, by obtaining real-time channel state data, it is ensured that the current channel condition can be responded to in a timely manner. By predicting the change trend of the future channel state, a basis is provided for making the optimal resource allocation result. In addition, by using a pre-trained channel state prediction model to predict the future channel state, the accuracy is higher. By combining real-time channel state data and predicted channel state data, dynamic, flexible, and accurate adjustment of resource allocation is realized to adapt to different service requirements. Communication resources corresponding to the resource allocation result are dynamically allocated for the vehicle, and the vehicle is used to communicate according to the communication resources, thereby improving the utilization rate of communication resources, reducing latency, enhancing communication reliability, and meeting the requirements for real-time performance and reliability in the autonomous driving scenario; in addition, by dynamically allocating resources, it is possible to better adapt to the dynamic changes of the wireless channel state and the complex communication environment.

[0076] In some embodiments, determining a resource allocation result corresponding to the target service data type for the vehicle according to real-time channel state data and predicted channel state data includes: inputting the real-time channel state data and the predicted channel state data into a pre-constructed deep reinforcement learning model to obtain a resource allocation result corresponding to the target service data type for the vehicle. The state space of the deep reinforcement learning model includes real-time channel state data, predicted channel state data, and service data type, and the action space of the deep reinforcement learning model includes at least one of a spectrum allocation strategy, a transmission power allocation strategy, and a modulation and coding allocation strategy.

[0077] Among them, the deep reinforcement learning (DRL) model can be understood as a trained model. For example, after constructing the deep reinforcement learning model, algorithms such as Deep Q-Network (DQN) or Proximal Policy Optimization (PPO) are used to train the deep reinforcement learning model. The DRL model combines the technologies of deep learning and reinforcement learning and is used to solve complex decision-making and control problems.

[0078] When constructing a DRL model, the state space is defined to include real-time channel state data, predicted channel state data, and service data types. If the state space is represented as S, then , where represents real-time channel state data, represents predicted channel state data, represents service data types. can be understood as a set, containing data of multiple vehicles. For example, contains real-time channel state data of multiple networked vehicles. Correspondingly, contains predicted channel state data of the multiple networked vehicles, contains service data types of the multiple networked vehicles. The action space is defined to include at least one decision objective.

[0079] In one example, the action space includes a spectrum allocation strategy, or a transmit power allocation strategy, or a modulation and coding allocation strategy. When determining the resource allocation result corresponding to the target service data type of a vehicle, only a single-dimensional decision objective is considered. For example, optimizing spectrum allocation alone may not take into account the impact of power adjustment on the modulation and coding scheme. For the DRL model, it is prone to local optimality.

[0080] In another example, the action space includes any two decision objectives among the spectrum allocation strategy, the transmit power allocation strategy, and the modulation and coding allocation strategy. When determining the resource allocation result corresponding to the target service data type of a vehicle, two-dimensional decision objectives are considered. For example, when the channel state quality deteriorates, switching the frequency band and reducing the transmit power simultaneously can better adapt to the change of the channel state.

[0081] In yet another example, the action space includes three decision objectives: the spectrum allocation strategy, the transmit power allocation strategy, and the modulation and coding allocation strategy. When determining the resource allocation result corresponding to the target service data type of a vehicle, for example, considering that the high-frequency band has a large bandwidth but is easily blocked, it is necessary to combine high power and low modulation and coding (such as QPSK) to ensure reliability; the low-frequency band has a wide coverage but limited bandwidth, which is suitable for high modulation and coding (such as 64-QAM) to improve the rate. By unifying the spectrum allocation strategy, the transmit power allocation strategy, and the modulation and coding allocation strategy into the action space, the decision-making process is simplified. At the same time, through multi-dimensional joint decision-making, the best balance point can be found, avoiding local optimality, maximizing the utilization efficiency of each resource, and enhancing the adaptability to the dynamic environment.

[0082] If the action space is represented as A, then , where represents the spectrum allocation strategy, represents the transmit power allocation strategy, represents the modulation and coding allocation strategy. It can be understood as a set that contains allocation strategies for multiple vehicles. For example, it contains spectrum allocation strategies for multiple networked vehicles, it contains transmit power allocation strategies for multiple networked vehicles, and it contains modulation and coding allocation strategies for multiple networked vehicles. For each networked vehicle, its spectrum allocation strategy, transmit power allocation strategy, and modulation and coding allocation strategy are dynamically updated and changed.

[0083] Exemplarily, real-time channel state data and predicted channel state data are used as the input of the DRL model. The DRL model outputs the optimal resource allocation result of the target service data type under this channel state according to the real-time channel state data and the predicted channel state data. For example, when the channel state is good, the modulation and coding allocation strategy is to allocate high-order modulation (such as 64-QAM) to improve the transmission efficiency. When the channel state is poor, the modulation and coding allocation strategy is to allocate low-order modulation (such as QPSK) to improve the transmission reliability.

[0084] In some embodiments, the reward function of the deep reinforcement learning model is determined based on multi-objective rewards of spectrum utilization rate, transmission delay, and communication loss. The reward function is used to evaluate the impact degree of the resource allocation result on vehicle communication; and / or, the resource allocation method is applied to a cloud server, and the deep reinforcement learning model is deployed on an edge computing node, and the cloud server and the edge computing node work together.

[0085] The vehicle communication network resource allocation method provided by the embodiments of this application, when implemented, may include the following three implementation manners:

[0086] In one implementation manner, the reward function of the deep reinforcement learning model is determined based on multi-objective rewards of spectrum utilization rate, transmission delay, and communication loss.

[0087] Among them, the design of the reward function is the core of whether the DRL model can achieve the expected goal, and it is used to evaluate the effect of the resource allocation result on vehicle communication. The reward function is determined based on multiple objectives. The larger the reward value, the more optimal balance point the DRL model finds in the multi-objective conflict.

[0088] The spectrum utilization rate measures the amount of effective data (bps) transmitted within a unit bandwidth (Hz), reflecting the resource usage efficiency. The communication loss includes path loss and interference loss, reflecting the attenuation and interference effects of the signal during transmission, and affecting the reliability of communication. The transmission delay involves the time from data sending to receiving. Low delay is crucial for real-time communication in autonomous driving.

[0089] When designing the reward function, it is necessary to consider the possible trade - off relationships among spectrum utilization, communication loss, and transmission delay. For example, increasing spectrum utilization may increase interference, thereby increasing communication loss; or, reducing transmission delay may require more resources, affecting spectrum utilization. Therefore, the reward function is obtained by linearly weighting spectrum utilization, communication loss, and transmission delay. If the reward function is denoted as R, then , where is the spectrum utilization, is the weight of the spectrum utilization, is the transmission delay, is the weight of the transmission delay, is the communication loss, is the weight of the communication loss.

[0090] It should be noted that , and can be dynamically adjusted based on the Pareto optimization algorithm, which can solve the problem of trade - off among multiple conflicting objectives (spectrum utilization, transmission delay, and communication loss). Among them, the initial weights can be set according to the historical proportions of spectrum utilization, transmission delay, and communication loss. For example, if the spectrum utilization is large, is set small, and if the spectrum utilization is small, is set large.

[0091] Optionally, the design of the reward function can also consider fairness, and determine the reward function based on multi - objective rewards of spectrum utilization, transmission delay, communication loss, and fairness. Correspondingly,

[0092] ,

[0093] where is fairness, is the weight of fairness.

[0094] Through dynamic weight adjustment, the system can flexibly respond to the changing communication environment and find the real - time optimal balance point in multi - objective conflicts. In addition, the reward function comprehensively considers spectrum utilization, transmission delay, and communication loss, ensuring that the resource allocation results meet multi - dimensional requirements, achieving comprehensive optimization and improving the overall performance.

[0095] In another implementation, the resource allocation method is applied to the cloud server, and the deep reinforcement learning model is deployed on the edge computing node. The cloud server and the edge computing node work together.

[0096] Exemplarily, the cloud server includes an edge computing node and a central control node. Among them, the DRL model is deployed on the edge computing node. The central control node sends the acquired real-time channel state data of the vehicle and the predicted channel state data to the edge computing node in real time. The edge computing node performs dynamic resource allocation to determine the resource allocation result, and the central control node dynamically allocates the communication resources corresponding to the resource allocation result to the vehicle. The cloud server may include multiple edge computing nodes, and multiple edge computing nodes complete the determination of the resource allocation result in parallel. For example, the training and optimization of the DRL model are completed by edge computing node a and edge computing node b, and the determination of the resource allocation result is completed by edge computing node c and edge computing node d. The edge computing node and the central control node work together to complete the resource allocation method.

[0097] By combining the edge computing node with distributed computing, the decision-making process of the DRL model is accelerated, the data transmission delay is reduced, and the real-time requirement of vehicle communication is met.

[0098] In another implementation, the reward function of the deep reinforcement learning model is determined based on multi-objective rewards of spectrum utilization rate, transmission delay, and communication loss. Moreover, the resource allocation method is applied to the cloud server, the deep reinforcement learning model is deployed on the edge computing node, and the cloud server and the edge computing node work together.

[0099] Through dynamic weight adjustment, the system can flexibly respond to the changing communication environment and find the real-time optimal balance point in multi-objective conflicts. In addition, the reward function comprehensively considers the spectrum utilization rate, transmission delay, and communication loss to ensure that the resource allocation result meets multi-dimensional requirements, realizes comprehensive optimization, and improves the overall performance. By combining the edge computing node with distributed computing, the decision-making process of the DRL model is accelerated, the data transmission delay is reduced, and the real-time requirement of vehicle communication is met.

[0100] In some embodiments, the deep reinforcement learning model is optimized in the following manner: storing the reward value corresponding to the resource allocation result, as well as the state space data and action space data corresponding to the reward value, in a buffer; sampling from the buffer to obtain sampling data; and optimizing the model parameters of the deep reinforcement learning model according to the sampling data.

[0101] Among them, for the calculation of the reward value, in one way, the vehicle terminal calculates the spectrum utilization rate, transmission delay, and communication loss based on the vehicle's communication data, and feeds back the spectrum utilization rate, transmission delay, and communication loss to the DRL model in the cloud server. Then, the reward value is obtained according to the reward function and sent to the cloud server in real time; in another way, the cloud server obtains the vehicle's communication data in real time, calculates the spectrum utilization rate, transmission delay, and communication loss corresponding to the communication data, and then obtains the reward value according to the reward function. The communication data can be obtained by parsing the vehicle OBU (On-Board Unit) log to obtain data such as transmission power, reception power, transmission rate, and bandwidth. Further, the state space data, action space data, and reward value are stored in the experience replay buffer, and data is randomly sampled from the experience replay buffer to update the DRL model parameters.

[0102] In the embodiments of the present application, through continuous feedback, the DRL model is optimized online to improve the model performance, and then the resource allocation result is continuously optimized to better adapt to the dynamically changing channel environment and service types.

[0103] In some embodiments, the deep reinforcement learning model is obtained through model compression, and the model compression includes pruning and / or quantization.

[0104] Among them, pruning reduces the size and computational complexity of the DRL model by removing unimportant weights or neurons in the neural network. For example, removing those weights that have little impact on the output of the DRL model. Common methods include threshold-based pruning and sparsity-based pruning. Pruning can significantly reduce the number of parameters and computational amount of the DRL model, while having little impact on the performance of the DRL model.

[0105] Quantization is to convert the floating-point representation (such as 32-bit floating-point number) of the DRL model into a low-precision representation (such as 8-bit integer) to reduce the storage requirement and computational complexity of the DRL model. Quantization can significantly reduce the storage requirement of the DRL model and accelerate the inference process.

[0106] In one implementation, the deep reinforcement learning model is obtained after pruning and quantization. Exemplarily, pruning compression is first performed to reduce the number of parameters of the model, providing a smaller model basis for quantization, and then quantization compression of the model is performed on the basis of pruning. It should be noted that pruning and quantization can also be performed alternately to gradually optimize the model, and the order of pruning and quantization can be adjusted according to specific circumstances.

[0107] In another implementation, the deep reinforcement learning model is obtained after pruning.

[0108] In yet another implementation, the deep reinforcement learning model is obtained after quantization.

[0109] In the embodiments of the present application, a deep reinforcement learning model is obtained through model compression, which reduces the model inference time while ensuring the performance of the deep reinforcement learning model, meeting the real-time requirements of the autonomous driving scenario.

[0110] In some embodiments, communication resources corresponding to a resource allocation result are dynamically allocated to a vehicle, including: determining a target channel for the resource allocation result corresponding to a target service data type based on a pre-set correspondence between channels and resource allocations; dynamically allocating the target channel to the vehicle, where the target channel is used to transmit service data corresponding to the target service data type.

[0111] Exemplarily, the correspondence between channels and resource allocations is pre-set in the cloud server. For example, the correspondence between channels and resource allocations is stored in the form of a table. When the DRL model outputs the resource allocation result corresponding to the target service data type, the cloud server determines the target channel corresponding to the resource allocation result by looking up the table internally. The target channel and the target service data type are both sent to the vehicle. The vehicle extracts the configuration parameters of the target channel, adjusts its own communication parameters, and communicates with the cloud server, or other surrounding vehicles, or a third-party platform, etc. through the target channel to transmit service data corresponding to the target service data type.

[0112] In the embodiments of the present application, a target channel for the resource allocation result corresponding to a target service data type is determined through the correspondence between channels and resource allocations, and communication resources corresponding to the resource allocation result are dynamically allocated to the vehicle. The vehicle is used to communicate according to the communication resources, thereby improving the utilization rate of communication resources, reducing latency, enhancing communication reliability, and meeting the real-time and reliability requirements in the autonomous driving scenario.

[0113] In some embodiments, inputting real-time channel state data into a pre-trained channel state prediction model includes: preprocessing the real-time channel state data to obtain processed channel state data, where the preprocessing includes data cleaning and normalization processing; inputting the processed channel state data into the pre-trained channel state prediction model.

[0114] Among them, data cleaning is not limited to including handling missing values, handling outliers, deduplication, and data consistency, etc. The normalization processing can scale the real-time channel state data to a specific range, improving the processing efficiency and performance of the channel state prediction model. The normalization processing is, for example, min-max normalization, standard normalization.

[0115] Exemplarily, for each type of data in the real-time channel state data, such as signal-to-noise ratio, bandwidth, delay, etc., data cleaning is first performed, and then normalization processing is performed on the cleaned data. When performing normalization processing, the following formula is used for min-max normalization processing:

[0116]

[0117] Among them, is the channel state data after cleaning, is the minimum value in the channel state data after cleaning, is the maximum value in the channel state data after cleaning, is the normalized data (i.e., the processed channel state data).

[0118] Finally, input the processed channel state data into the pre-trained channel state prediction model.

[0119] In the embodiments of the present application, by performing data cleaning and normalization processing on real-time channel state data, the integrity and consistency of the data are ensured, the processing efficiency and performance of the channel state prediction model are improved, and thus the accuracy of predicting channel state data is improved, providing a basis for making the optimal resource allocation result subsequently.

[0120] In some embodiments, the channel state prediction model can be trained in the following manner: Obtain training samples, where the training samples include the historical channel state data of the vehicle (assuming a certain historical moment is T, obtain the historical channel state data N of n time slots relative to moment T and the future channel state data M of m time slots); Input the historical channel state data N into the channel state prediction model for channel state prediction to obtain the predicted value N' of the historical channel state data N; Use the mean square error as the loss function to determine the loss value of the predicted value N' relative to the future channel state data M; According to the loss value and the gradient descent method, continuously iterate and optimize the model parameters of the channel state prediction model.

[0121] In summary, the present application has at least the following advantages:

[0122] First, by obtaining real-time channel state data, it is ensured that the current channel condition can be responded to in a timely manner. By predicting the change trend of the future channel state, a basis is provided for making the optimal resource allocation result. In addition, by predicting the future channel state through the pre-trained channel state prediction model, the accuracy is higher.

[0123] Second, by combining real-time channel state data and predicted channel state data, a deep reinforcement learning model is used to achieve dynamic, flexible, and accurate adjustment of resource allocation to adapt to different service requirements. Dynamically allocate the communication resources corresponding to the resource allocation result for the vehicle, and the vehicle is used to communicate according to the communication resources, thereby improving the utilization rate of communication resources, reducing latency, enhancing communication reliability, and meeting the requirements for real-time and reliability in the autonomous driving scenario; In addition, by dynamically allocating resources, it can better adapt to the dynamic changes of the wireless channel state and the complex communication environment.

[0124] III. Design of Multi-Objective Reward Function: The reward function comprehensively considers spectrum utilization, latency, fairness, and energy consumption to ensure that the resource allocation strategy meets multi-dimensional requirements, achieves comprehensive optimization, and improves overall performance.

[0125] IV. Design of Efficient Action Space: Unify the spectrum allocation strategy, transmit power allocation strategy, and modulation and coding allocation strategy into the action space to simplify the decision-making process. At the same time, through multi-dimensional joint decision-making, the best balance point can be found to avoid local optimality, maximize the utilization efficiency of each resource, and enhance the adaptability to dynamic environments.

[0126] V. Online Learning and Adaptive Update: Continuously optimize the DRL model through continuous feedback to improve the model performance, and then continuously optimize the resource allocation results to better adapt to the dynamically changing channel environment and service types.

[0127] VI. Obtain a deep reinforcement learning model through model compression. On the premise of ensuring the performance of the deep reinforcement learning model, reduce the model inference time, lower the complexity of the model algorithm, and meet the real-time requirements of the autonomous driving scenario.

[0128] VII. Accelerate the decision-making process of the DRL model by combining edge computing nodes and distributed computing, reduce data transmission latency, and meet the real-time requirements of vehicle communication.

[0129] Figure 3 The structural schematic diagram of the vehicle communication network resource allocation device provided for the application embodiment is as Figure 3 shown. The vehicle communication network resource allocation device 30 provided in this embodiment includes: an acquisition module 31, an input module 32, a determination module 33, and an allocation module 34. Among them:

[0130] The acquisition module 31 is used to acquire the real-time channel state data of the vehicle;

[0131] The input module 32 is used to input the real-time channel state data into the pre-trained channel state prediction model to obtain the predicted channel state data corresponding to the real-time channel state data;

[0132] The determination module 33 is used to determine the resource allocation result corresponding to the target service data type for the vehicle according to the real-time channel state data and the predicted channel state data;

[0133] The allocation module 34 is used to dynamically allocate the communication resources corresponding to the resource allocation result for the vehicle, and the vehicle communicates based on the communication resources.

[0134] In a possible implementation, the determining module 33 is specifically configured to: input real-time channel state data and predicted channel state data into a pre-constructed deep reinforcement learning model to obtain a resource allocation result corresponding to a target service data type for the vehicle. The state space of the deep reinforcement learning model includes real-time channel state data, predicted channel state data, and service data types, and the action space of the deep reinforcement learning model includes at least one of a spectrum allocation policy, a transmit power allocation policy, and a modulation and coding allocation policy.

[0135] In a possible implementation, the reward function of the deep reinforcement learning model is determined based on multi-objective rewards of spectrum utilization rate, transmission delay, and communication loss. The reward function is used to evaluate the influence degree of the resource allocation result on vehicle communication; and / or, the resource allocation method is applied to a cloud server, the deep reinforcement learning model is deployed on an edge computing node, and the cloud server and the edge computing node cooperate.

[0136] In a possible implementation, the deep reinforcement learning model is optimized in the following manner:

[0137] Store the reward value corresponding to the resource allocation result, as well as the state space data and action space data corresponding to the reward value, in a buffer; sample from the buffer to obtain sampled data; and optimize the model parameters of the deep reinforcement learning model according to the sampled data.

[0138] In a possible implementation, the deep reinforcement learning model is obtained through model compression, and the model compression includes pruning and / or quantization.

[0139] In a possible implementation, the allocation module 34 is specifically configured to: determine a target channel for the resource allocation result corresponding to the target service data type based on a pre-set correspondence between channels and resource allocations; and dynamically allocate the target channel for the vehicle, and the target channel is used to transmit service data corresponding to the target service data type.

[0140] In a possible implementation, the input module 32 is specifically configured to: preprocess the real-time channel state data to obtain processed channel state data, and the preprocessing includes data cleaning and normalization processing; and input the processed channel state data into a pre-trained channel state prediction model.

[0141] The vehicle communication network resource allocation device 30 provided in this embodiment can execute the method provided in the above method embodiment, and its implementation principle and technical effect are similar, and will not be elaborated here in this embodiment.

[0142] Figure 4 This is a schematic structural diagram of the device provided in the embodiment of the present application. As Figure 4As shown, the device 40 provided in this embodiment includes: at least one processor 41 and a memory 42. Optionally, the device 40 further includes a communication component 43. Among them, the processor 41, the memory 42, and the communication component 43 are connected through a bus 44.

[0143] In a specific implementation process, at least one processor 41 executes the computer-executable instructions stored in the memory 42, so that at least one processor 41 executes the above-mentioned method.

[0144] For the specific implementation process of the processor 41, reference can be made to the above method embodiment, and its implementation principle and technical effects are similar, so they will not be elaborated here in this embodiment.

[0145] In the above embodiment, it should be understood that the processor can be a central processing unit (English: Central Processing Unit, abbreviated: CPU), or other general-purpose processors, digital signal processors (English: Digital Signal Processor, abbreviated: DSP), application-specific integrated circuits (English: Application Specific Integrated Circuit, abbreviated: ASIC), etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. The steps of the method disclosed in combination with the invention can be directly embodied as being executed by a hardware processor, or executed by a combination of hardware and software modules in the processor.

[0146] The memory may include a high-speed memory (Random Access Memory, RAM), and may also include a non-volatile memory (Non-volatile Memory, NVM), such as at least one disk memory.

[0147] The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience in representation, the bus in the drawings of this application is not limited to only one bus or one type of bus.

[0148] This embodiment of the present application also provides a computer program product, including a computer program, and when the computer program is executed by a processor, the above-mentioned method is implemented.

[0149] An embodiment of the present application further provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the above method.

[0150] The above-readable storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, a magnetic disk, or an optical disk. The readable storage medium can be any available medium accessible by a general-purpose or special-purpose computer.

[0151] An exemplary readable storage medium is coupled to a processor, enabling the processor to read information from and write information to the readable storage medium. Of course, the readable storage medium can also be a component of the processor. The processor and the readable storage medium can be located in an application specific integrated circuit (ASIC). Of course, the processor and the readable storage medium can also exist as discrete components in a device.

[0152] The division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces, and the indirect coupling or communication connection of devices or units can be in an electrical, mechanical, or other form.

[0153] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0154] In addition, in each embodiment of the present invention, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.

[0155] If a function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods of the various embodiments of the present invention. The aforementioned storage medium includes: USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs, etc., which are various media that can store program codes.

[0156] Those of ordinary skill in the art can understand that all or part of the steps of implementing the above method embodiments can be completed by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When this program is executed, it performs the steps including the above method embodiments; and the aforementioned storage medium includes: ROMs, RAMs, magnetic disks, or optical discs, etc., which are various media that can store program codes.

[0157] Finally, it should be noted that: After considering the specification and practicing the invention disclosed herein, those skilled in the art will easily think of other implementation manners of the present invention. The present invention is intended to cover any variations, uses, or adaptive changes of the present invention. These variations, uses, or adaptive changes follow the general principles of the present invention and include the common general knowledge or conventional technical means in the technical field of the present invention that are not disclosed in the present invention. It is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present invention is only limited by the appended claims.

Claims

1. A method for allocating vehicle communication network resources, characterized in that Including: Obtain the real-time channel state data of the vehicle; Input the real-time channel state data into a pre-trained channel state prediction model to obtain predicted channel state data corresponding to the real-time channel state data; Determine a resource allocation result corresponding to the target service data type for the vehicle according to the real-time channel state data and the predicted channel state data; Dynamically allocate communication resources corresponding to the resource allocation result for the vehicle, and the vehicle communicates based on the communication resources.

2. The method according to claim 1, wherein The determining the resource allocation result corresponding to the target service data type for the vehicle according to the real-time channel state data and the predicted channel state data includes: Input the real-time channel state data and the predicted channel state data into a pre-constructed deep reinforcement learning model to obtain a resource allocation result corresponding to the target service data type for the vehicle; Wherein, the state space of the deep reinforcement learning model includes the real-time channel state data, the predicted channel state data, and the service data type, and the action space of the deep reinforcement learning model includes at least one of a spectrum allocation strategy, a transmit power allocation strategy, and a modulation and coding allocation strategy.

3. The method according to claim 2, wherein The reward function of the deep reinforcement learning model is determined based on multi-objective rewards of spectrum utilization rate, transmission delay, and communication loss, and the reward function is used to evaluate the influence degree of the resource allocation result on vehicle communication; And / or, the method is applied to a cloud server, the deep reinforcement learning model is deployed on an edge computing node, and the cloud server and the edge computing node work together.

4. The method according to claim 2, wherein The deep reinforcement learning model is optimized in the following manner: Store the reward value corresponding to the resource allocation result, as well as the state space data and action space data corresponding to the reward value, into a buffer; Sample from the buffer to obtain sampled data; Optimize the model parameters of the deep reinforcement learning model according to the sampled data.

5. The method according to claim 2, wherein The deep reinforcement learning model is obtained through model compression, and the model compression includes pruning and / or quantization.

6. The method according to any one of claims 1 to 5, characterized in that The dynamically allocating communication resources corresponding to the resource allocation result for the vehicle includes: Based on a preset correspondence between channels and resource allocation, determine a target channel for the resource allocation result corresponding to the target service data type; Dynamically allocate the target channel for the vehicle, and the target channel is used to transmit service data corresponding to the target service data type.

7. The method according to any one of claims 1 to 5, characterized in that, The inputting the real-time channel state data into a pre-trained channel state prediction model includes: Preprocess the real-time channel state data to obtain processed channel state data, and the preprocessing includes data cleaning and normalization processing; Input the processed channel state data into a pre-trained channel state prediction model.

8. A vehicle communication network resource allocation device, characterized in that, Including: An acquisition module, configured to acquire the real-time channel state data of the vehicle; An input module, configured to input the real-time channel state data into a pre-trained channel state prediction model to obtain predicted channel state data corresponding to the real-time channel state data; A determination module, configured to determine a resource allocation result corresponding to a target service data type for the vehicle according to the real-time channel state data and the predicted channel state data; An allocation module, configured to dynamically allocate communication resources corresponding to the resource allocation result for the vehicle, and the vehicle communicates based on the communication resources.

9. A device, characterized in that, Including: A memory and a processor; The memory stores computer-executable instructions; The processor executes the computer-executable instructions stored in the memory, so that the processor executes the method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, Computer-executable instructions are stored in the computer-readable storage medium, and when the computer-executable instructions are executed, they are used to implement the method according to any one of claims 1 to 7.

Citation Information

Cited By

  • Vehicle-mounted spectrum resource allocation method and device, equipment and storage medium

    CN120980695A