Internet of vehicles cooperative perception resource allocation method based on ISAC

By building a multi-layer collaborative perception area and time-division dynamic frame structure in the Internet of Vehicles, combining on-board computing capabilities, and optimizing the resource allocation of ISAC devices, the problem of unbalanced bandwidth and computing resource utilization in collaborative perception of the Internet of Vehicles is solved, latency is reduced, and system performance is improved.

CN119255377BActive Publication Date: 2025-10-17CHONGQING UNIV OF POSTS & TELECOMM
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410988773.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-23
Publication Date
2025-10-17
Estimated Expiration
2044-07-23

AI Technical Summary

Technical Problem

In existing collaborative perception solutions for the Internet of Vehicles, bandwidth and computing power utilization are unbalanced, resulting in high latency and underutilized computing resources. In addition, ISAC devices interfere with the balance between perception and communication performance, affecting system performance.

Method used

An ISAC-based collaborative perception resource allocation method for Internet of Vehicles (IoV) is adopted. By constructing a multi-layer collaborative perception interest area and a time-division dynamic frame structure, combined with on-board computing capabilities, the allocation of perception and communication resources is optimized, and a multi-agent deep deterministic policy gradient algorithm in a hybrid action space is used to optimize resource allocation.

Benefits of technology

The execution delay of collaborative perception is reduced, system performance and resource utilization are improved, and a more efficient balance of perception and communication is achieved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119255377B_ABST
    Figure CN119255377B_ABST
Patent Text Reader

Abstract

The present invention relates to a method for allocating collaborative perception resources in an Internet of Vehicles (IoV) based on ISAC, and belongs to the field of mobile communication technology. The method comprises: constructing a multi-layer collaborative perception area of ​​interest scheme based on relative position; confirming a common perceiver based on the relative position of the perception area of ​​interest of the current vehicle and the road test unit and other IoV vehicles; constructing an ISAC time allocation scheme based on a time-division dynamic frame structure to obtain the total perception mutual information SMI of the radar sensing duration of the common perceiver, the average communication rate of the communication duration, and the execution delay of each processing mode; establishing a joint problem of perception task allocation, offloading, and computing resource allocation with the goal of minimizing the delay in completing the collaborative perception task; and using a multi-agent deep deterministic policy gradient algorithm in a hybrid action space to solve the optimization problem and obtain a resource allocation scheme. The present invention can optimize the resource allocation strategy for collaborative perception of the IoV and reduce the execution delay of collaborative perception.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of mobile communication, and relates to a vehicle networking cooperative sensing resource allocation method based on ISAC. BACKGROUND

[0002] The perception ability of a single connected automatic vehicle (CAV) has inherent limitations. The emergence of cooperative sensing technology based on vehicle networking can overcome this defect. By fusing the perception data of multiple CAVs and RSUs, the perception range can be greatly expanded and the perception accuracy can be improved. However, the existing single-layer cooperative sensing scheme still has some problems. First, the bandwidth and computing power utilization rate is saturated or insufficient. In early fusion, excessive raw data transmission may cause bandwidth saturation while computing power is not fully utilized. In late fusion, sending object data after local processing may cause computing power saturation while bandwidth is not fully utilized. Even in the middle fusion, under dynamic network conditions, the above problems still exist. Second, for a single CAV, the road condition information it is most concerned about is not the global map information but the road condition information within a relatively short range. Generating a complete global environment map requires processing a large amount of perception data, which puts a huge burden on the system's computing and communication, resulting in high overall delay. Therefore, a more flexible and efficient cooperative sensing scheme needs to be designed to improve system performance.

[0003] Integrated communication and sensing (ISAC) devices play a key role in cooperative sensing. These devices can simultaneously perform radar sensing and communication functions, providing the necessary hardware foundation for cooperative sensing. By coordinating the use of radar and communication resources, ISAC devices can achieve an effective balance between sensing and communication, providing important support for the performance optimization of the entire cooperative sensing system. However, existing ISAC devices have some problems in terms of balancing sensing and communication performance and utilizing local computing resources. Due to the interference between sensing and communication functions, it is difficult to achieve the best resource allocation between the two. At the same time, the introduction of vehicle-mounted computing power is not considered, resulting in high latency in the entire cooperative sensing process. By analyzing the impact of different duration allocation ratios on sensing and communication performance, and combining vehicle-mounted computing power, the optimal time allocation ratio range for CAVs and RSUs is obtained, which can reduce the latency of the entire cooperative sensing process. This provides a flexible and efficient solution for the application of ISAC devices in cooperative sensing scenarios. SUMMARY

[0004] Therefore, the present application aims to provide an ISAC-based cooperative perception resource allocation method for vehicle networking, which can optimize the cooperative perception resource allocation strategy for vehicle networking and reduce the execution delay of cooperative perception.

[0005] To achieve the above-mentioned purpose, the present application provides the following technical solutions.

[0006] An ISAC-based cooperative perception resource allocation method for vehicle networking includes the following steps:

[0007] S1, constructing a multi-layer cooperative perception area of interest scheme based on relative position in a vehicle networking scenario;

[0008] S2, confirming the common perceivers according to the relative positions of the perception area of interest of the current vehicle, road test units and other vehicle networking vehicles;

[0009] S3, constructing an integrated ISAC time allocation scheme based on a time division dynamic frame structure, obtaining the total perception mutual information SMI of the radar sensing duration of the common perceivers, the average communication rate of the communication duration and the execution delay of each processing mode;

[0010] S4, establishing a joint problem of perception task allocation, unloading and computing resource allocation aiming to minimize the cooperative perception task completion delay;

[0011] S5, using a multi-agent deep deterministic policy gradient algorithm with a hybrid action space to solve the optimization problem to obtain the resource allocation scheme with the minimum task completion delay.

[0012] Further, in step S1, the vehicle networking scenario includes a plurality of vehicle networking autonomous vehicles CAVs, a plurality of road test units RSUs and a plurality of mobile edge servers MESs, the CAVs collect in-vehicle and out-vehicle data through various sensors, a plurality of RSUs are deployed along the road and connected to each other through high-speed wired links, and each RSU is connected to the MES through a wired link;

[0013] The current CAV selects common perceivers according to the relative positions of the RSUs and other CAVs in the perception area of interest, and the processing of the current CAV or the common perceivers for the perception data includes local processing and unloading processing;

[0014] Let N={1,...,i,...,N} be the set of CAVs in the system, the perception area of interest be P PAC , P RSU represent the coverage of the RSU, the current vehicle v1 cannot perceive all the areas in P PAC , divide P PAC into P1, P f and P b , P1 is the area within the field of view of v1, Pf and P b are the occlusion areas before and after v1 respectively, then:

[0015] P PAC = P1+ P f + P b

[0016] The perception tasks of the three areas are defined as T1, T f , T b respectively, and the perception task of the CAV concerned area is represented as:

[0017] T PAC = T1+ T f + T b

[0018] The CAV performs the perception task in the manner of local processing or offloading processing.

[0019] Further, in step S2, the process of transferring the current vehicle from the coverage range of the current RSU to the coverage range of the next RSU is divided into the following five stages, and the selection range of the common perceiver in each stage is:

[0020] In the first stage, T PAC is completely in the coverage range of RSU j , and the common perceiver can select RSU j and the CAV in P b , P f ;

[0021]

[0022] V f represents the CAV before v1 in P PAC , V b represents the CAV after v1 in P PAC , V f and V b may be multiple or none, V V2V is the set of CAVs within the communication range of v1, U j represents the RSU j in which v1 is currently located, and U j+1 represents the RSU j+1 that v1 is about to enter;

[0023] In the second stage, P PAC of T f enters the coverage range of RSU j+1 , and the common perceiver can select RSU j , RSU j+1 and the CAV in Pb , P f of CAV;

[0024]

[0025] In the third stage, T PAC of P f enter the coverage of RSU j+1 , P b do not enter, their common perceivers can select RSU j , RSU j+1 and CAV in P f ;

[0026]

[0027] In the fourth stage, T PAC of P b partly enter the coverage of RSU j+1 , their common perceivers can select RSU j , RSU j+1 and CAV in P b , P f ;

[0028]

[0029] In the fifth stage, T PAC of P j+1 have completely entered the coverage of RSU j+1 , their common perceivers can select RSU b and CAV in P f ;

[0030]

[0031] The task allocation function is defined as y = TA(x), which indicates that the cooperative task in T PAC is allocated to y, and a group of common perceivers are selected in RSU and CAV to assist v1 to complete the perception task according to the relative position information; let Ω represent the selected common perceivers set, which satisfies:

[0032] Ω = TA(T1)∪TA ρ (T f )∪TA ρ (T b )

[0033] Wherein TA ρ (T1), TA ρ (T f ), TA ρ (T b) respectively represent the perceivers who can perceive T0, T f , T b , ρ = 1, 2, 3, 4, 5 represents the stage.

[0034] Further, in step S3, first, a adjustable frame format including S sensing subframes, P calculation subframes and Q communication subframes is constructed, N s subframes in each cycle, assuming that the duration of each subframe is τ s , the time allocation decision of each vehicle ISAC device is described as χ vi = {a i , b i , c i}, where v i ∈ Ω V , a i , b i , c i are the normalized sensing duration, calculation duration and sensing duration corresponding to v i , respectively, a i + b i + c i = 1, where,

[0035] During the radar detection process, the average sensing mutual information (SMI) of v i in the radar sensing duration a i is expressed as:

[0036]

[0037] Where B i is the bandwidth allocated to v i ; is the signal-to-interference-plus-noise ratio of v i , which is expressed as:

[0038]

[0039] In the formula, P i is the transmit power of v i , represents the path propagation gain, n r-rad and n r-com represent the radar interference and communication interference caused by other CAVs to the radar signal of v i ; N0 represents the thermal noise power spectral density.

[0040] Further, in step S3, during the communication process, the total communication data volume of v i in the communication duration c i is:

[0041]

[0042] v i During the communication duration c i The average communication rate is:

[0043]

[0044] in, Indicates v i The signal-to-noise ratio of the communication signal is expressed as:

[0045]

[0046] Where G t and G r Indicates antenna transmission gain and antenna receiving gain, represents the communication channel gain from i to j, and They are the radar interference and communication interference respectively.

[0047] Further, in step S3, for each execution delay, the perception task is described as L i ={s i ,d i ,τ i}, s i represents the amount of environmental information that needs to be extracted in the perception task, d i Indicates the number of CPU cycles required to process one bit of environmental information, τ i The maximum tolerable delay for the entire perception task is:

[0048] The execution delay of the vehicle's local processing of perception data is:

[0049]

[0050] Where, Indicates the CPU cycle frequency allocated to the local task, and defines the maximum CPU cycle frequency of CAV satisfy

[0051] The execution delay of unloading a vehicle to the server is:

[0052]

[0053] Where, RSU j The maximum CPU cycle frequency, μ vi express Assigned to v ia proportion of computing resources.

[0054] The execution delay on the RSU side is:

[0055]

[0056] wherein, denotes u j a proportion of computing resources allocated for task processing.

[0057] Further, in step S4, the cooperative perception task allocation, offloading and computing resource allocation optimization scheme is formulated as a minimum cooperative perception task processing delay minimization problem, and the optimization problem is expressed as:

[0058]

[0059] wherein is a vehicle offloading decision, is a vehicle ISAC device time allocation proportion, is a proportion of computing resources allocated to the vehicle by the MES, Ω V denotes a CAV selected as a co-perceptioner, Ω U denotes an RSU selected as a co-perceptioner;

[0060] C1 denotes selecting a co-perceptioner; C2 denotes v i and u j The total processing time of v i The sum of the normalized perception, communication and computing time proportions allocated is 1; C4 denotes that the amount of data communicated cannot exceed the amount of data perceived in the same cycle; C5 denotes that the actual required perception, communication and computing time is not greater than the allocated perception, communication and computing time; C6 denotes that the sum of the resource proportions allocated by the RSU providing offloading services is not greater than 1.

[0061] Further, in step S5, the optimization problem is converted into a Markov decision process, and the state space is set as:

[0062]

[0063] wherein b, g, and l respectively represent the remaining computing resource state of the MES, the channel gain of the communication link between the CAV and the MES, and the connection state between the CAV and the MES;

[0064] Action space:

[0065] The discrete action of the offloading decision β vi , the time proportion allocation and the continuous variable of the computing resource allocation proportion μ jointly serve as the mixed action of the CAV

[0066] Reward:

[0067]

[0068] where β is a constant, the reward increases as the delay decreases;

[0069] The goal of training the value network is to minimize the prediction error of the state-action value function, assuming the parameters of each CAV are θ = {θ1, θ2,..., θ i ..., θ N}, for each CAV i , the target value is calculated using the target network; the target action a' j is calculated as: j j (s') represents the target action of the jth agent in the next state s' j , j ∈ {1, 2,..., N}, then the target value is:

[0070] y i = r i + γQ' i (s', a i )

[0071] where Q' i is the target value network of the CAV i ;

[0072] The value network parameters θ i are updated by minimizing the mean square error loss:

[0073]

[0074] where B is a batch drawn from the experience replay area;

[0075] When training the policy network, assume the policy network parameters of each CAV are ω = {ω1, ω2,..., ω i ,... ω N}, for each CAV i , the policy network is updated using the policy gradient, which is:

[0076]

[0077] The policy network parameters ω i are updated using the gradient ascent method;

[0078] For the target network parameters of the value network and the policy network, soft update is used:

[0079] θ​i ' = tau theta i ' = (1-tau) theta i

[0080] omega i ' = tau omega i ' = (1-tau) omega i

[0081] where tau represents a soft update parameter.

[0082] Further, in step S5, the specific process of solving the optimization problem by using the multi-agent deep deterministic policy gradient algorithm of the hybrid action space includes:

[0083] S51: initialize the actor network and critic network of each CAV;

[0084] S52: in each iteration, initialize a random process for action exploration, and obtain the initial observation of each CAV, i.e. the initial state of the environment;

[0085] S53: in each time slot, each CAV selects an action according to the current policy and state and executes it;

[0086] S54: in each time slot, all CAVs interact with the environment to obtain their respective rewards and jump to the next state, and store the experience data in the experience replay pool;

[0087] S55: for each CAV, randomly sample a small batch of samples from the experience pool;

[0088] S56: for each CAV, calculate the target state value of the critic, calculate the loss function, and minimize the loss to update the critic network, calculate the policy gradient, and update the actor network;

[0089] S57: after completing the target network parameters of the soft update value network and policy network of each CAV, return to S53, otherwise return to S55;

[0090] S58: after completing each time slot, return to S52, otherwise return to S53;

[0091] S59: stop after completing the iteration, otherwise return to S52.

[0092] The beneficial effects of the present application are:

[0093] The present application aims at the problems of saturation or deficiency of bandwidth and computing resource utilization in the existing cooperative perception fusion scheme of Internet of Vehicles, large amount of global map data generated, and the design of ISAC failing to balance the performance of perception and communication, resulting in high delay of the system, and proposes an ISAC-based cooperative perception resource allocation scheme for Internet of Vehicles. First, in order to reduce the amount of transmitted data and reasonably utilize bandwidth and computing resources, a multi-layer cooperative perception area of interest scheme based on relative position is proposed. Second, an ISAC perception, communication and computing time allocation scheme based on time division dynamic frame structure is proposed, the influence of different time allocation ratios on the mutual information (MI) of perception and communication is analyzed, and the time allocation is performed according to the computing capacity of CAV. The present application can realize the optimization of cooperative perception resource allocation strategy for Internet of Vehicles, and reduce the execution delay of cooperative perception.

[0094] Other advantages, objects, and features of the present application will be understood by those skilled in the art from the following specification in conjunction with the accompanying drawings in which: BRIEF DESCRIPTION OF DRAWINGS

[0095] In order to make the objects, technical solutions and advantages of the present application clearer, the preferred embodiments of the present application will be described in detail below with reference to the accompanying drawings, in which:

[0096] Fig. 1 Schematic diagram of cooperative perception scenario;

[0097] Fig. 2 Schematic diagram of five stages in the process of vehicle transfer from the current RSU to the next RSU;

[0098] Fig. 3 Network structure diagram of HAS-MADDPG algorithm. DETAILED DESCRIPTION

[0099] The embodiments of the present application are described below through specific specific examples, and those skilled in the art can easily understand other advantages and effects of the present application from the disclosure of the present specification. The present application can also be implemented or applied through other different specific embodiments, and the details in the present specification can be modified or changed in various ways based on different views and applications without departing from the spirit of the present application. It should be noted that the diagrams provided in the following examples only illustrate the basic concept of the present application in a schematic manner, and the following examples and features in the examples can be combined with each other without conflict.

[0100] The drawings are only used for exemplary illustration, and the representation is only a schematic diagram, not a physical diagram, and cannot be understood as a limitation on the present application; in order to better illustrate the embodiments of the present application, some components of the drawings may be omitted, enlarged or reduced, and do not represent the size of the actual product; for those skilled in the art, it is understandable that some well-known structures and their descriptions in the drawings may be omitted.

[0101] The same or similar reference numerals in the drawings of the embodiments of the present application correspond to the same or similar components; in the description of the present application, it should be understood that if the terms "upper", "lower", "left", "right", "front", "back" and the like indicate the orientation or positional relationship shown in the drawings, only for the convenience of describing the present application and simplifying the description, and do not indicate or imply that the devices or elements referred to must have a particular orientation, be constructed and operated in a particular orientation, therefore the positional relationship described in the drawings is only used for exemplary illustration, and cannot be understood as a limitation on the present application, for those skilled in the art, the specific meaning of the above terms can be understood according to the specific circumstances.

[0102] Please refer to Figs. 1-3 , a vehicle networking collaborative perception resource allocation method based on ISAC, which can realize the optimization of vehicle networking collaborative perception resource allocation strategy and reduce the execution delay of collaborative perception. Specifically, it includes the following steps:

[0103] S1: In the vehicle networking scene, a multi-layer collaborative perception area of interest scheme based on relative position is proposed, the road condition information in the close range of the car is defined as the perception area of interest, and comprehensive perception of the area is realized; Fig. 1 The scene is a collaborative perception scene diagram.

[0104] The scene mainly consists of a vehicle networking autonomous vehicle CAV, a road test unit RSU and a mobile edge server MES. The CAV is equipped with various sensors, responsible for collecting data inside and outside the vehicle; a number of equidistant RSUs are deployed along the roadside, and the RSUs are connected to each other through high-speed wired links; the MES provides computing services for the system, and each RSU is connected to the MES through a wired link.

[0105] The multi-layer collaborative perception area of interest scheme based on relative position includes that the current CAV selects appropriate co-perceivers according to the relative positions of the RSUs and other CAVs in the perception area of interest, the current CAV and the co-perceivers perform data fusion, and the current CAV or the co-perceivers can choose local processing or offload processing when processing the perception data. Let N={1,...,i,...,N} be the set of CAVs in the system.

[0106] The relevant area is described as a perception area of interest (PAC) frame P PAC , PRSU represents the coverage of RSU, v1 cannot perceive P PAC in all areas due to the occlusion of buildings and other vehicles, the occluded areas of v1 include the front occlusion and the rear blind spot. Therefore, P PAC is divided into P1, P f and P b . P1 is the area within the field of view of v1, P f and P b are the front and rear occlusions of v1 respectively. Therefore, there are

[0107] P PAC = P1+P f +P b

[0108] When facing occlusion, vehicles and pedestrians in P f and P b areas cannot be perceived, which may cause serious traffic accidents, and for P1 area, if the computing power of the vehicle is insufficient, it is difficult to process the perception data, and V2I can also request RSU to help process. The perception tasks of the three areas are defined as T1, T f , T b , therefore, the perception task of the CAV perception concerned area can be expressed as:

[0109] T PAC =T1+T f +T b

[0110] When performing the perception task, the CAV can choose local processing or offloading processing.

[0111] S2: analyze the influence of relative position on the selection of common perceivers, the relative position refers to the relative position of RSU, other CAV and the perception concerned area.

[0112] Therefore, appropriate common perceivers need to be selected to complete the perception in P1, P f and P b areas, and provide v1 with a good field of view, task T1 can be executed by v1. Since the moving speed of CAV is very fast, it is easy to move out of the coverage range of the current RSU, therefore, the process of moving from the coverage range of the current RSU to the coverage range of the next RSU is discussed, as shown in Fig. 2 , the process is divided into five stages:

[0113] In the first stage, T PAC is completely in the coverage range of RSU j , its common perceivers can select RSU j and P b, P f CAV of T

[0114]

[0115] V f CAV before v1 in P PAC , V b CAV after v1 in P PAC , V f and V b may or may not have, V V2V is the set of CAVs in the communication range of v1, U j denotes the RSU in which v1 is currently located j , U j+1 denotes the RSU into which v1 is about to enter j+1 ;

[0116] In the second stage, T PAC P f part of the CAVs enter the coverage of RSU j+1 , the co-awareness can select RSU j , RSU j+1 and CAV in P b , P f ;

[0117]

[0118] In the third stage, T PAC P f enter the coverage of RSU j+1 , P b does not enter, the co-awareness can select RSU j , RSU j+1 and CAV in P f ;

[0119]

[0120] In the fourth stage, T PAC P b part of the CAVs enter the coverage of RSU j+1 , the co-awareness can select RSU j , RSU j+1 and CAV in P b , P f ;

[0121]

[0122] In the fifth stage, T PAC has completely entered the coverage of RSUj+1 The coverage range of RSU j+1 and P b CAVs in P f .

[0123]

[0124] The task allocation function is defined as y = TA(x), which indicates that T PAC is allocated to y. According to the relative position information, a group of co-sensors is selected in RSU and CAV to assist v1 to complete the perception task. Let Ω represent the selected set of co-sensors, which satisfies:

[0125] Ω = TA(T1)∪TA ρ (T f )∪TA ρ (T b )

[0126] In the formula, TA ρ (T1), TA ρ (T f ), TA ρ (T b ) respectively represent the sensors that can perceive T0, T f , T b , and ρ = 1, 2, 3, 4, 5 represents the stage.

[0127] S3: A time allocation scheme for integrated sensing and communication ISAC based on time division dynamic frame structure is proposed, and the influence of different duration allocation ratios on the mutual information of perception and communication is analyzed, so that each vehicle can obtain the best duration allocation ratio; Specifically, based on the time division frame structure and radar communication interference analysis, the influence of different duration allocation ratios on the perception MI and communication MI is analyzed, and the allocation of calculation time is set considering the calculation ability of CAV itself, to obtain the best duration allocation ratio of a CAV and reduce the delay of cooperative perception process. Including the following steps:

[0128] An S / P / Q adjustable frame format composed of S sensing subframes, P calculation subframes and Q communication subframes is designed. There are N s subframes in each cycle, and the duration of each subframe is τ s . The time allocation decision of each vehicle ISAC device is described as χ vi = {a i ,b i ,c i}, where v i ∈Ω V , a i ,b i,c i v i The corresponding normalized perception duration, calculation duration and communication duration, a i +b i +c i =1. The calculation duration means that the CAV perception task processing may be selected to be processed locally. If b i Indicates the calculation duration. If you choose to offload to MES processing, b i is 0.

[0129] Analyze the radar detection process. When v i While performing radar sensing functions, other devices may be performing radar sensing, computing, or communication functions, so v i It may be interfered by radar or communication signals from other devices. Its radar receiving signal can be described as:

[0130] y rad =s rad h rad +n r-com +n r-rad +n

[0131] Where s rad Represents the radar transmission signal, h rad represents the propagation gain of the radar response signal in the channel, n r-rad and n r-com Represents other CAV pairs v i The radar interference and communication interference caused by the radar signal, n represents the interference to the signal receiver (including internal thermal noise and other external information interference). Therefore, v i The signal to interference plus noise ratio (SINR) can be expressed as:

[0132]

[0133] Where P i It is v i The transmission power, represents the path propagation gain, N0 represents the thermal noise power spectral density, B i Represented as assigned to v i Bandwidth. During the duration, the radar interference can be expressed as:

[0134]

[0135] Where r j,r Indicates the duration of radar detection Whether j is in the radar sensing time and the communication interference it suffers can be expressed as:

[0136]

[0137] where r j,c denotes the radar sensing duration, whether j is in the communication duration, denotes the communication signal path loss from RSU j to v i during the communication duration. G t and G r denote the antenna transmit gain and antenna receive gain.

[0138] v i The average sensing mutual information (SMI) during the radar sensing duration a i can be expressed as:

[0139]

[0140] Analyzing the communication process, during the communication duration c i of v i , it will be interfered by other devices’ radar and communication. During the communication duration c i of v i , the communication received signal can be described as:

[0141] y com = s com h com + n r-com + n r-rad + n

[0142] Assuming v i communicates with RSU j, the signal and interference plus noise ratio (SINR) during the communication duration can be expressed as:

[0143]

[0144] denotes the communication channel gain from i to j, the communication interference during the duration can be expressed as:

[0145]

[0146] The radar interference can be expressed as:

[0147]

[0148] where d i,j denotes the distance between v i and j.

[0149] vi The total communication data volume of the process is: i The average communication rate is:

[0150]

[0151] v i The average communication rate is: i The average communication rate is:

[0152]

[0153] The perception data processing process is analyzed, and the perception task can be described by three terms L i = {s i , d i , τ i}. s i represents the amount of environmental information that needs to be extracted in the perception task, d i represents the number of CPU cycles required to process one bit of environmental information, and τ i represents the maximum tolerable delay for completing the entire perception task.

[0154] When the vehicle chooses to locally process the perception data, the execution delay is as follows:

[0155]

[0156] wherein t represents the time required to locally process the amount of environmental information, and t represents the time required to upload the processed data to the MES for data fusion. Since the size of the processed result is small, the transmission delay of the result delivery is ignored. When allocating time for perception, communication, and calculation, the communication allocation time can be considered as 0, i.e. a i +b i = 1.

[0157] represents the perception time of the amount of environmental information that needs to be extracted, and the expression is as follows:

[0158]

[0159] Define as the highest CPU cycle frequency of the CAV (i.e. CPU cycles per second), and the CPU frequency allocated to the local task is satisfies In theory, the processing delay of the local perception data task is as follows:

[0160]

[0161] wherein, is the CPU frequency used to process perception tasks on the vehicle.

[0162] In summary, the execution delay of local vehicle processing of perception data is as follows:

[0163]

[0164] When the vehicle offloads to the server to process the perception data, the execution delay is as follows:

[0165]

[0166] in represents the perception time of the amount of environmental information that needs to be extracted, Indicates the time required to upload the original data to MES, Indicates the time required for MES to process this amount of environmental information. It can be calculated by the following formula:

[0167]

[0168] definition RSU j The processing delay on the MES is expressed as:

[0169]

[0170] where μ vi express Assigned to v i The proportion of computing resources.

[0171] In summary, the execution delay of unloading a vehicle to the server is as follows:

[0172]

[0173] When processing data on the RSU side, the execution delay is as follows:

[0174]

[0175] in, represents the perception time of the amount of environmental information that needs to be extracted, Indicates the time required for MES to process this amount of environmental information. The processing delay on MES is expressed as:

[0176]

[0177] in Indicates u j The proportion of computing resources allocated to task processing.

[0178] In summary, the execution delay on the RSU side is:

[0179]

[0180] S4: A resource allocation scheme targeting at minimizing the cooperative perception task completion delay, constructing a joint problem of perception task allocation, offloading and computing resource allocation;

[0181] To provide fast, reliable and efficient vehicle networking services, improve safety and user experience, the time required for task completion is crucial. This embodiment formulates the cooperative perception task allocation, offloading and computing resource allocation optimization scheme as a problem of minimizing the cooperative perception task processing delay minimization. The optimization problem can be expressed as:

[0182]

[0183] wherein is the vehicle offloading decision, is the vehicle ISAC device time allocation ratio, is the computing resource ratio allocated by the MES for the vehicle, Ω V represents the CAV selected as a co-perception, Ω U represents the RSU selected as a co-perception. The constraint conditions are described in detail below. C1 represents the selection of co-perception, C2 represents v i and u j The total processing time cannot exceed the maximum tolerable delay, C3 represents the sum of the normalized perception, communication and computing time ratio allocated by v i is 1, C4 represents that the data volume of communication cannot exceed the data volume of perception in the same cycle, C5 represents that the actual required perception, communication and computing time is not greater than the allocated perception, communication and computing time, C6 represents that the sum of the resource ratio allocated by the RSU providing offloading service is not greater than 1.

[0184] S5: To solve the optimization problem, a multi-agent deep deterministic policy gradient algorithm with hybrid action space, HAS-MADDPG, is proposed to learn the resource allocation scheme with minimum task completion delay; Fig. 3 is the network structure diagram of HAS-MADDPG algorithm.

[0185] The optimization problem is converted into a Markov decision process, and the state space is set as:

[0186]

[0187] wherein b, g, and l represent the remaining computing resource state of the MES, the channel gain of the communication link between the CAV and the MES, and the connection state between the CAV and the MES, respectively.

[0188] Action space: with offloading decisions Time proportion allocation The discrete action and the continuous variable of the proportion of computing resource allocation μ jointly as the hybrid action of CAV

[0189] Reward:

[0190]

[0191] where β is a constant, which is to increase the reward when the delay decreases.

[0192] The goal of training the value network is to minimize the prediction error of the state-action value function. Assume the parameters of each CAV are θ = {θ1, θ2,..., θ i ,..., θ N}. For each CAV i , the target value is calculated using the target network. The target action a' j is calculated as: j j (s') represents the target action of the jth agent in the next state s' j , j ∈ {1, 2,..., N}. The target value is:

[0193] y i = r i + γQ' i (s', a i )

[0194] where Q' i is the target value network of CAV i .

[0195] The value network parameters θ i are updated by minimizing the mean square error loss:

[0196]

[0197] where B is a batch drawn from the experience replay area.

[0198] When training the policy network, assume the policy network parameters of each CAV are ω = {ω1, ω2,..., ω i ,..., ω N}. For each CAV i , the policy network is updated using the policy gradient. The policy gradient is:

[0199]

[0200] The policy network parameters ω i are updated using the gradient ascent method.​

[0201] For the target network parameters of the value network and the policy network, soft updates are used:

[0202] θ i ' = τθ i + (1-τ)θ i

[0203] ω i ' = τω i + (1-τ)ω i

[0204] τ represents a soft update parameter.

[0205] The HAS-MADDPG algorithm specifically includes the following steps:

[0206] S51: initialize the actor network and critic network of each CAV;

[0207] S52: in each iteration, initialize a random process for action exploration, and obtain the initial observation of each CAV, i.e., the initial state of the environment;

[0208] S53: in each time slot, each CAV selects an action according to the current policy and state and executes it;

[0209] S54: in each time slot, all CAVs interact with the environment to obtain their respective rewards and jump to the next state, and store the experience data in the experience replay pool;

[0210] S55: for each CAV, randomly sample a small batch of samples from the experience pool;

[0211] S56: for each CAV, calculate the target state value of the critic, calculate the loss function, and minimize the loss to update the critic network, calculate the policy gradient, and update the actor network;

[0212] S57: after completing each CAV, soft update the target network parameters of the value network and the policy network, and then return to S53, otherwise return to S55;

[0213] S58: after completing each time slot, return to S52, otherwise return to S53;

[0214] S59: after completing the iteration, stop, otherwise return to S52.

[0215] Finally, it is to be explained that the above embodiments are only used to illustrate the technical solutions of the present application but not to limit the present application. Although the present application is described in detail with reference to the preferred embodiments, it should be understood by those skilled in the art that the technical solutions of the present application can be modified or equivalently replaced without departing from the purpose and scope of the technical solutions, and all should be covered in the scope of the claims of the present application.

Claims

1. A method for allocating collaborative sensing resources in an ISAC-based vehicle network, characterized by: It includes the following steps: S1. Construct a multi-layer collaborative perception area of ​​interest solution based on relative position in the Internet of Vehicles scenario; S2. Determine the common perceiver based on the relative positions of the current vehicle's perception area of ​​interest, the road test unit, and other connected vehicle vehicles; S3. Construct a time allocation scheme for interawareness integration ISAC based on a time-division dynamic frame structure to obtain the total perceptual mutual information (SMI) of the radar sensing duration of the co-perceiver, the average communication rate of the communication duration, and the execution delay of each processing method; S4. Establish a joint problem of perception task allocation, offloading, and computational resource allocation with the goal of minimizing the delay in completing collaborative perception tasks; S5. Use a multi-agent deep deterministic policy gradient algorithm in a hybrid action space to solve the optimization problem and obtain a resource allocation solution with minimal task completion delay. In step S1, the Internet of Vehicles scenario includes several Internet of Vehicles (CAVs), several Road Test Units (RSUs), and several Mobile Edge Servers (MESs). The CAVs collect data inside and outside the vehicle through various sensors. Several RSUs deployed along the road are connected to each other via high-speed wired links. Each RSU is connected to the MES via a wired link. The current CAV selects a co-sensor based on the relative positions of the RSU and other CAVs in the perception area of ​​interest. The processing of the perception data by the current CAV or the co-sensor includes local processing and offload processing. Assume N = {1,...,i,...,N} is the set of CAVs in the system, and the perception area of ​​interest is P PAC , P RSU Indicates the coverage of RSU. The current vehicle v1 cannot perceive P PAC All areas in P PAC Divided into P1, P f and P b , P1 is the area within the visual field of v1, P f and P b are the occluded areas before and after v1 respectively, then: P PAC =P1+P f +P b The perception tasks of the three regions are defined as T1, T f , T b , the perception task of CAV perception area of ​​interest is expressed as: T PAC =T1+T f +T b The ways in which CAVs process perception tasks include local processing or offload processing.

2. The ISAC-based vehicle network collaborative perception resource allocation method according to claim 1, characterized in that: In step S2, the process of transferring the current vehicle from the coverage of the current RSU to the coverage of the next RSU is divided into the following five stages, and the selection range of the co-perceiver in each stage is: In the first stage, T PAC Fully in RSU j The coverage of which the co-perceivers choose RSU j and in P b , P f CAV; V f Indicates P PAC CAV before v1, V b Indicates P PAC CAV after v1, V V2V is the set of CAVs within the communication range of v1, U j Indicates the RSU range that v1 is currently in j , U j+1 Indicates that v1 is about to enter the RSU j+1 ; In the second stage, T PAC P f Partially enter RSU j+1 The coverage of the co-perceiver selects RSU j 、RSU j+1 and in P b , P f CAV; In the third stage, T PAC P f Access to RSUs j+1 The coverage of P b Not entered, its co-perceiver chooses RSU j 、RSU j+1 and in P f CAV; In the fourth stage, T PAC P b Partially enter RSU j+1 The coverage of which the co-perceivers choose RSU j 、RSU j+1 and in P b , P f CAV; In the fifth stage, T PAC has fully entered into RSU j+1 The coverage of its co-perceiver selects RSUj+1 and b , P f CAV; The task allocation function is defined as y = TA(x), indicating that T PAC The collaborative task in is assigned to y. Based on the relative position information, a group of co-perceivers are selected from RSU and CAV to assist v1 in completing the perception task; Let Ω denote the set of selected co-perceivers such that: Ω=TA(T1)∪TA ρ (T f )∪TA ρ (T b ) Where TA ρ (T1), TA ρ (T f ),TA ρ (T b ) respectively indicate the ability to perceive T0, T f 、T b For the perceiver, ρ=1,2,3,4,5 indicates the stage.

3. The ISAC-based vehicle network collaborative perception resource allocation method according to claim 2, characterized in that: In step S3, an adjustable frame format including S sensing subframes, P computing subframes and Q communication subframes is first constructed, with N subframes in each cycle. s subframes, and the duration of each subframe is τ s , the time allocation decision of each vehicle ISAC device is described as Among them, v i ∈Ω V , a i ,b i ,c i v i The corresponding normalized perception duration, calculation duration and communication duration, a i +b i +c i =1, where During radar detection, v i During the radar sensing duration a i The average perceptual mutual information SMI is expressed as: Among them, B i Represented as v assigned to i bandwidth; v i The signal and interference plus noise ratio is expressed as: Where P i It is v i The transmission power, represents the path propagation gain, n r-rad and n r-com Represents other CAV pairs v i Radar interference and communication interference caused by radar signals; N0 represents the thermal noise power spectral density.

4. The ISAC-based vehicle network collaborative perception resource allocation method according to claim 3, characterized in that: In step S3, during the communication process, v i During the communication duration c i The total communication data volume of the process is: v i During the communication duration c i The average communication rate is: in, Indicates v i The signal-to-noise ratio of the communication signal is expressed as: Where G t and G r Indicates antenna transmission gain and antenna receiving gain, represents the communication channel gain from i to j, and They are the radar interference and communication interference respectively.

5. The ISAC-based vehicle network collaborative perception resource allocation method according to claim 4, characterized in that: In step S3, for each execution delay, the perception task is described as L i ={s i ,d i ,τ i }, s i represents the amount of environmental information that needs to be extracted in the perception task, d i Indicates the number of CPU cycles required to process one bit of environmental information, τ i The maximum tolerable delay for the entire perception task is: The execution delay of the vehicle's local processing of perception data is: Where, f i v Indicates the CPU cycle frequency assigned to the local task, and defines the maximum CPU cycle frequency F of CAV i v , f i v Satisfy f i v ≤F i v ; The execution delay of unloading a vehicle to the server is: Where, RSU j The maximum CPU cycle frequency, express Assigned to v i The proportion of computing resources; The execution delay on the RSU side is: Where, Indicates u j The proportion of computing resources allocated to task processing.

6. The ISAC-based vehicle network collaborative perception resource allocation method according to claim 5, characterized in that: In step S4, the collaborative sensing task allocation, offloading and computing resource allocation optimization scheme is formulated as the problem of minimizing the collaborative sensing task processing delay. The optimization problem is expressed as: in For vehicle unloading decisions, Time allocation ratio for vehicle ISAC equipment, is the ratio of computing resources allocated by MES to vehicles, Ω V represents the CAV selected as the co-perceiver, Ω U represents the RSU selected as the co-perceiver; C1 represents the selection of common perceivers; C2 represents v i and u j The total processing time cannot exceed the maximum tolerable delay; C3 means v i The sum of the allocated normalized perception, communication, and computation time ratios is 1; C4 indicates that the amount of communication data cannot exceed the amount of perception data in the same cycle; C5 indicates that the actual required perception, communication, and computation time is no greater than the allocated perception, communication, and computation time; and C6 indicates that the sum of the resource ratios allocated to the RSU providing offloading services is no greater than 1.

7. The ISAC-based vehicle network collaborative perception resource allocation method according to claim 6, characterized in that: In step S5, the optimization problem is transformed into a Markov decision process, and the state space is set as: Where b, g, and l represent the remaining computing resource status of MES, the channel gain of the communication link between CAV and MES, and the connection status between CAV and MES, respectively; Action Space: Uninstall decision Time Proportion Allocation The discrete action of and the continuous variable of the computing resource allocation ratio μ are used as the hybrid action of CAV. award: Where β is a constant, when the delay decreases, the reward increases; The goal of training the value network is to minimize the prediction error of the state-action value function. Assume that the parameters of each CAV are θ = {θ1,θ2,...,θ i ...,θ N }, for each CAV i , use the target network to calculate the target value; calculate the target action a' j , where a' j =μ' j (s') represents the next state s' of the jth agent j The target action under ,j∈{1,2,...,N}, then the target value is: y i =r i +γQ′ i (s',a i ) where Q′ i For CAV i target value network; Update the value network parameters θ by minimizing the mean squared error loss i : Where B is a batch drawn from the experience replay area; When training the policy network, assume that the policy network parameters of each CAV are ω={ω1,ω2,...ω i ,...ω N }, for each CAV i , use the policy gradient to update the policy network, the policy gradient is: Update the policy network parameters ω using the gradient ascent method i ; For the target network parameters of the value network and the policy network, soft updates are used: θ′ i =τθ′ i +(1-τ)θ i oh i '=to i '+(1-t)ω i Where τ represents the soft update parameter.

8. The ISAC-based vehicle network collaborative perception resource allocation method according to claim 7, characterized in that: In step S5, the specific process of using the multi-agent deep deterministic policy gradient algorithm in the hybrid action space to solve the optimization problem includes: S51: Initialize the actor network and critic network of each CAV; S52: In each iteration, a random process is initialized for action exploration to obtain the initial observation of each CAV, i.e., the initial state of the environment; S53: At each time slot, according to the current strategy and state, each CAV selects an action and executes it; S54: In each time slot, all CAVs interact with the environment to obtain their respective rewards and jump to the next state, storing the experience data in the experience replay pool; S55: For each CAV, randomly sample a small batch of samples from the experience pool; S56: For each CAV, calculate the target state value of the critic, calculate the loss function, and minimize the loss to update the critic network, calculate the policy gradient, and update the actor network; S57: After completing each CAV, soft-update the target network parameters of the value network and the policy network, and then return to S53, otherwise return to S55; S58: Return to S52 after completing each time slot, otherwise return to S53; S59: Stop after the iteration is completed, otherwise return to S52.

Citation Information

Patent Citations

  • Multi-source heterogeneous sensing data fusion method for autonomous vehicle

    CN114912532A

  • Intelligent network connection automobile cooperative perception and data fusion method

    CN117956507A