Multi-agent dynamic access and cooperation method for air-space-ground vehicle networking integrated architecture

Through deep reinforcement learning and multi-agent access and collaboration methods of Transformer architecture, the problems of vehicle perception range limitation and communication delay in the integrated architecture of the aerospace, earth and vehicle network are solved, and intelligent dynamic access and data space-time alignment of multiple vehicles are realized, improving the communication reliability and resource utilization of the system.

CN120343580APending Publication Date: 2025-07-18NANJING UNIV
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510476922.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-16
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

There are problems with vehicle perception range limitations and communication delays in the integrated architecture of the aerospace and earth-to-vehicle network. Traditional access solutions are difficult to adapt to dynamically changing communication environments, resulting in high queuing delays and high interference, and inconsistent vehicle clocks and differences in transmission delays between ground and satellite networks lead to data asynchronousness.

Method used

The multi-agent dynamic access strategy based on deep reinforcement learning is adopted to select the optimal access object in real time, reduce queuing and transmission delays, and solve the spatial and temporal alignment of asynchronous data through a multi-agent collaborative processing framework based on Transformer. Combined with the semi-deterministic channel model to accurately capture channel characteristics, intelligent matching and data fusion of multiple vehicles are realized.

Benefits of technology

It significantly reduces the end-to-end communication delay, improves network resource utilization and system stability, ensures high reliability and low-latency communication in complex scenarios, adapts to network dynamic fluctuations, and solves the timing consistency problem of asynchronous data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120343580A_ABST
    Figure CN120343580A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-agent dynamic access and cooperation method for an air-space-ground Internet of Vehicles integrated architecture, provides a dynamic access decision scheme based on the air-space-ground Internet of Vehicles integrated network architecture, and aims to solve the problems of vehicle sensing range limitation and communication delay. Through multi-agent deep reinforcement learning, channel quality and load conditions of a ground network and a satellite network are evaluated in real time, and an optimal access mode is selected for a vehicle node. Meanwhile, in order to solve the problem of data asynchronization caused by communication delay between vehicles, clock inconsistency and transmission delay difference between a ground network and a satellite network, a Transform-based multi-agent cooperation mechanism is provided on the basis of a dynamic access architecture, time sequence coding is performed on asynchronous delay, sparse features of the vehicles are extracted, and the data asynchronization problem is solved. Effective processing of asynchronous data and accurate reconstruction of a global environment are realized at a core network end. According to the invention, more flexible and efficient dynamic access management and cooperation are realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of vehicle networking for autonomous driving, and relates to the theory of cooperative perception and the design of access management solutions. Specifically, it relates to a multi-agent dynamic access and cooperation method for an integrated space-air-ground-vehicle networking architecture. Background Art

[0002] As a key technical framework for the new generation of intelligent transportation systems, the integrated space-air-ground-vehicle networking architecture provides unprecedented communication and perception capabilities for autonomous driving. This architecture integrates satellite communication and ground infrastructure to build an all-round coverage and highly reliable communication network environment. In areas with uneven ground link coverage or in dense and complex road environments, the introduction of this architecture can significantly improve the safety and reliability of the system, and ensure the smooth communication between vehicles and between vehicles and the core network. At the same time, satellites can provide a larger communication range, which has irreplaceable advantages in the global environmental perception reconstruction of autonomous driving systems.

[0003] However, in actual deployment, it still faces two core challenges: the limitation of vehicle perception range and communication delay problems. Although the integrated space-air-ground-vehicle networking architecture expands communication coverage and compensates for ground communication, it also brings relatively high air interface delay, which will affect the real-time performance of global vehicle data processing. Traditional vehicle access schemes may face high queuing delays or high interference. In addition, the perception of a single vehicle is still restricted by physical conditions, especially in complex road environments, especially in complex scenarios outside the central vision such as intersections, curves, and highway merge points. Obstacles such as buildings and large vehicles will cause serious perception blind spots. Therefore, a dynamic multi-vehicle access management decision and cooperation scheme is crucial for the autonomous driving system under the integrated space-air-ground-vehicle networking architecture.

[0004] Generally, the access nodes of autonomous driving vehicles to ground facilities are determined by distance. However, in the complex scenarios of space-air-ground integration, simple distance decisions may face high queuing delays or high interference. Therefore, dynamic communication link monitoring and access management decisions are crucial for vehicle networking systems. It ensures communication reliability while minimizing the queuing delay and transmission delay of the system. Generally, traditional access optimization constructs a model based on a task priority to dynamically assign weights to vehicles to be accessed. However, such methods based on fixed thresholds and preset rules often lead to very high processing costs when facing complex scenarios and are difficult to complete in real time in reality. On the other hand, although deep learning can make decisions quickly, it is difficult to adapt to dynamic user numbers. The access management problem in the integrated space-air-ground-vehicle networking architecture requires an effective method that can be completed in real time and adopt dynamic users. Summary of the Invention

[0005] Objective of the Invention: The objective of the present invention is to provide a multi-agent dynamic access and cooperation method for an integrated architecture of space-air-ground-vehicle Internet of Things. On the one hand, the present invention solves the problem that the traditional fixed-threshold access method is difficult to adapt to the dynamically changing communication environment in a heterogeneous network environment through a dynamic access strategy based on deep reinforcement learning, and realizes the intelligent matching of vehicle nodes with the optimal access objects. On the other hand, aiming at the data asynchrony problem caused by inconsistent vehicle clocks and the transmission delay difference between the ground and satellite networks in the integrated space-air-ground architecture, the present invention adopts a multi-agent cooperation processing framework based on Transformer to effectively address the spatio-temporal alignment problem of asynchronous data.

[0006] Technical Solution: To achieve the above objective of the invention, the present invention adopts the following technical solutions:

[0007] In the first aspect, the present invention proposes a multi-agent dynamic access method for an integrated architecture of space-air-ground-vehicle Internet of Things, including the following steps:

[0008] After collecting the vehicle multi-modal perception data set, perform accessible node matching according to the data and trajectory uploaded by the vehicle, in combination with the distribution of ground base stations and satellites.

[0009] Adopt an access algorithm based on reinforcement learning to select access objects for multiple vehicle nodes in real time, reducing the queuing delay and transmission delay. The access algorithm based on reinforcement learning aims to minimize the overall queuing delay and transmission delay of the vehicle network. The state space includes the channel quality and node load of ground and satellite connectable nodes, as well as the size of the data to be pre-transmitted by the vehicle. The action space is the selection of access nodes. In the algorithm design, a two-level decision-making structure based on delay estimation is adopted. First, select the link type with the minimum expected total delay, and then select the relatively optimal node among the same type of links in combination with the calculated expected delay and potential handover cost.

[0010] Furthermore, the accessible node matching process includes collecting multi-modal data mainly based on (Light Detection And Ranging, LiDAR) of in-vehicle sensors to calculate the real-time position and motion state of the vehicle, and establishing a candidate access node set in combination with the distribution density of ground base stations and the coverage range of satellite networks.

[0011] Furthermore, in the process of constructing the channel, refer to the semi-deterministic channel modeling method to be able to establish the time-frequency dynamic fading characteristics in a large-scale vehicle Internet of Things scenario. The channel quality matrix in the state space is composed of the complex channel coefficients and delays calculated from the semi-deterministic channel model. In addition, a Boolean variable is introduced in the state space to represent the coverage of ground base stations.

[0012] Furthermore, the handover cost C of the link type typeThe switching cost C of nodes within the same type node are respectively:

[0013]

[0014] L t and L s are respectively the average loads of the ground and satellite networks, is the load of the selected node, γ and β are adjustment coefficients, and are respectively the basic costs of link type switching and node switching within the same type.

[0015] Furthermore, the reinforcement learning algorithm also includes a dynamic equilibrium mechanism, which realizes load dispersion on the premise of meeting the system delay requirements by real-time monitoring the resource utilization rates of each access node, and avoids frequent switching of access objects by vehicle nodes by setting reasonable switching costs. The reward function is designed as: where s k , a k are respectively the state space and action space of the kth vehicle node, are respectively the total delay and the switching cost, represents the load balance factor, which is quantified by evaluating the deviation between the current load of the selected node and the system average load, and ω1, ω2 and ω3 are weight coefficients.

[0016] On the second aspect, the present invention also proposes a multi-agent dynamic access and cooperation method for the integrated architecture of space-air-ground-vehicle network, including selecting the optimal access node for multiple vehicles according to the foregoing method;

[0017] Performing sparse feature extraction and time series encoding on multiple vehicles through a Transformer-based architecture to correct the asynchronous delay caused by uneven clock sampling of vehicles and different ground-satellite transmission links; wherein the asynchronous delay processing process matches the historical data of the regions of interest of different vehicles, and enhances the feature representation ability through a multi-layer perceptron to adapt to the Transformer architecture and ensure the time series consistency of multi-source data.

[0018] Furthermore, the Transformer-based architecture first performs time encoding on the original perception data, then uses a target detection algorithm to generate regions of interest and extract sparse feature representations, realizes cross-time-domain feature matching through a multi-head attention mechanism, and finally completes the bird's-eye view perspective conversion and multi-source feature fusion based on the matching results.

[0019] Furthermore, the Transformer architecture also includes a feature fusion optimization strategy, which adopts a multi-scale maximum fusion method and combines a reliability evaluation mechanism to dynamically allocate the weights of different data sources to improve the accuracy of the fusion result.

[0020] In a third aspect, the present invention further provides a computer system, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the computer program is loaded onto the processor, the steps of the foregoing various methods are implemented.

[0021] In a fourth aspect, the present invention further provides a computer program product, including a computer program. When the computer program is executed by a processor, the steps of the foregoing various methods are implemented.

[0022] Beneficial effects: Compared with traditional vehicle cooperation and access management methods, the advantages and positive effects of the present invention are as follows: The dynamic access decision method based on reinforcement learning realizes the intelligent dynamic access of multiple vehicles in a heterogeneous network environment by constructing an optimization goal of minimizing the overall system delay, significantly reducing the end-to-end communication delay; the dynamically balanced mechanism further introduced in the present invention effectively solves the system overhead problems caused by uneven access resource allocation and frequent handovers, significantly improving the network resource utilization rate; and the integrated asynchronous data processing architecture serves as a support framework to ensure the effective fusion of data from different transmission links. This comprehensive solution enhances the system's adaptability to dynamic fluctuations in network load and can maintain stable performance under different communication conditions and road environments, providing a feasible technical path for the large-scale deployment of connected autonomous driving. Description of the Drawings

[0023] Figure 1 It is a multi-vehicle communication scenario diagram under the integrated space-air-ground-vehicle network architecture described in the embodiments of the present invention.

[0024] Figure 2 It is a semi-deterministic channel model diagram adopted in the embodiments of the present invention.

[0025] Figure 3 It is a flow chart of asynchronous delay processing based on Transformer adopted in the embodiments of the present invention.

[0026] Figure 4 It is a flow chart of the feature fusion network in the embodiments of the present invention. Detailed Embodiments

[0027] In order to make the objectives, technical solutions, and advantages of the present invention clearer and more understandable, the following provides a detailed description of the embodiments of the present invention with reference to the accompanying drawings: These embodiments are implemented on the premise of the technical solutions of the present invention, and detailed implementation manners and specific operation processes are given. It should be understood that the specific examples described herein are only used to explain the present invention, but the protection scope of the present invention is not limited to the following embodiments.

[0028] This embodiment discloses a multi-agent dynamic access method for an integrated architecture of space-air-ground-vehicle Internet of Things, which mainly includes: after collecting the vehicle multi-modal perception data set, matching the accessible nodes according to the vehicle uploaded data and trajectory, and combining the distribution of ground base stations and satellites; adopting an access algorithm based on reinforcement learning to select the access object for multi-vehicle nodes in real time. The access algorithm based on reinforcement learning takes the minimization of the overall queuing delay and transmission delay of the vehicle network as the optimization goal. The state space includes the channel quality and node load of ground and satellite connectable nodes, as well as the size of the data volume to be pre-transmitted by the vehicle. The action space is the access node selection. In the algorithm design, a two-level decision-making structure based on delay estimation is adopted. First, the link type with the minimum expected total delay is selected, and then, combining the calculated expected delay and potential handover cost, the relatively optimal node is selected among the links of the same type.

[0029] Furthermore, on this basis, a multi-source asynchronous data collaborative reconstruction framework based on Transformer is also disclosed, constituting a multi-agent dynamic access and collaboration method for an integrated architecture of space-air-ground-vehicle Internet of Things. The multi-source asynchronous data collaborative reconstruction framework based on Transformer extracts sparse features and performs time series encoding on multiple vehicles through an architecture based on Transformer, and corrects the asynchronous delay caused by uneven clock sampling of vehicles and different transmission links between ground and satellites. The asynchronous delay processing process matches the historical data of the regions of interest of different vehicles, and enhances the feature representation ability through a multi-layer perceptron to adapt to the Transformer architecture and ensure the time series consistency of multi-source data.

[0030] Specifically, Figure 1Shows a multi-vehicle communication scenario under the integrated architecture of space-air-ground-vehicle network. First, the data collector collects the multi-modal data set of in-vehicle sensors through the vehicle network. This data set mainly consists of point cloud data (PCD) and planar images. After collecting the time-series data, the data is organized according to the distribution of road ground base stations and the coverage of low-Earth orbit satellites. Using geometric constraints, the vehicles with known motion states are matched to the road segments to obtain an original candidate set of access nodes, and the channel conditions between the vehicles and the nodes and the node loads are detected. Then, with the help of the access strategy, multi-agent access management based on deep reinforcement learning is realized. To minimize the overall system latency, this embodiment further introduces a dynamic equilibrium mechanism to monitor the node resource utilization rate in real time and sets a reasonable negative reward function to avoid frequent switching of access objects by vehicle nodes. During the reconstruction process at the core network side, for the asynchronous latency caused by irregular vehicle clocks and different data flows, referring to the asynchronous time-series perception architecture, based on the multi-perceptron algorithm and the object detection algorithm, an area of interest (ROI) is established for each vehicle and time-series ROI matching and sparse feature extraction are realized. The Transformer algorithm is used to generate a BEV flow with spatio-temporal correlation. A large number of simulation experiments on real data sets show that the architecture scheme proposed in this embodiment is feasible for global environment reconstruction and collaborative path planning in vehicle networks and can meet the requirements of low latency and high reliability for autonomous driving in complex scenarios of high-speed movement.

[0031] Channel model:

[0032] In the case of high vehicle mobility and complex propagation environments, the vehicle network channels often have significant specificities. Especially when satellites are introduced to participate in communication, the channels will change rapidly and produce high Doppler frequency shifts and multipath propagation effects. It is difficult for the usual statistical channel models to accurately represent these dynamic changes in real time. To accurately capture the multipath-Doppler propagation characteristics in complex environments and balance the scenario accuracy and scalability of large-scale vehicle network communication simulations, a semi-deterministic channel model is introduced. Figure 2The construction process of the semi-deterministic channel model is shown, and this model can be further analyzed from the perspectives of large-scale fading (LSF) and small-scale fading (SSF). First, LSF describes the average attenuation of the signal over a relatively long distance. The large-scale parameters here include the Rice K factor, delay spread, shadow fading, and angular domain parameters. Taking the Rice K factor KF as an example, it represents the ratio of the direct-path power to the scattered-path power, directly reflecting the relative intensities of the line-of-sight (LOS) and non-line-of-sight (NLOS) components in the channel. Due to the influence of environmental factors on the direct-path intensity, this ratio fluctuates with the positions of the transmitter and receiver. Therefore, in a dynamic time-varying channel, the smooth variation of large-scale parameters, that is, the spatial consistency of the parameters, is a key issue in LSF modeling. Specifically, the spatially correlated Rice K factor is calculated as follows:

[0033] KF(p) = KF μ + KF σ · X KF (p)

[0034] where X KF (p) is a normally distributed random variable related to the position vector p, and KF μ and KF σ represent the reference expected value and reference standard deviation of the Rice K factor respectively. These reference values are affected by the dependent variables in the scenario, including the carrier frequency f c , the three-dimensional distance d 3D between the transmitter and receiver, the elevation angle α of the transmitter and receiver, the rainfall intensity R, and the atmospheric absorption coefficient γ o . Therefore, the calculation of KF μ can be expressed as follows,

[0035] KF μ = KF0 + KF f log(10 · f c ) + KF d log(10 · d 3D ) + KF α log(10 · α) - rR - KF L γ o

[0036] where KF f , KF d , KF α , r, KF LThey are the influence degrees of frequency, distance, elevation angle, rainfall, and atmosphere on the Rice K factor, which are defined by the 3GPP protocol in different scenarios. Among them, rainfall and atmosphere are more targeted at the air-to-ground channel conditions, reflecting the fading change process of signals over long distances. The calculation of large-scale parameters such as shadow fading and delay spread is the same. This process ensures the authenticity and spatial consistency of the LSF in the dynamic vehicle networking environment, and jointly calculates the final channel coefficient with the SSF model after power scaling.

[0037] The SSF model describes the channel variations at the wavelength level, including fast time-varying fading and multi-path - Doppler effects. As Figure 2 shown, large obstacles such as buildings and hills are regarded as clusters, representing groups of scatterers that affect the signal path, while scatterers are individual objects that cause path reflection and diffraction. Assume there are multiple paths i, each represented by a set of variables [τ i , P i , ψ i , where τ i represents the delay, P i represents the power, and ψ i represents the phase of each path. The model updates the signal propagation path in real time according to the movement of the transmitter and receiver and the change of the cluster position. By calculating the total path length vector d i and the above variables, the delay and coefficient of signal propagation are determined, and the change relationship of the channel over time slots is also captured. The initial power is affected by the radiation and polarization of the antenna and the large-scale parameters as described above, and each scatterer in the NLOS channel introduces a set of additional phase and power information. Therefore, the normalized power of each path can be represented by a composite gain as follows:

[0038]

[0039] where g r is the composite expression of the transmit antenna power and the normalized large-scale parameter. It should be clear that the index i represents the i-th end-to-end complete propagation path in the NLOS channel, and n represents the total number of sub-paths of each end-to-end complete propagation path. For the non-line-of-sight (NLOS) scenario, each complete propagation path i can be further decomposed into a cascade of multiple physical sub-path segments, each of which is composed of the electromagnetic wave propagation between adjacent scatterers. Therefore, P i and ψ i represent the combined power and phase of the i-th complete propagation path respectively, and these parameters comprehensively consider the cumulative effects of all sub-path segments of this path, including the power attenuation and phase change occurring at each scattering point. At the same time, the phase information ψ ireflects the Doppler shift of the path - the more intense the relative movement between the transceiver ends, the greater the phase change and the higher the Doppler shift. This change value is affected by the path length d at the wavelength level i , the elevation angle θ of the sub - path i and the azimuth angle under the combined effect:

[0040]

[0041] In summary, the complex channel coefficient g of each sub - path i can be expressed as the combination of the power composite gain and the path propagation loss L that is not affected by spatial variations P :

[0042]

[0043] In addition, the delay τ of each sub - path under multipath effects i is also calculated based on the geometric model:

[0044]

[0045] where n sca represents the number of scatterers that the path passes through, c represents the speed of light, represents the path length of each reflection segment, and it can be intuitively seen that a shorter path length results in a shorter delay. The complex channel coefficient g i and the delay τ i encapsulate the unique channel characteristics of each sub - path, especially the delay and phase changes caused by the complex propagation environment and high - speed relative movement - that is, the multipath - Doppler effect. This model effectively enhances the authenticity and accuracy of the air - ground - space - vehicle network channel simulation.

[0046] Dynamic access strategy:

[0047] In the integrated architecture of the air - ground - space - vehicle network, traditional fixed - threshold access methods or access strategies based on task priorities are difficult to adapt to the dynamically changing communication environment. This embodiment proposes a multi - agent dynamic access decision - making framework based on deep reinforcement learning, which can select the optimal access method for vehicle nodes in real - time and significantly reduce the overall system delay. The dynamic access problem is modeled as a Markov decision process (MDP), represented by the five - tuple (S, A, P, R, γ), where S is the state space, A is the action space, P is the state transition probability, R is the reward function, and γ is the discount factor. For the k - th vehicle node, its state space can be represented as follows:

[0048]

[0049] where and respectively represent the channel quality matrices of ground and satellite connectable nodes, which are composed of complex channel coefficients g finally calculated in the semi-deterministic channel model i and delay τ i and can be expressed as where M represents the number of connectable ground base stations; it should be noted that N p represents the number of main propagation paths considered. To ensure the consistency of the state space, for all vehicle nodes, we select a fixed number of main propagation paths based on power sorting. Matrix element H k,m is a binary tuple H k,m ={G k,m , T k,m}, where G k,m and T k,m are respectively the complex channel coefficient vector containing all effective propagation paths and the corresponding delay vector:

[0050]

[0051] and respectively represent the load conditions of ground and satellite candidate nodes; to cope with the special situation where ground base stations have insufficient coverage in remote areas and only satellites can be connected, a Boolean variable c k is introduced to represent the coverage of ground base stations; q k represents the current state of the vehicle, including the size of the data to be pre-transmitted. In a more refined modeling process, a priority sorting can also be introduced for each vehicle, and the priority and data volume can be jointly modeled into q k through a series of constraints on data transmission, so as to more accurately optimize the system queuing delay. The action space of vehicle nodes adopts a two-level decision-making structure based on delay estimation. The first layer distinguishes satellite nodes and ground nodes and selects the link type based on end-to-end delay estimation A very intuitive understanding is that the premise for data to flow to the satellite is definitely because the ground load is too high or there is no stable ground coverage, otherwise it is impossible to balance the air interface delay of the satellite. The solution of this embodiment is to make a preliminary screening of the queuing delay and transmission delay according to the vehicle and node conditions. Calculate the expected total delay for each ground candidate node m as follows:

[0052]

[0053] where and respectively represent the queuing delay and transmission delay of each ground candidate node. Similarly, calculate the expected total delay for each satellite candidate node n as follows:

[0054]

[0055] Among them, and respectively represent the queuing delay and transmission delay of each satellite candidate node. Select the link type with the minimum expected total delay Among them, the queuing delay is determined by the current node load and vehicle task status, while the transmission delay mainly considers the data volume and channel conditions. Taking the calculation of the ground link as an example:

[0056]

[0057] where B is the available bandwidth of the node, is the signal-to-noise ratio under the current channel condition, and α is the weight coefficient of the influence of the load on the transmission delay. Such a calculation comprehensively considers the influence of the vehicle task volume and node load on the queuing delay and transmission delay. Since this embodiment focuses on a relatively long-term dynamic change process, after determining the vehicle data flow direction, it is necessary to select the relatively optimal node among the same type of links by combining the calculated expected delay and potential handover cost:

[0058]

[0059] where λ is the weight coefficient of the handover cost. To further improve the system performance, this embodiment designs a dynamic handover cost mechanism, considering different types of handover costs:

[0060]

[0061] The handover cost also changes dynamically, depending on the load status of the current system, to prevent frequent handovers and resource preemption:

[0062]

[0063] where, L t and L s are the average loads of the ground and satellite networks respectively, is the load of the selected node, γ and β are adjustment coefficients, represents the basic cost of link type handover, that is, the minimum system overhead required to switch from the ground network to the satellite network or vice versa; Represents the basic cost of in - type node switching, that is, the minimum system overhead for switching from one node to another in the same - type network. Through this decision - making method based on end - to - end delay estimation, the system can comprehensively consider queuing delay and transmission delay at the first - level decision, achieving more accurate link selection. After determining whether to use the terrestrial network or the satellite network, the system enters the second - level decision, that is, to select the optimal access node in the same - type network. At this time, not only the delay performance of the node is considered, but also the handover cost is comprehensively evaluated to avoid the system overhead caused by frequent handovers. This hierarchical decision - making mechanism avoids the defects of simple threshold decision - making and can adaptively select the optimal link according to the actual network conditions and channel conditions. For example, when the load of the terrestrial link is high, resulting in an increase in queuing delay but still lower than the high transmission delay of the satellite link, the system will continue to select the terrestrial link in the first - level decision and then find a node with relatively lower load in the terrestrial network in the second - level decision; when the load of the terrestrial link is too high, resulting in the total delay exceeding that of the satellite link, the system will automatically switch to the satellite link in the first - level decision and then select the best satellite node in the second - level decision to achieve global optimality. To achieve the goal of minimizing the overall system delay, the reward function is designed as the negative value of the weighted sum of queuing delay, transmission delay, and handover cost, and the importance of each index is balanced by appropriately selecting the weight coefficients.

[0064]

[0065] Represents the load - balancing factor, which is used to promote the overall load balancing of the system; ω1, ω2, and ω3 are weight coefficients. Specifically, the load - balancing factor quantifies the impact of the access decision on the system load distribution by evaluating the deviation between the current load of the selected node a k and the average load of the system:

[0066] B k,ak = ξ·(L ak -L avg )

[0067] where L ak is the current load of the selected node, L avg is the average load level of the currently accessible nodes, and ξ is an adjustable coefficient. This design can effectively prevent the phenomenon that multiple vehicles switch to nodes with lower load simultaneously, resulting in instantaneous congestion of the node. By introducing this balance factor into the reward function, the system can, while meeting the delay requirements, achieve an even distribution of vehicles among multiple access nodes, improving the overall network resource utilization efficiency. This embodiment adopts a centralized training and distributed execution framework to implement multi - vehicle collaborative access decision - making, uses deep reinforcement learning to adaptively adjust the access strategy, and the optimization goal is to maximize the expected cumulative reward. In summary, the system state update formula is:

[0068] St+1 = {s 1,t+1 , s 2,t+1 , …, s n,t+1} = f(S t , A t , Z t )

[0069] Where S t represents the system state at time t, A t = {a 1,t , a 2,t , …, a n,t} represents the set of actions of all vehicles, Z t represents the environmental random factors, and f represents the state transition function. In this embodiment, an end-to-end deep reinforcement learning framework is adopted for training and implementation. The model is trained based on the Multi-Agent Proximal Policy Optimization (MAPPO) algorithm. In this algorithm, each agent adopts an Actor-Critic network architecture. The shared value network (Critic) is responsible for state evaluation and knowledge sharing, and the separate policy network (Actor) is responsible for generating actions for each vehicle, effectively balancing the collaboration and independent decision-making capabilities among multiple vehicle nodes. In the network structure design, the value network uses a multi-layer perceptron to process multi-source heterogeneous information in the state space (including the channel quality matrix, node load conditions, etc.). The complex matrix-like state input is extracted by a convolutional layer and then input into a fully connected layer to effectively process high-dimensional features and output the overall state value estimation; the policy network is designed for the independent decision-making needs of each vehicle and also generates an action probability distribution after receiving the state input. Although the decision-making process is logically divided into two stages: link type selection and node selection, in the model construction, it is integrated into a unified neural network architecture to ensure the integrity and collaborative optimization of decision-making.

[0070] The design method of this embodiment can achieve more accurate network selection and resource allocation, effectively balance the loads of the terrestrial and satellite networks, reduce the overall system delay, and improve communication reliability and user experience. This method is especially suitable for complex scenarios with dynamically changing network conditions and diverse service requirements, providing an efficient and intelligent access management solution for the integrated air-ground-space vehicle networking autonomous driving system.

[0071] Asynchronous delay processing:

[0072] In the integrated architecture of the air-ground-space vehicle network, due to the transmission delay differences between the terrestrial network and the satellite network, and the clock asynchronization among multiple vehicles, serious asynchronous problems exist in the data received by the core network. This asynchrony can lead to information mismatch in cooperative perception, seriously affecting the basis of overall cooperation. To address this challenge, as Figure 3As shown in the figure, this embodiment proposes an asynchronous delay processing method based on Transformer, referring to the design concept of the CoBEVFlow architecture, which effectively solves the data asynchrony problem caused by uneven clock sampling of vehicles and different transmission links.

[0073] The core idea of this embodiment is to align asynchronous cooperation messages sent by multiple agents by compensating for and estimating the motion trajectory. Specifically, based on the BEV flow, asynchronous perception features can be redistributed to appropriate positions, thereby reducing the impact of asynchrony. This processing method has two main advantages: First, it can process asynchronous cooperation messages sent at irregular and continuous timestamps without discretization; Second, through the BEV flow, the system only needs to transmit the original perception features instead of generating new perception features, thus avoiding introducing additional noise.

[0074] For the original perception data collected by vehicle k at time t To reduce the data processing burden, this embodiment first applies an object detection algorithm to generate regions of interest (ROIs) and extract sparse features to represent potential moving objects in the scene:

[0075]

[0076] where Φ roi_gen (·) is an ROI generation network with a detection decoder structure based on a convolutional neural network, and each element represents a detected ROI, which contains class confidence, position, size, and orientation information respectively. By thresholding the class confidence and applying non-maximum suppression, the spatial boundaries of the ROI set are obtained. Based on this ROI set, a binary mask H∈R {H×W} is generated, where H and W represent the height and width of the feature map respectively, and are consistent with the original perception data . In this binary mask, the values inside the ROI are 1 and the others are 0. Then a sparse feature map is obtained:

[0077]

[0078] where the symbol ⊙ represents the Hadamard product. This operation sets all feature values outside the ROI region in the original feature map to 0 and keeps the feature values inside the ROI region unchanged, thereby generating a sparse feature map. In this way, while retaining the features inside the ROI, they can be sent out as cooperation messages together with the ROI set. Then, in order to effectively process asynchronous data with irregular time intervals, a time encoding mechanism is introduced:

[0079]

[0080] where TE(t) is the time encoding function, is a feature matrix, N is the number of detected ROIs, and D is the feature dimension. Each row corresponds to the feature representation of an ROI, containing the semantic and spatial information of that ROI. This representation converts the original dense feature map into a more compact feature set, significantly reducing the computational burden of subsequent processing. This encoding method can effectively represent the time order of data and maintain the relative relationship between different timestamps, facilitating the subsequent processing of asynchronous data.

[0081] After obtaining the sparse feature map and the information with time series encoded features, in order to obtain the ROI information within the target timestamp, it is necessary to generate the motion trajectory by tracking the positions of each ROI at different time points. For the r-th ROI detected by the q-th agent, extract its position and orientation attributes in the historical frames:

[0082]

[0083] where is the set of 2D center coordinates and azimuth. As Figure 4 shown, based on these irregularly sampled trajectory data, the multi-head attention mechanism (MHA) is used to predict the position and orientation of each ROI at the target timestamp :

[0084]

[0085] where is the time encoding of the target timestamp, MLP(·) is the encoding function of the irregular historical sequence, and U p is the set of time encodings corresponding to each time point in the historical feature sequence V q,r . Specifically, the query of MHA is the time encoding of the target timestamp, and both the key and the value are the sum of the historical feature sequence and its corresponding time encoding:

[0086]

[0087] K = MLP(V q,r ) + U p

[0088] V = MLP(V q,r ) + U p

[0089] The formula for multi-head attention is:

[0090] MultiHead(Q, K, V) = Concat(head1, head2, …, head h )W O

[0091] head j = Attention(QW j Q , KW j K , VW j V )

[0092]

[0093] where is the time encoding of the target timestamp, which is a high-dimensional vector representation generated by the trigonometric function time encoding mechanism TE(t) described above. h is the number of attention heads, and d i is the dimension of the key vector. W j Q , W j k , W j V are learnable projection matrices that map queries, keys, and values to appropriate dimensional spaces for parallel computing respectively; W O is the output projection matrix used to map the result after concatenating multiple heads back to the original feature space. The dot product scaling factor is used to normalize the attention scores and prevent the problem of gradient vanishing. By calculating the similarity between the query and the key and performing a weighted sum of the values, multi-head attention can capture the correlations between features at different time points and effectively handle the temporal alignment problem of asynchronous data. After obtaining the predicted positions and directions of each ROI at the target timestamp, combined with the previously calculated ROI space set, in this embodiment, the motion vectors of each grid cell are calculated through affine transformation, and the features in the asynchronous feature map are moved to the estimated positions to achieve spatio-temporal alignment and fusion of features, thereby constructing a complete BEV flow graph:

[0094]

[0095] where t j represents the j-th timestamp of the k-th vehicle node, represents the i-th timestamp of the n-th vehicle node (usually the ego vehicle), and this formula represents the BEV flow graph mapping matrix from the k-th vehicle node at timestamp t j to the n-th vehicle node at timestamp . This flow graph is used to move the features in the asynchronous feature map to the estimated positions to achieve spatio-temporal alignment of features:

[0096]

[0097] where represents the k-th vehicle node at timestamp t jThe sparse feature map collected at represents the feature map of the k-th vehicle node at timestamp t j mapped to the timestamp to obtain the estimated feature map. This process relocates the features in each grid cell according to the predicted motion vector, thus solving the problem of feature misalignment caused by time delay. For feature fusion and weight assignment, this embodiment refers to the multi-agent collaborative perception framework proposed by Wei et al. [Asynchrony-Robust Collaborative Perception via Bird's Eye View Flow, Advances in Neural Information Processing Systems], adopts the multi-scale maximum fusion method, and combines the reliability evaluation mechanism to dynamically assign the weights of different data sources to improve the accuracy of the fusion result.

[0098] This feature alignment and fusion method based on BEV flow has two main advantages compared with traditional methods: one is that it can process asynchronous data in the continuous time domain without the need for discretization; the other is that it avoids the noise that may be introduced by generating new features through motion-guided feature relocation, improving the perception accuracy and robustness.

[0099] This embodiment forms a complementary mechanism between the above-mentioned asynchronous time delay processing method based on Transformer and the previous dynamic access strategy: the dynamic access strategy minimizes the overall communication time delay of the system by intelligently selecting the optimal communication link; while the asynchronous time delay processing mechanism solves the problem of inconsistent time axes between vehicles that still exists even under the optimal access selection. This two-layer optimization design enables the system to establish a reliable spatio-temporal alignment relationship between different agents, provides a basis for global environmental perception reconstruction, and completely solves the time delay and asynchronous challenges in the space-air-ground integrated architecture. Compared with traditional methods, this embodiment can not only process asynchronous data in the continuous time domain, but also avoid introducing additional noise through motion-guided feature relocation, significantly improving the adaptability and perception accuracy of the system in complex dynamic scenarios.

[0100] Simulation experiment settings:

[0101] To verify the effectiveness of the method in this embodiment, a set of simulation experiment environments are designed based on the DAIR-V2X dataset. This dataset is a real-world vehicle-road collaborative perception dataset, which contains rich collaboration scenarios between vehicles and roadside units. All frames in the dataset are from actual traffic scenarios, with a sampling frequency of 10Hz and complete 3D annotation information. The perception range of the experimental area is set as x ∈ [-100.8m, +100.8m], y ∈ [-40m, +40m], the voxel size h = w = 0.4m, and the feature map size is H = 200, W = 504.

[0102] Based on this, an integrated space-air-ground network architecture is constructed, including ground base stations distributed in the experimental area and low-orbit satellites covering this area. The ground network uses a carrier frequency of 5.9GHz, a bandwidth of 10MHz, and a transmission delay range of 10 - 30 milliseconds; the satellite network uses Ka-band communication, a bandwidth of 5MHz, and a transmission delay range of 50 - 300 milliseconds to simulate the network characteristics under different communication conditions. Background traffic and random interference are introduced into the network environment to truly reflect the load changes and instability of the communication system.

[0103] The dynamic access decision framework is implemented based on the MAPPO algorithm, adopting a centralized training and distributed execution paradigm. This method is based on the Actor-Critic network architecture, where the Actor network generates the access decision strategy of the vehicle, and the Critic network evaluates the global state value and guides policy optimization. The two-level decision-making mechanism of link type-node selection in this system is only an abstract concept for problem modeling. In the actual implementation of reinforcement learning, a unified policy optimization method is adopted to perform end-to-end training on the entire decision space. The feature extraction network uses a PointPillars encoder compatible with the DAIR-V2X dataset to generate a BEV feature map in standard format. In terms of asynchronous data processing, we implement a 6-layer Transformer encoder, configure 8 attention heads, and the hidden layer dimension is 256. The BEV flow generation network realizes temporal prediction based on the multi-head attention mechanism, supports a maximum of 10-frame historical data, and can effectively process asynchronous information under different delay conditions.

[0104] Through this experimental environment, this embodiment can evaluate the performance characteristics of the proposed method in the integrated space-air-ground vehicle networking architecture, especially the feasibility in challenging scenarios such as ultra-dense network environments and remote low-coverage areas. The theoretical model shows that the integrated architecture proposed in this embodiment based on dynamic access strategies and asynchronous delay processing has potential advantages in terms of end-to-end communication delay and global environment reconstruction. The modular design characteristics of this architecture make it applicable to various network conditions and traffic scenarios, providing a feasible technical path for the large-scale deployment of the integrated space-air-ground vehicle network.

[0105] An embodiment of the present invention also discloses a computer system, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the computer program is loaded onto the processor, the steps of the foregoing various methods are implemented.

[0106] An embodiment of the present invention also discloses a computer program product, including a computer program. When the computer program is executed by a processor, the steps of the foregoing various methods are implemented.

[0107] The foregoing is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A multi-agent dynamic access method for an integrated architecture of space-air-ground-vehicle Internet of Things, characterized in that It includes the following steps: After collecting the multi-modal perception data set of the vehicle, according to the data and trajectory uploaded by the vehicle, combined with the distribution of ground base stations and satellites, perform accessible node matching; Adopt an access algorithm based on reinforcement learning to select access objects for multi-vehicle nodes in real time. The algorithm aims to minimize the overall queuing delay and transmission delay of the vehicle network. The state space includes the channel quality and node load of ground and satellite connectable nodes, as well as the size of the data volume to be pre-transmitted by the vehicle. The action space is the selection of access nodes. The algorithm design adopts a two-level decision-making structure based on delay estimation. First, select the link type with the minimum expected total delay, and then combine the calculated expected delay and potential handover cost to select the relatively optimal node among links of the same type.

2. The multi-agent dynamic access method for an integrated architecture of space-air-ground-vehicle Internet of Things according to claim 1, wherein: The accessible node matching process includes collecting multi-modal data mainly based on LiDAR from in-vehicle sensors to calculate the real-time position and motion state of the vehicle, and establishing a candidate access node set in combination with the distribution density of ground base stations and the coverage range of the satellite network.

3. The multi-agent dynamic access method for an integrated architecture of space-air-ground-vehicle Internet of Things according to claim 1, wherein: Adopt a semi-deterministic channel model. The channel quality matrix in the state space is composed of the complex channel coefficients and delays calculated in the semi-deterministic channel model; a Boolean variable is also introduced in the state space to represent the coverage of the ground base station.

4. The multi-agent dynamic access and cooperation method for an integrated architecture of space-air-ground-vehicle Internet of Things according to claim 1, characterized in that: The switching cost C of the link type type and the switching cost C of the nodes within the same type node are respectively as follows: L t and L s are the average loads of the terrestrial and satellite networks respectively, L ak is the load of the selected node, and γ and β are adjustment coefficients, and are the basic costs of link type switching and intra-type node switching respectively.

5. The multi-agent dynamic access method for an integrated architecture of space-air-ground-vehicle Internet of Things according to claim 1, characterized in that: The reinforcement learning algorithm also includes a dynamic equilibrium mechanism, and the reward function where s k , a k are the state space and action space of the k-th vehicle node respectively, are the total delay and handover cost respectively, represents the load balancing factor, which is quantified by evaluating the deviation of the current load of the selected node from the system average load ω1, ω2, and ω3 are weight coefficients.

6. A multi-agent dynamic access and collaboration method for an integrated architecture of space-air-ground-vehicle Internet of Things, characterized in that, It includes: Select the optimal access node for multiple vehicles according to the multi-agent dynamic access method for an integrated space-air-ground-vehicle network architecture according to any one of claims 1-5; Perform sparse feature extraction and time series encoding for multiple vehicles through a Transformer-based architecture to correct the asynchronous delay caused by uneven clock sampling of vehicles and different ground-satellite transmission links; among which, the asynchronous delay processing process matches the historical data of the regions of interest of different vehicles, and enhances the feature representation ability through a multi-layer perceptron to adapt to the Transformer architecture and ensure the time series consistency of multi-source data.

7. A multi-agent dynamic access and cooperation method for an integrated architecture of space-air-ground-vehicle Internet of Things according to claim 6, characterized in that: The Transformer-based architecture first performs time encoding on the original perception data, then uses an object detection algorithm to generate regions of interest and extract sparse feature representations, realizes cross-time-domain feature matching through a multi-head attention mechanism, and finally completes the bird's-eye view perspective conversion and multi-source feature fusion based on the matching results.

8. A multi-agent dynamic access and cooperation method for an integrated architecture of space-air-ground-vehicle Internet of Things according to claim 7, characterized in that: The Transformer-based architecture also includes a feature fusion optimization strategy, which adopts a multi-scale maximum fusion method and combines a reliability evaluation mechanism to dynamically allocate the weights of different data sources to improve the accuracy of the fusion result.

9. A computer system, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the computer program is loaded into the processor, it implements the steps of the method according to any one of claims 1-8.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the method according to any one of claims 1-8.

Citation Information

Cited By

  • Ground-non-ground fusion network resource allocation method and system based on upper confidence bound algorithm

    CN121262660A

  • A ground-non-ground fusion network resource allocation method and system based on an upper confidence bound algorithm

    CN121262660B