An event-stream-assisted link adaptation method and system for vehicle-to-everything (V2X) networks

By employing an event-stream-assisted link adaptation method, which utilizes deep learning and high-frequency, low-redundancy sensing data from event cameras, the problem of real-time acquisition of channel changes in vehicle-to-everything (V2X) networks is solved, thereby improving coding rate performance and spectral efficiency.

CN121056964BActive Publication Date: 2026-01-06WUHAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511598522.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-11-04
Publication Date
2026-01-06
Estimated Expiration
2045-11-04

AI Technical Summary

Technical Problem

In the context of vehicle-to-everything (V2X) scenarios, traditional link adaptive methods cannot obtain channel changes in real time, leading to a reduction in spectrum efficiency.

Method used

An event-stream-assisted link adaptation method is adopted. A lightweight spatiotemporal depth estimation algorithm is designed using deep learning. High-frequency, low-redundancy sensing data captured by event cameras is used to construct a real-time mapping relationship between channel and environmental features. Based on deep learning, a lightweight spatiotemporal depth estimation algorithm is designed to achieve dynamic sensing of the distance between base stations and user equipment by utilizing the spatiotemporal correlation of event streams.

Benefits of technology

It improves coding rate performance, overcomes the time delay defect of traditional feedback mechanisms, and increases spectral efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121056964B_ABST
    Figure CN121056964B_ABST
Patent Text Reader

Abstract

This invention discloses an event-stream-assisted link adaptation method and system for vehicle-to-everything (V2X) networks, belonging to the field of V2X link adaptation. The method includes: acquiring event stream data; inputting the event stream data into a trained depth estimation model to obtain the environmental depth at the current time step; calculating the distance between the base station and the user equipment (UE) based on the event stream data and the environmental depth at the current time step; and estimating the channel state and calculating the coding rate at the current time step based on the distance between the base station and the UE to complete link adaptation. This invention utilizes the spatiotemporal correlation of event streams to achieve dynamic perception of the distance between the base station (BS) and the user equipment (UE), overcoming the time delay defect of traditional feedback mechanisms. It also overcomes the problem of inaccurate channel acquisition caused by outdated feedback information through visual information, reducing model computational complexity while improving coding rate performance, thereby improving spectral efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of vehicle-to-everything (V2X) link adaptation, specifically relating to an event-flow-assisted link adaptation method and system for V2X. Background Technology

[0002] The Internet of Vehicles (IoV) enables traffic efficiency optimization and other functions through communication between vehicles, between vehicles and people, and between vehicles and infrastructure, forming the foundation of smart city construction. With the continuous development of communication technology and the increasing intelligence of vehicles, the number of communication devices in IoV is expanding rapidly, and the applications supported by these devices place higher demands on communication latency and reliability. In mobile communication systems, especially in IoV scenarios, vehicle terminals and user equipment (UEs) carrying base stations (BSs) have higher mobility, and wireless channels exhibit highly time-varying characteristics. Link adaptation technology can adaptively adjust the modulation and coding scheme according to time-varying channel conditions to improve the system's spectral efficiency. OLLA (Open Loop Link Adaptation) is a widely used link adaptation algorithm. In IoV environments where UEs and base stations are moving or rapidly changing, the current channel has already changed significantly by the time the transmitter receives channel information. Algorithms based on the OLLA algorithm struggle to obtain timely and accurate channel information. Achieving highly reliable, low-latency communication using link adaptation in time-varying channel environments is a major challenge.

[0003] Numerous researchers have conducted studies on the convergence of link adaptation, system power consumption, link adaptation algorithms based on the physical layer or link layer, and the use of machine learning methods. Unlike traditional static scenarios, especially in IoV communication systems, wireless channels exhibit highly time-varying characteristics. This makes current link adaptation methods unable to acquire channel changes in real time, often sacrificing spectral efficiency to meet communication requirements. Summary of the Invention

[0004] The purpose of this invention is to address the reduced spectral efficiency of IoV (Inter-vehicle Vehicle) in high-reliability, low-latency scenarios due to inaccurate and non-real-time channel feedback. It provides an event-stream-assisted link adaptation method for vehicular networks. Considering the limited computing and storage resources of vehicle BS (Base Station), this invention innovatively introduces dynamic event streams captured by event cameras into the link adaptation framework. It constructs a real-time mapping relationship between channel and environmental features using high-frequency, low-redundancy sensing data. A lightweight spatiotemporal depth estimation algorithm based on deep learning is designed, utilizing the spatiotemporal correlation of event streams to achieve dynamic perception of the distance between the BS and UE (User Equipment). This overcomes the time delay defect of traditional feedback mechanisms and overcomes the problem of inaccurate channel acquisition caused by outdated traditional feedback information through visual information. While reducing model computational complexity, it improves coding rate performance, thereby enhancing spectral efficiency.

[0005] According to one aspect of this specification, an event-flow-assisted link adaptation method for vehicle-to-everything (V2X) networks is provided, comprising:

[0006] S1. Obtain event stream data;

[0007] S2. Input the event stream data into the trained depth estimation model to obtain the environment depth at the current time step; wherein, the training of the depth estimation model includes:

[0008] Construct an event stream training dataset;

[0009] The depth estimation model is constructed, including: an event stream encoder, which extracts features from event stream data to obtain multi-scale features; a spatiotemporal consistency module, which performs spatiotemporal context modeling based on multi-scale features and outputs low-resolution features; and a decoder, which transforms low-resolution features into high-resolution depth maps.

[0010] The depth estimation model is trained on the constructed event stream training dataset, and the trained depth estimation model is output.

[0011] S3. Calculate the distance between the base station and the user equipment based on event stream data and the environmental depth at the current time step;

[0012] S4. Based on the distance between the base station and the user equipment, estimate the channel state and calculate the coding rate at the current time step to complete link adaptation.

[0013] Furthermore, the construction of the event stream encoder includes:

[0014] Reduce the size of the input event stream data;

[0015] We capture multi-scale context in event stream data after scaling down and enhance long-range dependencies through an attention mechanism to obtain multi-scale features.

[0016] Furthermore, the construction of the spatiotemporal consistency module includes:

[0017] Deformable convolutions are used to spatiotemporally align the multi-scale features of the input and output a weight vector.

[0018] The weight vector is multiplied by the features at the current time step and the features at the previous time step, and then fused to obtain the fused features;

[0019] Multiply the weight vector by the fused feature to obtain the low-resolution feature at the current time step.

[0020] Further, S3 includes:

[0021] The target detection model is used to obtain the upper left and lower right corner coordinates of the user equipment at the current time step. Combined with the environmental depth at the current time step, the distance between the base station and the user equipment is calculated.

[0022] Further, S4 includes:

[0023] A small-scale fading model is established based on the Rician channel fading model, and the small-scale fading and large-scale fading are fused to obtain the channel; wherein, the large-scale fading is determined by the distance between the base station and the user equipment.

[0024] By combining event stream data, the signal-to-noise ratio of the channel is calculated to obtain the channel quality;

[0025] Based on channel quality and combined with finite block length theory, the coding rate is calculated to achieve link adaptation.

[0026] Furthermore, the method also includes constructing a scale-invariant distance-weighted loss function, expressed as:

[0027]

[0028] in, Indicates position The weighting coefficients, It is the actual depth value. It is the predicted depth value. This indicates the number of pixels within the valid area.

[0029] Furthermore, acquiring event stream data also includes the following preprocessing steps:

[0030] Convert event stream data into a structured representation and separate it according to the number of time bins;

[0031] Aggregate the polarity values ​​of events within each timeframe to generate a three-dimensional tensor.

[0032] According to one aspect of this specification, an event-flow-assisted link adaptive system for vehicle-to-everything (V2X) networks is provided, comprising:

[0033] The data acquisition module is used to acquire event stream data;

[0034] An environmental depth calculation module is used to input event stream data into a trained depth estimation model to obtain the environmental depth at the current time step; wherein, the training of the depth estimation model includes:

[0035] Construct an event stream training dataset;

[0036] The depth estimation model is constructed, including: an event stream encoder, which extracts features from event stream data to obtain multi-scale features; a spatiotemporal consistency module, which performs spatiotemporal context modeling based on multi-scale features and outputs low-resolution features; and a decoder, which transforms low-resolution features into high-resolution depth maps.

[0037] The depth estimation model is trained on the constructed event stream training dataset, and the trained depth estimation model is output.

[0038] The distance calculation module is used to calculate the distance between the base station and the user equipment based on event stream data and the environmental depth at the current time step;

[0039] The link adaptive module is used to obtain the large-scale fading based on the distance between the base station and the user equipment, fuse the large-scale fading with the small-scale fading to obtain the channel, and calculate the coding rate at the current time step.

[0040] According to one aspect of this specification, a vehicle terminal carrying a base station is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of the event flow-assisted link adaptation method for vehicle networking.

[0041] According to one aspect of this specification, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the steps of the event-flow-assisted link adaptation method for vehicle-to-everything (V2X) networks.

[0042] Compared with the prior art, the beneficial effects of the present invention are:

[0043] 1. This invention addresses the issue of reduced spectral efficiency in IoV scenarios due to inaccurate and non-real-time channel feedback. It proposes a deep learning-based event-stream-assisted multi-user link adaptive framework for vehicle networking, overcoming the time delay defect of traditional feedback mechanisms. By using visual information, it overcomes the problem of inaccurate channel acquisition caused by outdated traditional feedback information. This reduces the computational complexity of the model while improving coding rate performance, thereby enhancing spectral efficiency.

[0044] 2. This invention addresses the highly dynamic channel environment and limited computing resources in IoV scenarios by proposing a lightweight depth estimation algorithm based on event streams with spatiotemporal dependence, which accurately and quickly calculates the distance between the BS and UE.

[0045] 3. The depth estimation model proposed in this embodiment of the invention performs well in terms of lightweighting, and the proposed link adaptation method has excellent coding rate performance, which can effectively improve spectral efficiency and open up new avenues for link adaptation schemes. Attached Figure Description

[0046] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0047] Figure 1 This is a flowchart of a method according to an embodiment of the present invention;

[0048] Figure 2 This is a schematic diagram of the network model according to an embodiment of the present invention;

[0049] Figure 3 This is a schematic diagram of a scenario according to an embodiment of the present invention;

[0050] Figure 4 This is a schematic diagram illustrating the training changes of the Loss value of the network model in an embodiment of the present invention.

[0051] Figure 5 This is a comparison chart of the coding rate results in the simulation experiment of this invention. Detailed Implementation

[0052] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0053] like Figure 1As shown, this embodiment of the invention provides an event-stream-assisted link adaptation method for vehicle-to-everything (V2X) networks, comprising: S1, acquiring event stream data; S2, inputting the event stream data into a trained depth estimation model to obtain the environmental depth at the current time step; wherein, the training of the depth estimation model includes: constructing an event stream training dataset; building the depth estimation model, comprising: an event stream encoder for extracting features from the event stream data to obtain multi-scale features; a spatiotemporal consistency module for performing spatiotemporal context modeling based on multi-scale features and outputting low-resolution features; and a decoder for converting low-resolution features into a high-resolution depth map; S3, calculating the distance between the base station and the user equipment based on the event stream data and the environmental depth at the current time step; S4, obtaining the large-scale fading based on the distance between the base station and the user equipment, fusing the large-scale fading with the small-scale fading to obtain the channel, and calculating the coding rate at the current time step to complete the link adaptation.

[0054] Specifically, event streams (environmental information) are acquired from mobile vehicle base stations equipped with event cameras and preprocessed. Unlike traditional RGB cameras, event cameras asynchronously detect pixel-level brightness changes, outputting data only when an event occurs. This reduces redundant information caused by the fixed frame rate of traditional RGB cameras, lowering the resource consumption of the vehicle base station. Furthermore, it can capture the dynamic changes of fast-moving targets in an IoV environment in real time with a microsecond-level temporal resolution. The data format of the event camera also differs from traditional RGB; it consists of four tuples, as follows:

[0055] (1)

[0056] in, An event stamp indicating when an event occurred. , The x and y axes represent the occurrence of the event. The polarity is represented by -1 indicating decreased brightness and +1 indicating increased brightness. This parameter is not calculated but is part of the output format of the event camera. These data streams are inherently unordered because they represent events occurring independently on each pixel. Discrete and asynchronous events change over time, forming event stream data. However, discrete event streams cannot be directly passed to the method proposed in this invention. This invention converts asynchronous sparse event streams into a structured representation, first by setting the time window length. Extracting time period The event set within each time chamber is divided into time chambers, and finally the polarity values ​​of the events within each time chamber are aggregated. Generate a three-dimensional tensor, in which Indicates the beginning of time.

[0057] Specifically, after preprocessing, the event stream enters the depth estimation module to calculate the environmental depth at the current time step. Considering the actual needs of mobile vehicular base stations (BSs) in IoV scenarios, computational resources are limited, and they are highly sensitive to short-range accuracy. This invention proposes a depth estimation network architecture, including an event stream encoder, a spatiotemporal consistency module, and a decoder, such as... Figure 2 As shown in the diagram. The event encoder encodes the raw event data and extracts features. The spatiotemporal consistency module models the spatiotemporal context to improve the spatiotemporal consistency of depth estimation; this module reduces distance estimation errors between the BS and UE caused by dynamic environmental changes. The decoder recovers a high-resolution depth map from the low-resolution feature representation.

[0058] Specifically, event camera data processed by voxel grids is fed into the event encoder for feature extraction. Subsequently, depth features are extracted progressively through three progressive scale feature extraction modules. The feature extraction module consists of a dilated convolution and a feature interaction module. The former expands the receptive field by dynamically adjusting the dilation rate to capture multi-scale context, while the latter introduces an attention mechanism to enhance long-range dependency modeling. The encoder stitches together a low-resolution version of the original image at each downsampling to preserve details and employs a stochastic depth decay strategy to optimize training stability, ultimately outputting multi-scale features that maintain lightweight design while ensuring both local accuracy and global consistency.

[0059] Specifically, independent processing of each time slot can lead to depth jumps, which in turn cause jumps in the distance estimation between the BS and UE, ultimately affecting the link adaptive performance. Therefore, a deformable convolution is used in the spatiotemporal consistency module to align spatiotemporal features. Furthermore, a channel attention mechanism is used to adaptively fuse the current input and temporal hidden states. The spatiotemporal consistency module is further subdivided into a spatiotemporal alignment module, a channel attention module, and a spatiotemporal feature fusion module. First, the features of the 128 channels extracted by the event stream encoder at the current time are... Compared to the previous time step, the 128-channel hidden state The data is concatenated and input into the spatiotemporal alignment module. Its core principle lies in using deformable convolutions to obtain a larger receptive field while maintaining the size of the output feature map. Its expression is shown below:

[0060] (2)

[0061] Unlike ordinary convolution, This is the current output position. It is the feature map at the location The value is a weighted sum of several points obtained by offset sampling in the input feature map, where N represents the number of sampling points in the convolution kernel and the index of the nth sampling point in the convolution kernel. It refers to the relative positions of the standard convolution sampling points. It is the predicted spatial offset, which is dynamically calculated from the input feature map. It is the first in the convolution kernel The weight of each position, This is the current output position. Offset This is achieved by compressing the input of the spliced ​​features back to 128 channels via convolution, and then obtaining it through additional convolutions. Finally, It is obtained by using a 3-fold deformable convolution aligned with the state of the previous time step. The above process can be expressed as:

[0062] (3)

[0063] in, This represents the convolution operation. Represents deformable convolution. This represents the aligned feature map. This represents the feature map used to calculate the offset. The feature map representing the current time step. This represents element-wise addition of vectors. In IoV, more attention needs to be paid to user experience (UE) information such as vehicles and pedestrians, rather than background information. Therefore, channel attention is added after the spatiotemporal alignment module to enhance key features and suppress useless features. Features are compressed using the following expression:

[0064] (4)

[0065] in, Indicates the first Each channel is located in Feature map, and Indicates the height and width of the feature map. Indicates channel The global description is as follows: After compression, the dimensionality is reduced through convolution, then ReLU activation is applied, followed by another convolution operation to restore it to 256 channels. Finally, the activation function outputs a weight vector. . The acquisition process is represented as follows:

[0066] (5)

[0067] in, It is the current frame Aligned features The joint feature map is formed by splicing. This represents the channel attention mechanism. In the spatiotemporal feature fusion module, the weight vector... The features are multiplied by the original features at the current time step and the hidden state points at the previous time step, respectively, to perform fusion. The fused features are then fed into the edge enhancement module. The above process can be represented as:

[0068] (6)

[0069] (7)

[0070] (8)

[0071] in, and Spatial attention maps are generated for the current time step and the aligned historical features, respectively. and These represent the convolution operations for modeling the current and historical data, respectively. This represents element-wise multiplication of vectors. This represents the element-wise addition of a vector. and The two branches are then weighted separately, concatenated, and integrated into a new representation through convolution. Edge enhancement module uses The high response values ​​at the boundaries are generated by convolution and then input into the Sigmoid function. Finally, the weight vector and the fused features are multiplied by a dot product to obtain the features at the current time step, which are then input into the decoder. The decoder consists of an upsampling module and an estimation layer. It converts the 128-channel low-resolution feature map output by the spatiotemporal coherence module into a single-channel depth matrix with the same resolution as the input event bin. The upsampling module consists of cubic bilinear interpolation upsampling combined with convolution operations, restoring spatial resolution through progressive upsampling while refining features using convolution operations. The estimation layer consists of a single convolutional layer with a kernel size of [missing value]. It maps the 32-channel features output by the upsampling module to a single-channel depth matrix.

[0072] Specifically, the closer the distance between the BS and the UE, the greater the impact of the distance estimation error on the channel; conversely, the greater the distance, the smaller the impact of the error. Traditional depth estimation typically uses scale-invariant loss as the objective function, but this design cannot predict accuracy at close range. Therefore, this embodiment of the invention designs a scale-invariant distance-weighted loss function, the expression of which is:

[0073] (9)

[0074] in, Indicates position The weighting coefficients are used to emphasize the importance of proximity estimates. It is the actual depth value. It is the predicted depth value. This indicates the number of pixels within the valid area. This represents the distance-weighted term. It is a small constant used to prevent errors caused by division by zero. Because the channel characteristics are more sensitive to depth estimation errors of nearby objects, it needs to be assigned a higher weight. Ensure that the average of these weights is 1, then divide them by the average to obtain... , This represents the distance-weighted term. This represents the average of the distance-weighted terms. It is a location The weighting coefficients before processing become... To maintain consistency when processing data of different scales, embodiments of the present invention use... This represents the logarithmic error term with constant scale.

[0075] Specifically, the precise distance between the BS and the UE is calculated based on the event flow and environmental depth information. The UE is detected within the field of view using a target detection model. This model outputs the coordinates of the UE's top-left and bottom-right corners at the current time step, which, combined with the environmental depth at the current time step, yields the real-time distance between the BS and the UE.

[0076] Specifically, the channel is estimated and the coding rate at the current time step is calculated to complete link adaptation. For example... Figure 3 As shown, this embodiment of the invention pertains to a downlink system in an IoV scenario, where a vehicle equipped with an event camera acts as a mobile BS. The event camera, computing device, and vehicle are connected via a wired connection, so transmission latency and losses are not considered. The BS travels on urban roads, providing services to UEs (pedestrians and other vehicles) on the roads. The channel is determined by both large-scale and small-scale fading. Large-scale fading is primarily determined by the distance between the BS and the UE, significantly impacting the coverage and signal strength of the communication system. Large-scale fading can be expressed as:

[0077] (10)

[0078] in, The loss is calculated at a reference distance of 1m. This is the path loss index. The distance between the vehicle's base station (BS) and the user equipment (UE) is given. Signal power decreases with distance. Small-scale fading is investigated using the Rician channel fading model. The Rician model consists of line-of-sight and non-line-of-sight scattering, and is expressed as follows:

[0079] (11)

[0080] in, , where is the Rice factor, representing the ratio of direct path power to scattered path power. This represents the line-of-sight distance component, which is a complex exponential term. This represents the non-linear line-of-sight component, with each element following a complex Gaussian distribution. Finally, the channel after fusing large-scale and small-scale fading can be expressed as:

[0081] (12)

[0082] The signal-to-noise ratio (SNR) can only be calculated after the channel is obtained from the information obtained from the event stream to obtain the channel quality. The calculation method is as follows:

[0083] (13)

[0084] in, For transmission power, This represents noise power. In IoV scenarios, low-latency communication uses short block transmission. Based on the finite block length theory, the coding rate is calculated to achieve link adaptation. The coding rate calculation is expressed as follows:

[0085] (14)

[0086] in, It is Shannon capacity. It is a finite block length correction term. It is channel dispersion, It is the block length. Gauss The inverse function of the function This represents the bit error rate. The more accurate the distance between the BS and UE calculated from the environment, the more accurate the channel estimation, the more accurate the coding rate calculation, and ultimately, the better the communication reliability.

[0087] Specifically, to verify the performance of the event-stream-assisted link adaptation method for vehicle-to-everything (V2X) proposed in this invention, a set of experiments are conducted. For training the depth estimation module, this embodiment uses the DSEC dataset. DSEC is an event camera dataset collected during real-world vehicle driving, where vehicles move freely in urban areas without a fixed tracking target. The model is built using the PyTorch framework and trained on a server equipped with an NVIDIA GeForce RTX 4090 with 24 GB of VRAM. To simulate the channel propagation characteristics in wireless communication systems, for the NLOS part, this invention selects the CDL-C channel model provided by the 3GPP organization, suitable for dense urban areas. Specific parameters are as follows: noise power of -45 dBm, transmission power of 110 dBm, Rice factor of 3 dB, path loss coefficient of 3.5, and path loss of -30 dB at a baseline distance of 1 m. Regarding training parameters, the batch size is 4, the number of epochs is 100, and the initial learning rate is... The learning rate scheduler used is OneCycleLR, and the optimizer is AdamW.

[0088] Specifically, the embodiments of the present invention are in Figure 4 The convergence and training performance of the proposed method are demonstrated. Overall, the loss value decreases with increasing training epochs, indicating continuous learning and optimization of the model. In the initial training phase (first 30 epochs), the loss value decreases rapidly, indicating that the model can quickly adapt to the data at this stage. After 30 epochs, the rate of decrease in the loss value slows significantly and stabilizes in subsequent training, eventually settling at a low level, indicating convergence in model efficiency. The model in this invention has 4.33M parameters. The number of model parameters directly affects the model's storage space; fewer parameters mean less terminal storage space required. Therefore, the proposed method demonstrates its ability to be deployed in vehicle-mounted business units (BSs) with limited storage and computing resources.

[0089] Specifically, the coding rate simulation results are as follows: Figure 5 As shown in the figure, the experimental results demonstrate that the coding rate achieves the highest performance under ideal channel information, representing the upper limit of performance. However, this is unattainable in real-world environments and serves only as a benchmark for method performance. The link-adaptive framework proposed in this paper can obtain the real-time distance between the user equipment and the base station in a highly dynamic environment, thereby acquiring more accurate channel information to select the coding rate. Compared to traditional algorithms, our link-adaptive framework improves coding rate performance by 23.95%, bringing it closer to the ideal upper limit.

[0090] The implementation of the various embodiments of the present invention is based on programmed processing through a device with processor functionality. Therefore, in practical engineering, the technical solutions and functions of the various embodiments of the present invention are encapsulated into various modules. Based on this reality, and building upon the above embodiments, the embodiments of the present invention provide an event-flow-assisted link adaptive system for vehicle-to-everything (V2X) networks. This system is used to execute an event-flow-assisted link adaptive method for V2X networks as described in the above method embodiments.

[0091] The system includes: a data acquisition module for acquiring event stream data; an environmental depth calculation module for inputting the event stream data into a trained depth estimation model to obtain the environmental depth at the current time step; wherein, the training of the depth estimation model includes: constructing an event stream training dataset; building the depth estimation model, including: an event stream encoder for extracting features from the event stream data to obtain multi-scale features; a spatiotemporal consistency module for performing spatiotemporal context modeling based on multi-scale features and outputting low-resolution features; a decoder for converting low-resolution features into high-resolution depth maps; training the depth estimation model on the constructed event stream training dataset and outputting the trained depth estimation model; a distance calculation module for calculating the distance between the base station and the user equipment based on the event stream data and the environmental depth at the current time step; and a link adaptation module for estimating the channel state and calculating the coding rate at the current time step based on the distance between the base station and the user equipment to complete link adaptation.

[0092] The event-stream-assisted link adaptive system for vehicle-to-everything (V2X) provided in this invention addresses the spectral efficiency reduction problem caused by inaccurate and non-real-time channel feedback in high-reliability, low-latency scenarios. It employs several modules to construct a real-time mapping relationship between channel and environmental features using high-frequency, low-redundancy sensing data. A lightweight spatiotemporal depth estimation algorithm based on deep learning is designed, and the spatiotemporal correlation of event streams enables dynamic perception of the distance between the BS and UE. This overcomes the time delay defect of traditional feedback mechanisms and overcomes the problem of inaccurate channel acquisition caused by outdated traditional feedback information through visual information. While reducing model computational complexity, it improves coding rate performance, thereby enhancing spectral efficiency.

[0093] Based on the same inventive concept as the foregoing embodiments, this embodiment of the invention also provides a vehicle terminal carrying a base station, including a memory and a processor. The memory is used for computer programs, and the processor is used for executing computer-executable instructions to implement an event-flow-assisted link adaptation method for vehicle networking as proposed in the above embodiments.

[0094] This invention also provides a computer-readable storage medium storing a computer program thereon. When executed by a processor, this program overcomes the problem of reduced spectral efficiency caused by inaccurate and non-real-time channel feedback in high-reliability, low-latency scenarios, overcomes the time delay defects of traditional feedback mechanisms, and overcomes the problem of inaccurate channel acquisition caused by outdated traditional feedback information through visual information. While reducing the computational complexity of the model, it improves coding rate performance, thereby improving spectral efficiency.

[0095] The storage medium can be any non-volatile storage device such as a hard disk, solid-state drive, flash drive, or optical disk, used to store computer program code and necessary data files. The stored computer program includes: a data acquisition module, an environmental depth calculation module, a distance calculation module, and a link adaptation module.

[0096] Finally, it should be noted that the above specific embodiments are merely representative examples of the present invention. Obviously, the present invention is not limited to the above specific embodiments and many variations are possible. Any simple modifications, equivalent changes, and alterations made to the above specific embodiments based on the technical essence of the present invention should be considered within the protection scope of the present invention.

Claims

1. A method for link adaptation assisted by event stream for vehicle-to-everything, characterized in that, The method comprises the following steps: S1, obtaining event stream data; S2, inputting the event stream data into a trained depth estimation model to obtain the environmental depth at the current time step; wherein the training of the depth estimation model comprises: constructing an event stream training data set; building a depth estimation model, which comprises: an event stream encoder for feature extraction of the event stream data to obtain multi-scale features; a spatio-temporal consistency module for spatio-temporal context modeling based on the multi-scale features to output low-resolution features; and a decoder for converting the low-resolution features into a high-resolution depth map; training the depth estimation model on the constructed event stream training data set to output the trained depth estimation model; S3, calculating the distance between the base station and the user equipment based on the event stream data and the environmental depth at the current time step; S4, estimating the channel state and calculating the coding rate at the current time step based on the distance between the base station and the user equipment to complete link adaptation.

2. The method of claim 1, wherein, The construction of the event stream encoder comprises: reducing the size of the input event stream data; capturing multi-scale contexts in the size-reduced event stream data and enhancing long-range dependencies through an attention mechanism to obtain multi-scale features.

3. The method of claim 1, wherein, The construction of the spatio-temporal consistency module comprises: using deformable convolution to perform spatio-temporal alignment on the input multi-scale features and output a weight vector; multiplying the weight vector with the features at the current time and the features at the previous time, respectively, and fusing them to obtain fused features; point-multiplying the weight vector with the fused features to obtain the low-resolution features at the current time.

4. The method of claim 1, wherein, The S3 comprises: using a target detection model to obtain the top-left corner coordinates and the bottom-right corner coordinates of the user equipment at the current time step, and combining the environmental depth at the current time step to calculate the distance between the base station and the user equipment.

5. The method of claim 1, wherein, The S4 comprises: establishing a small-scale fading based on a Rician channel fading model, fusing the small-scale fading and a large-scale fading to obtain a channel; wherein the large-scale fading is determined by the distance between the base station and the user equipment; combining the event stream data to calculate the signal-to-noise ratio of the channel to obtain the channel quality; based on the channel quality, combining the finite block length theory to calculate the coding rate to complete link adaptation.

6. The method of claim 1, wherein, The method further comprises constructing a scale-invariant distance weighted loss function, the expression of which is: , wherein, a weighting factor, is a true depth value, is a predicted depth value, is a predicted depth value, denotes the number of pixels within the valid region.

7. The method of claim 1, wherein, The acquisition of the event stream data further comprises the following preprocessing process: converting the event stream data into a structured representation and separating them according to time bins; aggregating the polarity values of events in each time bin to generate a three-dimensional tensor. 8.A system for vehicular internet of things (V-IoT) event stream assisted link adaptation, comprising: The method comprises the following steps: a data acquisition module for acquiring event stream data; an environmental depth calculation module for inputting the event stream data into a trained depth estimation model to obtain the environmental depth at the current time step; wherein the training of the depth estimation model comprises: constructing an event stream training data set; building a depth estimation model, which comprises: an event stream encoder for feature extraction of the event stream data to obtain multi-scale features; a spatio-temporal consistency module for spatio-temporal context modeling based on the multi-scale features to output low-resolution features; and a decoder for converting the low-resolution features into a high-resolution depth map; training a depth estimation model on the constructed event stream training dataset, and outputting the trained depth estimation model; a distance calculation module configured to calculate the distance between the base station and the user equipment based on the event stream data and the environment depth at the current time step; a link adaptation module configured to estimate the channel state and calculate the coding rate at the current time step based on the distance between the base station and the user equipment to complete link adaptation. 9.A vehicle terminal carrying a base station, comprising a memory and a processor, the memory storing a computer program, characterized in that, The processor implements the steps of the vehicle Internet-oriented event stream assisted link adaptation method of any one of claims 1-7 when executing the computer program.

10. A computer-readable storage medium having stored thereon a computer program, characterized in that, The computer program is executed by the processor to implement the steps of the vehicle Internet-oriented event stream assisted link adaptation method of any one of claims 1-7.

Citation Information

Patent Citations

  • Depth estimation method based on laser radar and event camera fusion

    CN114359744A

  • Scene recognition method and system based on fusion event camera

    CN116188930A