Internet of vehicles channel prediction method based on multi-modal fusion and related equipment
By combining weather scenarios in urban road scenarios to obtain multimodal data, data enhancement and deep fusion are performed, and channel prediction is used to use the Transformer model to perform channel prediction, the problems of limited single-modal perception capability and insufficient data quality in vehicle network channel prediction are solved, and high-precision and efficient channel prediction are achieved.
Patent Information
- Application Number
- CN202510826145.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-19
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-06-19
AI Technical Summary
In the prior art, the Internet of Vehicle channel prediction method has problems such as limited single-modal perception capability, poor quality of original data, and insufficient robustness in complex scenarios.
The Internet of Vehicle Channel Prediction Method based on multimodal fusion is adopted. By constructing urban road scenarios, multimodal primitive data are obtained in combination with weather scenarios, data enhancement, feature extraction and deep fusion are performed, and channel prediction is performed using the Transformer model.
It improves the robustness and accuracy of channel prediction, ensures stability and reliability in multiple scenarios, improves the computing efficiency of the model, and provides technical guarantees for real-time channel prediction.
Smart Images

Figure CN120342527A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of vehicle - to - everything (V2X) communication, and specifically to a V2X channel prediction method based on multimodal fusion and related devices. Background Technique
[0002] With the rapid development of V2X technology, the importance of channel prediction in improving the efficiency and reliability of vehicle - road collaborative communication has become increasingly prominent. Thanks to policy support and the gradual improvement of technical standards, V2X communication systems are moving towards the ultra - high frequency band and multimodal fusion. Currently, with the implementation of 5G technology and the forward - looking advancement of 6G communication, V2X has gradually acquired the ability to support the efficient transmission and perception of multimodal data. At the same time, the popularization of intelligent sensors (such as lidar, millimeter - wave radar, etc.) and edge computing has continuously enhanced the feasibility of incorporating multimodal sensing data into the design of communication systems, further promoting the research and application of channel prediction technology.
[0003] Currently, many studies have explored the potential and advantages of incorporating sensing data into channel prediction. In the prior art, the research on V2X channel prediction based on various multimodal enhanced data mainly focuses on the following aspects: 1. Channel prediction methods based on visual sensor data Visual sensors (such as cameras, etc.) capture visible or infrared light in the environment to generate two - dimensional image or video data. Its technical characteristics are that it can provide rich scene information, such as object shape, color, and texture, and is suitable for the perception of static and dynamic environments. In channel prediction, environmental features (such as buildings, vegetation, vehicles, etc.) in the data are usually used to infer signal propagation paths and attenuation characteristics.
[0004] 2. Channel prediction methods based on optical sensor data Optical sensors (such as lidar, etc.) emit laser beams and receive reflected signals to generate high - precision three - dimensional point cloud data. Its technical characteristics are that it can provide accurate distance and shape information and is suitable for high - resolution environmental modeling. In channel prediction, the environmental structure (such as buildings, obstacles, terrain, etc.) in the point cloud data is usually used to model signal propagation paths and reflection characteristics.
[0005] 3. Channel prediction methods based on positioning data Location data usually comes from the Global Positioning System (GPS), Inertial Measurement Unit (IMU), or integrated positioning system, providing information about the device's location, speed, and direction. Its technical feature is the ability to provide accurate location and motion information, suitable for channel prediction in dynamic environments. In channel prediction, the device's location and motion trajectory are usually utilized, combined with map data or environmental models, to infer the signal propagation path and dynamic changes.
[0006] 4. Channel Prediction Method Based on Modal Fusion The modal fusion method integrates multi-source heterogeneous data such as vision, optics, and location, comprehensively utilizes the complementary characteristics of different modalities, and improves the robustness and accuracy of channel prediction. Its technical features are reflected in the use of a deep learning framework to achieve feature alignment and joint modeling of heterogeneous data. However, different multi-modal enhanced data still have limitations in terms of environmental adaptability and diversity. First, the perception ability of vision sensors significantly decreases in low-light, extreme weather, or occlusion scenarios, and they cannot directly provide distance information, relying on complex algorithms for indirect derivation, increasing computational complexity and uncertainty. Second, optical sensors are easily affected by weather and the characteristics of reflective surfaces, with large amounts of data, high computational resource requirements, and expensive hardware costs, restricting their large-scale application. Third, in occluded environments (such as urban canyons, tunnels, etc.) or indoor scenarios, the positioning accuracy of GPS may drop significantly or even fail due to satellite signal shielding or attenuation. Fourth, existing research has not effectively integrated multi-source heterogeneous information, resulting in limited adaptability to diverse traffic conditions. Summary of the Invention
[0007] In order to overcome the defects of the above-mentioned existing technologies, the purpose of the present invention is to provide a vehicle-to-everything (V2X) channel prediction method and related devices based on multi-modal fusion to solve the technical problems of limited single-modal perception ability, poor quality of original data, and insufficient robustness in complex scenarios existing in the prior art.
[0008] The present invention is implemented through the following technical solutions: In the first aspect, the present invention provides a vehicle-to-everything (V2X) channel prediction method based on multi-modal fusion, including: Construct the urban road scene of the vehicle-to-everything (V2X) channel, and obtain known time frame data and unknown time frame data to be predicted by combining the weather scene in the urban road scene; Perform data enhancement, data vector representation, and location information processing on the known time frame data and unknown data to be predicted to obtain multi-modal vectors, and perform feature extraction and splicing processing on the multi-modal vectors to obtain multi-modal feature vectors; Input the multi-modal feature vectors into the Transformer model, and perform deep fusion between modalities through the multi-head attention mechanism and the feed-forward neural network of the Transformer model to obtain joint feature vectors; Perform non-linear mapping on the joint feature vectors through a multi-layer perceptron module, and output the channel index probability distribution, and predict the vehicle network channel based on the channel index probability distribution.
[0009] Preferably, in the urban road scenario, when obtaining the known time frame data and the unknown time frame data to be predicted by combining the weather scenario, the weather scenario includes sunny scenario, rainy scenario and snowy scenario; The known time frame data includes channel index values, communication signal time frame serial numbers, position coordinates of vehicles and RSUs, RGB images, LiDAR point clouds; The unknown time frame data to be predicted includes communication signal time frame serial numbers, position coordinates of vehicles and RSUs, RGB images, LiDAR point clouds.
[0010] Preferably, data augmentation includes image augmentation and point cloud augmentation; Among them, image augmentation includes brightness adjustment, contrast enhancement, gamma correction, Gaussian blur and background masking; The process of point cloud augmentation includes background filtering by referring to a two-dimensional bird's-eye view, mapping the building positions into the point cloud data, removing the point clouds in the building position areas, and concentrating the remaining point clouds in the areas of vehicle movement and road structure to reduce the interference of static objects on the model; The data vector representation includes image vector representation and point cloud vector representation; Among them, the process of image vector representation includes vectorizing the picture data in the RGB three-channel format and using it as the model input after normalization processing; The process of point cloud vector representation includes first dividing the area of interest into grids to achieve effective vectorization representation of the point cloud data; setting the area of interest to limit the point cloud range, and counting the density information in each grid and using it as the vector representation; secondly, setting a maximum density threshold, when the number of points in the grid exceeds the maximum density threshold, fixing it to this maximum density threshold to suppress the interference of outliers on the vector representation and the model performance; finally, normalizing the number of points in all grids and using it as the model input; The position information processes the spatial consistency of the multi-modal data, aligns it through a unified coordinate system standardization method, and uses it as the model input after normalization processing.
[0011] Preferably, the specific process of extracting and splicing features of the multi-modal vectors to obtain the multi-modal feature vectors is as follows: The multi-modal feature vectors are respectively input into the ResNet34 module and the ResNet18 module for feature extraction, and the multi-modal vectors are concatenated to obtain multi-modal feature vectors.
[0012] Preferably, the Transformer model includes an encoder and a decoder. The encoder performs multi-level abstraction and expression on the input multi-modal feature vectors through the multi-head attention mechanism and the feed-forward neural network, and then combines with the target context through the decoder to output a joint feature vector.
[0013] Furthermore, the specific process of the multi-head attention mechanism is as follows: Project the input multi-modal feature vectors into multiple subspaces, calculate the attention weights according to the multiple subspaces respectively, and establish the relationships between multiple feature dimensions; Generate query, key, and value vectors independently for each attention according to the relationships between multiple feature dimensions, calculate the weights through scaled dot-product attention, and concatenate the multi-head results into a unified output according to the obtained weights; Among them, the scaled dot-product attention scales the multi-head results by measuring the correlation through the dot product of the query and key vectors, which is used to avoid numerical instability caused by high-dimensional features.
[0014] Preferably, the joint feature vectors are subjected to non-linear mapping through a multi-layer perceptron module. The multi-layer perceptron module includes multiple fully connected layers and a Softmax activation function, which is used for dimensionality reduction and outputting a normalized probability distribution; the calculation formula of the Softmax activation function is as follows:
[0015] Where x is a vector of length K, and the output is a probability distribution.
[0016] In a second aspect, the present invention also provides a vehicle-to-everything (V2X) channel prediction system based on multi-modal fusion, including: A multi-modal raw data acquisition module, which is used to construct an urban road scenario of the V2X channel, and obtain known time frame data and unknown time frames to be predicted by combining weather scenarios in the urban road scenario; A data processing and conversion module, which is used to perform data augmentation, data vector representation, and location information processing on the known time frame data and the unknown data to be predicted to obtain multi-modal vectors, and perform feature extraction and concatenation processing on the multi-modal vectors to obtain multi-modal feature vectors; A model processing module, which is used to input the multi-modal feature vectors into the Transformer model, and perform deep fusion between modalities through the multi-head attention mechanism and the feed-forward neural network of the Transformer model to obtain joint feature vectors; A channel prediction module is used to perform non - linear mapping on the joint feature vector through a multi - layer perceptron module, and output a channel index probability distribution, and predict the vehicle - to - everything (V2X) channel based on the channel index probability distribution.
[0017] In a third aspect, the present invention further provides a mobile terminal, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the method for predicting V2X channels based on multi - modal fusion as described above is implemented.
[0018] In a fourth aspect, the present invention further provides a computer - readable storage medium. The computer - readable storage medium stores a computer program, and when the computer program is executed by a processor, the method for predicting V2X channels based on multi - modal fusion as described above is implemented.
[0019] Compared with the prior art, the present invention has the following beneficial technical effects: The present invention provides a method for predicting V2X channels based on multi - modal fusion. By constructing an urban road scenario and combining weather scenarios in the urban road scenario to obtain multi - modal raw data, the efficient fusion ability of the Transformer model and the optimization of the data enhancement strategy are realized under different traffic densities and various weather conditions, improving the robustness of channel prediction in multi - scenarios. The quality of the raw data is significantly improved through the pre - processing process, providing reliable input for the high - precision prediction of the model. Based on the Transformer model, the deep fusion and dynamic modeling of multi - modal data are realized, and the computational efficiency of the model is significantly improved through the parallel computing characteristics, enabling it to quickly process large - scale multi - modal data and providing technical support for real - time channel prediction. The present invention solves the problem of dynamic modeling of wireless channels at the access network (AN) level in network slicing. By constructing an urban road scenario, combining weather scenarios to obtain multi - modal data, and performing data enhancement, feature extraction, and deep fusion, accurate prediction of channel states is achieved. The present invention is connected before and after with access network security authentication, jointly serving the integrity of the slice end - to - end architecture, ensuring the stability and security of network slicing in the access link.
[0020] Furthermore, the present invention performs multi-modal data fusion through a Transformer model, significantly improving the accuracy of channel prediction. Traditional single-modal methods often perform inadequately in complex scenarios. For example, visual information cannot capture depth, lidar data lacks semantic understanding, and position information is also difficult to independently reflect dynamic changes. The Transformer model, through multi-modal data fusion, fully utilizes the advantages of each modality: visual information provides rich spatial details, lidar data supplements depth and speed information, and position information dynamically reflects scene changes. This fusion method overcomes the limitations of single-modal methods, making channel prediction more accurate and reliable.
[0021] Furthermore, the present invention effectively solves the problem of insufficient quality of the original data through a data augmentation strategy, providing higher-quality input for channel prediction. The original multi-modal data often has quality defects. For example, images are vulnerable to illumination and background interference, lidar point clouds have more noise, and position information also lacks consistency. These problems directly affect the effect of model training and prediction accuracy. To address these deficiencies, the present invention adopts a series of optimization measures: improving the quality of images and masking the background to reduce irrelevant interference and enhance image quality; masking, downsampling, and noise filtering lidar point clouds to enhance the robustness of the data; and performing unified coordinate alignment on position information to ensure the spatial consistency of multi-modal data. Through these optimizations, the present invention effectively solves the problem of poor quality of the original data, providing cleaner and more consistent input data for the model.
[0022] Furthermore, the present invention conducts systematic experiments on the proposed channel prediction method under various weather conditions and different traffic flow densities. Traditional channel prediction models are usually tested in idealized environments, ignoring the diverse actual application scenarios, resulting in a significant decline in their performance in complex environments. The present invention designs an experimental environment covering different weather conditions (such as sunny, rainy, snowy, etc.) and traffic densities (such as medium density, high density, etc.), comprehensively verifying the robustness and stability of the model under various actual conditions. The experimental results show that the proposed method can maintain excellent prediction performance in multiple scenarios, fully demonstrating the robustness of the method in complex real-world environments. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] Figure 1 It is a flowchart of the vehicle-to-everything (V2X) channel prediction method based on multi-modal fusion in an embodiment of the present invention; Figure 2 It is a scene diagram of the vehicle-to-infrastructure (V2I) channel prediction in an embodiment of the present invention; Figure 3 It is an algorithm model diagram of the V2I channel prediction in an embodiment of the present invention; Figure 4 It is a diagram showing the multi-modal original data in an embodiment of the present invention; Figure 5 It is a display diagram of RGB image data enhancement processing in an embodiment of the present invention; Figure 6 It is a display diagram of LiDAR point cloud data enhancement processing in an embodiment of the present invention; Figure 7 It is a visualization schematic diagram of the input process from multi-modal data to Transformer in an embodiment of the present invention; Figure 8 It is a schematic diagram of the basic architecture of Transformer and the multi-head attention mechanism in an embodiment of the present invention; Figure 9 It is a schematic diagram of the accuracy value comparison of different experimental schemes under different weather and traffic flows in an embodiment of the present invention; Figure 10 It is a schematic diagram of a vehicle-to-everything (V2X) channel prediction system based on multi-modal fusion in an embodiment of the present invention; In the figure: 1. Multi-modal raw data acquisition module; 2. Data processing and conversion module; 3. Model processing module; 4. Channel prediction module. Detailed implementation manners
[0024] In order to enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0025] The purpose of the present invention is to provide a vehicle-to-everything (V2X) channel prediction method and related devices based on multi-modal fusion to solve the technical problems of limited single-modal perception ability, poor quality of raw data, and insufficient robustness in complex scenarios existing in the prior art.
[0026] The terms involved in the present invention are explained as follows: 1. Transformer model: Self-attention transformation model; 2. RGB image: Red-Green-Blue image; 3. LiDAR point cloud: Light Detection And Ranging point cloud; 4. ResNet module: Residual Network module; 5. GPS: Global Positioning System; 6: IMU: Inertial Measurement Unit; 7: V2I communication: Vehicle-to-Infrastructure communication; 8: Air-Sim: Air-Sim is an open source cross-platform simulator based on a game engine, which can be used for physical and visual simulation of robots such as drones and unmanned vehicles; 9: WaveFarer: WaveFarer is a high-fidelity radar simulator that accounts for multipath and scattering from structures and vehicles in the environment surrounding the radar system, as well as critical atmospheric and scattering effects at frequencies up to and beyond 100GHz; its applications include simulating automotive driving scenarios, indoor sensors, and far-field radar cross sections; 10: Wireless InSite: Wireless InSite is a prediction tool for understanding wireless coverage, channel multipath, and data throughput for 5G, 6G, and WiFi networks. Advanced accelerated 3D ray tracing and fast ray-based surrogacy methods, along with empirical models, enable efficient and accurate prediction of multipath channel characteristics in indoor, urban, and rural environments. Dynamic scenario modeling of vehicles and pedestrians captures time-varying attenuation, while frequency sweeps include broadband effects. The communication analysis feature applies multi-input, multi-output algorithms to channel prediction to estimate the coverage and throughput of wireless networks.
[0027] The present invention is further described in detail below in conjunction with the accompanying drawings: Example 1 See also Figure 1 In one embodiment of the present invention, a vehicle network channel prediction method based on multimodal fusion is provided, comprising: Step 1: construct an urban road scene of the Internet of Vehicles channel, and obtain known time frame data and unknown time frame data to be predicted in the urban road scene in combination with the weather scene; Specifically, according to Figure 2 As shown in the figure, the urban road scenarios of the constructed Internet of Vehicles channel cover experimental environments with different weather (such as sunny, rainy, snowy, etc.) and traffic densities (such as medium density, high density, etc.), which fully verifies the robustness and stability of the model under various practical conditions.
[0028] In the medium-density scenario, the scenario contains 10 cars, 3 buses and 11 RSUs; in the high-density scenario, the number of vehicles increases to 15 cars and 6 buses, while keeping the number of RSUs at 11. Raw data is collected by sensing devices in three weather scenarios of sunny, rainy and snowy, with 1000 time frames for each scenario, including known time frames (75%) and unknown data to be predicted (25%). Each known time frame contains the signal-to-noise ratio of the first three channel index values of V2I communication, the sequence number of the communication signal time frame, the position coordinates of the vehicle and the RSU, the RGB image, and the LiDAR point cloud. Each unknown time frame to be predicted only contains the sequence number of the communication signal time frame, the position coordinates of the vehicle and the RSU, the RGB image, and the LiDAR point cloud, as Figure 4 shown.
[0029] Step 2: Perform data augmentation, data vector representation, and position information processing on the known time frame data and the unknown data to be predicted to obtain multi-modal vectors, and perform feature extraction and splicing processing on the multi-modal vectors to obtain multi-modal feature vectors; Specifically, according to Figure 3 shown, in the preprocessing of the multi-modal raw data to obtain multi-modal augmented data, the preprocessing process includes image augmentation, image vector representation, point cloud augmentation, point cloud vector representation, and preprocessing of position information.
[0030] Specifically, according to Figure 5 shown, the specific process of image augmentation is as follows: Image improvement techniques include a variety of methods such as brightness adjustment, contrast enhancement, gamma correction, Gaussian blur, and background masking. These processing means can significantly improve the clarity and consistency of the image under different weather conditions (such as sunny, rainy, and snowy days). At the same time, the main purpose of background masking is to cover buildings and other irrelevant background information to reduce the interference of invalid features on the model, thereby enhancing the model's attention to roads, vehicles, and dynamic traffic targets. If a pixel area is completely within the masked range, the area is directly cropped to further reduce the computational burden and optimize the data quality.
[0031] Specifically, the specific process of image vector representation is as follows: The picture data is vector-represented in the RGB three-channel format and used as the model input after normalization. The data of each channel is normalized to reduce the impact of numerical differences between different channels on model training, thereby improving the stability and effect of training. During the normalization process, each channel is processed with the commonly used fixed values of [0.229, 0.224, 0.225]. This method not only adapts to common models but also provides a more stable input data distribution for subsequent feature extraction. The processed RGB data is input into the model as the basic data representation for further feature learning, effectively supporting the channel prediction task in complex scenarios.
[0032] The specific process of point cloud enhancement is as follows: Background filtering maps the building positions into the point cloud data by referring to the 2D bird's-eye view and removes the point clouds in these areas, ensuring that the remaining point cloud information is mainly concentrated in key areas such as vehicle movement and road structure, thereby effectively reducing the interference of static objects to the model. In addition, in order to reduce the data scale while maintaining the integrity of the geometric features of the point cloud, the research introduces a downsampling operation to reduce the point cloud density by 20%, improve the data processing efficiency and reduce the computational overhead. At the same time, a random noise perturbation between -0.4 and 0.4 is appropriately introduced to simulate the measurement uncertainty in the actual scenario and enhance the robustness of the model, as Figure 6 shown.
[0033] The specific process of point cloud vector representation is as follows: The area of interest is divided into grids to achieve an effective vector representation of the point cloud data. The point cloud range is restricted by setting the area of interest, and the density information within each grid is counted and used as the vector representation. During this process, to prevent excessive local points from causing outliers, a maximum density threshold is set. When the number of points in the grid exceeds the threshold, it is fixed at this threshold to suppress the interference of outliers to the vector representation and the model performance. Finally, the number of points in all grids is normalized to ensure that the vector representation has a consistent data distribution, thereby improving the convergence and prediction performance of the model. The above steps effectively convert the point cloud data into usable vectors, laying a data foundation for subsequent multi-modal fusion and dynamic modeling.
[0034] The preprocessing of the location information focuses on the spatial consistency of multi-modal data and is aligned through a unified coordinate system standardization method. Specifically, the positions of the vehicle and the RSU are mapped into the coordinate system with the RSU as the origin, and the coordinates are normalized in combination with the farthest detection distance of the sensing device. Normalization is achieved by dividing the actual coordinate value by the farthest distance of the sensing device, making the data distributed within the range of 0 to 1 and eliminating the influence of the absolute coordinate value on model training.
[0035] After all multi-modal enhanced data is converted into vector representations, the RGB image vector is input into the ResNet34 module, and the LiDAR point cloud vector is input into the ResNet18 module for feature extraction, and then the multi-modal vectors are concatenated. Through the residual connection mechanism, the ResNet module can effectively capture local and global characteristics, and at the same time gradually compress the high-dimensional feature representation to make it more compact and have strong semantic expression ability, as Figure 7 shown.
[0036] Step 3, input the multi-modal feature vectors into the Transformer model, and perform in-depth fusion between modalities through the multi-head attention mechanism and the feed-forward neural network of the Transformer model to obtain the joint feature vectors, as Figure 8 shown; Specifically, according to Figure 3 As shown, the Transformer first linearly projects the input features to generate query vectors (Q), key vectors (K), and value vectors (V). Subsequently, the attention weights are calculated through the dot product operation between the query vectors and the key vectors, and the value vectors are weighted and summed based on the weights, thereby realizing the deep fusion of modal features and the aggregation of associated information. This multi-modal fusion process is iterated 4 times in the Transformer model, and each iteration further improves the abstract expression ability of the features and the depth of interactive information. Finally, the fused RGB images, LiDAR point clouds, and coordinate features are represented as a 512×1 joint feature vector, forming an enhanced multi-modal feature representation.
[0037] Step 4, the joint feature vector is non-linearly mapped through a multi-layer perceptron module, and the channel index probability distribution is output, and the vehicle-to-everything (V2X) channel is predicted based on the channel index probability distribution.
[0038] Specifically, according to Figure 3 As shown, in the channel prediction stage, the joint feature vector is non-linearly mapped through a multi-layer perceptron (MLP) module. The MLP module includes multiple fully connected layers and a Softmax activation function for dimensionality reduction and outputting a normalized probability distribution.
[0039] The calculation formula of the Softmax activation function includes:
[0040] where x is a vector of length K, and the output is a probability distribution.
[0041] Finally, the algorithm maps the joint feature vector to a 128×1 channel index probability distribution, representing the prediction result of the target channel. The specific meaning of the channel index is to evenly divide the 180-degree plane at the RSU and road level into 128 parts, and the channel index represents one of the directions. For example, the index values of 0 and 127 represent the leftmost and rightmost directions respectively. The channel index can accurately represent the communication direction and further improve the performance of the communication system.
[0042] To verify the beneficial effects of the proposed Transformer model in the channel prediction task, a series of simulation experiments were carried out based on a public dataset The proposed method in this embodiment was experimentally compared with Scheme 1 that only uses location information and Scheme 2 that uses location information and image data.
[0043] Data source: The dataset uses Air-Sim and WaveFarer to collect multi-modal perception data and Wireless InSite to collect communication data. The dataset realizes a communication-aware integrated system by deeply integrating and precisely aligning AirSim, WaveFarer, and WirelessInSite. In addition, The dataset also covers various weather conditions, multiplexing frequency bands, and different times of the day.
[0044] Hardware environment and software platform: The hardware environment uses an Intel Xeon W-3235 12-core processor and an RTX 3090 GPU, and the programming language used is Python 3.9.
[0045] In terms of dataset selection and experimental settings, in order to ensure the reliability of experimental results and the scientific nature of model evaluation, the dataset adopts a division strategy of training set, validation set, and test set. Specifically, the training set accounts for 72%, and the validation set and test set account for 8% and 20% respectively. In the model training stage, the training set is used to optimize the model, and the validation set is used to monitor the model performance and perform hyperparameter tuning to avoid overfitting. In the testing stage, the generalization performance of the model is finally evaluated through the test set, so as to verify its prediction ability on unseen data.
[0046] To simulate the actual vehicle networking scenario, different vehicle densities and RSU distributions are set to reflect the typical urban road communication environment. In the medium-density scenario, the scenario contains 10 cars, 3 buses, and 11 RSUs; in the high-density scenario, the number of vehicles increases to 15 cars and 6 buses, while keeping the number of RSUs at 11. Through these two different density scenario settings, the adaptability of the model under different traffic conditions can be comprehensively evaluated. In addition, to simulate real road conditions, two weather scenarios of rainy day and snowy day are also introduced. In the rainy day scenario, the rainfall is set to 50 mm / hr and the road humidity is 100%; in the snowy day scenario, the snowfall is set to 10 mm / hr and the road humidity is also 100%. At the same time, to more accurately reflect the impact of bad weather on channel propagation, the experiment also models the physical space and electromagnetic space parameters under different weather conditions. For example, in rainy and cloudy weather, the ambient temperature is set to 22.2 °C and the humidity is 100%; in the snowy day scenario, the temperature is set to -10 °C and the humidity is 20%.
[0047] In the communication parameter settings, the experiment adopted an integrated design of communication and sensing to ensure the accuracy of channel prediction and the efficiency of model training. Specifically, each RSU and vehicle were respectively equipped with 128 and 32 antenna units to support high-precision beamforming and channel prediction. In terms of frequency band selection, the experiment simultaneously simulated the communication performance of the Sub-6 GHz band and the millimeter wave band. In the Sub-6 GHz band, the carrier frequency was set to 5.9 GHz and the communication bandwidth was 20 MHz; in the millimeter wave band, the carrier frequency was set to 28 GHz and the communication bandwidth was 2 GHz. Through the comprehensive comparison of the communication performance of different frequency bands, the experiment further verified the adaptability and performance of the model in a multi-band environment.
[0048] The summary of the simulation experiment parameter settings is shown in Table 1: Table 1 Channel Prediction Simulation Parameter Settings Based on the Transformer Model
[0049] According to Figure 9 As shown, the comparison of the channel prediction accuracy of the three schemes under different weather and traffic flow scenarios, including Comparative Scheme 1, Comparative Scheme 2, and the multi-modal method of the present invention. The experimental results show that Scheme 1 that only uses location information has limited performance in complex scenarios. For example, the accuracy in the snow day medium density scenario is only 0.58, and it is only 0.63 in the sunny day high density scenario. When combined with camera features, the accuracy of Scheme 2 has improved. For example, it has increased from 0.69 to 0.78 in the sunny day medium density scenario. The scheme proposed in this paper not only realizes the deep fusion of features based on multi-modal data, but also combines targeted data augmentation strategies, enabling the model to show significant performance advantages in all scenarios. In the sunny day high density scenario, the accuracy reaches 0.91, and in the snow day medium density scenario, the accuracy is even improved to 0.98. These results fully demonstrate the robustness and superiority of the V2I channel prediction strategy based on the Transformer model proposed in this paper under different traffic flows and adverse weather conditions, which is superior to other existing schemes.
[0050] Table 2 shows the time performance of channel prediction based on the Transformer model under different input data types and augmentation strategies. Through the analysis of input data types and model running time, the real-time performance of the model in multi-modal data fusion and expansion scenarios can be comprehensively evaluated. First, for the original input data, which includes vehicle position, RSU position, a single RGB image, and a LiDAR point cloud data, the size of a single sample is 5.31 MB. Under this condition, the training time for a single sample of the model is 0.399 ms, and the inference time for a single sample is 2.247 ms. In contrast, the augmented input data includes more RGB images (8) and LiDAR point cloud data (2), and the size of a single sample is reduced to 4.14 MB. This data compression mainly benefits from the optimization of information redundancy by the data augmentation strategy, while reducing the sample complexity. Further analyze the impact of data augmentation on the model's time performance. After data augmentation, the training time for a single sample slightly increases from 0.399 ms of the original data to 0.466 ms, while the inference time for a single sample decreases from 2.247 ms to 2.073 ms. This phenomenon indicates that the data augmentation strategy significantly improves the model inference efficiency while slightly increasing the training time, reducing the computational overhead during inference, especially showing good performance in high-dynamic scenarios that require real-time prediction. The experimental results show that the data augmentation strategy effectively reduces the inference time by optimizing the input data structure while ensuring data quality, providing strong support for the efficient real-time prediction of the model in multi-modal data scenarios.
[0051] Table 2 Time Performance of the Channel Prediction Model Based on the Transformer Model
[0052] In summary, the experimental results verify that this embodiment effectively overcomes the problems of multi-modal heterogeneity and channel dynamics by fully exploiting the spatio-temporal correlation and semantic features in multi-modal data, achieving high-precision and high-robustness channel prediction.
[0053] In summary, in this embodiment, the Transformer model realizes the deep fusion and dynamic modeling of multi-modal data through its multi-head attention mechanism, position encoding ability, and parallel computing characteristics. Among them, the multi-head attention mechanism can effectively capture the spatial characteristics of visual information, the depth and velocity characteristics of lidar data, and the dynamic change characteristics of position data, providing a basis for the efficient fusion of multi-modal data. The position encoding ability enhances the model's modeling ability for dynamic temporal scenarios by introducing sequence position information, enabling it to better adapt to the changes in complex scenarios. The parallel computing characteristics significantly improve the computational efficiency of the model, enabling it to quickly process large-scale multi-modal data, providing technical support for real-time channel prediction.
[0054] For image data, the input quality of RGB images is optimized through quality improvement and background masking techniques, and vector representation is performed in the RGB three-channel format. After normalization, it is used as the input to the model. For LiDAR point cloud data, strategies such as background filtering, downsampling, and random noise perturbation are adopted, and the region of interest is divided into grids in the vector representation process to achieve an effective vectorized representation of the point cloud data. The position information is aligned through a unified coordinate system standardization method to ensure the spatial consistency of multi-modal data. These data augmentation strategies significantly improve the quality of the original data and provide reliable input for the high-precision prediction of the model.
[0055] This embodiment exhibits excellent robustness in multiple scenarios, mainly due to the efficient fusion ability of the Transformer model and the optimization of data augmentation strategies. The experiments cover different traffic densities (such as medium density, high density, etc.) and various weather conditions (such as sunny, rainy, snowy, etc.). When the traffic density is high, the channel prediction accuracy is likely to decrease due to vehicle occlusion. However, through multi-modal data fusion in this embodiment, the complementary advantages of visual, lidar, and position information are fully utilized to effectively compensate for the vehicle occlusion problem. At the same time, for adverse weather, the data augmentation strategy significantly improves the quality of the input data through measures such as optimizing image brightness, contrast, and lidar point cloud noise reduction, enabling the model to maintain stable performance in complex weather environments. These technical advantages together ensure the high robustness of this embodiment in multiple scenarios.
[0056] Embodiment 2 According to Figure 10 as shown, the present invention also provides a vehicle-to-everything (V2X) channel prediction system based on multi-modal fusion, including: A multi-modal raw data acquisition module 1 for constructing an urban road scenario of a V2X channel and obtaining known time frame data and unknown time frames to be predicted in combination with a weather scenario in the urban road scenario; A data processing and conversion module 2 for performing data augmentation, data vector representation, and position information processing on the known time frame data and unknown data to be predicted to obtain multi-modal vectors, and performing feature extraction and splicing processing on the multi-modal vectors to obtain multi-modal feature vectors; A model processing module 3 for inputting the multi-modal feature vectors into a Transformer model and performing deep fusion between modalities through the multi-head attention mechanism and feed-forward neural network of the Transformer model to obtain joint feature vectors; A channel prediction module 4 for performing non-linear mapping on the joint feature vectors through a multi-layer perceptron module and outputting a channel index probability distribution to predict the V2X channel based on the channel index probability distribution.
[0057] Embodiment 3 The present invention also provides a mobile terminal, including a memory, a processor, and a computer program stored in the memory and executable on the processor, such as a vehicle-to-internet channel prediction program based on multimodal fusion.
[0058] When the processor executes the computer program, the vehicle-to-internet channel prediction method based on multimodal fusion described above is implemented, for example: Construct the urban road scenario of the vehicle-to-internet channel, and obtain the known time frame data and the unknown time frame data to be predicted by combining the weather scenario in the urban road scenario; Perform data enhancement, data vector representation, and location information processing on the known time frame data and the unknown data to be predicted to obtain a multimodal vector, and perform feature extraction and splicing processing on the multimodal vector to obtain a multimodal feature vector; Input the multimodal feature vector into the Transformer model, and perform deep fusion between modalities through the multi-head attention mechanism and the feed-forward neural network of the Transformer model to obtain a joint feature vector; Perform non-linear mapping on the joint feature vector through a multi-layer perceptron module, and output the channel index probability distribution, and predict the vehicle-to-internet channel based on the channel index probability distribution.
[0059] Alternatively, when the processor executes the computer program, the functions of each module in the above system are implemented, for example: Multimodal raw data acquisition module 1, which is used to construct the urban road scenario of the vehicle-to-internet channel, and obtain the known time frame data and the unknown time frame data to be predicted by combining the weather scenario in the urban road scenario; Data processing and conversion module 2, which is used to perform data enhancement, data vector representation, and location information processing on the known time frame data and the unknown data to be predicted to obtain a multimodal vector, and perform feature extraction and splicing processing on the multimodal vector to obtain a multimodal feature vector; Model processing module 3, which is used to input the multimodal feature vector into the Transformer model, and perform deep fusion between modalities through the multi-head attention mechanism and the feed-forward neural network of the Transformer model to obtain a joint feature vector; Channel prediction module 4, which is used to perform non-linear mapping on the joint feature vector through a multi-layer perceptron module, and output the channel index probability distribution, and predict the vehicle-to-internet channel based on the channel index probability distribution.
[0060] Exemplarily, the computer program may be divided into one or more modules / units, which are stored in the memory and executed by the processor to implement the present invention. The one or more modules / units may be a series of computer program instruction segments capable of performing specific functions, and these instruction segments are used to describe the execution process of the computer program in the mobile terminal.
[0061] For example, the computer program may be divided into a multi-modal raw data acquisition module 1, a data processing and conversion module 2, a model processing module 3, and a channel prediction module 4; The specific functions of each module are as follows: The multi-modal raw data acquisition module 1 is used to construct the urban road scenario of the vehicle-to-everything (V2X) channel, and obtain the known time frame data and the unknown time frame data to be predicted by combining the weather scenario in the urban road scenario; The data processing and conversion module 2 is used to perform data augmentation, data vector representation, and location information processing on the known time frame data and the unknown data to be predicted to obtain multi-modal vectors, and perform feature extraction and splicing processing on the multi-modal vectors to obtain multi-modal feature vectors; The model processing module 3 is used to input the multi-modal feature vectors into the Transformer model, and perform deep fusion between modalities through the multi-head attention mechanism and the feed-forward neural network of the Transformer model to obtain joint feature vectors; The channel prediction module 4 is used to perform non-linear mapping on the joint feature vectors through a multi-layer perceptron module, and output the channel index probability distribution, and predict the vehicle-to-everything (V2X) channel based on the channel index probability distribution.
[0062] The mobile terminal may be a computing device such as a desktop computer, a notebook, a palm computer, and a cloud server. The mobile terminal may include, but is not limited to, a processor and a memory.
[0063] The processor may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The processor is the control center of the mobile terminal and connects various parts of the entire mobile terminal through various interfaces and circuits.
[0064] The memory can be used to store the computer programs and / or modules. By running or executing the computer programs and / or modules stored in the memory, and by calling the data stored in the memory, the processor implements various functions of the mobile terminal.
[0065] The memory may mainly include a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function (such as a sound playback function, an image playback function, etc.); the data storage area can store data created according to the use of the mobile phone (such as audio data, phone book, etc.). In addition, the memory may include high-speed random access memory, and may also include non-volatile memory, such as a hard disk, memory, plug-in hard disk, Smart Media Card (SMC), Secure Digital (SD) card, Flash Card, at least one magnetic disk storage device, flash device, or other volatile solid-state storage devices.
[0066] Embodiment 4 The present invention also provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, it implements the method for predicting a vehicle networking channel based on multimodal fusion.
[0067] If the modules / units integrated in the mobile terminal are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium.
[0068] Based on such understanding, all or part of the processes in the above method of the present invention can also be completed by instructing relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the above-mentioned aggregated reinforcement learning resource scheduling method can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc.
[0069] The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electrical carrier signal, telecommunication signal, and software distribution medium, etc.
[0070] It should be noted that the content included in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.
[0071] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the above embodiments, those of ordinary skill in the art should understand that: modifications or equivalent replacements can still be made to the specific implementation manners of the present invention, and any modification or equivalent replacement without departing from the spirit and scope of the present invention should be covered within the protection scope of the present invention.
Claims
1. A vehicle networking channel prediction method based on multimodal fusion, characterized in that Including: Construct the urban road scenario of the vehicle networking channel, and obtain the known time frame data and the unknown to-be-predicted time frame data by combining the weather scenario in the urban road scenario; Perform data augmentation, data vector representation, and location information processing on the known time frame data and the unknown to-be-predicted data to obtain a multi-modal vector, and perform feature extraction and splicing processing on the multi-modal vector to obtain a multi-modal feature vector; Input the multi-modal feature vector into the Transformer model, and perform deep fusion between modalities through the multi-head attention mechanism and the feed-forward neural network of the Transformer model to obtain a joint feature vector; Perform non-linear mapping on the joint feature vector through a multi-layer perceptron module, and output the channel index probability distribution, and predict the vehicle networking channel based on the channel index probability distribution.
2. The method for predicting a vehicle network channel based on multimodal fusion according to claim 1, wherein In the process of obtaining the known time frame data and the unknown to-be-predicted time frame data by combining the weather scenario in the urban road scenario, the weather scenario includes sunny scenario, rainy scenario, and snowy scenario; The known time frame data includes channel index value, communication signal time frame serial number, position coordinates of vehicles and RSUs, RGB images, and LiDAR point clouds; The unknown to-be-predicted time frame data includes communication signal time frame serial number, position coordinates of vehicles and RSUs, RGB images, and LiDAR point clouds.
3. A vehicle networking channel prediction method based on multimodal fusion according to claim 1, wherein, The data augmentation includes image augmentation and point cloud augmentation; Among them, the image augmentation includes brightness adjustment, contrast enhancement, gamma correction, Gaussian blur, and background masking; The process of point cloud augmentation includes background filtering by referring to a two-dimensional bird's-eye view, mapping the building positions into the point cloud data, removing the point clouds in the building position areas, and concentrating the remaining point clouds in the areas of vehicle movement and road structure; The data vector representation includes image vector representation and point cloud vector representation; Among them, the process of image vector representation includes vectorizing the picture data in the RGB three-channel format and using it as the model input after normalization processing; The process of point cloud vector representation includes first dividing the area of interest into grids to achieve effective vectorization representation of the point cloud data; setting the area of interest to limit the point cloud range, and counting the density information in each grid and using it as the vector representation; secondly, setting a maximum density threshold, and when the number of points in the grid exceeds the maximum density threshold, fixing it to this maximum density threshold; finally, normalizing the number of points in all grids and using it as the model input; The location information processes the spatial consistency of the multi-modal data, aligns it through a unified coordinate system standardization method, and uses it as the model input after normalization processing.
4. A vehicle networking channel prediction method based on multimodal fusion according to claim 1, wherein The specific process of performing feature extraction and splicing processing on the multi-modal vector to obtain a multi-modal feature vector is as follows: Input the multi-modal feature vector into the ResNet34 module and the ResNet18 module respectively for feature extraction, and splice the multi-modal vectors to obtain a multi-modal feature vector.
5. A method for predicting a vehicle network channel based on multimodal fusion according to claim 1, characterized in that The Transformer model includes an encoder and a decoder. The encoder performs multi-level abstraction and expression on the input multi-modal feature vectors through the multi-head attention mechanism and the feed-forward neural network, and then combines with the target context through the decoder to output a joint feature vector.
6. The method for predicting a vehicle networking channel based on multimodal fusion according to claim 5, wherein The specific process of the multi-head attention mechanism is as follows: Project the input multi-modal feature vectors into multiple subspaces, calculate the attention weights respectively according to the multiple subspaces, and establish the relationships between multiple feature dimensions; Generate query, key, and value vectors independently for each attention according to the relationships between multiple feature dimensions, calculate the weights through scaled dot-product attention, and splice the multi-head results into a unified output according to the obtained weights; Among them, the scaled dot-product attention scales the multi-head results by measuring the correlation through the dot product of the query and key vectors.
7. A method for predicting vehicle - to - everything (V2X) channel based on multimodal fusion according to claim 1, characterized in that, The joint feature vector is non-linearly mapped through a multi-layer perceptron module. The multi-layer perceptron module includes multiple fully connected layers and a Softmax activation function, and is used for dimensionality reduction and outputting a normalized probability distribution; The calculation formula of the Softmax activation function includes: where x is a vector of length K, and the output is a probability distribution.
8. A vehicle networking channel prediction system based on multimodal fusion, characterized in that, includes: A multi-modal raw data acquisition module, which is used to construct the urban road scene of the vehicle-to-everything (V2X) channel, and obtain the known time frame data and the unknown time frame data to be predicted by combining the weather scene in the urban road scene; A data processing and conversion module, which is used to perform data augmentation, data vector representation, and location information processing on the known time frame data and the unknown data to be predicted to obtain multi-modal vectors, and perform feature extraction and splicing processing on the multi-modal vectors to obtain multi-modal feature vectors; A model processing module, which is used to input the multi-modal feature vectors into the Transformer model, and perform deep fusion between modalities through the multi-head attention mechanism and the feed-forward neural network of the Transformer model to obtain a joint feature vector; A channel prediction module, which is used to perform non-linear mapping on the joint feature vector through a multi-layer perceptron module, and output a channel index probability distribution, and predict the vehicle-to-everything (V2X) channel according to the channel index probability distribution.
9. A mobile terminal, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the vehicle-to-everything (V2X) channel prediction method based on multi-modal fusion according to any one of claims 1-7.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the vehicle-to-everything (V2X) channel prediction method based on multi-modal fusion according to any one of claims 1-7.
Citation Information
Patent Citations
Wireless environment reconstruction method driven by multi-modal data
CN117579203A
Internet of vehicles channel path loss prediction method, system and device, and storage medium
CN118155167A
Point cloud segmentation method and system, medium, equipment and information data processing terminal
CN118967698A
Transform and end point induction-based multi-modal trajectory prediction method for automatic driving vehicle
CN119975390A
Road monitoring multi-mode sensing method and system adapting to dynamic environment
CN120047899A
Cited By
Internet of vehicles sensing data transmission method and device
CN120935533A
An internet of vehicles perception data transmission method and device
CN120935533B
Vehicle condition estimation method, device and equipment based on multi-modal data fusion, storage medium and product
CN121010770A
Enhanced sense integration method and system fused with machine vision
CN121037783A