Internet of vehicles beam prediction method and device based on multi-modal large model

By combining multimodal large models with vehicle camera images and positioning data, ray tracing and scatterer projection models are constructed to predict vehicle-to-everything (V2X) beam index and coherence time. This solves the problem of insufficient accuracy and robustness of traditional beam training methods in complex environments, and improves communication efficiency and stability.

CN120165733BActive Publication Date: 2026-01-06XIAMEN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510212337.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-25
Publication Date
2026-01-06
Estimated Expiration
2045-02-25

AI Technical Summary

Technical Problem

Traditional vehicle-to-everything (V2X) beamforming methods face challenges such as insufficient positioning accuracy, lack of robustness, and beam inaccuracy in fast-moving and complex environments, resulting in limitations on communication accuracy and transmission rate.

Method used

A multimodal large model-based approach is adopted. By acquiring vehicle camera images, positioning data, and road segment 3D maps, ray tracing and scatterer projection models are constructed. Combined with channel characteristics and positioning information, beam index and beam coherence time are predicted, and a beam prediction model is constructed to improve beam accuracy and communication efficiency.

Benefits of technology

It enhances beam accuracy in vehicle-to-everything (V2X) environments, improves communication transmission rates and energy efficiency, enhances system adaptability and robustness, reduces pilot overhead, and ensures stable communication under both line-of-sight and non-line-of-sight paths.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120165733B_ABST
    Figure CN120165733B_ABST
Patent Text Reader

Abstract

Embodiments of the present application provide a vehicle networking beam prediction method and device based on a multi-modal large model. The method comprises: obtaining a camera image of a target vehicle, positioning data, base station positions of a road section and a three-dimensional road map of the road section; inputting the positioning data, the base station positions and the three-dimensional road map into a ray tracing model to obtain propagation characteristics of a wireless communication signal; inputting the propagation characteristics and the camera image into a scatterer projection model to obtain scatterer characteristics in the image; optimizing a multi-modal large model according to a channel characteristic data set containing the camera image and the scatterer characteristics, and constructing a scatter heat map, and then obtaining a beam prediction data set for training of a beam prediction model for training. The technical solution of the present application can effectively extract channel characteristics in the camera image, accurately predict beam indexes and beam coherence times in a codebook by combining the channel characteristics and positioning information, and improve the communication transmission rate and energy efficiency of the system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the fields of communication and computer technology, and more specifically, to a method and apparatus for beam prediction in vehicle networking based on a multimodal large model. Background Technology

[0002] With the development of applications such as autonomous driving and vehicle-to-everything (V2X) communication, the demand for data communication speeds in vehicles is becoming increasingly stringent, and millimeter-wave communication can enhance the communication capabilities of vehicles. Beamforming technology, as one of the key technologies in millimeter-wave communication, utilizes directional beam focusing to transmit signals, thereby increasing the signal power at the receiving end and ensuring communication quality.

[0003] However, the rapid mobility and complex environment (such as buildings, obstacles, and other vehicles) in the Internet of Vehicles (IoV) present numerous challenges to traditional beam training methods. Current vision-assisted beam prediction solutions, which utilize the relative position information of vehicles obtained from onboard camera images and combine it with positioning information to assist base station beam prediction, address the issues of insufficient positioning accuracy and the susceptibility of base station cameras to obstruction, effectively reducing beam training overhead. However, this end-to-end beam prediction method based on onboard images often suffers from insufficient accuracy and robustness in situations with heavy traffic and base station obstruction. Furthermore, a fixed beam coherence time (the time during which the beam remains constant) can lead to prolonged beam inaccuracies or significant transmission rate losses in highly mobile and dynamic IoV scenarios. Summary of the Invention

[0004] The embodiments of this application provide a beam prediction method and apparatus for vehicle-to-everything (V2X) networks based on a multimodal large model. This method can effectively extract channel features from vehicle camera images to at least a certain extent, and accurately predict beam indexes and beam coherence times in the codebook by combining channel features and positioning information. This addresses the shortcomings of insufficient adaptability and robustness in different V2X environments, enhances beam accuracy under line-of-sight and non-line-of-sight switching in millimeter-wave communication, and improves the communication transmission rate and energy efficiency of the system.

[0005] Other features and advantages of this application will become apparent from the following detailed description, or may be learned in part from practice of this application.

[0006] According to one aspect of the embodiments of this application, a beam prediction method for vehicle-to-everything (V2X) networks based on a multimodal large model is provided, comprising:

[0007] Acquire camera images, location data, base station locations on the road segment where the target vehicle is located, and a 3D map of the road segment;

[0008] The target vehicle's location data, base station location, and road segment 3D map are used as inputs to a pre-built ray tracing model, so that the ray tracing model outputs the propagation characteristics of the target vehicle's wireless communication signal during driving.

[0009] The propagation characteristics of the wireless communication signal and the camera image are input into a pre-constructed scatterer projection model to obtain the scatterer features in the image;

[0010] The pre-built multimodal large model is trained under supervision based on the channel feature dataset composed of the camera images and the scatterer features to minimize the prediction error of the scatterer coordinates and intensity, thereby obtaining the channel estimation large model.

[0011] A scattering heatmap is constructed based on the channel feature dataset, and the scattering heatmap, localization, and triplet data composed of the optimal beam are segmented according to the beam coherence time to obtain a beam prediction dataset.

[0012] The pre-built beam prediction model is trained based on the beam prediction dataset to obtain the target beam prediction model for use in vehicle-to-everything (V2X) beam prediction.

[0013] According to one aspect of the embodiments of this application, a vehicle-to-everything (V2X) beam prediction device based on a multimodal large model is provided, comprising:

[0014] The acquisition module is used to acquire camera images, location data, base station locations on the road segment, and a 3D map of the road segment of the target vehicle.

[0015] The ray tracing module is used to take the target vehicle's positioning data, base station location, and road segment 3D map as input to a pre-built ray tracing model, so that the ray tracing model outputs the propagation characteristics of the target vehicle's wireless communication signal during driving.

[0016] The projection module is used to input the propagation characteristics of the wireless communication signal and the camera image into a pre-constructed scatterer projection model to obtain the scatterer features in the image;

[0017] The optimization module is used to supervise the training of a pre-built multimodal large model based on the channel feature dataset composed of the camera images and the scatterer features, so as to minimize the prediction error of the scatterer coordinates and intensity and obtain the channel estimation large model.

[0018] The construction module is used to construct a scattering heatmap based on the channel feature dataset, and to segment the scattering heatmap, localization and the triplet data composed of the optimal beam according to the beam coherence time to obtain the beam prediction dataset.

[0019] The processing module is used to train a pre-built beam prediction model based on the beam prediction dataset to obtain a target beam prediction model for use in vehicle-to-everything (V2X) beam prediction.

[0020] According to one aspect of the embodiments of this application, a computer-readable medium is provided having a computer program stored thereon, which, when executed by a processor, implements the vehicle-to-everything (V2X) beam prediction method based on a multimodal large model as described in the above embodiments.

[0021] According to one aspect of the embodiments of this application, an electronic device is provided, including: one or more processors; and a storage device for storing one or more programs, which, when executed by the one or more processors, cause the one or more processors to implement the vehicle-to-everything (V2X) beam prediction method based on a multimodal large model as described in the above embodiments.

[0022] According to one aspect of the embodiments of this application, a computer program product or computer program is provided, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the vehicle-to-everything (V2X) beam prediction method based on a multimodal large model provided in the above embodiments.

[0023] In some embodiments of this application, the technical solutions are obtained by acquiring camera images, positioning data, base station locations on the road segment, and a 3D map of the road segment of the target vehicle. The positioning data, base station locations, and 3D map of the road segment of the target vehicle are used as inputs to a pre-built ray tracing model, so that the ray tracing model outputs the propagation characteristics of the wireless communication signal of the target vehicle during its travel on the road segment. The propagation characteristics of the wireless communication signal and the camera image are input into a pre-built scatterer projection model to obtain the scatterer features in the image. The pre-built multimodal large model is optimized based on the channel feature dataset composed of the camera image and the scatterer features to obtain a channel estimation large model. A scattering heatmap is constructed based on the channel feature dataset, and the triplet data composed of the scattering heatmap, positioning, and optimal beam is segmented according to the beam coherence time to obtain a beam prediction dataset. The pre-built beam prediction model is trained based on the beam prediction dataset to obtain a target beam prediction model for use in vehicle-to-everything (V2X) beam prediction. In this way, the scatterer features in the vehicle camera images can be effectively extracted, and the beam index and beam coherence time in the codebook can be accurately predicted by combining the scatterer features and positioning information. This solves the shortcomings of insufficient adaptability and robustness in different vehicle networking environments, and improves the communication transmission rate and energy efficiency of the system.

[0024] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and do not limit this application. Attached Figure Description

[0025] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application. It is obvious that the drawings described below are merely some embodiments of this application, and those skilled in the art can obtain other drawings based on these drawings without any inventive effort. In the drawings:

[0026] Figure 1 A flowchart illustrating a beam prediction method for vehicle-to-everything (V2X) networks based on a multimodal large model according to an embodiment of this application is shown.

[0027] Figure 2 A schematic diagram of signal ray propagation according to an embodiment of this application is shown;

[0028] Figure 3 A schematic diagram of a scattering point projection according to an embodiment of this application is shown;

[0029] Figure 4 A schematic diagram of the structure of a multimodal large model according to an embodiment of this application is shown;

[0030] Figure 5 A schematic diagram of beam coherence time according to an embodiment of this application is shown;

[0031] Figure 6 A schematic diagram of the structure of a beam prediction model according to an embodiment of this application is shown;

[0032] Figure 7 A block diagram of a vehicle-to-everything (V2X) beam prediction device based on a multimodal large model according to an embodiment of this application is shown;

[0033] Figure 8 A schematic diagram of the structure of a computer system suitable for implementing the electronic device of the present application is shown. Detailed Implementation

[0034] Exemplary embodiments will now be described more fully with reference to the accompanying drawings. However, these exemplary embodiments can be implemented in many forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided to make this application more comprehensive and complete, and to fully convey the concept of the exemplary embodiments to those skilled in the art.

[0035] Furthermore, the described features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. Numerous specific details are provided in the following description to give a thorough understanding of embodiments of this application. However, those skilled in the art will recognize that the technical solutions of this application can be practiced without one or more of the specific details, or other methods, components, apparatuses, steps, etc., can be employed. In other instances, well-known methods, apparatuses, implementations, or operations are not shown or described in detail to avoid obscuring various aspects of this application.

[0036] The block diagrams shown in the accompanying drawings are merely functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities can be implemented in software, in one or more hardware modules or integrated circuits, or in different network and / or processor devices and / or microcontroller devices.

[0037] The flowcharts shown in the accompanying drawings are merely illustrative and do not necessarily include all content and operations / steps, nor do they necessarily have to be performed in the described order. For example, some operations / steps can be broken down, while others can be combined or partially combined; therefore, the actual execution order may change depending on the specific circumstances.

[0038] Figure 1 A flowchart illustrating a beam prediction method for vehicle-to-everything (V2X) networks based on a multimodal large model according to an embodiment of this application is shown.

[0039] Please refer to Figure 1 The beam prediction method for vehicle-to-everything (V2X) networks based on a multimodal large model includes at least steps S110 to S160, which are detailed below:

[0040] In step S110, the camera image, positioning data, base station location of the road segment, and 3D map of the road segment of the target vehicle are acquired.

[0041] In this embodiment, the millimeter-wave base station can be deployed along roadsides, such as streetlights and utility poles. The camera images can be acquired by an onboard camera on the target vehicle, and these images can include multi-view images with non-overlapping perspectives, such as images from the front and rear of the target vehicle.

[0042] The target vehicle's positioning data may include latitude and longitude, as well as the vehicle's inertial acceleration and axial acceleration. Latitude and longitude can be obtained through GPS or the BeiDou satellite positioning system, while the vehicle's inertial acceleration and axial acceleration can be obtained from the vehicle's own attitude sensors. In one example, after acquiring the target vehicle's camera images and positioning data, data preprocessing can be performed, including but not limited to one or more of data cleaning, data standardization, and data synchronization.

[0043] The location of the base station and the three-dimensional map of the road segment where the target vehicle is located can be obtained in advance by those skilled in the art.

[0044] In step S120, the positioning data of the target vehicle, the location of the base station, and the three-dimensional map of the road segment are used as inputs to a pre-built ray tracing model, so that the ray tracing model outputs the propagation characteristics of the wireless communication signal of the target vehicle during driving.

[0045] In this embodiment, the ray tracing model can be implemented using the channel simulation software Wireless InSite. For example... Figure 2 As shown, the propagation process of millimeter-wave wireless signals involves multiple scattering paths, which can be divided into line-of-sight (LOS) paths and non-line-of-sight (NLOS) paths. The LOS path is the direct link between the base station and the vehicle. The NLOS path occurs when the direct link between the base station and the vehicle is blocked by an obstacle, and the signal propagates only through reflection, refraction, or scattering. The point where the signal changes its propagation direction after passing through the surface of the scattering object is called the scattering point. When unobstructed, the base station can be considered the scattering point with the strongest scattering intensity. The millimeter-wave channel can be represented by a geometric channel model as follows:

[0046]

[0047] Where, α p Let be the complex gain coefficient of the p-th scattering path. and The horizontal and vertical angles originating from the scattering path. and The horizontal and vertical angles reached by the scattering path. Let a be the conjugate transpose of the transmit antenna steering vector. r This is the guide vector for the receiving antenna. The number of antennas at the base station and vehicle end is N = N2 x N y For a Uniform Planar Array (UPA), the antenna steering vector can be expressed as:

[0048]

[0049] in, It represents the Kronecker product.

[0050] In one example, the propagation characteristics of the wireless communication signal include, but are not limited to, the angle of arrival for each scattering path, the scattering path distribution, and the location and intensity of scattering points in the environment. Based on the channel information, the optimal beam pair between the base station and the vehicle is obtained by scanning the codebook; this step can be represented as...

[0051]

[0052] Where W is the optimal beam pair, H is the channel between the target vehicle and the base station, and B i and C j These are the beams in the base station codebook B and the vehicle codebook C, respectively.

[0053] In step S130, the propagation characteristics of the wireless communication signal and the camera image are input into a pre-constructed scatterer projection model to obtain the scatterer features in the image.

[0054] In this embodiment, a scatterer projection model is pre-constructed. The propagation characteristics of the wireless communication signal and the camera image are input into the scatterer projection model so that the scatterer projection model projects the scattering point according to the angle of arrival of the wireless communication signal to obtain the scatterer features in the image. The scatterer features include the coordinates of the projection point and the scattering intensity.

[0055] In one example, based on Figure 3 The projection of the scattering point shown can be represented by the following model:

[0056]

[0057] Where W and H are the width and height of the camera image, respectively, and 2β is the horizontal field of view of the camera. and θ p,i Let x be the azimuth and elevation angles of the p-th scattering path in the i-th camera image, respectively. p and y p These are the coordinates of the projected point in the camera image.

[0058] The scattering intensity is obtained from the ray tracing model. After thresholding and normalization, the scattering intensity can be expressed as:

[0059]

[0060] Where, α p Let α be the original reflection coefficient of the p-th scattering path, and ε be the intensity threshold used to filter out scattering paths with lower intensity. max | 2 α represents the maximum scattering intensity across all scattering paths. ′ p Let be the normalized scattering intensity at the p-th projection point.

[0061] In this embodiment, except for the line-of-sight path, each scattering path undergoes at least one scattering. For scattering paths with multiple scatterings, only the scattering point closest to the vehicle end is taken, and the intensity of this scattering point is defined as the normalized scattering intensity α of the path.′ p .

[0062] Please continue to refer to this. Figure 1 In step S140, a pre-built multimodal large model is trained under supervision based on the channel feature dataset composed of the camera image and the scatterer features, so as to minimize the prediction error of the scatterer coordinates and intensity, and obtain the channel estimation large model.

[0063] In this embodiment, those skilled in the art can pre-construct a large multimodal model, such as Figure 4 As shown, this multimodal large model can include a depth estimation network, a visual encoder, a learnable connector, and a large language model. The depth estimation network, such as Depth-anything or ZoeDepth, is used to calculate the spatial depth information of objects in an image, thereby obtaining the distance from objects in the environment to the vehicle to be predicted. The visual encoder is used to extract object category features from camera images; commonly used visual encoders undergo extensive image-text feature alignment, including CLIP and SigLIP. The learnable connector, such as Q-Former or MLP (Multilayer Perceptron), is responsible for bridging the gap between image and text modalities, enabling the large language model to parse the input spatial depth and object category features. The large language model, such as Chatgpt, Llama, or Qwen models, is used to estimate channel features based on spatial depth information and object category features, combined with cue words.

[0064] In this embodiment, the large model determines the location and altitude information of base stations in the road segment by retrieving data from a database based on vehicle positioning data. Then, the large model can be guided to predict the main scattering path of the millimeter-wave channel in the vehicle-to-everything (V2X) environment from multi-view images, using ray propagation as an example. Based on the type of scatterers and spatial depth information in the multi-view images, the location and intensity of scattering points in the scattering path are estimated. Finally, the large model outputs the scatterer features of each view image in natural language.

[0065] In one example, digital twins can be used to simulate real-world road environments and traffic flow, expanding the channel feature dataset. A large model identifies road segments based on vehicle location data, and deployment is accelerated by training with data from the digital twin world and fine-tuning with limited real-world data.

[0066] Next, a channel feature dataset is constructed, which includes multi-view camera images and scatterer features; such as Figure 4As shown, the deep estimation network, visual encoder, connector, or large language model in the multimodal large model are fine-tuned using a channel feature dataset, or any combination of these three components. Fine-tuning methods include, but are not limited to, full parameter fine-tuning, LoRA, and QLoRA, to obtain the original channel estimation large model. Cue words supplement the large model with knowledge, guide thinking, and provide instructions; in this embodiment, cue words include background knowledge, vehicle perception data, and questions. The output of the multimodal large model includes the coordinates and scattering intensity of each scattering point in the camera image from each viewpoint.

[0067] In one example, the original large-scale channel estimation model can be lightweighted, including but not limited to quantization, sparsification, knowledge distillation, low-rank decomposition, and parameter sharing. The lightweight large-scale channel estimation model is then deployed to a vehicle to predict scatterer features in camera images.

[0068] In step S150, a scattering heatmap is constructed based on the channel estimation dataset, and the triplet data consisting of the scattering heatmap, localization, and optimal beam is segmented according to the beam coherence time to obtain a beam prediction dataset.

[0069] In this embodiment, each sample in the beam prediction dataset includes a scattering heatmap and a localization sequence, with a length of M. t The labels of the beam prediction dataset include the optimal beam W in the next beam coherence time. t+1 and time length M t+1 Beamcoherence time refers to the length of time that the base station and the vehicle maintain a beam pair, including the beam alignment phase and the communication phase. During the beam alignment phase, the base station and the vehicle predict and adjust the beam according to the channel environment, which can be expressed as follows: Figure 5 As shown, from Figure 5 As can be seen, using a fixed beam coherence time requires frequent updates, which causes the original communication time to be used for beam alignment, reducing the transmission rate, as seen in time period ①. Furthermore, the beam may become inaccurate, as seen in time period ②.

[0070] Setting the beam coherence time to an integer multiple of the vehicle data acquisition interval can be expressed as follows:

[0071] T = M t T s M = 1, ..., M max

[0072] Among them, M t M represents the number of data acquisitions within the coherence time of the t-th beam. max For the maximum number of times, T s This refers to the interval for vehicle data collection.

[0073] In one example, the size of the scattering heatmap is obtained from camera images in the channel feature dataset; the scattering intensity is normalized using the radiative function and then expanded to obtain the scattering heatmap. This step can be represented as:

[0074]

[0075] Where I is the pixel value (maximum value is 255) of the point with coordinates (x, y) in the heatmap, and (x p ,y p ) represents the coordinates of the p-th scattering point projected onto the heat map, and σ represents the extended range of the projection point when the scattering point is projected onto the heat map.

[0076] In step S160, the pre-built beam prediction model is trained based on the beam prediction dataset to obtain the target beam prediction model for use in vehicle-to-everything (V2X) beam prediction.

[0077] In one embodiment, such as Figure 6 As shown, the beam prediction model can include a spatiotemporal feature extractor, a localization feature extractor, a beam classification network, and a beam coherence temporal classification network. The spatiotemporal feature extractor can be a ConvLSTM, PredRNN, or SwinLSTM network, used to extract spatiotemporal features from multi-view camera image sequences within the beam coherence time. The localization feature extractor can be a Convolutional Neural Network (CNN) or a Transformer, used to extract vehicle localization features. Feature fusion can be achieved through feature vector multiplication and an attention mechanism. The beam classification network and the beam coherence temporal classification network can be fully connected neural networks, used to classify and output classification results based on the fused features obtained by fusing the outputs of the spatiotemporal feature extractor and the localization feature extractor. It should be noted that, as... Figure 6 As shown, the output of the beam prediction model includes the beam index and time length in the next beam coherence time, where the time length is defined as the sampling interval of the sensing data.

[0078] Once the target beam prediction model is trained, it can be deployed in vehicles for vehicle-to-everything (V2X) beam prediction.

[0079] In some embodiments of this application, after obtaining the target beam prediction model, the method further includes:

[0080] The channel estimation model is used to predict the corresponding channel features based on multi-view camera images of the vehicle to be predicted.

[0081] The base station retrieves the coordinates and intensity of the projection point in the camera image from each viewpoint based on the channel characteristics of the vehicle to be predicted.

[0082] A scattering heat map is constructed based on the coordinates and intensity of the projection point. The positioning data and scattering heat map of the vehicle to be predicted are then input into the target beam prediction model so that the target beam prediction model can predict the beam index and time length within the next beam coherence time.

[0083] In this embodiment, the large-scale channel estimation model deployed on the vehicle can predict its channel characteristics based on camera images from each viewpoint. The output format can be that the coordinates of the projection point j in the image from viewpoint i are (x...). j ,y j The intensity is S. The base station can retrieve the above output results according to regular matching to obtain the coordinates and intensity of the projection point in the camera image of each viewpoint; then, it constructs a scattering heat map based on the coordinates and intensity of the projection point, and inputs the positioning data of the vehicle to be predicted and the scattering heat map into the target beam prediction model to obtain the beam index and time length in the next beam coherence time.

[0084] Specifically, a large channel estimation model is used to predict the corresponding channel features based on multi-view camera images of the vehicle to be predicted. This includes: using a depth estimation network to extract spatial depth information from the camera images to obtain the distance from objects in the camera images to the vehicle to be predicted; using a visual encoder to extract object category features from the camera images; using a learnable connector to map image features including spatial depth information and object category features to a large language model to bridge the gap between image and text modalities; and using prompt words to guide the large language model to output channel features in the camera images containing the location and intensity of scattering points based on the angle of ray propagation.

[0085] Thus, the vehicle-to-everything (V2X) beam prediction method based on a multimodal large model provided in this application utilizes multi-view camera images of the vehicle to predict scatterer features, including the coordinates and intensity of scattering points. These features help the base station accurately understand the channel characteristics in the environment, thereby achieving accurate beam prediction. Compared with the traditional codebook scanning method, the camera image-based method in this application significantly reduces pilot overhead. Compared with deploying cameras at the base station, it can operate normally even when there is obstruction in the line-of-sight path, enhancing the accuracy of the vision-based beam prediction method under switching between line-of-sight and non-line-of-sight paths. Simultaneously, leveraging the powerful generalization capability of the multimodal large model, the model can be quickly deployed in different communication environments, improving the robustness of the system in extracting channel features based on visual images. Furthermore, the beam prediction model predicts the next beam coherence time length and the optimal beam index by combining a multi-view scattering heatmap constructed based on channel features with the scattering heatmap within the current beam coherence time and the vehicle positioning sequence. Compared with the traditional fixed beam coherence time scheme, this effectively reduces the risk of beam misalignment and significantly improves the communication transmission rate.

[0086] The following describes an embodiment of the apparatus described in this application, which can be used to execute the vehicular network beam prediction method based on a multimodal large model described in the above embodiments of this application. For details not disclosed in the apparatus embodiments of this application, please refer to the embodiments of the vehicular network beam prediction method based on a multimodal large model described in the above embodiments of this application.

[0087] Figure 7 A block diagram of a vehicle-to-everything (V2X) beam prediction device based on a multimodal large model according to an embodiment of this application is shown.

[0088] Reference Figure 7 As shown, a vehicle-to-everything (V2X) beam prediction device based on a multimodal large model according to an embodiment of this application includes:

[0089] The acquisition module is used to acquire camera images, location data, base station locations on the road segment, and a 3D map of the road segment of the target vehicle.

[0090] The ray tracing module is used to take the target vehicle's positioning data, base station location, and road segment 3D map as input to a pre-built ray tracing model, so that the ray tracing model outputs the propagation characteristics of the target vehicle's wireless communication signal during driving.

[0091] The projection module is used to input the propagation characteristics of the wireless communication signal and the camera image into a pre-constructed scatterer projection model to obtain the scatterer features in the image;

[0092] The optimization module is used to supervise the training of a pre-built multimodal large model based on the channel feature dataset composed of the camera images and the scatterer features, so as to minimize the prediction error of the scatterer coordinates and intensity and obtain the channel estimation large model.

[0093] The construction module is used to construct a scattering heatmap based on the channel feature dataset, and to segment the scattering heatmap, localization and the triplet data composed of the optimal beam according to the beam coherence time to obtain the beam prediction dataset.

[0094] The processing module is used to train a pre-built beam prediction model based on the beam prediction dataset to obtain a target beam prediction model for use in vehicle-to-everything (V2X) beam prediction.

[0095] In one embodiment, after obtaining the target beam prediction model, the processing module is further configured to:

[0096] The channel estimation model is used to predict the corresponding channel features based on multi-view camera images of the vehicle to be predicted.

[0097] The base station retrieves the coordinates and intensity of the projection point in the camera image from each viewpoint based on the channel characteristics of the vehicle to be predicted.

[0098] A scattering heat map is constructed based on the coordinates and intensity of the projection point. The positioning data and scattering heat map of the vehicle to be predicted are then input into the target beam prediction model so that the target beam prediction model can predict the beam index and time length within the next beam coherence time.

[0099] Figure 8 A schematic diagram of the structure of a computer system suitable for implementing the electronic device of the present application is shown.

[0100] It should be noted that, Figure 8 The computer system of the electronic device shown is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments of this application.

[0101] like Figure 8 As shown, the computer system includes a Central Processing Unit (CPU) 801, which can perform various appropriate actions and processes based on programs stored in Read-Only Memory (ROM) 802 or programs loaded from storage portion 808 into Random Access Memory (RAM) 803, such as performing the methods described in the above embodiments. The RAM 803 also stores various programs and data required for system operation. The CPU 801, ROM 802, and RAM 803 are interconnected via a bus 804. An Input / Output (I / O) interface 805 is also connected to the bus 804.

[0102] The following components are connected to I / O interface 805: an input section 806 including a keyboard, mouse, etc.; an output section 807 including a cathode ray tube (CRT), liquid crystal display (LCD), etc., and speakers, etc.; a storage section 808 including a hard disk, etc.; and a communication section 809 including a network interface card such as a LAN (Local Area Network) card, modem, etc. The communication section 809 performs communication processing via a network such as the Internet. A drive 810 is also connected to I / O interface 805 as needed. A removable medium 811, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed on drive 810 as needed so that computer programs read from it can be installed into storage section 808 as needed.

[0103] Specifically, according to embodiments of this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program including a computer program for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication section 809, and / or installed from removable medium 811. When the computer program is executed by central processing unit (CPU) 801, it performs various functions defined in the system of this application.

[0104] It should be noted that the computer-readable medium shown in the embodiments of this application can be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), flash memory, optical fiber, portable compact disc read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this application, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In this application, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying a computer-readable computer program. The transmitted data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. The computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The computer program contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to wireless, wired, etc., or any suitable combination thereof.

[0105] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. Each block in a flowchart or block diagram may represent a module, segment, or portion of code, which contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram or flowchart, and combinations of blocks in a block diagram or flowchart, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0106] The units described in the embodiments of this application can be implemented in software or hardware, and the described units can also be located in a processor. The names of these units do not necessarily limit the specific unit itself.

[0107] In another aspect, this application also provides a computer-readable medium, which may be included in the electronic device described in the above embodiments; or it may exist independently and not assembled into the electronic device. The computer-readable medium carries one or more programs, which, when executed by the electronic device, cause the electronic device to perform the methods described in the above embodiments.

[0108] It should be noted that although several modules or units for the device used to perform actions have been mentioned in the detailed description above, this division is not mandatory. In fact, according to the embodiments of this application, the features and functions of two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.

[0109] Through the above description of the embodiments, those skilled in the art will readily understand that the exemplary embodiments described herein can be implemented by software or by combining software with necessary hardware. Therefore, the technical solutions according to the embodiments of this application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, USB flash drive, external hard drive, etc.) or on a network, including several instructions to cause a computing device (such as a personal computer, server, touch terminal, or network device, etc.) to execute the method according to the embodiments of this application.

[0110] Other embodiments of this application will readily occur to those skilled in the art upon consideration of the specification and practice of the embodiments disclosed herein. This application is intended to cover any variations, uses, or adaptations of this application that follow the general principles of this application and include common knowledge or customary techniques in the art not disclosed herein.

[0111] It should be understood that this application is not limited to the precise structure described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this application is limited only by the appended claims.

Claims

1. A method for vehicle-to-everything beam prediction based on a multi-modal large model, characterized in that, The method comprises the following steps: acquiring camera images, positioning data, base station positions of a road segment and a three-dimensional map of the road segment of a target vehicle; inputting the positioning data, base station positions and three-dimensional map of the road segment of the target vehicle into a pre-constructed ray tracing model as inputs of the ray tracing model, so that the ray tracing model outputs propagation characteristics of wireless communication signals of the target vehicle during driving; inputting the propagation characteristics of the wireless communication signals and the camera images into a pre-constructed scatterer projection model to obtain scatterer features in the images; performing supervised training on a pre-constructed multi-modal large model according to a channel feature data set composed of the camera images and the scatterer features to minimize prediction errors of scatterer coordinates and intensities, thereby obtaining a channel estimation large model; constructing a scatter heat map according to the channel feature data set, and dividing a three-tuple data set composed of the scatter heat map, positioning and optimal beams according to a beam coherence time, thereby obtaining a beam prediction data set; training a pre-constructed beam prediction model according to the beam prediction data set to obtain a target beam prediction model for vehicle networking beam prediction; using the channel estimation large model to perform prediction based on multi-view camera images of a vehicle to be predicted, thereby obtaining corresponding channel features; retrieving, by a base station, the channel features of the vehicle to be predicted to obtain coordinates and intensities of projection points in camera images of each view; constructing a scatter heat map according to the coordinates and intensities of the projection points, and inputting positioning data and the scatter heat map of the vehicle to be predicted into the target beam prediction model, so that the target beam prediction model predicts beam indexes and time lengths within a next beam coherence time.

2. The method of claim 1, wherein, The scatterer features include coordinates of scatter points and scatter intensities. Then, constructing a scatter heat map according to the channel feature data set comprises: calculating according to camera images in the channel feature data set to obtain a size of the scatter heat map; normalizing the scatter intensities by using a radiation function and extending to obtain the scatter heat map.

3. The method of claim 1, wherein, The beam prediction model comprises a space-time feature extractor, a positioning feature extractor, a beam classification network and a beam coherence time classification network, wherein the space-time feature extractor is used to perform space-time feature extraction on a multi-view camera image sequence within a beam coherence time, the positioning feature extractor is used to extract vehicle positioning features, and the beam classification network and the beam coherence network are used to classify and output classification results according to fused features fused from output results of the space-time feature extractor and output results of the positioning feature extractor.

4. The method of claim 1, wherein, The multi-modal large model comprises a depth estimation network, a visual encoder, a learnable connector and a large language model. Then, using the channel estimation large model to perform prediction based on multi-view camera images of a vehicle to be predicted to obtain corresponding channel features comprises: extracting spatial depth information in the camera images by using the depth estimation network to obtain distances of objects in the camera images to the vehicle to be predicted; using the visual encoder to extract object category features in the camera images; The image features including spatial depth information and object category features are mapped to the large language model by using the learnable connector to bridge the gap between image and text modalities. The large language model is guided by a prompt word to output channel features containing the positions and intensities of scattering points in the camera image based on the angle of ray propagation. 5.A vehicular internet-of-things beam prediction device based on a multi-modal large model, characterized in that, The method comprises the following steps: An acquisition module is configured to acquire a camera image of a target vehicle, positioning data, base station positions of a road segment, and a three-dimensional map of the road segment; A ray tracing module is configured to input the positioning data of the target vehicle, the base station positions, and the three-dimensional map of the road segment into a pre-constructed ray tracing model as inputs, so that the ray tracing model outputs propagation characteristics of wireless communication signals of the target vehicle during driving; A projection module is configured to input the propagation characteristics of the wireless communication signals and the camera image into a pre-constructed scatterer projection model to obtain scatterer features in the image; An optimization module is configured to supervise training of a pre-constructed multi-modal large model based on a channel feature dataset composed of the camera image and the scatterer features, so as to minimize prediction errors of scatterer coordinates and intensities and obtain a channel estimation large model; A construction module is configured to construct a scatter heat map based on the channel feature dataset, and to split a three-tuple dataset composed of the scatter heat map, positioning, and an optimal beam based on a beam coherence time to obtain a beam prediction dataset; A processing module is configured to train a pre-constructed beam prediction model based on the beam prediction dataset to obtain a target beam prediction model for vehicle networking beam prediction; The channel estimation large model is used to make a prediction based on multi-view camera images of a vehicle to be predicted to obtain corresponding channel features; A base station searches based on the channel features of the vehicle to be predicted to obtain coordinates and intensities of projection points in camera images of each view; A scatter heat map is constructed based on the coordinates and intensities of the projection points, and positioning data and the scatter heat map of the vehicle to be predicted are input into the target beam prediction model, so that the target beam prediction model predicts a beam index and a time length within a next beam coherence time.

6. A computer readable medium having stored thereon a computer program, characterized in that, The computer program is executed by a processor to implement the method of any one of claims 1-4.

7. An electronic device, comprising: The method comprises the following steps: One or more processors; A storage device is configured to store one or more programs, which, when executed by the one or more processors, cause the one or more processors to implement the method of any one of claims 1-4.

Citation Information

Patent Citations

  • Beam alignment method and device based on terminal visual perception

    CN116094559A

  • Millimeter wave beam prediction method of industrial wireless system assisted by multi-modal environment semantics

    CN119402048A