Internet of vehicles beam prediction method and device based on multi-modal large model

Through the Internet of Vehicle Beam Prediction method based on multimodal large model, the channel characteristics in the vehicle environment are extracted and beam index and coherence time are predicted, which solves the problem of insufficient accuracy and robustness in the traditional method, and improves the performance of Internet of Vehicles communication.

CN120165733AActive Publication Date: 2025-06-17XIAMEN UNIV

Patent Information

Application Number
CN202510212337.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-25
Publication Date
2025-06-17
Estimated Expiration
2045-02-25

AI Technical Summary

Technical Problem

Traditional Internet of Vehicle beam prediction methods have problems with insufficient accuracy and robustness in fast mobility and complex environments, and fixed beam coherence time may lead to beam misalignment or transmission rate loss.

Method used

The Internet of Vehicle beam prediction method based on multimodal large model is adopted. By obtaining the vehicle's camera image, positioning data and road segment information, combining the ray tracing model and the scatterer projection model, channel characteristics are extracted and supervised and trained to predict the beam index and beam coherence time.

Benefits of technology

It improves the adaptability and robustness of Internet of Vehicles beam prediction, enhances the beam accuracy under non-horizontal switching of millimeter wave communication, and improves the communication transmission rate and energy efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120165733A_ABST
    Figure CN120165733A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides an Internet of Vehicles beam prediction method and device based on a multi-modal large model. The method comprises the following steps: acquiring a camera image and positioning data of a target vehicle, a base station position of a road section and a three-dimensional map of the road section; inputting the positioning data, the position of the base station and the three-dimensional map of the road section into a ray tracing model to obtain propagation characteristics of a wireless communication signal; inputting the propagation characteristics and the camera image into a scatterer projection model to obtain scatterer characteristics in the image; and optimizing the multi-modal large model and constructing a scattering thermodynamic diagram according to a channel feature data set containing camera images and scatterer features, and further obtaining a beam prediction data set for beam prediction model training for training. According to the technical scheme, the channel characteristics in the camera image can be effectively extracted, the beam index and the beam coherence time in the codebook can be accurately predicted in combination with the channel characteristics and the positioning information, and the communication transmission rate and the energy efficiency of the system are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the fields of communication and computer technologies. Specifically, it relates to a vehicle-to-everything (V2X) beam prediction method and apparatus based on a multimodal large model. Background Art

[0002] With the development of applications such as autonomous driving and vehicle-to-everything (V2X), the demand for data communication rate by vehicles is becoming increasingly stringent, and millimeter-wave communication can provide guarantee for vehicle communication. As one of the key technologies of millimeter-wave communication, beamforming technology uses directional beams to focus signal transmission to increase the signal power at the receiving end and ensure communication quality.

[0003] However, the fast mobility and complex environment (such as buildings, obstacles, and other vehicles) in vehicle-to-everything (V2X) networks pose many problems to traditional beam training methods. In the current technical solutions of vision-assisted beam prediction, the beam prediction scheme that obtains the relative position information of a vehicle from in-vehicle camera images and jointly uses positioning information to assist the base station can solve the problems of insufficient positioning accuracy and accuracy, and the vulnerability of the camera from the base station to occlusion, effectively reducing the beam training overhead. However, this end-to-end beam prediction method based on in-vehicle images often faces problems of insufficient accuracy and robustness in cases of heavy traffic and base station occlusion. In addition, the fixed beam coherence time (the time during which the beam remains unchanged) may lead to long-term beam misalignment or a large loss in transmission rate for the vehicle-to-everything (V2X) scenario with strong mobility changes. Summary of the Invention

[0004] Embodiments of the present application provide a vehicle-to-everything (V2X) beam prediction method and apparatus based on a multimodal large model, which can, at least to a certain extent, effectively extract channel features in in-vehicle camera images, and jointly use the channel features and positioning information to accurately predict the beam index and beam coherence time in the codebook, solve the disadvantages of insufficient adaptability and robustness in different vehicle-to-everything (V2X) environments, enhance the beam accuracy under the handover between line-of-sight and non-line-of-sight of millimeter-wave communication, and improve the communication transmission rate and energy efficiency of the system.

[0005] Other features and advantages of the present application will become apparent through the following detailed description, or be learned in part through the practice of the present application.

[0006] According to one aspect of the embodiments of the present application, there is provided a vehicle-to-everything (V2X) beam prediction method based on a multimodal large model, including:

[0007] Obtaining a camera image, positioning data, the location of a base station on a road section, and a three-dimensional map of the road section of a target vehicle;

[0008] Use the positioning data of the target vehicle, the base station location, and the three-dimensional road map as the input of a pre-constructed ray tracing model, so that the ray tracing model outputs the propagation characteristics of the wireless communication signal during the driving of the target vehicle;

[0009] Input the propagation characteristics of the wireless communication signal and the camera image into a pre-constructed scatterer projection model to obtain the scatterer characteristics in the image;

[0010] Supervise and train a pre-constructed multi-modal large model according to the channel characteristic data set composed of the camera image and the scatterer characteristics to minimize the prediction error of the scatterer coordinates and intensity, and obtain a channel estimation large model;

[0011] Construct a scatter heat map according to the channel characteristic data set, and segment the triple data composed of the scatter heat map, positioning, and the optimal beam according to the beam coherence time to obtain a beam prediction data set;

[0012] Train a pre-constructed beam prediction model according to the beam prediction data set to obtain a target beam prediction model for vehicle-to-everything (V2X) beam prediction.

[0013] According to one aspect of the embodiments of the present application, a V2X beam prediction device based on a multi-modal large model is provided, including:

[0014] An acquisition module, configured to acquire the camera image, positioning data, base station location of the road section where the target vehicle is located, and the three-dimensional road map;

[0015] A ray tracing module, configured to use the positioning data of the target vehicle, the base station location, and the three-dimensional road map as the input of a pre-constructed ray tracing model, so that the ray tracing model outputs the propagation characteristics of the wireless communication signal during the driving of the target vehicle;

[0016] A projection module, configured to input the propagation characteristics of the wireless communication signal and the camera image into a pre-constructed scatterer projection model to obtain the scatterer characteristics in the image;

[0017] An optimization module, configured to supervise and train a pre-constructed multi-modal large model according to the channel characteristic data set composed of the camera image and the scatterer characteristics to minimize the prediction error of the scatterer coordinates and intensity, and obtain a channel estimation large model;

[0018] A construction module, configured to construct a scatter heat map according to the channel characteristic data set, and segment the triple data composed of the scatter heat map, positioning, and the optimal beam according to the beam coherence time to obtain a beam prediction data set;

[0019] A processing module, configured to train a pre-constructed beam prediction model according to the beam prediction data set to obtain a target beam prediction model for vehicle-to-everything (V2X) beam prediction.

[0020] According to one aspect of the embodiments of the present application, there is provided a computer-readable medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the V2X beam prediction method based on a multimodal large model as described in the above embodiments.

[0021] According to one aspect of the embodiments of the present application, there is provided an electronic device, including: one or more processors; a storage device for storing one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors are caused to implement the V2X beam prediction method based on a multimodal large model as described in the above embodiments.

[0022] According to one aspect of the embodiments of the present application, there is provided a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the V2X beam prediction method provided in the above embodiments.

[0023] In the technical solutions provided in some embodiments of the present application, by acquiring a camera image, positioning data, base station positions of a section where a target vehicle is located, and a three-dimensional map of the section, the positioning data, base station positions, and three-dimensional map of the section of the target vehicle are used as inputs of a pre-constructed ray tracing model, so that the ray tracing model outputs the propagation characteristics of wireless communication signals during the driving process of the target vehicle on the section. The propagation characteristics of the wireless communication signals and the camera image are input into a pre-constructed scatterer projection model to obtain scatterer characteristics in the image. The pre-constructed multimodal large model is optimized according to a channel feature data set composed of the camera image and the scatterer characteristics to obtain a channel estimation large model. A scatter heat map is constructed according to the channel feature data set, and the triple data composed of the scatter heat map, positioning, and optimal beam is segmented according to the beam coherence time to obtain a beam prediction data set. The pre-constructed beam prediction model is trained according to the beam prediction data set to obtain a target beam prediction model for V2X beam prediction. In this way, the scatterer characteristics in the in-vehicle camera image can be effectively extracted, and the beam index and beam coherence time in the codebook can be accurately predicted by combining the scatterer characteristics and positioning information, solving the disadvantages of insufficient adaptability and robustness in different V2X environments, and improving the communication transmission rate and energy efficiency of the system.

[0024] It should be understood that the above general description and the following detailed description are merely exemplary and explanatory, and do not limit this application. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] The accompanying drawings are incorporated herein and constitute a part of this specification, showing embodiments consistent with this application, and are used together with the specification to explain the principles of this application. Obviously, the drawings in the following description are only some embodiments of this application, and those of ordinary skill in the art can obtain other drawings based on these drawings without creative efforts. In the drawings:

[0026] Figure 1 shows a schematic flow diagram of a vehicle networking beam prediction method based on a multimodal large model according to an embodiment of this application;

[0027] Figure 2 shows a schematic diagram of signal ray propagation according to an embodiment of this application;

[0028] Figure 3 shows a schematic diagram of scatter point projection according to an embodiment of this application;

[0029] Figure 4 shows a schematic structural diagram of a multimodal large model according to an embodiment of this application;

[0030] Figure 5 shows a schematic diagram of beam coherence time according to an embodiment of this application;

[0031] Figure 6 shows a schematic structural diagram of a beam prediction model according to an embodiment of this application;

[0032] Figure 7 shows a block diagram of a vehicle networking beam prediction device based on a multimodal large model according to an embodiment of this application;

[0033] Figure 8 shows a schematic structural diagram of a computer system of an electronic device suitable for implementing the embodiments of this application. DETAILED DESCRIPTION

[0034] Example embodiments will now be described more fully with reference to the accompanying drawings. However, the example embodiments can be implemented in various forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this application will be more complete and comprehensive, and will fully convey the concept of the example embodiments to those skilled in the art.

[0035] In addition, the described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments. In the following description, numerous specific details are provided to give a thorough understanding of the embodiments of the present application. However, those skilled in the art will realize that the technical solutions of the present application can be practiced without one or more of the specific details, or other methods, components, devices, steps, etc. may be adopted. In other cases, well-known methods, devices, implementations, or operations are not shown or described in detail to avoid obscuring aspects of the present application.

[0036] The block diagrams shown in the drawings are only functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities can be implemented in software form, or implemented in one or more hardware modules or integrated circuits, or implemented in different networks and / or processor devices and / or microcontroller devices.

[0037] The flowcharts shown in the drawings are only illustrative and do not necessarily include all the contents and operations / steps, nor are they necessarily executed in the described order. For example, some operations / steps can be decomposed, while some operations / steps can be combined or partially combined, so the actual execution order may change according to the actual situation.

[0038] Figure 1 The flowchart showing the method for predicting vehicle-to-everything (V2X) beams based on a multimodal large model according to an embodiment of the present application is presented.

[0039] Please refer to Figure 1 , the method for predicting V2X beams based on a multimodal large model at least includes steps S110 to S160, which are introduced in detail as follows:

[0040] In step S110, the camera image, positioning data, base station location of the section where the target vehicle is located, and the three-dimensional map of the section are obtained.

[0041] In this embodiment, millimeter-wave base stations can be deployed on street lamps and utility poles beside the road. The camera image can be obtained by an in-vehicle camera on the target vehicle, and the camera image can include multi-view images with non-overlapping image views, such as images of the front view and rear view of the target vehicle.

[0042] The positioning data of the target vehicle may include longitude and latitude, as well as vehicle inertial acceleration and axle acceleration, where the longitude and latitude can be obtained through GPS or the Beidou satellite positioning system, and the vehicle inertial acceleration and axle acceleration can be obtained by the vehicle's own attitude sensors. In one example, after obtaining the camera image and positioning data of the target vehicle, data preprocessing can be performed, including but not limited to one or more of data cleaning, data normalization, and data synchronization.

[0043] The base station location of the road section where the target vehicle is located and the three-dimensional road map can be pre-acquired by those skilled in the art.

[0044] In step S120, the positioning data of the target vehicle, the base station location, and the three-dimensional road map are used as the inputs of a pre-constructed ray tracing model, so that the ray tracing model outputs the propagation characteristics of the wireless communication signal during the driving process of the target vehicle.

[0045] In this embodiment, the ray tracing model can be implemented by the channel simulation software Wireless InSite. As Figure 2 shown, the channel of the millimeter-wave wireless signal propagation process includes multiple scattering paths, which can be divided into the line-of-sight path and the non-line-of-sight path. The line-of-sight (LOS) path is the direct link between the base station and the vehicle, and the non-line-of-sight (NLOS) path is that the direct link between the base station and the vehicle is blocked by obstacles, and the signal propagates only through reflection, refraction, or scattering. The point where the signal changes the propagation direction after passing through the surface of the scatterer is the scattering point. When not blocked by obstacles, the base station can be regarded as the scatterer with the strongest scattering intensity. The millimeter-wave channel can be represented by the geometric channel model as:

[0046]

[0047] where, α p is the complex gain coefficient of the p-th scattering path, and are the horizontal angle and the elevation angle from which the scattering path departs, and are the horizontal angle and the elevation angle at which the scattering path arrives, is the conjugate transpose of the transmitting antenna steering vector, a r is the receiving antenna steering vector. The antennas at the base station and the vehicle ends are uniform planar arrays (UPAs) with the number N = N x N y . The antenna steering vector can be expressed as:

[0048]

[0049] where, represents the Kronecker product.

[0050] In an example, the propagation characteristics of the wireless communication signal include but are not limited to the angle of arrival of each scattering path, the scattering path distribution, and the position and scattering intensity of the scattering points in the environment. According to the channel information, the optimal beam pair between the base station and the vehicle is obtained by scanning the codebook; this step can be expressed as

[0051]

[0052] Among them, W is the optimal beam pair, H is the channel between the target vehicle and the base station, and B i and C j are the beams in the base station codebook B and the vehicle codebook C, respectively.

[0053] In step S130, the propagation characteristics of the wireless communication signal and the camera image are input into a pre-constructed scatterer projection model to obtain the scatterer characteristics in the image.

[0054] In this embodiment, a scatterer projection model is pre-constructed, and the propagation characteristics of the wireless communication signal and the camera image are input into the scatterer projection model, so that the scatterer projection model projects the scatter points according to the angle of arrival of the wireless communication signal to obtain the scatterer characteristics in the image. Among them, the scatterer characteristics include the coordinates of the projection points and the scattering intensity.

[0055] In an example, based on Figure 3 the scatter point projection shown, the scatterer projection model can be expressed as:

[0056]

[0057] Among them, W and H are the width and height of the camera image respectively, 2β is the horizontal field of view angle of the camera, and θ p,i are the azimuth angle and elevation angle of the p-th scattering path in the i-th camera image respectively, x p and y p are the coordinates of the projection points in the camera image.

[0058] The scattering intensity is obtained from the ray tracing model. After threshold screening and normalization, the scattering intensity can be expressed as

[0059]

[0060] Among them, α p is the original reflection coefficient of the p-th scattering path, ε is the intensity threshold used to filter out the scattering paths with smaller intensities, |α max | 2 is the maximum scattering intensity among all scattering paths, and α ′ p is the normalized scattering intensity of the p-th projection point.

[0061] In this embodiment, except for the line-of-sight path, each scattering path undergoes at least one scattering. For the scattering paths with multiple scatterings, only the scatter point closest to the vehicle end is taken, and the intensity of this scatter point is defined as the normalized scattering intensity α of this path′ p 。

[0062] Please continue to refer to Figure 1 , in step S140, a pre-constructed multimodal large model is supervised and trained based on the channel feature dataset composed of the camera image and the scatterer characteristics to minimize the prediction error of the scatterer coordinates and intensity, and a channel estimation large model is obtained.

[0063] In this embodiment, those skilled in the art can pre-construct a multimodal large model, such as Figure 4 shown, the multimodal large model may include a depth estimation network, a visual encoder, a learnable connector, and a large language model. Among them, the depth estimation network may be Depth-anything, ZoeDepth, which is used to calculate the spatial depth information of the objects in the image, so as to obtain the distance from the objects in the environment to the vehicle to be predicted; the visual encoder is used to extract the object category features in the camera image, and common visual encoders include CLIP and SigLIP after a large number of text-image feature alignments; the learnable connector may be Q-Former, MLP (Multi-Layer Perceptron), which is responsible for bridging the gap between the two modalities of image and text, so that the large language model can parse the input spatial depth and object category features; the large language model may be Chatgpt, Llama, and Qwen models, which are used to estimate the channel features according to the spatial depth information and object category features and in combination with the prompt words.

[0064] In this embodiment, the large model determines the position and height information of the base stations in the road section by retrieving the database according to the vehicle positioning data. Then, the large model can be guided in thinking to predict the main scattering paths of the millimeter-wave channels in the vehicle networking environment from multiple perspectives at the angle of ray propagation. The position and intensity of the scatterers in the scattering path are estimated according to the category and spatial depth information of the scatterers in the multi-perspective image, and finally the large model outputs the scatterer characteristics of each perspective image in a natural language manner.

[0065] In one example, the digital twin method can be used to simulate the road section environment and traffic flow in the real world to expand the channel feature dataset. The large model determines the road section according to the vehicle positioning data, and speeds up the model deployment by training with the data in the twin world and fine-tuning with a small amount of data in the real world.

[0066] Next, a channel feature dataset is constructed, which includes multi-perspective camera images and scatterer characteristics; as Figure 4As shown, the depth estimation network, visual encoder, connector, or large language model in the multi-modal large model, or any combination of the three, is fine-tuned using the channel feature dataset. The fine-tuning methods include, but are not limited to, full-parameter fine-tuning, LoRA, and QLoRA, to obtain the original channel estimation large model. The prompt is the knowledge supplement, thinking guidance, and instruction requirements for the large model. In this embodiment, the prompt includes background knowledge, vehicle perception data, and questions. The output of the multi-modal large model includes the coordinates and scattering intensity of each scattering point in the camera image of each perspective.

[0067] In one example, the original channel estimation large model can be lightweighted, including but not limited to quantization, sparsification, knowledge distillation, low-rank decomposition, and parameter sharing. The lightweighted channel estimation large model is deployed to the vehicle side to predict the scattering body characteristics in the camera image.

[0068] In step S150, a scattering heat map is constructed according to the channel estimation dataset, and the triple data of the scattering heat map, positioning, and optimal beam is segmented according to the beam coherence time to obtain a beam prediction dataset.

[0069] In this embodiment, each sample of the beam prediction dataset includes a scattering heat map and a positioning sequence, with a length of M t ; the label of the beam prediction dataset includes the optimal beam W within the next beam coherence time t+1 and the time length M t+1 . The beam coherence time refers to the time length for the base station and the vehicle to maintain the beam pair, including the beam alignment phase and the communication phase. In the beam alignment phase, the base station and the vehicle predict and adjust the beam according to the channel environment, which can be represented as shown in Figure 5 As shown, from Figure 5 it can be seen that using a fixed beam coherence time requires frequent updates, which will cause the original communication time to be used for beam alignment, reducing the transmission rate, such as time period ①. In addition, the beam may face the situation of misalignment, such as time period ②.

[0070] Setting the beam coherence time to an integer multiple of the vehicle data acquisition interval can be represented as

[0071] T = M t T s , M = 1, …, M max

[0072] where M t is the number of data acquisitions within the t-th beam coherence time, M max is the maximum number, and T s is the vehicle data acquisition interval.

[0073] In one example, the size of the scattering heat map is obtained from the camera images in the channel feature dataset; the projection points are normalized using a radiation function and extended to obtain the scattering heat map, and this step can be expressed as:

[0074]

[0075] where I is the pixel value (maximum value is 255) of the point with coordinates (x, y) in the heat map, (x p , y p ) are the coordinates of the p-th scattering point, and σ is the extension range of the projection point.

[0076] In step S160, the pre-constructed beam prediction model is trained according to the beam prediction dataset to obtain a target beam prediction model for vehicle-to-internet-of-things (V2X) beam prediction.

[0077] In one embodiment, as Figure 6 shown, the beam prediction model may include a spatio-temporal feature extractor, a positioning feature extractor, a beam classification network, and a beam coherence time classification network. The spatio-temporal feature extractor may be a ConvLSTM, PredRNN, or SwinLSTM network for extracting spatio-temporal features from a multi-view camera image sequence within the beam coherence time; the positioning feature extractor may be a convolutional neural network (CNN) or a Transformer for extracting vehicle positioning features; the feature fusion method may be feature vector multiplication and an attention mechanism; the beam classification network and the beam coherence time classification network may be fully connected neural networks for classifying and outputting classification results based on the fused features obtained by fusing the output results of the spatio-temporal feature extractor and the positioning feature extractor. It should be noted that, as Figure 6 shown, the output of the beam prediction model includes the beam index and the time length within the next beam coherence time, where the time length is defined as the sampling interval of the sensing data.

[0078] After obtaining the trained target beam prediction model, it can be deployed in the vehicle terminal for V2X beam prediction.

[0079] In some embodiments of the present application, after obtaining the target beam prediction model, the method further includes:

[0080] Using the channel estimation large model to predict based on the multi-view camera images of the vehicle to be predicted to obtain corresponding channel features;

[0081] The base station retrieves according to the channel features of the vehicle to be predicted to obtain the coordinates and intensities of the projection points in the camera images of each view;

[0082] Construct a scattering heat map based on the coordinates and intensities of the projection points, and input the positioning data of the vehicle to be predicted and the scattering heat map into the target beam prediction model, so that the target beam prediction model predicts the beam index and time length within the next beam coherence time.

[0083] In this embodiment, the large channel estimation model deployed at the vehicle side can predict its channel characteristics based on the camera images from each perspective, and the format of its output result can be that the coordinate of the projection point j in the image of perspective i is (x j , y j ), and the intensity is S. The base station can retrieve the above output result according to regular matching to obtain the coordinates and intensities of the projection points in the camera images from each perspective; then construct a scattering heat map based on the coordinates and intensities of the projection points, and input the positioning data of the vehicle to be predicted and the scattering heat map into the target beam prediction model to obtain the beam index and time length within the next beam coherence time.

[0084] Specifically, use the large channel estimation model to predict based on the camera images from multiple perspectives of the vehicle to be predicted to obtain the corresponding channel characteristics, including: using a depth estimation network to extract the spatial depth information in the camera image to obtain the distance from the object in the camera image to the vehicle to be predicted; using a vision encoder to extract the object category characteristics in the camera image; using a learnable connector to map the image features including spatial depth information and object category characteristics to a large language model to bridge the gap between the two modalities of image and text; using prompt words to guide the large language model to output the channel characteristics including the position and intensity of the scattering points in the camera image based on the angle of ray propagation.

[0085] In this way, based on the vehicle-to-everything (V2X) beam prediction method based on a multimodal large model provided in this application embodiment, by using the camera images from multiple perspectives of the vehicle to predict the scatterer characteristics, including the coordinates and intensities of the scattering points, these characteristics help the base station accurately understand the channel characteristics in the environment, so as to achieve accurate beam prediction. Compared with the traditional codebook scanning method, the method based on camera images in this application embodiment can greatly reduce the pilot overhead; compared with the method of deploying cameras at the base station, it can work normally when the line-of-sight path is blocked, enhancing the accuracy of the vision-based beam prediction method when switching between line-of-sight and non-line-of-sight paths. At the same time, with the powerful generalization ability of the multimodal large model, the model can be quickly deployed in different communication environments, and the robustness of the system to extract channel characteristics based on visual images can be improved. In addition, the beam prediction model predicts the next beam coherence time length and the optimal beam index by combining the multi-perspective scattering heat map constructed based on the channel characteristics, the scattering heat map within the current beam coherence time, and the vehicle positioning sequence. Compared with the traditional fixed beam coherence time scheme, it can effectively reduce the risk of beam misalignment and significantly improve the transmission rate of communication.

[0086] The following describes the device embodiments of the present application, which can be used to execute the vehicle-to-everything (V2X) beam prediction method based on a multimodal large model in the above embodiments of the present application. For details not disclosed in the device embodiments of the present application, please refer to the embodiments of the V2X beam prediction method based on a multimodal large model above.

[0087] Figure 7 The block diagram of a V2X beam prediction device based on a multimodal large model according to an embodiment of the present application is shown.

[0088] Refer to Figure 7 As shown, a V2X beam prediction device based on a multimodal large model according to an embodiment of the present application includes:

[0089] An acquisition module, configured to acquire a camera image, positioning data, base station locations of the section where the target vehicle is located, and a three-dimensional map of the section;

[0090] A ray tracing module, configured to use the positioning data, base station locations, and three-dimensional map of the section of the target vehicle as inputs to a pre-constructed ray tracing model, so that the ray tracing model outputs the propagation characteristics of wireless communication signals during the driving process of the target vehicle;

[0091] A projection module, configured to input the propagation characteristics of the wireless communication signals and the camera image into a pre-constructed scatterer projection model to obtain scatterer characteristics in the image;

[0092] An optimization module, configured to perform supervised training on a pre-constructed multimodal large model according to a channel feature data set composed of the camera image and the scatterer characteristics, so as to minimize the prediction error of scatterer coordinates and intensities, and obtain a channel estimation large model;

[0093] A construction module, configured to construct a scatter heat map according to the channel feature data set, and segment a triple data composed of the scatter heat map, positioning, and optimal beam according to the beam coherence time to obtain a beam prediction data set;

[0094] A processing module, configured to train a pre-constructed beam prediction model according to the beam prediction data set to obtain a target beam prediction model for V2X beam prediction.

[0095] In one embodiment, after obtaining the target beam prediction model, the processing module is further configured to:

[0096] Use the channel estimation large model to perform prediction based on multi-view camera images of the vehicle to be predicted to obtain corresponding channel characteristics;

[0097] The base station retrieves according to the channel characteristics of the vehicle to be predicted, and obtains the coordinates and intensities of the projection points in the camera images of each perspective;

[0098] Construct a scattering heat map based on the coordinates and intensities of the projection points, and input the positioning data and the scattering heat map of the vehicle to be predicted into the target beam prediction model, so that the target beam prediction model predicts the beam index and time length within the next beam coherence time.

[0099] Figure 8 The structural schematic diagram of a computer system of an electronic device suitable for implementing the embodiments of the present application is shown.

[0100] It should be noted that Figure 8 The computer system of the shown electronic device is only an example, and should not bring any limitation to the functions and usage scope of the embodiments of the present application.

[0101] As Figure 8 shown, the computer system includes a central processing unit (CPU) 801, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 802 or the program loaded from the storage part 808 into the random access memory (RAM) 803, such as executing the method described in the above embodiments. In the RAM 803, various programs and data required for system operation are also stored. The CPU 801, ROM 802, and RAM 803 are connected to each other through a bus 804. The input / output (I / O) interface 805 is also connected to the bus 804.

[0102] The following components are connected to the I / O interface 805: an input part 806 including a keyboard, a mouse, etc.; an output part 807 including such as a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage part 808 including a hard disk, etc.; and a communication part 809 including a network interface card such as a LAN (Local Area Network) card, a modem, etc. The communication part 809 performs communication processing via a network such as the Internet. The drive 810 is also connected to the I / O interface 805 as needed. A removable medium 811, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 810 as needed, so that the computer program read from it can be installed into the storage part 808 as needed.

[0103] In particular, according to an embodiment of the present application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, an embodiment of the present application includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program contains a computer program for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication section 809, and / or installed from the removable medium 811. When the computer program is executed by the central processing unit (CPU) 801, various functions defined in the system of the present application are executed.

[0104] It should be noted that the computer-readable medium shown in the embodiments of the present application can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, the computer-readable storage medium can be any tangible medium that contains or stores a program, and the program can be used by or in combination with an instruction execution system, apparatus, or device. In the present application, the computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries a computer-readable computer program. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, and the computer-readable medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The computer program contained on the computer-readable medium can be transmitted by any suitable medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.

[0105] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present application. Among them, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the above-mentioned module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, as well as the combination of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0106] The units involved in the embodiments described in the present application can be implemented in software or in hardware, and the described units can also be provided in a processor. Among them, the names of these units do not constitute a limitation to the units themselves in some cases.

[0107] On the other hand, the present application also provides a computer-readable medium, which may be included in the electronic device described in the above embodiments; or may exist separately without being assembled into the electronic device. The above computer-readable medium carries one or more programs, and when the above one or more programs are executed by an electronic device, the electronic device implements the method described in the above embodiments.

[0108] It should be noted that although several modules or units of the devices for action execution are mentioned in the above detailed description, such a division is not mandatory. In fact, according to the embodiments of the present application, the features and functions of the two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.

[0109] Through the description of the above embodiments, those skilled in the art can easily understand that the example embodiments described here can be implemented by software or by a combination of software and necessary hardware. Therefore, the technical solutions according to the embodiments of the present application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, including several instructions to enable a computing device (such as a personal computer, a server, a touch terminal, or a network device, etc.) to execute the method according to the embodiments of the present application.

[0110] Other embodiments of the present application will be readily apparent to those skilled in the art upon consideration of the specification and practice of the embodiments disclosed herein. The present application is intended to cover any variations, uses, or adaptations of the present application, which follow the general principles of the present application and include known common knowledge or conventional technical means in the technical field not disclosed in the present application.

[0111] It should be understood that the present application is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present application is only limited by the appended claims.

Claims

1. A vehicle network beam prediction method based on a multi-modal large model, characterized in that: include: Obtain the target vehicle’s camera image, positioning data, base station location of the road section it is on, and a three-dimensional map of the road section; The positioning data of the target vehicle, the base station location and the three-dimensional map of the road section are used as inputs of a pre-built ray tracing model, so that the ray tracing model outputs the propagation characteristics of the wireless communication signal of the target vehicle during driving; Inputting the propagation characteristics of the wireless communication signal and the camera image into a pre-built scatterer projection model to obtain scatterer features in the image; Performing supervised training on a pre-built multimodal large model according to the camera image and the channel feature data set composed of the scatterer features to minimize the prediction error of the scatterer coordinates and intensity, thereby obtaining a channel estimation large model; Constructing a scattering heat map according to the channel feature data set, and segmenting the ternary data consisting of the scattering heat map, positioning, and optimal beam according to the beam coherence time to obtain a beam prediction data set; The pre-built beam prediction model is trained according to the beam prediction data set to obtain a target beam prediction model for use in Internet of Vehicles beam prediction.

2. The method according to claim 1, characterized in that The scatterer characteristics include the coordinates of the scattering points and the scattering intensity; Then constructing a scattering heat map according to the channel feature data set includes: Calculating based on the camera image in the channel feature data set to obtain the size of the scattering heat map; The projection points are normalized using a radiation function and expanded to obtain a scattering heat map.

3. The method according to claim 1, characterized in that The beam prediction model includes a spatiotemporal feature extractor, a positioning feature extractor, a beam classification network and a beam coherence time classification network, wherein the spatiotemporal feature extractor is used to extract spatiotemporal features of a multi-perspective camera image sequence within a beam coherence time, the positioning feature extractor is used to extract vehicle positioning features, and the beam classification network and the beam coherence network are used to classify and output classification results based on fused features obtained by fusing the output results of the spatiotemporal feature extractor and the output results of the positioning feature extractor.

4. The method according to claim 3, characterized in that After obtaining the target beam prediction model, the method further includes: Using the channel estimation large model to perform prediction based on multi-view camera images of the vehicle to be predicted, to obtain corresponding channel features; The base station searches according to the channel characteristics of the vehicle to be predicted to obtain the coordinates and intensity of the projection point in the camera image of each viewing angle; A scattering heat map is constructed according to the coordinates and intensities of the projection points, and the positioning data of the vehicle to be predicted and the scattering heat map are input into the target beam prediction model so that the target beam prediction model predicts the beam index and time length within the next beam coherence time.

5. The method according to claim 4, characterized in that The multimodal large model includes a depth estimation network, a visual encoder, a learnable connector and a large language model; The channel estimation model is used to make predictions based on the multi-view camera images of the vehicle to be predicted, and the corresponding channel features are obtained, including: Extracting spatial depth information from the camera image using the depth estimation network to obtain the distance from the object in the camera image to the vehicle to be predicted; Utilizing the visual encoder to extract object category features in a camera image; Mapping image features including spatial depth information and object category features to the large language model using the learnable connector to bridge the gap between image and text modalities; The large language model is guided by prompt words to output channel features including scattering point positions and intensities in the camera image based on the angle of ray propagation.

6. A vehicle network beam prediction device based on a multi-modal large model, characterized in that: include: An acquisition module is used to acquire the camera image, positioning data, base station location of the road section and three-dimensional map of the target vehicle; A ray tracing module, used to use the positioning data of the target vehicle, the base station location and the three-dimensional map of the road section as inputs of a pre-built ray tracing model, so that the ray tracing model outputs the propagation characteristics of the wireless communication signal of the target vehicle during driving; A projection module, used for inputting the propagation characteristics of the wireless communication signal and the camera image into a pre-built scatterer projection model to obtain scatterer features in the image; An optimization module, used for performing supervised training on a pre-built multimodal large model based on the camera image and the channel feature data set composed of the scatterer features, so as to minimize the prediction error of the scatterer coordinates and intensity, and obtain a channel estimation large model; A construction module, used to construct a scattering heat map according to the channel feature data set, and segment the triple data consisting of the scattering heat map, positioning and optimal beam according to the beam coherence time to obtain a beam prediction data set; The processing module is used to train a pre-built beam prediction model according to the beam prediction data set to obtain a target beam prediction model for use in Internet of Vehicles beam prediction.

7. A computer readable medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 5 is implemented.

8. An electronic device, characterized in that: include: one or more processors; A storage device for storing one or more programs, which, when executed by the one or more processors, enables the one or more processors to implement the method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Internet-of-vehicles wave beam real-time alignment method based on multi-modal information consciousness

    CN115412844A

  • Beam alignment method and device based on terminal visual perception

    CN116094559A

  • Millimeter wave beam prediction method based on multi-mode sensing data

    CN118509824A

  • Visual assistance millimeter wave beam prediction method for low-light environment

    CN119183120A

  • Millimeter wave beam prediction method of industrial wireless system assisted by multi-modal environment semantics

    CN119402048A

Cited By

  • Internet of vehicles channel prediction method based on multi-modal fusion and related equipment

    CN120342527A