Multi-modal large model-based beam prediction method, apparatus and device, and medium

Through fine-tuning training and feature fusion of multimodal large models, the training overhead and generalization capabilities of beam acquisition methods are solved, and efficient and accurate beam prediction is achieved, which is suitable for mobile scenarios such as the Internet of Vehicles.

CN120389773AActive Publication Date: 2025-07-29BEIJING UNIV OF POSTS & TELECOMM
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510743049.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-05
Publication Date
2025-07-29
Estimated Expiration
2045-06-05

AI Technical Summary

Technical Problem

In the prior art, beam acquisition methods lead to increased training overhead and lack the ability to generalize new scenarios, making it difficult to meet the real-time needs of mobile scenarios such as the Internet of Vehicles.

Method used

A multimodal large model is used for one-time fine-tuning training and second-time fine-tuning training. Using position information and multi-view image features in the target multimodal environment data, multi-modal features are obtained through text embedding and image embedding fusion, and mapped to the beam codebook to select the beam index with the greatest probability.

Benefits of technology

It improves the accuracy of beam prediction, reduces training overhead, and improves the generalization ability of new scenarios, meeting the real-time needs of mobile scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120389773A_ABST
    Figure CN120389773A_ABST
Patent Text Reader

Abstract

The invention provides a beam prediction method and device based on a multi-modal large model, equipment and a medium. The method comprises the following steps: performing primary fine tuning training on a beam prediction model based on the multi-modal large model by using a first multi-modal environment-beam data set; performing secondary fine tuning training on the beam prediction model based on the multi-modal large model after the primary fine tuning training by using the small sample training set; according to the multi-modal large model-based beam prediction model after the secondary fine tuning training and the target multi-modal environment data of the target terminal at the first moment, obtaining a corresponding beam of the target terminal at the first moment; the beam prediction model based on the multi-modal large model is used for obtaining text embedding features and image embedding features according to target multi-modal environment data, obtaining multi-modal features related to beam selection, mapping the multi-modal features to probability distribution of all beams in a beam codebook, and selecting a beam corresponding to a beam index with the maximum probability; the scheme of the invention has relatively high prediction accuracy and relatively strong generalization ability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of wireless communication technologies, and particularly to a beam prediction method, apparatus, device, and medium based on a multimodal large model. Background Art

[0002] In 5G (5th Generation Mobile Communication Technology), beamforming and Massive MIMO (Massive Multiple Input Multiple Output) technologies are key technologies for solving the high path loss problem in high frequency bands. By forming narrow beams, the base station can concentrate its transmission energy along the desired direction, thereby increasing the strength of the received signal. Currently, in mobile scenarios such as vehicle-to-everything (V2X), due to the rapid change of the vehicle's surrounding environment, the optimal beam direction is constantly updated. There are two methods for obtaining the optimal beam, namely the beam scanning method and the beam prediction method based on deep learning. Among them, the beam scanning method needs to traverse all beams in the predefined codebook, and frequent training leads to huge time and resource overheads, and it is difficult to meet the real-time requirements in mobile scenarios. Although the beam prediction method based on deep learning can significantly reduce the training overhead, due to the limited scale of model parameters, it can often only adapt to specific environments and lacks the generalization ability for new scenarios. Summary of the Invention

[0003] At least one embodiment of this application provides a beam prediction method, apparatus, device, and medium based on a multimodal large model, so as to solve the problems in the prior art that the optimal beam acquisition method will lead to an increase in training overhead and a lack of generalization ability for new scenarios.

[0004] To solve the above technical problems, this application is implemented as follows:

[0005] In a first aspect, an embodiment of this application provides a beam prediction method based on a multimodal large model, including:

[0006] Performing a first fine-tuning training on a beam prediction model based on a multimodal large model by using a first multimodal environment-beam data set, to obtain a beam prediction model based on a multimodal large model after the first fine-tuning training; wherein, the first multimodal environment-beam data set is obtained by matching each multimodal environment data in the first multimodal environment data set with the beam index of the maximum received signal power at the corresponding position;

[0007] Performing a second fine-tuning training on the beam prediction model based on a multimodal large model after the first fine-tuning training by using a small sample training set, to obtain a beam prediction model based on a multimodal large model after the second fine-tuning training;

[0008] Based on the beam prediction model based on the multimodal large model after secondary fine-tuning and the acquired target multi-modal environment data of the target terminal at the first moment, obtain the beam corresponding to the target terminal at the first moment;

[0009] Among them, the beam prediction model based on the multimodal large model is used to obtain text embedding features according to the position information in the target multi-modal environment data and obtain image embedding features according to the multi-view images in the target multi-modal environment data, and based on the text embedding features and the image embedding features, obtain multi-modal features related to beam selection, and map the multi-modal features to the probability distribution of all beams in the predefined beam codebook, and select the beam corresponding to the beam index with the highest probability as the beam corresponding to the target terminal at the first moment.

[0010] Optionally, in the beam prediction method based on the multimodal large model, the method further includes:

[0011] Match each multi-modal environment data in the second multi-modal environment data set with the beam index of the maximum received signal power at the corresponding position to obtain a second multi-modal environment-beam data set;

[0012] Select multi-modal environment-beam data with different proportions from the second multi-modal environment-beam data set respectively to construct different small sample training sets.

[0013] Optionally, in the beam prediction method based on the multimodal large model, the method further includes:

[0014] According to the position information and multi-view images of each terminal in the first wireless propagation scenario collected at every preset time interval, obtain the first multi-modal environment data set; and,

[0015] According to the position information and multi-view images of each terminal in the second wireless propagation scenario collected at every preset time interval, obtain the second multi-modal environment data set, where the first wireless propagation scenario and the second wireless propagation scenario are different wireless propagation scenarios.

[0016] Optionally, in the beam prediction method based on the multimodal large model, the beam prediction model based on the multimodal large model includes:

[0017] An input preprocessing module for converting the position information in the target multi-modal environment data into position text and stitching the multi-view images in the target multi-modal environment data into a panoramic image;

[0018] A text embedding module for processing the position text to obtain text embedding features;

[0019] An image encoding module, configured to extract features from the panoramic image to obtain image embedding features;

[0020] A large model module, configured to fuse the image embedding features and the text embedding features to obtain fused features, and extract multi-modal features related to beam selection from the fused features;

[0021] An output mapping module, which maps the multi-modal features to the probability distribution of all beams in the beam codebook, selects the beam index with the highest probability, and obtains a beam according to the beam index.

[0022] Optionally, in the beam prediction method based on a multi-modal large model, a first multi-modal environment dataset is obtained according to the position information and multi-view images of each terminal in the first wireless propagation scenario collected at preset time intervals, including:

[0023] Construct the first wireless propagation scenario, which includes scatterers and moving terminals;

[0024] At preset time intervals, the position information of each terminal is collected through the sensors of each terminal, and the multi-view images of each terminal are collected through the multi-view image collectors of each terminal;

[0025] For each terminal, the position information is associated with the corresponding multi-view image to obtain the multi-modal environment data of the terminal;

[0026] According to the multi-modal environment data of each terminal in the first wireless propagation scenario, the first multi-modal environment dataset is obtained.

[0027] Optionally, in the beam prediction method based on a multi-modal large model, the method further includes:

[0028] Perform ray tracing simulation on the positions of each terminal at preset time intervals to generate multi-path channel state information;

[0029] Using the multi-path channel state information and a predefined beam codebook, calculate the received signal power corresponding to each beam in the beam codebook;

[0030] Obtain the beam index corresponding to the maximum received signal power for the positions of each terminal at preset time intervals.

[0031] Optionally, in the beam prediction method based on a multimodal large model, a fine-tuning training is performed on the beam prediction model based on the multimodal large model using a first multimodal environment-beam dataset to obtain a beam prediction model based on the multimodal large model after the first fine-tuning training, including:

[0032] Select a training set from the first multimodal environment-beam dataset;

[0033] Use the training set and LoRA (Low-Rank Adaptation) technology to perform fine-tuning training on the weights of the beam prediction model based on the multimodal large model to obtain a beam prediction model based on the multimodal large model after the first fine-tuning training.

[0034] Optionally, in the beam prediction method based on a multimodal large model, after performing a first fine-tuning training on the beam prediction model based on the multimodal large model using the first multimodal environment-beam dataset to obtain a beam prediction model based on the multimodal large model after the first fine-tuning training, the method further includes at least one of the following:

[0035] Use the accuracy rate to evaluate the consistency between the beam index obtained by the beam prediction model based on the multimodal large model after the first fine-tuning training and the true index;

[0036] Evaluate the complexity of the beam prediction model based on the multimodal large model after the first fine-tuning training according to at least one of the inference time, video memory occupancy, and computing resource consumption.

[0037] In a second aspect, an embodiment of the present application further provides a beam prediction device based on a multimodal large model, including:

[0038] A first training module for performing a first fine-tuning training on a beam prediction model based on a multimodal large model using a first multimodal environment-beam dataset to obtain a beam prediction model based on the multimodal large model after the first fine-tuning training; wherein, the first multimodal environment-beam dataset is obtained by matching each multimodal environment data in the first multimodal environment dataset with the beam index of the maximum received signal power at the corresponding position;

[0039] A second training module for performing a second fine-tuning training on the beam prediction model based on the multimodal large model after the first fine-tuning training using a small-sample training set to obtain a beam prediction model based on the multimodal large model after the second fine-tuning training;

[0040] A prediction module for obtaining the beam of the target terminal at the first moment according to the beam prediction model based on the multimodal large model after the second fine-tuning training and the acquired target multimodal environment data of the target terminal at the first moment.

[0041] Among them, the beam prediction model based on the multimodal large model is used to convert the position information in the target multimodal environment data into text embedding features and convert the multi-view images in the target multimodal environment data into image embedding features, and obtain multimodal features related to beam selection according to the text embedding features and the image embedding features, and map the multimodal features to the probability distribution of all beams in a predefined beam codebook, and select the beam corresponding to the beam index with the highest probability as the beam of the target terminal at the first moment.

[0042] In a third aspect, an embodiment of the present application further provides a beam prediction device based on a multimodal large model, including: a processor, a memory, and a program stored on the memory and executable on the processor, and when the program is executed by the processor, it implements the beam prediction method based on the multimodal large model as described in the first aspect.

[0043] In a fourth aspect, an embodiment of the present application further provides a readable storage medium, on which a program is stored, and when the program is executed by a processor, it implements the beam prediction method based on the multimodal large model as described in the first aspect.

[0044] In a fifth aspect, an embodiment of the present application further provides a computer program product, including computer instructions, and when the computer instructions are executed by a processor, they implement the beam prediction method based on the multimodal large model as described in the first aspect.

[0045] Compared with the prior art, the embodiments of the present application provide a beam prediction method, apparatus, device and medium based on a multimodal large model. The method includes: performing a first fine-tuning training on a beam prediction model based on a multimodal large model using a first multimodal environment-beam dataset to obtain a beam prediction model based on the multimodal large model after the first fine-tuning training; wherein, the first multimodal environment-beam dataset is obtained by matching each multimodal environment data in the first multimodal environment dataset with the beam index of the maximum received signal power at the corresponding position; performing a second fine-tuning training on the beam prediction model based on the multimodal large model after the first fine-tuning training using a small-sample training set to obtain a beam prediction model based on the multimodal large model after the second fine-tuning training; obtaining the beam corresponding to the target terminal at the first moment according to the beam prediction model based on the multimodal large model after the second fine-tuning training and the collected target multimodal environment data of the target terminal at the first moment; the beam prediction model based on the multimodal large model is used to obtain text embedding features according to the position information in the target multimodal environment data and image embedding features according to the multi-view images in the target multimodal environment data, and according to the text embedding features and image embedding features, obtain multimodal features related to beam selection, and map the multimodal features to the probability distribution of all beams in a predefined beam codebook, and select the beam corresponding to the beam index with the maximum probability as the beam corresponding to the target terminal at the first moment. In this way, the beam prediction accuracy can be improved through the target multimodal environment data and the beam prediction model based on the multimodal large model, and the beam prediction model based on the multimodal large model can avoid the problems of increased training overhead and lack of generalization ability for new scenarios. Description of the Drawings

[0046] By reading the detailed description of the preferred embodiments below, various other advantages and benefits will become clear to those of ordinary skill in the art. The drawings are only for the purpose of illustrating the preferred embodiments and are not considered to be a limitation of the present application. And throughout the drawings, the same reference numerals are used to represent the same components. In the drawings:

[0047] Figure 1 It is a flowchart of the beam prediction method based on the multimodal large model according to the embodiment of the present application;

[0048] Figure 2 It is a schematic diagram of the architecture of the beam prediction model based on the multimodal large model according to the embodiment of the present application;

[0049] Figure 3 It is a flowchart of one implementation manner of the beam prediction method based on the multimodal large model according to the embodiment of the present application;

[0050] Figure 4Schematic diagram of the beam prediction device based on the multimodal large model according to the embodiments of the present application;

[0051] Figure 5 Hardware block diagram of the beam prediction device based on the multimodal large model according to the embodiments of the present application. Detailed implementation manners

[0052] The terms "first", "second", etc. in the present application are used to distinguish similar objects, rather than to describe a specific order or sequence. It should be understood that such terms can be interchanged under appropriate circumstances so that the embodiments of the present application can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first" and "second" are generally of the same category, and the number of objects is not limited. For example, the first object can be one or multiple. In addition, "or" in the present application means at least one of the connected objects. For example, "A or B" covers three scenarios, namely, Scenario 1: including A and not including B; Scenario 2: including B and not including A; Scenario 3: including both A and B. The character " / " generally indicates an "or" relationship between the associated objects before and after.

[0053] It should be noted that the above device provided by the embodiments of the present application can implement all the method steps implemented by the above method embodiments and can achieve the same technical effects. Therefore, the same parts and beneficial effects as those in the method embodiments will not be specifically described in this embodiment.

[0054] Please refer to Figure 1 , the embodiments of the present application provide a beam prediction method based on a multimodal large model, including:

[0055] Step 101: Perform a first fine-tuning training on the beam prediction model based on the multimodal large model by using the first multimodal environment-beam data set, and obtain the beam prediction model based on the multimodal large model after the first fine-tuning training; wherein, the first multimodal environment-beam data set is obtained by matching each multimodal environment data in the first multimodal environment data set with the beam index of the maximum received signal power at the corresponding position;

[0056] Step 102: Perform a second fine-tuning training on the beam prediction model based on the multimodal large model after the first fine-tuning training by using the small sample training set, and obtain the beam prediction model based on the multimodal large model after the second fine-tuning training;

[0057] Step 103: Obtain the beam corresponding to the target terminal at the first moment according to the beam prediction model based on the multimodal large model after the second fine-tuning training and the collected target multimodal environment data of the target terminal at the first moment;

[0058] Among them, the beam prediction model based on the multimodal large model is used to obtain text embedding features according to the position information in the target multimodal environment data, obtain image embedding features according to the multi-view images in the target multimodal environment data, obtain multimodal features related to beam selection according to the text embedding features and the image embedding features, map the multimodal features to the probability distributions of all beams in a predefined beam codebook, and select the beam corresponding to the beam index with the highest probability as the beam corresponding to the target terminal at the first moment.

[0059] It should be noted that inputting the target multimodal environment data of the target terminal at the first moment into the beam prediction model based on the multimodal large model can output the beam corresponding to the target terminal at the first moment. The beam corresponding to the target terminal at the first moment is the optimal beam, which is the beam corresponding to the beam index with the highest probability in the beam codebook.

[0060] In an embodiment of the present application, optionally, before performing one fine-tuning training on the beam prediction model based on the multimodal large model using the first multimodal environment-beam data set, the above method further includes:

[0061] Construct the first wireless propagation scenario, where the first wireless propagation scenario includes scatterers and moving terminals;

[0062] At each preset time interval, collect the position information of each terminal through the sensor of each terminal, and collect the multi-view images of each terminal through the multi-view image collector of each terminal;

[0063] For each terminal, associate each piece of position information with the corresponding multi-view image to obtain the multimodal environment data of the terminal;

[0064] According to the multimodal environment data of each terminal in the first wireless propagation scenario, obtain the first multimodal environment data set.

[0065] The scatterers in the embodiments of the present application can be buildings, and the moving terminals can be vehicles. Therefore, the first wireless propagation scenario can be understood as a 3D vehicle networking simulation scenario. In addition to including scatterers and moving terminals, the first wireless propagation scenario can also include lanes, and multiple vehicles are driving on different lanes. By collecting multi-view images and position information regularly through in-vehicle cameras and positioning devices, multimodal environment data containing information about the scatterers around the vehicles is obtained, providing data support for subsequent beam prediction. Specifically, obtaining the multimodal environment data of the terminal includes the following steps:

[0066] Use a 3D modeling software to build an urban scene map. The scene size is, for example, 200 meters × 200 meters, including 8 lanes and 4 building complexes. Different-sized and -typed buildings are placed in each building complex to simulate the scatterer distribution in a real communication environment. The built urban scene map is the first wireless propagation scene;

[0067] Import the built first wireless propagation scene into the autonomous driving simulator CARLA to simulate the movement of vehicles and the acquisition of sensor data. For example, set 3 small cars to drive at different speeds on different lanes to simulate a dynamic vehicle-to-everything (V2X) communication environment. Install 6 on-vehicle cameras on the top of each vehicle, respectively facing 0°, 60°, 120°, 180°, 240°, and 300° to collect multi-view images and achieve 360° panoramic perception. The image acquisition interval is 0.02 seconds, and the image resolution is 800 × 600 pixels. At the same time, record the three-dimensional position information of the vehicle at each moment according to the positioning device. Each vehicle collected data of 1891 sample points, forming multi-modal environment data containing multi-view images and position information.

[0068] In the embodiment of the present application, optionally, the above method further includes:

[0069] Perform ray tracing simulation on the position of each terminal at preset time intervals to generate multi-path channel state information;

[0070] Use the multi-path channel state information and a predefined beam codebook to calculate the received signal power corresponding to each beam in the beam codebook;

[0071] Obtain the beam index corresponding to the maximum received signal power at the position of each terminal at preset time intervals.

[0072] In the embodiment of the present application, use the ray tracing software Wireless Insite to perform ray tracing simulation on the position of each terminal in the first wireless propagation scene at each sampling moment to generate multi-path channel state information, and calculate the received signal power in combination with a predefined beam codebook to obtain the best beam index, that is, the beam index of the maximum received signal power, and pair it with the multi-modal environment data to construct a first multi-modal environment-beam data set, providing data support for subsequent beam prediction model training. Specifically, constructing the first multi-modal environment-beam data set includes the following steps:

[0073] Import the first wireless propagation scenario into the ray tracing software Wireless Insite. Set the simulation frequency band to 28 GHz. The transmitting antenna is a uniform planar array with 64 antennas, which is fixed at a position 6 meters above the ground beside the lane. The receiving antenna is a single antenna, which is fixed at a position 1.5 meters above the ground on the top of the vehicle. OFDM (Orthogonal Frequency Division Multiplexing) is used for information transmission at the transmitter and receiver ends. The downlink channel state information of the k-th subcarrier can be expressed as follows:

[0074]

[0075] where, α l , τ l and ψ l are the attenuation, time delay and phase shift of the l-th path respectively; f k is the frequency of the k-th subcarrier; L is the number of multipaths; θ l and φ l represent the azimuth angle and elevation angle of the l-th path respectively; a(θ l , φ l ) is the steering vector of the transmitting antenna array. After the ray tracing simulation is completed, the downlink channel state information of each subcarrier is calculated according to the multipath channel state information.

[0076] Use a predefined beam codebook with 64 beams to calculate the received signal power corresponding to each beam with the obtained multipath channel state information, and obtain the best beam from the beam index corresponding to the maximum received signal power, which can be expressed as follows:

[0077]

[0078] where, f * is the best beamforming vector; F is the predefined beam codebook; N s is the number of subcarriers; f is a beamforming vector in the predefined beam codebook. Pair the best beam index with the multimodal environment data at the same position according to the position of the receiving antenna to construct the first multimodal environment-beam dataset.

[0079] In an embodiment of the present application, optionally, the beam prediction model based on the multimodal large model includes:

[0080] An input preprocessing module, configured to convert the position information in the target multimodal environment data into position text, and stitch the multi-view images in the target multimodal environment data into a panoramic image;

[0081] A text embedding module for processing the location text to obtain text embedding features;

[0082] An image encoding module for extracting features from the panoramic image to obtain image embedding features;

[0083] A large model module for fusing the image embedding features and the text embedding features to obtain fused features, and extracting multi-modal features related to beam selection from the fused features;

[0084] An output mapping module that maps the multi-modal features to the probability distribution of all beams in the beam codebook, selects the beam index with the highest probability, and obtains the beam according to the beam index.

[0085] Figure 2 This is a schematic diagram of the architecture of the beam prediction model based on the multi-modal large model according to the embodiments of the present application. As Figure 2 shown, the embodiments of the present application provide a beam prediction model based on a multi-modal large model. The beam prediction model based on the multi-modal large model can be constructed based on the DeepSeek Janus-Pro-1B model. An input preprocessing module for location information and multi-view images is designed to convert location information into location text and stitch multi-view images into a panoramic image. The location text is converted into text embedding features through a tokenizer and a high-dimensional mapper in the text embedding module, and the panoramic image is converted into image embedding features through an encoder and an aligner in the image encoding module. The text embedding features and the image embedding features are used to achieve multi-modal fusion and key feature extraction through a decoder in the large model module to obtain multi-modal features. An output mapping module is designed to map the multi-modal features output by the large model to the probability distribution of each beam in a predefined beam codebook, and the beam corresponding to the beam index with the highest probability is the best beam as the beam corresponding to the target terminal at the first moment. Specifically, the processing process of the beam prediction model based on the multi-modal large model is as follows:

[0086] Design an input preprocessing module. Given that the DeepSeek Janus-Pro-1B model requires the input image to have a fixed size, the sizes of 6 multi-view images are uniformly scaled to the same size and then stitched horizontally into a single panoramic image I ∈ R 384×384×3 , which is used as the input of the image encoding module. The location information of the vehicle is converted from a numerical value to a natural language description to obtain location text T, which is used as the input of the text embedding module.

[0087] Use the text embedding module to process the location text T, and convert it into a discrete sequence through a tokenizer where L pDenote the length of the sequence, and map the discrete sequence into an embedding vector in a high-dimensional continuous space through a high-dimensional mapper to obtain text embedding features where d m is the embedding dimension of the large model module.

[0088] Use the image encoding module to extract features from the panoramic image I, divide the image into 576 image patches of the same size and embed them into a high-dimensional space to obtain image patch embedding features where d v represents the embedding dimension of the image encoder. Subsequently, extract the semantic information of the image through an encoder composed of 24 encoder blocks, and then project the output of the encoder to a space with dimension d m to obtain image embedding features

[0089] Use the large model module to fuse the image embedding feature I e and the text embedding feature T e to establish the relationship between the two modalities through a decoder composed of 24 decoder blocks, and extract the key features related to beam selection to obtain the multi-modal features output by the large model

[0090] Design an output mapping module to project the high-dimensional output of the large model to the probability distribution of all beams in a predefined beam codebook. By calculating the average value, simplify the output multi-modal feature L of the large model module o to a matrix Subsequently, the matrix L m obtains the final output P of the entire model through an MLP (Multilayer Perceptron) with three hidden layers, that is, the probability distribution of the best beam index. Specifically, this process can be expressed as the following formula:

[0091] P = Softmax(MLP(L m ))

[0092] In the embodiments of the present application, optionally, use the first multi-modal environment-beam dataset to perform a fine-tuning training on the beam prediction model based on the multi-modal large model to obtain a beam prediction model based on the multi-modal large model after the first fine-tuning training, including:

[0093] Select a training set from the first multi-modal environment-beam dataset;

[0094] Use the training set and LoRA technology to fine-tune the weights of the beam prediction model based on the multi-modal large model to obtain a beam prediction model based on the multi-modal large model after the first fine-tuning training.

[0095] In the embodiment of the present application, the beam prediction model based on the multimodal large model is fine-tuned and trained using the constructed first multimodal environment-beam dataset. The LoRA technique is used to fine-tune the attention weights in the image encoding module, and only the partial layers participating in task learning in the image encoding module and the large model module are unfrozen to reduce the computational overhead. At the same time, a suitable loss function is selected to optimize the difference between the predicted beam and the true value. Specifically, fine-tuning and training the beam prediction model based on the multimodal large model using the first multimodal environment-beam dataset once includes:

[0096] The entire first multimodal environment-beam dataset is divided into a training set, a validation set, and a test set according to the ratio of 7:1:2. Among them, the training set includes 3971 sample data, the validation set includes 567 sample data, and the test set includes 1135 sample data.

[0097] The LoRA technique is used to fine-tune the weights of the query matrix, key matrix, and value matrix in the self-attention part of the image encoding module. The pre-trained weight matrix is where d v is the embedding dimension of the image encoding module. Specifically, the process of fine-tuning the weights of the beam prediction model based on the multimodal large model using the LoRA technique can be expressed by the following formula:

[0098]

[0099] where the matrix and the matrix contain trainable parameters; r and α represent the rank and scaling factor of LoRA respectively. Usually, r is much smaller than d v , the matrix B is initialized as a zero matrix, and the matrix A is initialized using a random Gaussian distribution. During the training process of the beam prediction model based on the multimodal large model, only the weights of the normalization layer and the root mean square normalization layer in the image encoding module and the large model module are unfrozen to save computational resources.

[0100] The cross-entropy function is used as the loss function during the model training process to minimize the difference between the predicted beam index and the true index during the model training process. The cross-entropy function can be expressed by the following formula:

[0101]

[0102] where M is the number of beams in the codebook; is the predicted best beam probability distribution; P is the true best beam probability distribution. Here, the best beam is the beam with the highest probability. This model is trained using the Adam optimizer, with a batch size of 10, a learning rate of 0.0001, and a total of 200 training epochs.

[0103] In an embodiment of the present application, optionally, after using the first multi-modal environment-beam dataset to perform a first fine-tuning training on the beam prediction model based on the multi-modal large model to obtain the beam prediction model based on the multi-modal large model after the first fine-tuning training, the method further includes at least one of the following:

[0104] Using the accuracy rate to evaluate the consistency between the beam index obtained by the beam prediction model based on the multi-modal large model after the first fine-tuning training and the true index;

[0105] Evaluating the complexity of the beam prediction model based on the multi-modal large model after the first fine-tuning training according to at least one of the inference time, video memory occupancy, and computing resource consumption.

[0106] It should be noted that the beam prediction model based on the multi-modal large model in the embodiment of the present application is comprehensively evaluated from two dimensions of accuracy rate and resource consumption. The accuracy of the model is evaluated by comparing the consistency between the predicted beam and the true beam, and the model complexity is analyzed from aspects such as inference time and video memory occupancy to verify its deployment feasibility in the actual system. Specifically, the comprehensive evaluation of the beam prediction model based on the multi-modal large model includes:

[0107] Using the Top-K accuracy, the accuracy of the model is evaluated by calculating the probability that the correct best beam index appears within the top k highest predicted probabilities. The Top-K accuracy can be expressed as the following formula:

[0108]

[0109] Among them, D is the total number of samples, and I(·) is an indicator function. If the true beam index is within the top k highest predicted probability ranges, it returns 1, otherwise it returns 0. The Top-1 accuracy rate of the beam prediction model based on the multi-modal large model in the embodiment of the present application quickly reaches about 97.5% after 20 rounds of training, and remains stable during the subsequent training process, and finally reaches 98.1%.

[0110] Using the required computing power and model inference time to evaluate the complexity of the model. The required computing power of the beam prediction model based on the multi-modal large model in the embodiment of the present application is 9.2 TFLOPs, and the inference time for each sample data is 8.66 ms.

[0111] In an embodiment of the present application, optionally, the above method further includes:

[0112] Matching each multi-modal environment data in the second multi-modal environment dataset with the beam index of the maximum received signal power at the corresponding position to obtain the second multi-modal environment-beam dataset;

[0113] Select multi-modal environment-beam data with different proportions from the second multi-modal environment-beam dataset to construct different small-sample training sets.

[0114] It should be noted that the second multi-modal environment dataset can be obtained by using the method for obtaining the first multi-modal environment dataset as described above, which will not be elaborated here. The second multi-modal environment-beam dataset is a dataset collected based on a second wireless propagation scenario other than the first wireless propagation scenario, and is used to evaluate the small-sample generalization ability of the model. This second wireless propagation scenario can be a real scenario.

[0115] In the embodiments of the present application, the generalization ability of the beam prediction model based on the multi-modal large model is evaluated under small-sample conditions. Different proportions of small-sample training sets are constructed, and on this basis, the model is fine-tuned and trained. Its performance is evaluated through the accuracy index to verify the adaptability of the model under low data volume conditions in real scenarios. Specifically, the beam prediction model based on the multi-modal large model after the first fine-tuning training is secondarily fine-tuned and trained using the small-sample training set, including:

[0116] Randomly select 10%, 20%, and 30% of the samples from the second multi-modal environment-beam dataset collected from the second wireless propagation scenario to construct different small-sample training sets respectively. For each data volume setting, the generalization ability of the model is measured by evaluating the accuracy of the model.

[0117] The model is fine-tuned and trained under the above different small-sample conditions, and the Top-1 accuracy and Top-3 accuracy are used as evaluation indicators. When the sample proportion is 10%, the Top-1 accuracy and Top-3 accuracy of the beam prediction model in the embodiments of the present application are 60.2% and 79.9% respectively. When the sample proportion is 20%, the Top-1 accuracy and Top-3 accuracy of the model are 68.7% and 85.1% respectively. When the sample proportion is 30%, the Top-1 accuracy and Top-3 accuracy of the model are 72.7% and 92.4% respectively.

[0118] Figure 3 It is a schematic flowchart of one implementation manner of the beam prediction method based on the multi-modal large model described in the embodiments of the present application. As Figure 3 shown, the method includes:

[0119] Step 301: Construct the first wireless propagation scenario and obtain the first multi-modal environment data. Specifically, construct a wireless propagation scenario map, plan the positions of lanes and building complexes in this scenario, and arrange diverse scatterers in each building complex, including static or dynamic objects with different geometric shapes and materials, to reconstruct the first wireless propagation environment. Place multiple terminals moving at different speeds on different lanes in the constructed wireless propagation scenario, and install a certain number of cameras at the top of the terminals to collect multi-view images at fixed time intervals. Save the position information and multi-view images of each terminal as the first multi-modal environment data.

[0120] Step 302: Construct the first multi-modal environment-beam dataset. Specifically, use a ray-tracing channel simulation tool, place a base station as the transmitter at an appropriate position, and arrange the receiving antennas at the positions where the terminals are located at corresponding times. After the ray-tracing simulation is completed, save the channel information, including multi-path channel state information and path loss, etc. According to the predefined beam codebook and the multi-path information collected by ray-tracing, calculate the received signal power corresponding to each beamforming vector, save the index corresponding to the beamforming vector with the maximum received signal power, and associate the best beam index with the environmental information at the same position to form the first multi-modal environment-beam dataset.

[0121] Step 303: Establish a beam prediction model based on a multi-modal large model. Specifically, design an input preprocessing module according to the architecture of the multi-modal large model, resize and splice multiple input pictures into a panoramic picture as the input of the image coding module, and at the same time convert the position information from a numerical value into a position text as the input of the text embedding module. Use the tokenizer of the text embedding module to convert the position text into a discrete sequence, and map the discrete sequence into a high-dimensional continuous representation to obtain the text embedding feature containing the terminal position information. Use the encoder of the image coding module to extract the environmental semantic information in the panoramic picture, including environmental features related to channel propagation such as occlusion and scatterer distribution, to obtain the image embedding feature that can reflect the wireless propagation environment. Use the decoder of the large model module to fuse different modal information from the text embedding feature and the image embedding feature, and extract the key features that affect beam selection. Design an output mapping module according to the output of the large model module and the number of beams in the predefined beam codebook, map the high-dimensional output of the large model to the probability distribution of all beams in the beam codebook, and the beam corresponding to the index with the highest probability is the best beam.

[0122] Step 304: Fine-tune and train the beam prediction model based on the multi-modal large model using the first multi-modal environment-dataset. Specifically, divide the first multi-modal environment-dataset into a training set, a validation set, and a test set according to an appropriate ratio. Use the LoRA technique to fine-tune the weights of the multi-head self-attention part in the image encoding module, and only unfreeze the weights of specific layers in the image encoding module and the large model module during the training process to save computing resources. Select an appropriate loss function to minimize the difference between the predicted beam index and the true value during model training.

[0123] Step 305: Evaluate the accuracy and complexity of the beam prediction model based on the multi-modal large model. Specifically, use the accuracy rate to evaluate the consistency between the beam index output by the model and the true best beam to measure its prediction accuracy. Evaluate the complexity of the model from aspects such as inference time, video memory occupancy, and computing resource consumption to verify its deployment feasibility in an actual system.

[0124] Step 306: Evaluate the few-shot generalization ability of the beam prediction model based on the multi-modal large model based on the second multi-modal environment-beam dataset. Specifically, randomly select a small amount of data, such as 10%, 20%, 30%, etc., from the second multi-modal environment-beam dataset to construct few-shot training sets with different data volumes. Under the corresponding few-shot conditions, fine-tune and train the model, and use the accuracy rate to evaluate its performance to verify the generalization ability of the model under low data volume conditions.

[0125] In summary, by adopting the beam prediction method based on the multi-modal large model described in the embodiments of the present application, the position information and multi-view images collected by the terminal are fully utilized, and joint modeling is performed through the multi-modal large model architecture, enabling the model to understand and represent the characteristics of complex communication environments. The environmental characteristics mainly include factors such as the distribution of obstacles and building structures that have a significant impact on the signal propagation path. Through the deep fusion of multi-view images and position information, the adaptability of the beam prediction model to environmental changes is effectively improved, overcoming the limitations of single-modal in terms of representation ability. Moreover, the embodiments of the present application adopt a multi-modal large model architecture based on DeepSeek Janus-Pro-1B and combine it with the LoRA technology, enabling the model to more accurately capture the key environmental characteristics affecting beam selection while retaining the pre-trained knowledge. Through the fusion of multi-view images and position information, as well as the powerful knowledge transfer ability of the large model, the discrimination performance of the best beam is enhanced, making the model have higher prediction accuracy. In addition, the embodiments of the present application perform transfer learning by means of parameter fine-tuning based on the large model, enabling the model to still maintain good prediction performance when facing new scenarios with limited data volume. This small-sample adaptation ability significantly enhances the practicality of the model, meets the communication requirements in variable environments such as real vehicle-to-everything (V2X), and effectively alleviates the problem of the dependence of current deep learning-based beam prediction methods on large-scale data.

[0126] Please refer to Figure 4 , the embodiments of the present application further provide a beam prediction device based on a multi-modal large model, including:

[0127] The first training module 401 is configured to perform a first fine-tuning training on a beam prediction model based on a multi-modal large model by using a first multi-modal environment-beam data set, and obtain a beam prediction model based on the multi-modal large model after the first fine-tuning training; wherein, the first multi-modal environment-beam data set is obtained by matching each multi-modal environment data in the first multi-modal environment data set with the beam index of the maximum received signal power at the corresponding position;

[0128] The second training module 402 is configured to perform a second fine-tuning training on the beam prediction model based on the multi-modal large model after the first fine-tuning training by using a small-sample training set, and obtain a beam prediction model based on the multi-modal large model after the second fine-tuning training;

[0129] The prediction module 403 obtains the beam of the target terminal at the first moment according to the beam prediction model based on the multi-modal large model after secondary fine-tuning training and the acquired target multi-modal environment data of the target terminal at the first moment; wherein, the beam prediction model based on the multi-modal large model is used to convert the position information in the target multi-modal environment data into text embedding features and convert the multi-view images in the target multi-modal environment data into image embedding features, and according to the text embedding features and the image embedding features, obtain multi-modal features related to beam selection, and map the multi-modal features to the probability distribution of all beams in the predefined beam codebook, and select the beam corresponding to the beam index with the highest probability as the beam of the target terminal at the first moment.

[0130] Optionally, for the beam prediction device based on the multi-modal large model, wherein the device further includes:

[0131] An obtaining module, configured to match each multi-modal environment data in the second multi-modal environment data set with the beam index of the maximum received signal power at the corresponding position, and obtain a second multi-modal environment-beam data set;

[0132] A construction module, configured to respectively select different proportions of multi-modal environment-beam data from the second multi-modal environment-beam data set to construct different small sample training sets.

[0133] Optionally, for the beam prediction device based on the multi-modal large model, wherein the device further includes:

[0134] A first acquisition module, configured to obtain a first multi-modal environment data set according to the position information and multi-view images of each terminal in the first wireless propagation scenario collected at preset time intervals; and,

[0135] According to the position information and multi-view images of each terminal in the second wireless propagation scenario collected at intervals of the preset time duration, obtain the second multi-modal environment data set, where the first wireless propagation scenario and the second wireless propagation scenario are different wireless propagation scenarios.

[0136] Optionally, for the beam prediction device based on the multi-modal large model, wherein the beam prediction model based on the multi-modal large model includes:

[0137] An input preprocessing module, configured to convert the position information in the target multi-modal environment data into position text, and splice the multi-view images in the target multi-modal environment data into a panoramic image;

[0138] A text embedding module, configured to process the position text to obtain text embedding features;

[0139] An image encoding module for extracting features from the panoramic image to obtain image embedding features;

[0140] A large model module for fusing the image embedding features and the text embedding features to obtain fused features, and extracting multi-modal features related to beam selection from the fused features;

[0141] An output mapping module that maps the multi-modal features to the probability distribution of all beams in the beam codebook, selects the beam index with the maximum probability, and obtains the beam according to the beam index.

[0142] Optionally, in the beam prediction device based on a multi-modal large model, the first acquisition module is specifically configured to:

[0143] Construct the first wireless propagation scenario, where the first wireless propagation scenario includes scatterers and moving terminals;

[0144] At each interval of the preset duration, collect the position information of each terminal through the sensors of each terminal, and collect the multi-view images of each terminal through the multi-view image collectors of each terminal;

[0145] For each terminal, associate each position information with the corresponding multi-view image to obtain the multi-modal environment data of the terminal;

[0146] According to the multi-modal environment data of each terminal in the first wireless propagation scenario, obtain the first multi-modal environment data set.

[0147] Optionally, in the beam prediction device based on a multi-modal large model, the device further includes:

[0148] A generation module for performing ray tracing simulation on the positions of each terminal at each interval of the preset duration to generate multi-path channel state information;

[0149] A calculation module for calculating the received signal power corresponding to each beam in the beam codebook by using the multi-path channel state information and a predefined beam codebook;

[0150] A second acquisition module for obtaining the beam index corresponding to the maximum received signal power at the positions of each terminal at each interval of the preset duration.

[0151] Optionally, in the beam prediction device based on a multi-modal large model, the first training module is specifically configured to:

[0152] Select a training set from the first multi-modal environment-beam data set;

[0153] Fine-tune the weights of the beam prediction model based on the multi-modal large model using the training set and low-rank adaptation technology to obtain the beam prediction model based on the multi-modal large model after the first fine-tuning training.

[0154] Optionally, for the beam prediction device based on the multi-modal large model, the device further includes at least one of the following:

[0155] The first evaluation module is used to evaluate the consistency between the beam index obtained by the beam prediction model based on the multi-modal large model after the first fine-tuning training and the true index using the accuracy rate;

[0156] The second evaluation module is used to evaluate the complexity of the beam prediction model based on the multi-modal large model after the first fine-tuning training according to at least one of the inference time, video memory occupancy, and computing resource consumption.

[0157] It should be noted that the above device provided in the embodiments of the present application can implement all the method steps implemented in the above method embodiments, and can achieve the same technical effects. The same parts and beneficial effects as those in the method embodiments in this embodiment will not be specifically described here.

[0158] The embodiments of the present application also provide a beam prediction device based on a multi-modal large model, as Figure 5 shown, including:

[0159] A processor 501, a memory 502, a transceiver 503, and a program or instruction stored on the memory 502 and executable on the processor 501; when the processor 501 executes the program or instruction, it implements each process of the above method embodiment of the beam prediction method based on the multi-modal large model, and can achieve the same technical effects. To avoid repetition, it will not be elaborated here.

[0160] The transceiver 503 is used to receive and send data under the control of the processor 501.

[0161] Among them, in Figure 5Among them, the bus architecture may include any number of interconnected buses and bridges, specifically, various circuits represented by one or more processors represented by processor 501 and a memory represented by memory 502 are linked together. The bus architecture may also link together various other circuits such as peripheral devices, voltage regulators, and power management circuits, which are well known in the art, and thus will not be further described herein. The bus interface provides an interface. The transceiver 503 may be multiple components, that is, including a transmitter and a receiver, and provides a unit for communicating with various other devices on a transmission medium. For different user devices, the user interface 504 may also be an interface capable of externally connecting and internally connecting required devices, and the connected devices include but are not limited to a keypad, a display, a speaker, a microphone, a joystick, etc.

[0162] The processor 501 is responsible for managing the bus architecture and general processing, and the memory 502 may store data used by the processor 501 when performing operations.

[0163] The embodiment of the present application also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements each process of the above embodiment of the beam prediction method based on a multimodal large model and can achieve the same technical effect. To avoid repetition, it will not be elaborated here. Among them, the computer-readable storage medium is, for example, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc, etc.

[0164] The embodiment of the present application also provides a computer program product, including computer instructions. When the computer instructions are executed by a processor, it implements each process of the above embodiment of the beam prediction method based on a multimodal large model and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.

[0165] It should be noted that in this article, the term "including", "containing in", or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article, or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such a process, method, article, or device. Without more limitations, an element defined by the statement "including one..." does not exclude the existence of another identical element in the process, method, article, or device including that element.

[0166] Through the description of the above embodiments, those skilled in the art can clearly understand that the above method of the embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases the former is a better implementation. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. The computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes several instructions for causing a terminal (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of the present application.

[0167] The embodiments of the present application have been described above in conjunction with the accompanying drawings. However, the present application is not limited to the above specific implementation manners. The above specific implementation manners are merely illustrative rather than restrictive. Under the inspiration of the present application, those of ordinary skill in the art can also make many forms without departing from the purpose of the present application and the scope protected by the claims, and all of them belong to the protection scope of the present application.

Claims

1. A beam prediction method based on a multimodal large model, characterized in that Including: Performing a first fine-tuning training on a beam prediction model based on a multimodal large model using a first multimodal environment-beam dataset to obtain a beam prediction model based on the multimodal large model after the first fine-tuning training; wherein, the first multimodal environment-beam dataset is obtained by matching each multimodal environment data in the first multimodal environment dataset with the beam index of the maximum received signal power at the corresponding position. Performing a second fine-tuning training on the beam prediction model based on the multimodal large model after the first fine-tuning training using a small sample training set to obtain a beam prediction model based on the multimodal large model after the second fine-tuning training. Obtaining the beam corresponding to the target terminal at the first moment according to the beam prediction model based on the multimodal large model after the second fine-tuning training and the target multimodal environment data of the target terminal collected at the first moment. Wherein, the beam prediction model based on the multimodal large model is used to obtain text embedding features according to the position information in the target multimodal environment data and image embedding features according to the multi-view images in the target multimodal environment data, and obtain multimodal features related to beam selection according to the text embedding features and the image embedding features, and map the multimodal features to the probability distribution of all beams in a predefined beam codebook, and select the beam corresponding to the beam index with the highest probability as the beam corresponding to the target terminal at the first moment.

2. The beam prediction method based on a multimodal large model according to claim 1, wherein, The method further includes: Matching each multimodal environment data in the second multimodal environment dataset with the beam index of the maximum received signal power at the corresponding position to obtain a second multimodal environment-beam dataset. Selecting different proportions of multimodal environment-beam data from the second multimodal environment-beam dataset respectively to construct different small sample training sets.

3. The beam prediction method based on a multimodal large model according to claim 2, wherein The method further includes: Obtaining the first multimodal environment dataset according to the position information and multi-view images of each terminal in the first wireless propagation scenario collected at every preset time interval; and Obtaining the second multimodal environment dataset according to the position information and multi-view images of each terminal in the second wireless propagation scenario collected at every preset time interval, where the first wireless propagation scenario and the second wireless propagation scenario are different wireless propagation scenarios.

4. The beam prediction method based on a multimodal large model according to claim 1, wherein The beam prediction model based on the multimodal large model includes: An input preprocessing module for converting the position information in the target multimodal environment data into position text and stitching the multi-view images in the target multimodal environment data into a panoramic image. A text embedding module for processing the position text to obtain text embedding features. An image encoding module for extracting features from the panoramic image to obtain image embedding features. A large model module for fusing the image embedding features and the text embedding features to obtain a fused feature and extracting multimodal features related to beam selection from the fused feature. An output mapping module for mapping the multimodal features to the probability distribution of all beams in the beam codebook, selecting the beam index with the highest probability, and obtaining a beam according to the beam index.

5. The beam prediction method based on a multimodal large model according to claim 3, wherein Obtain the first multi-modal environment dataset based on the position information and multi-view images of each terminal in the first wireless propagation scenario collected at preset time intervals, including: Construct the first wireless propagation scenario, which includes scatterers and moving terminals; At each preset time interval, collect the position information of each terminal through the sensors of each terminal, and collect the multi-view images of each terminal through the multi-view image collectors of each terminal; For each terminal, associate each position information with the corresponding multi-view image to obtain the multi-modal environment data of the terminal; Obtain the first multi-modal environment dataset according to the multi-modal environment data of each terminal in the first wireless propagation scenario.

6. The beam prediction method based on a multimodal large model according to claim 1, wherein The method further includes: Perform ray tracing simulation on the positions of each terminal at preset time intervals to generate multi-path channel state information; Use the multi-path channel state information and a predefined beam codebook to calculate the received signal power corresponding to each beam in the beam codebook; Obtain the beam index corresponding to the maximum received signal power at the positions of each terminal at preset time intervals.

7. The beam prediction method based on a multimodal large model according to claim 1, wherein Perform one-time fine-tuning training on the beam prediction model based on the multi-modal large model using the first multi-modal environment-beam dataset to obtain the beam prediction model based on the multi-modal large model after one-time fine-tuning training, including: Select a training set from the first multi-modal environment-beam dataset; Use the training set and low-rank adaptation technology to fine-tune the weights of the beam prediction model based on the multi-modal large model to obtain the beam prediction model based on the multi-modal large model after one-time fine-tuning training.

8. The beam prediction method based on a multimodal large model according to claim 1, wherein After performing one-time fine-tuning training on the beam prediction model based on the multi-modal large model using the first multi-modal environment-beam dataset to obtain the beam prediction model based on the multi-modal large model after one-time fine-tuning training, the method further includes at least one of the following: Use the accuracy rate to evaluate the consistency between the beam index obtained by the beam prediction model based on the multi-modal large model after one-time fine-tuning training and the true index; Evaluate the complexity of the beam prediction model based on the multi-modal large model after one-time fine-tuning training according to at least one of the inference time, video memory occupancy, and computing resource consumption.

9. A beam prediction device based on a multimodal large model, characterized in that, Include: The first training module is used to perform one-time fine-tuning training on the beam prediction model based on the multi-modal large model using the first multi-modal environment-beam dataset to obtain the beam prediction model based on the multi-modal large model after one-time fine-tuning training; wherein, the first multi-modal environment-beam dataset is obtained by matching each multi-modal environment data in the first multi-modal environment dataset with the beam index of the maximum received signal power at the corresponding position; The second training module is used to perform secondary fine-tuning training on the beam prediction model based on the multi-modal large model after one-time fine-tuning training using a small sample training set to obtain the beam prediction model based on the multi-modal large model after secondary fine-tuning training; A prediction module, which obtains the beam of the target terminal at the first moment according to the beam prediction model based on the multimodal large model after secondary fine-tuning training and the collected target multimodal environment data of the target terminal at the first moment. Among them, the beam prediction model based on the multimodal large model is used to convert the position information in the target multimodal environment data into text embedding features and convert the multi-view images in the target multimodal environment data into image embedding features, and obtain multimodal features related to beam selection according to the text embedding features and the image embedding features, and map the multimodal features to the probability distribution of all beams in a predefined beam codebook, and select the beam corresponding to the beam index with the highest probability as the beam of the target terminal at the first moment.

10. A beam prediction device based on a multimodal large model, characterized in that, It includes: A processor, a memory, and a program stored on the memory and executable on the processor. When the program is executed by the processor, it implements the beam prediction method based on the multimodal large model according to any one of claims 1 to 8.

11. A readable storage medium, characterized in that, A program is stored on the readable storage medium. When the program is executed by the processor, it implements the beam prediction method based on the multimodal large model according to any one of claims 1 to 8.

12. A computer program product, characterized in that, It includes computer instructions. When the computer instructions are executed by the processor, it implements the beam prediction method based on the multimodal large model according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Beam tracking method and device, equipment and storage medium

    CN117560046A

  • Millimeter wave beam tracking method based on contrast learning

    CN119154979A

  • Unmanned aerial vehicle beam prediction method and device based on multi-modal information intelligent screening mechanism

    CN119962337A