Visual assistance millimeter wave beam prediction method for large view scene
Through object detection technology, an interference information is filtered, and a neural network model based on background filtering is built, which solves the problem of low beam prediction accuracy in large-view scenarios, and improves the stability and reliability of the millimeter wave communication system.
Patent Information
- Application Number
- CN202510611214.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-13
- Publication Date
- 2025-07-04
AI Technical Summary
In large-view scenarios, in millimeter wave communication systems, the image contains a large amount of interference information, making it difficult for deep learning networks to learn effective features, and the beam prediction accuracy is reduced, which affects system stability and reliability.
The object detection algorithm is used to identify the target equipment in the image, and the bounding box coordinates are used to filter interference information, and a neural network model based on background filtering is built to predict the optimal beam.
By accurately identifying the target equipment, the complexity of model training is reduced, the beam prediction accuracy is improved, the system stability and reliability are enhanced, the dynamic matching of channel state changes is achieved, and data transmission reliability is ensured.
Smart Images

Figure CN120263249A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of wireless communication and relates to a vision-assisted millimeter-wave beam prediction method for large-view scenarios. Background Art
[0002] Millimeter waves have characteristics such as large bandwidth and low latency, and can achieve efficient data transmission. They have become a key technology to meet the real-time high-speed communication requirements of 5G and 6G. Since millimeter-wave communication has a short transmission distance and is easily blocked, it is very likely to reduce the reliability and stability of communication. Therefore, it is necessary to form narrow beams through beam management to achieve long-distance wireless communication and ensure the reliability of millimeter-wave communication. However, the traditional beam-scanning-based management method has huge beam training overhead and resource waste. To solve this problem, a wireless communication scheme assisted by visual perception data can learn the environmental change rules in the communication environment through a deep learning framework, and avoid training overhead while predicting the optimal communication beam.
[0003] Visual data can effectively reflect the change rules of wireless communication users in the communication environment, which is helpful for realizing beamforming of the optimal path and selecting the optimal communication beam. Therefore, visual data is widely used in millimeter-wave beam prediction. However, the image captured by the base station has a wide coverage range, and the pixel volume of the terminal device in the image shrinks sharply, making it more difficult to extract image feature information, thus increasing the difficulty of beam prediction. In addition, there are multiple types of interference information in the captured large-view scene image, which will generate feature information similar to that of the user equipment, affecting the parameter optimization performance of the network model and reducing the accuracy of the optimal beam prediction, thereby affecting the stability and reliability of the system. Therefore, how to improve the accuracy of beam prediction in large-view scenarios has become an important problem to be solved urgently. Summary of the Invention
[0004] In view of this, the purpose of the present invention is to provide a vision-assisted millimeter-wave beam prediction method for large-view scenarios. In large-view scenarios, there is a lot of interference information in the image, and it is difficult for the deep learning network model to learn effective features, resulting in reduced prediction. The present invention uses object detection technology to detect the target device in the image, reduce the model training complexity, and improve the prediction accuracy.
[0005] To achieve the above purpose, the present invention provides the following technical solutions:
[0006] A visual-aided millimeter-wave beam prediction method for large-view scenarios. Aiming at the problem that there is a lot of interference information in the large-view scenario images, making it difficult for deep learning network models to learn effective features and resulting in reduced prediction, a beam prediction method based on object detection is proposed. The object detection algorithm is used to detect the target devices in the large-view images, and the interference information in the images is filtered according to the bounding box coordinates of the object detection. Then, a neural network model is built based on the image data filtered by the background to predict the optimal beam. The method specifically includes the following steps:
[0007] S1: Obtain the parameters of the millimeter-wave wireless communication system, construct a millimeter-wave wireless communication system model, and define the system optimal beam selection strategy;
[0008] S2: Define the optimization problem of visual-aided millimeter-wave beam prediction, construct a visual-aided beam prediction scheme in the large-view scenario, use the neural network framework to learn the mapping relationship between the images of the communication scenario and the optimal communication beam, predict the optimal beam vector, achieve the maximization of the received power, and ensure reliable data transmission;
[0009] S3: Collect the visual data of the millimeter-wave wireless communication scenario and obtain the corresponding optimal beam vectors in the codebook, preprocess the collected image data, construct a millimeter-wave beam prediction model based on the image data, and define the loss function and model optimizer, etc.;
[0010] S4: Use the object detection algorithm to extract the bounding box coordinates of the object targets in the collected images, and then calculate the coverage range of the target devices in the images according to the bounding box coordinates, discard the image information outside the coverage range of the target devices, obtain the image data only containing the information of the target devices, use the obtained image data to train the time series neural network model, update the network parameters, and after the loss function tends to be stable, obtain the optimal beam prediction network for the large-view scenario, and then use the optimal prediction network to achieve the optimal beam selection.
[0011] Furthermore, in step S1, when constructing the millimeter-wave wireless communication system model, it specifically includes: defining the communication system as a millimeter-wave communication between a base station and a mobile user, equipping the base station with a camera and a millimeter-wave phased array, using the camera to capture the real-time images of the communication scenario for training the network model, the phased array contains a ULA antenna array, there are M antennas in the antenna array, and weights are assigned to the antenna array through a predefined codebook to achieve beamforming. The codebook is expressed as where represents the beamforming vector in the codebook, represents a complex vector with a dimension of M×1, and Q represents the total number of vectors in the codebook; the system adopts orthogonal frequency division multiplexing (OFDM) technology to transmit signals in parallel on K subcarriers, and the communication channel corresponding to each subcarrier is expressed as h k, for k = 1, 2, ..., K, if the base station sends a downlink signal to the user at the t-th moment, the signal received by the user is expressed as:
[0012]
[0013] where y k [t] represents the received signal, v k [t] represents the interference signal, and x represents the transmitted data signal.
[0014] Furthermore, in step S1, the system optimal beam selection strategy is defined, specifically including: the optimization objective of the system is to select the optimal communication beam. At each moment t, the beamforming vector w q [t] is selected from the beam vector codebook to maximize the average received SNR of the system, as shown in the following formula:
[0015]
[0016] where w q [t]* represents the optimal beamforming vector.
[0017] Furthermore, step S2 specifically includes the following steps:
[0018] S21: For the optimization objective proposed in step S1, the image of the communication scenario is used to optimize the beam selection. Let represent the RGB image taken at time t, where W, H, and C represent the width, height, and color channel number of the image respectively. Then, for the image data X[t] at the t-th moment, the beam prediction optimization problem is defined as constructing a mapping function that can predict the optimal communication beam w★[t] through the image data samples;
[0019] S22: The image data in continuous time can accurately reflect the mapping relationship between the image and the optimal communication beam. Find the optimal mapping function. For the time series of m RGB images the beam mapping function can be expressed as:
[0020]
[0021] where, represents the mapping function of the network model, θ w represents the optimization parameter of the network model, and w[τ] represents the optimal beam predicted at the τ-th moment; since there is a one-to-one correspondence between the beam vectors in the codebook and their indices, a mapping function from the image to the beam index can be learned more simply, and the original formula can be expressed as:
[0022]
[0023] Among them, represents the corresponding model mapping function, represents the predicted beam index;
[0024] S23: For an image dataset The optimal mapping function needs to satisfy having the largest number of correct predictions on all samples in the dataset D, as shown in the following formula:
[0025]
[0026] Among them, represents the optimal mapping function, f Θ (.) represents the set of all mapping functions, N represents the number of samples in the dataset, s n represents the true index label, represents the predicted optimal beam index, and represents the probability that the model predicts correctly under the input X n below.
[0027] Furthermore, in step S3, a millimeter wave beam prediction model is constructed, specifically including the following steps:
[0028] S31: Use a base station equipped with a camera to capture real-time large view scene images. The images are in RGB format and contain multiple feature information in the communication environment. The base station selects the optimal beamforming vector w from a preset beam codebook * , and the codebook contains multiple beamforming vectors to cover the entire communication scene;
[0029] S32: Construct a beam prediction neural network framework. Use the YOLOv7 detection model to identify the target device in the image, obtain the target bounding box coordinates, calculate the specific position and coverage range of the target device based on the bounding box coordinates, and then perform background filtering to obtain image data containing only the target device information;
[0030] S33: Input the image data obtained through filtering into the LSTM network model. Each time step of the LSTM model uses a fully connected layer to map the features to the beamforming vector space, and calculates the probability distribution of each beamforming vector through the Softmax function. The training of the model uses the cross-entropy loss function, and the optimization goal is to maximize the probability of correctly predicting the beamforming vector. The cross-entropy loss function is expressed as:
[0031]
[0032] Among them, n is the number of training samples; P(sn |X n ) is the probability that the model predicts the beam vector index s n given the image X. To further improve the generalization ability of the model, the Adam optimizer is used to dynamically adjust the model parameters to ensure that the model converges quickly during training. n
[0033] Furthermore, step S4 specifically includes the following steps:
[0034] S41: Use the YOLOv7 object detection neural network model to identify the target objects in the large-view scene image, and obtain the center coordinates (x c , y c ) of the normalized target bounding box center, as well as the width and height (w, h) of the bounding box, and the predicted class. The predicted output class is set to daily communication users (such as cars, pedestrians, etc.) to avoid detecting redundant classes, reduce the model complexity, and improve the model performance. The specific loss function is as follows:
[0035] L = L loc + L conf + L cls
[0036] where L represents the total model loss, L loc represents the bounding box loss, which is used to optimize the center coordinates and size of the predicted box; L conf represents the confidence loss, which measures the degree of matching between the predicted bounding box and the true target; L cls represents the classification loss; the bounding box loss function uses the CIOU loss.
[0037] S42: Calculate the actual position and coverage range of the target device in the image according to the bounding box output by the object detection. The pixel height of the image data is height, and the pixel width is width. Since the output bounding box coordinates are in normalized form, the actual size of the bounding box in the image is:
[0038]
[0039] where represents the de-normalized center coordinates of the bounding box, and (w pixel , h pixel ) represents the de-normalized width and height of the bounding box;
[0040] Furthermore, the upper-left coordinates (x min , y min ) and the lower-right coordinates (x max , y max ) of the bounding box can be calculated as follows:
[0041]
[0042] Based on the actual coordinates of the vertices of the bounding box in the image data, determine the specific position and coverage range of the target device in the image, and then implement background filtering to discard the remaining complex interference background information, and only retain the key image information of the target device;
[0043] S43: The image after target detection and background filtering is input into the LSTM time series network. The hidden state of the LSTM is initialized with all zeros. The entire prediction network contains m LSTM prediction units. Each unit is a feature extraction module composed of two convolutional layers and two fully connected layers. A layer of LSTM network is connected behind this module. The image data at each moment is input into the corresponding LSTM prediction unit at that moment to achieve feature extraction; and the prediction unit at the last moment realizes the optimal beam prediction through the connected classifier, and uses the cross-entropy loss as the final loss function;
[0044] S44: Train the time series neural network model, update the network parameters. After the loss function tends to be stable, obtain the optimal beam prediction network for the large-view scenario, and then use the optimal prediction network to achieve the optimal wave speed selection.
[0045] The beneficial effects of the present invention are as follows:
[0046] (1) Efficient filtering of interference information and feature focusing
[0047] By introducing the YOLOv7 target detection technology, accurately identify the target devices (such as cars, pedestrians) in the image, and use the bounding box coordinates to eliminate the complex background interference information in the large-view scenario, and only retain the key visual features of the target device. This innovation solves the problems of low model learning efficiency and difficult feature extraction caused by redundant image information in traditional methods, significantly reduces the complexity of model training, and at the same time improves the beam prediction accuracy.
[0048] (2) Temporal dynamic modeling and prediction performance optimization
[0049] Adopt the LSTM time series neural network to model the continuous time series image data after background filtering, and capture the time evolution law of the beam vector. Through the cascaded design of multiple convolutional and fully connected modules, combined with the cross-entropy loss function and the Adam dynamic optimization strategy, realize the dynamic matching of beam selection and channel state change. This method improves the model prediction accuracy and maximizes the system average received signal-to-noise ratio (SNR), ensuring the stability of 5G / 6G millimeter-wave communication and the reliability of data transmission.
[0050] (3) Balance between system complexity and generalization ability
[0051] By limiting the target detection categories (only objects related to communication users), redundant calculations are avoided, and the number of model parameters is reduced. At the same time, the CIOU loss function is used to optimize the bounding box positioning accuracy, and the Softmax probability distribution is used to enhance the classification robustness. This design takes into account both the algorithm efficiency and generalization ability, enabling the model to converge quickly in complex scenarios and adapt to the position changes of mobile devices, providing a feasibility guarantee for the deployment in actual communication scenarios.
[0052] Other advantages, objectives, and features of the present invention will, to some extent, be described in the subsequent specification, and to some extent, will be obvious to those skilled in the art based on an examination of the following text, or can be learned from the practice of the present invention. The objectives and other advantages of the present invention can be achieved and obtained through the following specification. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] In order to make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be described in detail preferably with reference to the accompanying drawings, where:
[0054] Figure 1 It is a framework diagram of a vision-assisted millimeter-wave beam prediction based on object detection for large-view scenarios according to the present invention;
[0055] Figure 2 It is a flowchart of a vision-assisted millimeter-wave beam prediction method based on object detection according to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0056] The following specific examples illustrate the embodiments of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the diagrams provided in the following embodiments only illustrate the basic concept of the present invention schematically. Without conflict, the following embodiments and the features in the embodiments can be combined with each other.
[0057] Among them, the drawings are only for illustrative purposes, showing only schematic diagrams, not physical diagrams, and should not be construed as a limitation to the present invention; for better illustrating the embodiments of the present invention, some components in the drawings will be omitted, enlarged, or reduced, and do not represent the dimensions of actual products; for those skilled in the art, it is understandable that some well-known structures and their descriptions in the drawings may be omitted.
[0058] In the accompanying drawings of the embodiments of the present invention, the same or similar reference numerals correspond to the same or similar components; in the description of the present invention, it should be understood that if there are terms such as "upper", "lower", "left", "right", "front", "rear", etc. indicating the orientation or positional relationship based on the orientation or positional relationship shown in the drawings, it is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, the terms describing the positional relationship in the drawings are only for illustrative purposes and cannot be construed as a limitation of the present invention. For those of ordinary skill in the art, the specific meanings of the above terms can be understood according to specific circumstances.
[0059] Please refer to Figures 1 to 2 , the present invention is directed to a millimeter-wave wireless communication system in a large-view scenario, which uses visual perception data of the communication scenario to assist the base station in selecting the optimal communication beam to ensure a highly reliable connection with mobile users. In view of the large amount of interference information contained in the large-view image and the relatively small pixel volume of the device, the target detection technology is used to identify the target device in the image, and background filtering is performed based on the output bounding box coordinates of the target detection to discard the redundant interference information and obtain image data that only retains important device information. Then, these image data are input into the final temporal neural network model to achieve optimal millimeter-wave beam selection.
[0060] Figure 1 It is a schematic diagram of the beam prediction network architecture based on target detection. First, the YOLO-v7 target detection neural network is used to perform target detection on the input image data. The YOLO-v7 network can quickly and efficiently detect multi-scale target devices in the image and output the bounding box coordinates of each device. According to these bounding box coordinates, the specific position and range of the target device in the image can be further calculated, and this information is used for background filtering. The purpose of background filtering is to remove the irrelevant information in the image and only retain the image information containing the target device, so as to obtain more effective and target-device-focused image data. Finally, the image data containing only the target device information is input into the LSTM temporal neural network to predict the optimal beam vector.
[0061] Figure 2 It is a flowchart of the visual-aided millimeter-wave beam prediction method based on target detection according to the present invention. This method is for a millimeter-wave communication system. It collects the image data of the communication environment in the large-view scenario through a camera installed on the base station, constructs a millimeter-wave beam prediction system model, uses the target detection algorithm and background filtering strategy to obtain the image data that only contains the target device information, and then uses the temporal neural network model to learn the obtained image data and predict the optimal communication beam. The specific steps are as follows:
[0062] V1 - V5: Fetch and preprocess the relevant data sets for beam prediction, use the YOLOv7 object detection technology to identify the target devices in the image, obtain the bounding box coordinates of the target devices, calculate the specific positions and coverage ranges of the target devices in the image based on the bounding box coordinates, discard the interference image information outside the coverage range, and obtain the image data containing only the information of the target devices.
[0063] V6 - V8: Construct the image sequence segment data set for final beam prediction, build the LSTM time series neural network model, and set the model training hyperparameters, model loss function, optimizer, training termination conditions, etc.
[0064] V9 - V13: Train the network model, input the image time series data after background filtering into the LSTM network model for training, update the model parameters, retain the network model parameters with the best prediction accuracy during the training process, and verify the actual performance of the model after the training ends.
[0065] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the purpose and scope of the present technical solution, and they should all be covered within the scope of the claims of the present invention.
Claims
1. A visual-aided millimeter-wave beam prediction method for large-view scenarios, characterized in that The method specifically includes the following steps: S1: Obtain the parameters of the millimeter-wave wireless communication system, construct a millimeter-wave wireless communication system model, and define the optimal beam selection strategy of the system; S2: Define the optimization problem of vision-assisted millimeter-wave beam prediction, construct a vision-assisted beam prediction scheme in a large-view scenario, use a neural network framework to learn the mapping relationship between the image of the communication scenario and the optimal communication beam, predict the optimal beam vector, and achieve maximum received power; S3: Collect the visual data of the millimeter-wave wireless communication scenario and obtain the corresponding optimal beam vector in the codebook, preprocess the collected image data, construct a millimeter-wave beam prediction model based on the image data, and define the loss function and the model optimizer; S4: Use the object detection algorithm to extract the bounding box coordinates of the object target in the collected image, and then calculate the coverage range of the target device in the image based on the bounding box coordinates, discard the image information outside the coverage range of the target device, obtain the image data containing only the information of the target device, use the obtained image data to train the time series neural network model, update the network parameters, and after the loss function tends to be stable, obtain the optimal beam prediction network for the large-view scenario, and then use the optimal prediction network to achieve optimal beam selection.
2. The visual assistance millimeter wave beam prediction method according to claim 1, characterized in that In step S1, a millimeter-wave wireless communication system model is constructed, which specifically includes: defining the communication system as a millimeter-wave communication between a base station and a mobile user. The base station is equipped with a camera and a millimeter-wave phased array. The camera captures real-time images of the communication scenario for training the network model. The phased array contains a ULA antenna array with M antennas. Weights are assigned to the antenna array through a predefined codebook to achieve beamforming. The codebook is represented as where represents the beamforming vector in the codebook, represents a complex vector with dimension M×1, and Q represents the total number of vectors in the codebook. The system adopts orthogonal frequency division multiplexing technology to transmit signals in parallel on K subcarriers. The communication channel corresponding to each subcarrier is represented as h k , k = 1, 2,..., K. If the base station sends a downlink signal to the user at the t-th moment, the signal received at the user end is represented as: where y k [t] represents the received signal, v k [t] represents the interference signal, and x represents the transmitted data signal.
3. The visual-aid millimeter-wave beam prediction method according to claim 2, wherein In step S1, a system-optimal beam selection strategy is defined, which specifically includes: the optimization objective of the system is to select the optimal communication beam. At each moment t, a beamforming vector w q [t] is selected from the beam vector codebook to maximize the average received SNR of the system, as shown in the following equation: where, w q [t] * represents the optimal beamforming vector.
4. The visual-aided millimeter-wave beam prediction method according to claim 3, wherein Step S2 specifically includes the following steps: S21: For the optimization objective proposed in step S1, use the images of the communication scenario to optimize beam selection, and let denote the RGB image captured at time t, where W, H, and C represent the width, height, and number of color channels of the image respectively. Then, for the image data X[t] at the t-th moment, the beam prediction optimization problem is defined as constructing a mapping function that can predict the optimal communication beam w * [t]; S22: The image data in continuous time can accurately reflect the mapping relationship between the image and the optimal communication beam. Search for the optimal mapping function for a time series of m RGB images The beam mapping function is expressed as: Among them, represents the mapping function of the network model, and θ w represents the optimization parameter of the network model, and w[τ] represents the optimal beam at the τ-th moment of prediction; since there is a one-to-one correspondence between the beam vectors and their indices in the codebook, learning a mapping function from the image to the beam index, the original formula is expressed as: Among them, represents the corresponding model mapping function, represents the predicted beam index; S23: For an image dataset The optimal mapping function needs to satisfy having the largest number of correct predictions on all samples in the dataset D, that is, the probability that the model predicts the correct label on the entire dataset is the largest, as shown in the following formula: Among them, represents the optimal mapping function, f Θ (.) represents the set of all mapping functions, N represents the number of samples in the dataset, s n represents the true index label, represents the predicted optimal beam index, and represents the probability that the model predicts correctly under the input X n and below.
5. The visual-aid millimeter-wave beam prediction method according to claim 4, wherein In step S3, to construct the millimeter-wave beam prediction model, it specifically includes the following steps: S31: Use a base station equipped with a camera to capture a real-time large-view scene image. The image is in RGB format and contains multiple feature information in the communication environment. The base station selects the optimal beamforming vector w from a preset beam codebook according to the captured image information, where the codebook contains multiple beamforming vectors * to cover the entire communication scene; S32: Construct a beam prediction neural network framework, use the YOLOv7 detection model to identify the target device in the image, obtain the target bounding box coordinates, calculate the specific position and coverage range of the target device based on the bounding box coordinates, and then perform background filtering to obtain the image data containing only the information of the target device; S33: Input the image data obtained through filtering into the LSTM network model. In each time series of the LSTM model, use a fully connected layer to map the features to the beamforming vector space, and calculate the probability distribution of each beamforming vector through the Softmax function. The training of the model uses the cross-entropy loss function, and the optimization goal is to maximize the probability of correctly predicting the beamforming vector.
6. The visual-aided millimeter wave beam prediction method according to claim 5, wherein In step S33, the cross-entropy loss function is expressed as: where n is the number of training samples; P(s n |X n ) is the probability that the model predicts the beam vector index s n given the image X n . The Adam optimizer is used to dynamically adjust the model parameters to ensure that the model converges quickly during training.
7. The visual-aid millimeter-wave beam prediction method according to claim 5 or 6, characterized in that Step S4 specifically includes the following steps: S41: Use the YOLOv7 object detection neural network model to identify the target objects in the large-view scene image, and obtain the center coordinates (x c , y c ) and width and height (w, h) of the bounding box in the normalized form, as well as the predicted category, and set the predicted output category to daily communication users; the specific loss function is as follows: L = L loc + L conf + L cls Among them, L represents the total model loss, L loc represents the bounding box loss, which is used to optimize the center coordinates and size of the predicted box; L conf represents the confidence loss, which measures the degree of matching between the predicted bounding box and the true target; L cls represents the classification loss; S42: Calculate the actual position and coverage range of the target device in the image according to the bounding box output by the object detection. The pixel height of the image data is height, and the pixel width is width. The actual size of the bounding box in the image is: Among them, represents the center coordinates of the bounding box after denormalization, and (w pixel , h pixel ) represents the width and height of the bounding box after denormalization; Further calculate the upper left corner coordinates (x min , y min ) and the lower right corner coordinates (x max , y max ) as follows: Based on the actual coordinates of the vertices of the bounding box of the image data, determine the specific position and coverage range of the target device in the image; S43: The image after target detection and background filtering is input into the LSTM time series network. The hidden state of the LSTM is initialized with all zeros. The entire prediction network contains m LSTM prediction units. Each unit is a feature extraction module composed of two convolutional layers and two fully connected layers. A layer of LSTM network is connected behind this module. The image data at each moment is input into the corresponding LSTM prediction unit at that moment to achieve feature extraction; while the prediction unit at the last moment realizes the optimal beam prediction through the connected classifier, and the cross-entropy loss is used as the final loss function; S44: Train the time series neural network model and update the network parameters. After the loss function tends to be stable, obtain the optimal beam prediction network for the large view scenario, and then use the optimal prediction network to achieve the optimal wave speed selection.
8. The visual-aided millimeter wave beam prediction method according to claim 7, wherein In step S41, the bounding box loss function adopts the CIOU loss.