Millimeter-wave beam tracking method based on contrastive learning

By using a multimodal data fusion neural network model and utilizing visual and lidar data for beam tracking, the problem of insufficient single-modal data is solved, achieving efficient and stable beam tracking and reducing latency and overhead in millimeter-wave communication systems.

WO2026091167A1PCT designated stage Publication Date: 2026-05-07CHONGQING UNIV OF POSTS & TELECOMM
View PDF 6 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
CHONGQING UNIV OF POSTS & TELECOMM
Filing Date
2024-11-08
Publication Date
2026-05-07

AI Technical Summary

Technical Problem

In millimeter-wave communication systems, single-mode data cannot fully capture the characteristics of complex scenes, resulting in high beam acquisition overhead. Existing methods are unable to effectively reduce communication latency and computational overhead.

Method used

A multimodal data fusion neural network model based on contrastive learning is adopted. Visual and lidar data are used for feature extraction and temporal modeling. The contrastive learning mechanism optimizes the synergistic effect between different modes to achieve beam tracking.

Benefits of technology

It improves the accuracy and stability of beam tracking, reduces the latency and computational overhead of millimeter-wave communication systems, and adapts to the low-latency requirements of multi-sensor fusion frameworks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024130674_07052026_PF_FP_ABST
    Figure CN2024130674_07052026_PF_FP_ABST
Patent Text Reader

Abstract

The present invention belongs to the field of wireless communications and relates to a millimeter-wave beam tracking method based on contrastive learning. The method is oriented to a millimeter-wave communication system. In combination with visual and LiDAR information in an environment that are collected by a plurality of sensors installed at a millimeter-wave base station, a machine learning model is used to perform feature processing on data of different modalities, and a multi-modal fusion strategy based on contrastive learning is used to enhance the synergistic effect between data of various modalities, so as to obtain a method for real-time prediction of the current beam and tracking of a plurality of future beams. In the present invention, multi-modal data fusion and contrastive learning are combined for beam tracking, such that the accuracy and stability of beam tracking can be effectively ensured, thereby ensuring the communication efficiency of a millimeter-wave communication system.
Need to check novelty before this filing date? Find Prior Art

Description

A millimeter-wave beam tracking method based on contrastive learning Technical Field

[0001] This invention belongs to the field of wireless communication and relates to a millimeter-wave beam tracking method based on contrastive learning. Background Technology

[0002] With the rapid development of artificial intelligence and communication technologies, next-generation mobile communication networks are beginning to utilize millimeter waves and even higher frequency bands to meet the growing demands of mobile services. Compared to the low-frequency signals used in traditional communication systems, millimeter wave signals suffer from more severe propagation loss. Due to the shorter wavelength of millimeter waves, a large number of antenna arrays and narrow directional beams are required to ensure sufficiently high received signal strength. However, applying traditional beamforming methods based on training a large number of antenna array vectors to the millimeter wave band incurs significant beamforming overhead. Therefore, it is necessary to find new methods to reduce this overhead in order to decrease communication latency and power consumption in millimeter wave systems.

[0003] Machine learning-based methods offer new insights into addressing beam management and channel assessment overhead in high-frequency communication systems. Information acquired from environmental sensors, such as user location, images of user locations captured by base station cameras, LiDAR point clouds, and radar data, can be used for beam prediction and tracking in millimeter-wave communication systems. However, single-modal data can only extract information from a single data source, which may lead to information loss or inadequacy, especially in complex scenarios where single-modal data cannot fully capture scene features. Therefore, developing a millimeter-wave beam tracking method suitable for multimodal data fusion, leveraging the complementarity of data from different modalities to compensate for the shortcomings of single-modal data and provide more accurate and stable prediction and tracking, has become a significant challenge.

[0004] Summary of the Invention

[0005] In view of this, the purpose of this invention is to provide a millimeter-wave beam tracking method based on contrastive learning, which utilizes the complementarity of data from different modes to compensate for the deficiencies of single-mode data and provide more accurate and stable prediction and tracking.

[0006] To achieve the above objectives, the present invention provides the following technical solution:

[0007] A millimeter-wave beam tracking method based on contrastive learning is proposed. This method constructs a multimodal data fusion neural network model for millimeter-wave communication beam tracking, trains the neural network using contrastive learning, optimizes the synergistic effect between different modes, and uses the fused features for beam tracking. It can predict the current optimal beam in real time and simultaneously predict the optimal index of multiple future optimal beams.

[0008] The method specifically includes the following steps:

[0009] S1: Obtain the model parameters of the multi-sensor millimeter-wave communication system, construct a model of the beam tracking problem of multimodal sensing-assisted millimeter-wave communication, and transform the beam tracking problem into an optimization problem based on machine learning;

[0010] S2: Construct a multimodal data fusion neural network model for millimeter-wave communication beam tracking, including: a feature extraction module, a recurrent neural network module, and a fusion classification module; this model acquires sensing data of different modalities from multiple sensors equipped in the system, uses the feature extraction module to extract features from the acquired multimodal data, and uses the recurrent neural network module to model the multimodal temporal features. The fusion classification module uses a multi-head attention mechanism to fuse the extracted features.

[0011] S3: Use a contrastive learning mechanism to train a multimodal data fusion neural network model, optimize the synergistic effect between different modalities, use the fused features for beam tracking, and predict the best index of multiple current and future optimal beams in real time.

[0012] Furthermore, step S1 specifically includes the following steps:

[0013] S11: Consider a millimeter-wave communication system model surrounding a single mobile user equipment (MUE), where a base station provides services to the MUE. The base station includes N antennas, a lidar sensor, and an RGB camera (visual data sensor) to provide awareness of the surrounding environment. Assume the base station has only one antenna and uses a predefined beamforming codebook F = {f1, f2, ..., f...}. q ,...,f Q}, where Q is the total number of beamforming vectors in the codebook, f q ∈F is the beamforming vector in the codebook;

[0014] S12: At time step t, if each millimeter-wave user u is beamformed by the base station using a beamforming vector f q [t] service, channel The signal from the base station to the user is then represented as:

[0015] Wherein, E|x[t]| 2 =P T P T The average transmit power, To transmit the complex number symbol, (·) H denoted as conjugate transpose, and n[t] represents a noise sample extracted from a complex Gaussian distribution;

[0016] S13: In a single-user scenario, multimodal beam tracking is performed using lidar and visual information. Based on available sensing data up to time t-1, the base station aims to predict the next... The goal of this task is to select the optimal beam for the upcoming steps, specifically the optimal beam from time t to time (t+α-1); the objective is to select the optimal beam from the candidate beams accessible in the codebook F. To maximize the received signal power; this problem can be mathematically represented as:

[0017] S14: Select the optimal beam index from a pre-determined codebook F, which is typically associated with higher acquisition overhead; the signal is concentrated in a certain direction in space, which is the mechanism for guiding narrow beams; the spatial dimension of the scene is divided into many potentially overlapping sectors using beam vectors, and each sector is assigned a unique value; therefore, the machine learning task can be transformed into a classification task using a pre-determined codebook, where the user's real-time location in the visual and LiDAR scene determines which beam index in the codebook to assign; under specified codebook constraints, the optimal beam can be identified individually by matching codebook indices. At time t, the optimal beam index can be expressed as:

[0018] Where |Q| represents the cardinality of F, and the optimal beam index is denoted as g[t]. It should be noted that finding the optimal beam is the same as finding the optimal beam index under codebook constraints; therefore, the beam tracking problem can be expressed as:

[0019] Where P{·∣·} represents the conditional probability. The o represents the predicted optimal beam index for time step t. t-α It is auxiliary information collected before the time step (t-α+1), which contains some detailed information about the optimal beam at time t;

[0020] S15: The RGB image (visual perception information) obtained at time step t is defined as... Where H, W, and C represent the image height, width, and number of color channels, respectively; (Definition) This refers to the user information obtained by the lidar sensor at time step t, where D represents the angle in the lidar's perceived field of view, and each angle has a corresponding distance value; the present invention aims to utilize the input visual information x u and lidar information u Tracking the optimal beam index To serve user u, this invention proposes to learn a function f using machine learning methods. Θ(S), where S=(x) u ,l u The observed visual and lidar information is fed into this function, which outputs probabilities P∈{p1,...,p}. Q The index of the most likely element in codebook F Formally, it can be expressed as:

[0021] in, Indicates the optimal beam index. Indicates the beam index value;

[0022] S16: To utilize these visual and lidar data for beam tracking, the objective of this invention is to develop a function that can predict the optimal future beam starting from time step t based on the sensing information collected up to time step t-1; let S t,i ={S[t-i+1],...,S[t]} represents the visual and lidar data sequences, where i represents the time step within the observation window; then the visual and lidar-assisted beam tracking optimization problem can be expressed as:

[0023] S17: To accomplish this beam tracking task, this invention uses a prediction function, denoted as f. Θ (S t,i ), where a set of model parameters is denoted as Θ; using a labeled dataset Train the estimation function, where S represents the visual and LiDAR input pair. This represents the actual optimal beam index; following standard machine learning conventions, the goal of this task is to increase the probability of accurately predicting the optimal beam and improve the overall success probability; mathematically, it can be represented as:

[0024] Furthermore, step S2 specifically includes the following steps:

[0025] S21: The feature extraction module processes the original input data. Since this data usually contains irrelevant information that does not contribute to the beam tracking task, the feature extraction module aims to reduce the dimensionality of the data by extracting key features, thereby minimizing the feature space. The main goal of this module is to identify and capture the basic features that are crucial to the subsequent beam tracking process. By eliminating unnecessary information, the training process is made more stable, and the efficiency of beam prediction and tracking is improved.

[0026] For the raw image data, the feature extraction module uses object detection to identify and annotate key objects in the image, such as vehicles. Given that beam tracking tasks require rapid processing of visual data over time, YOLOv4 is particularly well-suited for this application due to its real-time detection capabilities and low computational complexity. YOLOv4 processes the input visual (RGB) image, generating bounding boxes for all detected objects in the image, represented as... Each bounding box vector S bbox =[x c ,y c ,w,h] T These represent the x-center, y-center, width, and height of the detected object, respectively; therefore, the bounding box obtained from the original image is used as a key feature for the beam tracking task.

[0027] LiDAR sensors generate two-dimensional point clouds of their surroundings by emitting and receiving laser pulses; raw LiDAR data often contains a lot of redundant or noisy information; the main goal of feature extraction is to extract key geometric features relevant to the task, effectively reducing the dimensionality of the data; this feature extraction process is crucial for supporting subsequent beam tracking tasks.

[0028] S22: For visual features (bounding boxes) and LiDAR features, this invention uses an embedding layer to... Transform into Will Transform into To embed the prior beam index metric, this invention employs a trainable lookup table containing the embedding vector |Q|, where the embedding vector... Corresponding to beam index q;

[0029] S23: The data processed by the feature extraction module enters the recurrent neural network module, which uses a long short-term memory network model to process time series data. The long short-term memory network model can effectively capture the time dependencies in the input data, which is crucial for beam tracking tasks. By processing features from vision and lidar data, the long short-term memory network processes the relationships between time steps and facilitates accurate prediction of future beams.

[0030] S24: The data processed by the recurrent neural network module enters the fusion classification module, which concatenates the feature vectors of visual (RGB) data output by the long short-term memory network and the temporal features of the LiDAR data to form a multimodal fusion feature vector. The fused feature vector is passed through a multi-head attention mechanism. This mechanism can assign different attention weights to different features, thereby enhancing the model's attention to key information and ignoring irrelevant information. Finally, the feature vector is input into a fully connected network for classification prediction. The classifier is responsible for predicting the current and future optimal beam, i.e., outputting the optimal beam index.

[0031] Furthermore, step S3 specifically includes:

[0032] In this invention, a contrastive learning mechanism is introduced to enhance the feature alignment capability between multimodal data (visual RGB images and LiDAR data) and improve the beam tracking accuracy of millimeter-wave communication systems. Specifically, this invention uses a normalized temperature-scale cross-entropy loss function as the core of the contrastive learning, which is defined as follows:

[0033] Among them, NT-Xent(z i ,z j ) represents sample z i With z j The loss, sim(z) i ,z j ) represents sample z i With z j The similarity, where τ is the temperature coefficient;

[0034] During model training, feature representations are first extracted from visual (RGB) images and LiDAR data, respectively. To introduce contrastive learning, the visual (RGB) features and corresponding LiDAR features of each data sample are used as positive pairs. The cosine similarity of these positive pairs is calculated using a normalized temperature-scaled cross-entropy loss function, and scaling is performed using a temperature parameter to ensure stable training at different feature scales. Furthermore, the feature pairs of all other samples in the same batch during machine learning training are used as negative pairs. The goal of contrastive learning is to maximize the similarity between positive pairs and minimize the similarity between negative pairs, thereby enhancing the consistency between features of different modalities.

[0035] During training, the model's total loss function consists of several parts: classification loss, classification loss of visual (RGB) features and LiDAR features, and normalized temperature-scale cross-entropy contrastive learning loss. This combined strategy not only improves classification performance but also effectively aligns multimodal feature representations, thereby achieving higher tracking accuracy in millimeter-wave communication beam tracking tasks. By introducing contrastive learning, the model can better fuse multimodal data to achieve accurate beam tracking. Thus, a multimodal data fusion millimeter-wave communication beam tracking method based on contrastive learning is obtained.

[0036] The beneficial effects of this invention are as follows:

[0037] (1) By combining visual and lidar data acquired by multiple sensors of millimeter-wave base stations and utilizing a multimodal fusion strategy of contrastive learning, this invention can significantly improve the accuracy and stability of beam tracking, thereby enhancing the communication efficiency of millimeter-wave communication systems.

[0038] (2) This invention transforms the beam tracking problem into an optimization problem, and uses machine learning to perform real-time calculation and updating of beam selection, reducing the communication latency and computational overhead of millimeter-wave systems, and adapting to the low-latency requirements of multi-sensor fusion frameworks. This enables the tracking method to be deployed in large-scale millimeter-wave communication systems.

[0039] Other advantages, objectives, and features of the invention will be set forth in part in the description which follows, and in part will be apparent to those skilled in the art from the following examination, or may be learned from practice of the invention. The objectives and other advantages of the invention can be realized and obtained through the following description. Attached Figure Description

[0040] To make the objectives, technical solutions, and advantages of the present invention clearer, the preferred embodiments of the present invention will be described in detail below with reference to the accompanying drawings, wherein:

[0041] Figure 1 is a schematic diagram of the multi-sensor millimeter-wave communication system model described in this invention;

[0042] Figure 2 is a schematic diagram of the multimodal data fusion neural network model proposed in this invention;

[0043] Figure 3 is a flowchart of the multimodal data fusion beam tracking based on contrastive learning according to the present invention. Detailed Implementation

[0044] The following specific examples illustrate the implementation of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and various details in this specification can be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the illustrations provided in the following embodiments are only schematic representations of the basic concept of the present invention. Unless otherwise specified, the following embodiments and features can be combined with each other.

[0045] Referring to Figures 1-3, this invention provides a millimeter-wave beam tracking method based on multimodal sensing assistance. In a multi-sensor millimeter-wave communication system model, key parameters of the system are acquired, and sensing data from different modalities are obtained from the system's multiple sensors. A multimodal sensing-assisted millimeter-wave communication beam tracking problem model is constructed, transforming the beam tracking problem into a machine learning-based optimization problem. Feature extraction and temporal modeling are performed by establishing a multimodal data fusion neural network model. A multi-head attention mechanism is used to fuse extracted features, and a contrastive learning mechanism is used to train the multimodal data fusion neural network model, optimizing the synergistic effect between different modalities. The fused features are then used for beam tracking, enabling real-time prediction of the optimal beam at current and future times, thereby reducing latency and computational overhead while ensuring millimeter-wave communication efficiency.

[0046] Figure 1 is a schematic diagram of a multi-sensor millimeter-wave communication system model. Figure 1 shows a millimeter-wave communication system model surrounding a single mobile user device (MDP), where the base station provides services to the MDP. The base station contains N antennas, a lidar sensor, and an RGB camera (visual data sensor) to provide perception of the surrounding environment. Assuming the base station has only one antenna, the base station uses a predefined beamforming codebook F = {f1, f2, ..., f...} Q}, where Q is the total number of beamforming vectors in the codebook, f q ∈F is the beamforming vector in the codebook.

[0047] At time step t, if each millimeter-wave user u is beamformed by the base station using a beamforming vector f q [t] service, channel The signal from the base station to the user is then represented as:

[0048] Wherein, E|x[t]| 2 =P T P T The average transmit power, To transmit complex numbers, n[t] represents a noise sample extracted from a complex Gaussian distribution.

[0049] Figure 2 is a schematic diagram of a multimodal data fusion neural network model. This model mainly consists of three parts: a feature extraction module, a recurrent neural network module, and a fusion classification module. First, the feature extraction module processes the raw input data. The main goal of this module is to identify and capture fundamental features crucial for the subsequent beam tracking process.

[0050] The data processed by the feature extraction module enters the recurrent neural network module, which uses a long short-term memory network to process time series data. The long short-term memory network model can effectively capture the time dependencies in the input data.

[0051] The data processed by the recurrent neural network module enters the fusion and classification module. The feature vectors of the image output from the long short-term memory network and the temporal features of the LiDAR data are concatenated to form a multimodal fusion feature vector. This fused feature vector is passed through a multi-head attention mechanism. This mechanism assigns different attention weights to different features, thereby enhancing the model's focus on key information and ignoring irrelevant information. Finally, the feature vector is input into a fully connected network for classification prediction. The classifier is responsible for predicting the current and future optimal beam, i.e., outputting the optimal beam index.

[0052] Figure 3 is a flowchart of the multimodal data fusion beam tracking process based on the contrastive learning mechanism of the present invention. The process specifically includes the following steps:

[0053] V1~V5: Obtain the model parameters of the multi-sensor millimeter-wave communication system, construct a model of the beam tracking problem of multimodal sensing-assisted millimeter-wave communication, and transform the beam tracking problem into an optimization problem based on machine learning.

[0054] V6–V12: Construct a multimodal data fusion neural network model for millimeter-wave communication beam tracking. The neural network extracts features from the acquired multimodal data, pairing image features and LiDAR features as positive and negative pairs. A recurrent neural network is used to model the multimodal temporal features, employing a normalized temperature-scale cross-entropy loss function for comparative learning to maximize positive similarity and minimize negative similarity. The fusion classification module is used to fuse the temporal multimodal feature vectors for training.

[0055] V13~V15: Calculate the loss function and update the model parameters by combining classification loss and contrastive learning loss.

[0056] V16~V17: After the training termination condition is met, neural network parameters are generated, and the multi-sensor millimeter-wave communication system performs beam tracking based on the fully trained beam tracking network.

[0057] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.

Claims

1. A millimeter-wave beam tracking method based on contrastive learning, characterized in that, The method specifically includes the following steps: S1: Obtain the model parameters of the multi-sensor millimeter-wave communication system, construct a model of the beam tracking problem of multimodal sensing-assisted millimeter-wave communication, and transform the beam tracking problem into an optimization problem based on machine learning; S2: Construct a multimodal data fusion neural network model for millimeter-wave communication beam tracking, including: a feature extraction module, a recurrent neural network module, and a fusion classification module; this model acquires sensing data of different modalities from multiple sensors equipped in the system, uses the feature extraction module to extract features from the acquired multimodal data, and uses the recurrent neural network module to model the multimodal temporal features. The fusion classification module uses a multi-head attention mechanism to fuse the extracted features. S3: Use a contrastive learning mechanism to train a multimodal data fusion neural network model, optimize the synergistic effect between different modalities, use the fused features for beam tracking, and predict the best index of multiple current and future optimal beams in real time.

2. The millimeter-wave beam tracking method based on contrastive learning according to claim 1, characterized in that, Step S1 specifically includes the following steps: S11: Consider a millimeter-wave communication system model surrounding a single mobile user equipment (Mobile User Equipment), where a base station provides services to the Mobile User Equipment. The base station contains N antennas, a lidar sensor, and a visual data sensor to provide awareness of the surrounding environment. Assume the base station has only one antenna and uses a predefined beamforming codebook F = {f1, f2, ..., f...}. q ,...,f |Q| }, where Q is the total number of beamforming vectors in the codebook, f q ∈F is the beamforming vector in the codebook; S12: At time step t, if each millimeter-wave user u is beamformed by the base station using a beamforming vector f q [t] Service, Channel h u [t]∈£ N×1 The signal from the base station to the user is then represented as: Wherein, E|x[t]| 2 =P T P T Let x[t]∈£ be the average transmit power, and x[t]∈£ be the transmission complex symbol (·). H denoted as conjugate transpose, and n[t] represents a noise sample extracted from a complex Gaussian distribution; S13: In a single-user scenario, multimodal beam tracking is performed using lidar and visual information. Based on available sensing data up to time t-1, the base station aims to predict the next... The goal of this task is to select the optimal beam for the upcoming steps, specifically the optimal beam from time t to time (t+α-1); the objective is to select the optimal beam from the candidate beams accessible in the codebook F. To maximize the received signal power; mathematically, this is expressed as: S14: Select the optimal beam index from a pre-determined codebook F, concentrating the signal in a certain direction within space; divide the spatial dimension of the scene into many potentially overlapping sectors using beam vectors, and assign a unique value to each sector; thus, the machine learning task can be transformed into a classification task using a pre-determined codebook, where the user operates in both vision and LiDAR... The real-time location in the scene determines which beam index in the codebook should be assigned; under specified codebook constraints, the optimal beam is identified individually by matching codebook indexes. At time t, the optimal beam index is expressed as: Where |Q| represents the cardinality of F, and the optimal beam index is denoted as g[t]. It should be noted that finding the optimal beam is the same as finding the optimal beam index under codebook constraints; therefore, the beam tracking problem can be expressed as: Where P{·∣·} represents the conditional probability. The o represents the predicted optimal beam index for time step t. t-α It is auxiliary information collected before the time step (t-α+1), which contains some detailed information about the optimal beam at time t; S15: The visual perception information obtained at time step t is defined as... Where H, W, and C represent the image height, width, and number of color channels, respectively; (Definition) This represents the user information obtained by the LiDAR sensor at time step t, where D represents the angle in the LiDAR's perceived field of view, and each angle has a corresponding distance value; using the input visual information x u and lidar information u Tracking the optimal beam index To serve user u; to achieve this, a function f is learned using machine learning methods. Θ (S), where S=(x) u ,l u The observed visual and lidar information is fed into this function, which outputs probabilities P∈{p1,...,p}. Q The index of the most likely element in codebook F In terms of form, it is expressed as: in, Indicates the optimal beam index. Indicates the beam index value; S16: To perform beam tracking using visual and lidar data, the goal is to develop a function that predicts the optimal future beam starting from time step t, based on the sensing information collected up to time step t-1; let S... t,i ={S[t-i+1],...,S[t]} represents the visual and lidar data sequences, where i represents the time step within the observation window; then the visual and lidar-assisted beam tracking optimization problem is expressed as: S17: To complete the beam tracking task, a prediction function is used, denoted as f. Θ (S t,i ), where a set of model parameters is denoted as Θ; using a labeled dataset Train the estimation function, where S represents the visual and LiDAR input pair. This represents the actual optimal beam index; following standard machine learning conventions, the goal of this task is to increase the accuracy of predicting the optimal beam. The probability of a bundle, and to increase the overall probability of success; mathematically speaking, it can be expressed as:

3. The millimeter-wave beam tracking method based on contrastive learning according to claim 2, characterized in that, Step S2 specifically includes the following steps: S21: The feature extraction module processes the original input data; the feature extraction module reduces the dimensionality of the data by extracting key features, thereby minimizing the feature space; For the raw image data, the feature extraction module uses object detection to identify and annotate key objects in the image. Specifically, it uses YOLOv4 to process the input visual image, generating bounding boxes for all detected objects in the image, represented as... Each bounding box vector S bbox =[x c ,y c ,w,h] T These represent the x-center, y-center, width, and height of the detected object, respectively; therefore, the bounding box obtained from the original image is used as a key feature for the beam tracking task. LiDAR sensors generate two-dimensional point clouds of their surrounding environment by emitting and receiving laser pulses; the goal of feature extraction is to extract key geometric features relevant to the task. S22: For visual features and LiDAR features, use an embedding layer to... Transform into Will Transform into To embed the prior beam index metric, a trainable lookup table containing the embedding vectors |Q| is applied, where the embedding vectors... Corresponding to beam index q; S23: The data processed by the feature extraction module enters the recurrent neural network module, which uses a long short-term memory network model to process time series data; by processing features from visual and lidar data, the long short-term memory network processes the relationship between time steps and facilitates accurate prediction of future beams; S24: The data processed by the recurrent neural network module enters the fusion classification module, which concatenates the feature vectors of the visual and LiDAR data time features output by the long short-term memory network to form a multimodal fusion feature vector; the fused feature vector is passed through a multi-head attention mechanism; finally, the feature vector is input into a fully connected network for classification prediction; the classifier is responsible for predicting the current and future optimal beam, that is, outputting the optimal beam index.

4. The millimeter-wave beam tracking method based on contrastive learning according to claim 3, characterized in that, Step S3 specifically includes: introducing a contrastive learning mechanism to enhance the feature alignment capability between multimodal data and improve the beam tracking accuracy of millimeter-wave communication systems; specifically, the normalized temperature-scale cross-entropy loss function is used as the core of the contrastive learning, and its definition is as follows: Among them, NT-Xent(z i ,z j ) represents sample z i With z j The loss, sim(z) i ,z j ) represents sample z i With z j The similarity, where τ is the temperature coefficient; During model training, feature representations are first extracted from visual images and LiDAR data, respectively. To introduce contrastive learning, the visual features and corresponding LiDAR features of each data sample are treated as positive pairs. The cosine similarity of these positive pairs is calculated using a normalized temperature-scaled cross-entropy loss function, and scaling is performed using a temperature parameter to ensure stable training at different feature scales. Furthermore, feature pairs of all other samples in the same batch during machine learning training are treated as negative pairs. The goal of contrastive learning is to maximize the similarity between positive pairs and minimize the similarity between negative pairs, thereby enhancing the consistency between features of different modalities. During training, the model's total loss function consists of several parts: classification loss, classification loss of visual features and LiDAR features, and normalized temperature-scale cross-entropy contrastive learning loss.

Citation Information

Patent Citations

  • Millimeter wave communication link blocking prediction method based on visual information fusion

    CN114845332A

  • Internet-of-vehicles wave beam real-time alignment method based on multi-modal information consciousness

    CN115412844A

  • Millimeter wave beam tracking method in microwave and millimeter wave heterogeneous distribution network scene

    CN116800321A

  • Millimeter wave beam tracking method based on multi-mode fusion

    CN117692033A

  • Fusion models for beam prediction

    US20240144087A1