Fusion model for beam prediction

Through machine learning models, the beam selection in wireless communication is optimized, and the problem of insufficient beam selection efficiency and accuracy in the prior art is solved, and higher communication robustness and throughput are achieved.

CN120077579APending Publication Date: 2025-05-30QUALCOMM INC
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202380074244.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-06-23
Filing Date
2023-08-24
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The prior art is difficult to effectively utilize multimodal data and situational awareness to optimize beam selection in wireless communications, especially in 6G systems, where beam selection efficiency and accuracy are insufficient in the face of complex environments and rapid changes.

Method used

Using a machine learning model, by accessing multiple data modes (such as image data, radar data, LIDAR data and position data), perform feature extraction and fusion, and using attention mechanisms to fuse features of different modes, ultimately generating an optimized wireless communication configuration.

Benefits of technology

By converging multimodal data, the most suitable radio frequency beam can be predicted and selected more accurately, improving the robustness and throughput of wireless communications, especially in the face of complex environments and rapid changes in 6G systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120077579A_ABST
    Figure CN120077579A_ABST
Patent Text Reader

Abstract

Certain aspects of the present disclosure provide techniques and apparatus for beam selection using machine learning. A plurality of data samples corresponding to a plurality of data modalities are accessed. A plurality of features is generated by performing feature extraction for each respective data sample of the plurality of data samples based at least in part on a respective modality of the respective data sample. The plurality of features are fused using one or more attention-based models, and a wireless communication configuration is generated based on processing the fused plurality of features using a machine learning model.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross - reference to related applications

[0002] This application claims priority to U.S. Patent Application No. 18 / 340,671, filed on Jun. 23, 2023, which claims the benefit and priority of U.S. Provisional Patent Application No. 63 / 381,408, filed on Oct. 28, 2022, and U.S. Provisional Patent Application No. 63 / 500,496, filed on May 5, 2023. The entire contents of each of these applications are incorporated herein by reference. Background of the disclosure

[0003] Aspects of the present disclosure relate to machine learning, and more particularly to using machine learning to provide improved beam selection (e.g., in wireless communication).

[0004] Wireless communication systems are widely deployed to provide various telecommunication services such as telephony, video, data, messaging, broadcasting, etc. The current and future demands for wireless communication networks continue to grow. For example, the sixth - generation (6G) systems are expected to support applications such as augmented reality, multisensory communication, and high - fidelity holograms. It is further expected that these systems serve an increasing number of devices while also achieving high standards with respect to performance. Summary of the disclosure

[0005] The systems, methods, and devices of the present disclosure each have several aspects, no single one of which is solely responsible for its desirable attributes. Without limiting the scope of the present disclosure as expressed by the appended claims, some features will now be briefly discussed. After considering this discussion, and particularly after reading the section entitled "Detailed Description", one will understand how the features of the present disclosure provide the advantages described herein.

[0006] Some aspects of the present disclosure provide a method (e.g., a processor - implemented method). The method generally includes: accessing a plurality of data samples corresponding to a plurality of data modalities; generating a plurality of features by performing feature extraction for each respective data sample of the plurality of data samples based at least in part on the respective modality of the respective data sample; using one or more attention - based models to fuse the plurality of features; and generating a wireless communication configuration based on processing the fused plurality of features using a machine - learning model.

[0007] In other aspects, provided are: a processing system configured to perform the foregoing methods and those described herein; non-transitory computer-readable media including instructions that, when executed by one or more processors of the processing system, cause the processing system to perform the foregoing methods and those described herein; a computer program product embodied on a computer-readable storage medium, the computer program product including code for performing the foregoing methods and those further described herein; and a processing system including components for performing the foregoing methods and those further described herein.

[0008] To achieve the foregoing and related purposes, one or more aspects include the features described comprehensively below and particularly pointed out in the claims. The following description and the drawings set forth certain illustrative features of these one or more aspects in detail. However, these features are indicative of only some of the various ways in which the principles of the various aspects may be employed. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] For a more particular description of the features briefly summarized above, reference may be had to some aspects illustrated in the accompanying drawings. It should be noted, however, that the drawings illustrate only certain typical aspects of the present disclosure and are therefore not to be considered as limiting its scope, as the specification may admit other equally effective aspects.

[0010] Figure 1 An example environment for using fusion-based machine learning to provide improved beam selection is depicted.

[0011] Figure 2 An example architecture for fusing and evaluating image data and location data to provide improved beam selection is depicted.

[0012] Figure 3 An example architecture for using light detection and ranging (LIDAR) data to provide improved beam selection is depicted.

[0013] Figure 4 An example architecture for using radar data to provide improved beam selection is depicted.

[0014] Figure 5 An example architecture for using fusion to provide improved beam selection is depicted.

[0015] Figure 6 An example architecture for using sequential fusion to provide improved beam selection is depicted.

[0016] Figure 7 An example workflow for using pre-training and scenario adaptation of simulation data is depicted.

[0017] Figure 8 is a flowchart depicting an example method for implementing improved beam selection through data modality fusion.

[0018] Figure 9 is a flowchart depicting an example method for pre-training and scenario adaptation.

[0019] Figure 10 is a flowchart depicting an example method for implementing improved wireless communication configuration using machine learning.

[0020] Figure 11 depicts an example processing system configured to perform various aspects of the present disclosure.

[0021] For ease of understanding, the same reference numerals have been used, where possible, to denote the same elements common to the figures. It is contemplated that elements disclosed in one aspect may be beneficially utilized in other aspects without specific recitation. Detailed Description

[0022] Aspects of the present disclosure provide apparatuses, methods, processing systems, and computer-readable media for using machine learning to drive, for example, improved wireless communication and other applications using beams.

[0023] In some aspects, techniques are disclosed for improving a wireless system (e.g., a 6G system) by using machine learning (ML) and / or artificial intelligence (AI) to leverage multimodal data and context awareness to provide improved communication (such as through more optimized beam selection). In some aspects, since different data modalities generally have significantly different characteristics, the fusion of these different data modalities involves target feature extraction and fusion operations using machine learning.

[0024] In some aspects, the best or otherwise desired wireless beam for communicating with a given wireless device or equipment (e.g., user equipment (UE)), (e.g., by a 6G base station), depends at least in part on the relative positioning between the transmitter and the receiver and the geometry of the environment. Additionally, beam selection can benefit from awareness of the surrounding environment and other contexts. In some aspects, beam selection can be performed based on a codebook, where each entry or code in the codebook corresponds to a beam (e.g., the direction from the transmitter) that covers a specific portion or region of physical space. In some aspects, the codebook generally can include a set of such beams that cover the entire angular space, and beams can be selected for various zones in physical space.

[0025] In some aspects of the present disclosure, a machine learning model is trained and used to predict or select a radio frequency (RF) beam that is predicted to be most suitable (or among the most suitable) for a given communication based on input data such as images, radar data, LIDAR data, Global Navigation Satellite System (GNSS) data, and the like. In some aspects, the performance of the machine learning model is enhanced by leveraging different data modalities, and a fusion module capable of using the available data modalities can be provided to achieve improved prediction performance. In some aspects, since each data modality has different characteristics, dedicated sub-modules (branches) are used to extract information from each of these data modalities. Additionally, in some aspects, beam prediction is consistent over time for a mobile UE, which can be particularly important when the UE is moving at high speed.

[0026] In some aspects, a fusion model for beam prediction in a multi-modal scenario is provided. In some aspects, the model is capable of fusing different data modalities to achieve improved performance in the beam prediction task. In some aspects, the fusion model includes one or more attention modules that can allow the model itself to determine how to fuse the modalities and utilize different modalities for each data point. In some aspects, each branch of the model is designed and dedicated to a single modality, ensuring that the branch can extract meaningful information from each modality. Additionally, in some aspects, the fusion model includes one or more recurrent modules that allow for the analysis of the temporal evolution of the UE / environment, thus allowing for robust prediction over time.

[0027] Example Environment for Fusion-Based Beam Selection

[0028] Figure 1 An example environment 100 for using fusion-based machine learning to provide improved beam selection is depicted.

[0029] In the illustrated example, base station 105 (e.g., a next-generation Node B (gNB)) is configured to collect, generate, and / or receive data 115 that provides an environmental context for communication. Generally speaking, data 115 may include data belonging to or associated with multiple types or modalities. For example, data 115 may include one or more modalities such as (but not limited to) image data 117A (e.g., captured by one or more cameras on or near base station 105), radio detection and ranging (radar) data 117B (e.g., captured by one or more radar sensors on or near base station 105), LIDAR data 117C (e.g., captured by one or more LIDAR sensors on or near base station 105), location data 117D (e.g., GNSS positioning coordinates of base station 105 and / or one or more other nearby objects such as UE 110), etc. As used herein, “UE” generally may refer to any wireless device or system capable of performing wireless communication (e.g., via base station 105), such as a cellular phone, a smart phone, a smart vehicle, a laptop computer, etc.

[0030] Although four specific modalities are depicted for clarity of concept, in some aspects, machine learning system 125 may use any number and variety of modalities to generate prediction beam 130. In some aspects, machine learning system 125 may selectively use or avoid using one or more of the modalities depending on the particular configuration. That is, in some aspects, machine learning system 125 may determine which modalities to use in generating prediction beam 130, potentially avoiding using one or more modalities (e.g., avoiding using LIDAR data 117C), even if those modalities are present. In some aspects, machine learning system 125 may assign or give a higher weight to one or more of the modalities than the remaining modalities. For example, machine learning system 125 may use four modalities, thereby giving a higher weight (e.g., twice the weight of the other modalities or some other factor) to one modality (e.g., image data).

[0031] Although the illustrated example depicts the base station 105 providing data 115 (e.g., after obtaining it from another source such as a nearby camera, vehicle, etc.), in some aspects, some or all of the data in data 115 may come directly from other sources (such as UE 110) (i.e., without passing through the base station 105). For example, the image data 117A may be generated using one or more cameras on or near the base station 105 (e.g., to visually identify the UE 110), and the location data 117D may be generated at least in part based on the GNSS location of the UE 110. For example, the location data 117D may indicate the relative position and / or orientation of the UE 110 with respect to the base station 105, as determined by the GNSS coordinates of the UE 110 and the known or determined location of the base station 105. One or more of these data 117A - 117D may be directly transmitted to the machine learning system 125.

[0032] In the illustrated example, the base station 105 may use a variety of configurations or parameters to control the beam 135 for communicating with the UE 110. For example, in multi - input multi - output (MIMO) aspects, the base station 105 may use or control various beamforming codebooks, dictionaries, phase shifters, etc. to change the focus of the beam 135 (e.g., change where the center or focus of the beam 135 is located). By appropriately manipulating or adjusting the beam 135 (e.g., selecting a specific beam), the base station 105 may be able to provide or improve communication with the UE 110. In the illustrated example, the UE 110 may be moving. In such mobile aspects (and particularly when the UE 110 is moving rapidly), appropriate beam selection can achieve significantly improved results.

[0033] Although a single base station 105 and a single UE 110 are depicted for conceptual clarity, in various aspects, there may be any number of base stations and UEs. Additionally, although a 6G base station 105 is depicted for conceptual clarity, in various aspects, the base station 105 may include or be configured to use any standard or technology such as 5G, 4G, WiFi, etc. to provide wireless communication.

[0034] In some aspects, the collected data 115 includes a time series or series. That is, within at least one of the modalities, there may be a set or sequence of data points. For example, the data 115 may include a series of images or frames from a video, a series of radar measurements, etc. In some aspects, the machine learning system 125 may process the data 115 as a sequence (e.g., generate a predicted beam 130 based on a timestamp or a sequence of data points, where each timestamp or data point may include data from multiple modalities) and / or discrete data points (e.g., generate a predicted beam 130 for each timestamp or data point, where each timestamp or data point may include data from multiple modalities).

[0035] In some aspects, before processing data 115 with one or more machine learning models, the machine learning system 125 may first perform various preprocessing operations. In some aspects, the machine learning system 125 may synchronize the data 115 for each modality. For example, for a given data point in a given modality (e.g., a given frame of the image data 117A), the machine learning system 125 may identify the corresponding data points in each other modality (e.g., radar captured at the same timestamp, LIDAR captured at the same timestamp, location data captured at the same timestamp, etc.).

[0036] In some aspects, the machine learning system 125 may perform extrapolation on the data 115 when appropriate. For example, if the location data 117D is only available for a subset of timestamps, the machine learning system 125 may use the available location data to determine or infer the relative movement of the UE (e.g., speed and / or direction), and use that movement to extrapolate and generate data points at the corresponding timestamps to match the other modalities in the data 115.

[0037] In some aspects, the machine learning system 125 processes the data 115 independently for each timestamp (e.g., generates the prediction beam 130 for each timestamp). In some aspects, the machine learning system 125 processes the data 115 from a time window. That is, instead of evaluating the data 115 corresponding to a single time, the machine learning system 125 may jointly evaluate N data points or timestamps to generate the prediction beam 130, where N may be a hyperparameter configured by a user or an administrator, or may be a learned value. For example, the machine learning system may use five data points from each modality (e.g., five images, five radar markers, etc.). In some aspects, the time interval between data points (e.g., whether the data points are one second apart, five seconds apart, etc.) may similarly be a hyperparameter configured by a user or an administrator, or may be a learned value.

[0038] In some aspects, the machine learning system 125 may calibrate or transform one or more modalities such as the location data 117D in the modality into a local reference frame to improve the generalization of the model. For example, the machine learning system 125 may transform the location information (e.g., the GNSS coordinates of the base station 105 and / or the UE 110) into a Cartesian coordinate system (e.g., (x,y) coordinates). The machine learning system 125 may then transform the location of the UE 110 from the global reference frame into a local reference frame, such as by subtracting the coordinates of the base station 105 from the coordinates of the UE 110. In some aspects, the machine learning system 125 may then transform this local location of the UE 110 (relative to the base station 105) into a radius r and an angle α relative to the base station 105 in polar coordinates.

[0039] In the illustrated example, the machine learning system 125 evaluates data 115 (e.g., independently or as a sequence of points) to generate predicted beam 130. As discussed above, the predicted beam 130 corresponds to the beam predicted to have the best or optimal characteristics for communicating with UE 110. As shown, this predicted beam 130 is provided to base station 105, which uses this predicted beam to select beam 135 for communication.

[0040] In some aspects, the process may be repeated (e.g., continuously, intermittently, or periodically) or until termination. For example, the machine learning system 125 may repeatedly evaluate data 115 continuously and / or periodically (e.g., every five seconds) when the data is received to generate new beam predictions, allowing the base station 105 to continue to select (optimal) beam 135 for communicating with UE 110. As discussed above, although a single UE 110 is depicted for clarity of concept, in some aspects, the machine learning system 125 may similarly evaluate the corresponding data 115 for each UE 110 and generate corresponding predicted beam 130 for each UE.

[0041] Generally, the machine learning model used by the machine learning system 125 may be trained and / or used by any suitable computing system, and the machine learning system 125 may be implemented by any suitable system. For example, the base station 105 itself may include or be associated with computing resources for training and / or implementing the machine learning system 125. In other aspects, the data 115 may be provided to a remote system for training and / or implementing the model. For example, a cloud system may train the model, and the trained model may be used by the cloud system or by another system (such as the base station 105 or the machine learning system 125) to generate the predicted beam 130. In other examples, some or all of the data 115 may be collected by UE 110 and / or provided to the UE to train the model (or a portion thereof) on the UE 110 and / or use the trained model to generate beam predictions.

[0042] In some aspects, the machine learning system 125 uses a fusion method to dynamically fuse data 115 from each modality, thereby enabling improved beam prediction that enhances communication robustness and throughput.

[0043] Example Architecture for Fusion-Based Beam Selection

[0044] Figure 2 An example architecture 200 for fusing and evaluating image data and location data to provide improved beam selection is depicted. In some aspects, the architecture 200 is used by a machine learning system such as Figure 1 the machine learning system 125.

[0045] Specifically, in the illustrated example, the architecture 200 is configured to evaluate the image data 205 and the location data 210 to generate a predicted beam 235 (e.g., an optimal beam selected or predicted for communication in terms of robustness, throughput, etc.). In some aspects, some or all of the architectures in the architecture 200 can be used as part of a fusion model rather than as an independent model. For example, as discussed in more detail below, the feature extraction and / or preprocessing performed in the feature extraction 215 and / or the preprocessing 220 can be used to provide feature extraction and / or preprocessing at the feature extraction 510A and / or the feature extraction 510D of Figure 5 and / or the feature extraction 610A and / or the feature extraction 610D of Figure 6 In some aspects, the image data 205 and the location data 210 correspond to Figure 1 the data 115 (e.g., corresponding to the image data 117A and the location data 117D respectively).

[0046] In the illustrated example, the image data 205 is first processed using the feature extraction 215. Generally speaking, the feature extraction 215 can include a variety of operations and techniques for extracting or generating features based on the image data 205. For example, in some aspects, the feature extraction 215 corresponds to one or more trained models, such as a feature extractor neural network trained to generate or extract output features based on the input image data 205. As an example, in some aspects, the feature extraction 215 can be performed using a pre-trained model (e.g., the ResNet50 model), where a multi-layer perceptron (MLP) is used as the final or output layer of the feature extraction 215 (rather than a conventional classifier layer). In some such aspects, the MLP can include a linear layer, followed by a batch normalization layer, followed by a second linear layer, followed by a second batch normalization layer, and finally a third linear layer. In various aspects, the specific architecture of the feature extraction 215 can vary depending on the particular implementation. Generally speaking, any architecture that receives an input image and outputs the extracted or generated features can be used.

[0047] Generally speaking, the feature extraction 215 can be used to generate or extract features for each data point in the image data 205. That is, for each frame or image, a corresponding set of features can be generated by the feature extraction 215. These features can then be provided as inputs to the machine learning model 225. In some aspects, these features are provided independently for each image frame (e.g., rather than a sequence). In some aspects, these features are provided as a sequence or time series, as discussed above. For example, if a sequence of features is used, any architecture that receives a sequence of input data points can be used, such as a gated recurrent unit (GRU) network, a recurrent neural network (RNN), a long short-term memory (LSTM) architecture, etc. can be used as the machine learning model 225.

[0048] In the illustrated example, the location data 210 (e.g., coordinates and / or movement of a base station and / or UE determined via, for example, Global Positioning Satellite (GPS) or other positioning technologies) is first processed using preprocessing 220. The preprocessing 220 may generally include any relevant operations or transformations. For example, in some aspects, the preprocessing 220 may include extrapolating any incomplete location data 210 to match the timestamp of the image data 205. For example, if there are five timestamps in the image data 205 and only a subset of these timestamps are available in the location data 210, the preprocessing 220 may include inferring or determining the UE location based on the trajectory of the UE (relative to the base station), as indicated in the existing location data 210. That is, given the trajectory of the UE relative to the base station location, the preprocessing 220 may include inferring the location of the UE at the next time step.

[0049] Generally speaking, the preprocessing 220 can be used to generate or extract features for each data point or timestamp. That is, for each timestamp used (e.g., for each data point in the image data 205), a corresponding set of features (e.g., relative position, angle, range, trajectory, etc.) can be generated or determined by the preprocessing 220. This data can then be provided (individually or as a sequence) as input to the machine learning model 225.

[0050] In some aspects, rather than providing the extracted image features and location data separately, the machine learning system may first combine the image features and location features. For example, at each timestamp, the machine learning system may concatenate, sum, or otherwise aggregate or fuse the image features and location features. In this way, the system generates a combined or concatenated set of features based on the image data 205 and the location data 210. This combined data (or sequence) can then be used as input to the machine learning model 225. In some aspects, if the data from one modality (e.g., image data) does not have corresponding data in the other modality for the same timestamp (e.g., there is no location data for that time), the image data may be discarded or otherwise less weighted compared to data that has a counterpart in the other modality.

[0051] In some aspects, if a recursive or time series model (such as a GRU) is used, the machine learning model 225 may include a set of nodes (e.g., one node for each timestamp in the input sequence). The input to a given node may include the features at a given timestamp in the data and the output of the previous node. That is, the first node may receive the features at the first timestamp to generate an output, and the second node may receive the features and the output of the first node at the second timestamp to generate a second output. This process may continue until the final node, where the output of the final node may be output by the machine learning model 225 or may be processed (e.g., using a classifier) to generate an output prediction.

[0052] In the illustrated example, the machine learning model 225 outputs data or features to a classifier 230 (e.g., an MLP). Generally, any classifier architecture may be used. For example, in some aspects, the classifier 230 may be an MLP that includes a linear layer, a layer for dropout, a non-linear layer (e.g., a rectified linear unit (ReLU) operation), and another linear layer.

[0053] As shown, the classifier 230 outputs a predicted beam 235. As discussed above, this predicted beam 235 may correspond to the beam that is predicted or expected to provide the best available communication to the UE, such as the best robustness, the highest throughput, etc.

[0054] Figure 3 An example architecture 300 for using LIDAR data to provide improved beam selection is depicted. In some aspects, the architecture 300 may be used by a machine learning system such as Figure 1 the machine learning system 125. In some aspects, in addition to Figure 2 the architecture 200, the architecture 300 may also be used. In other aspects, the architecture 300 may be used to replace Figure 2 the architecture 200.

[0055] Specifically, in the illustrated example, the architecture 300 is configured to evaluate LIDAR data 305 (e.g., Figure 1 the LIDAR data 117C) to generate a predicted beam 330 (e.g., the optimal beam selected or predicted for communication in terms of robustness, throughput, etc.). In some aspects, some or all of the architecture 300 may be used as part of a fusion model rather than as an independent model. For example, as discussed in more detail below, the feature extraction performed in the encoder 310 and the deep learning model 315 may be used to provide LIDAR feature extraction in Figure 5 the feature extraction 510C and / or Figure 6 the feature extraction 610C.

[0056] In the illustrated example, the LIDAR data 305 is first processed using feature extraction 308 (also referred to as an embedding network in some aspects) that includes an encoder 310 and a deep learning model 315. Generally speaking, feature extraction 308 can include a variety of operations and techniques for extracting or generating features based on the LIDAR data 305. For example, in some aspects, feature extraction 308 can correspond to one or more trained models that are trained to generate or extract output features or generate embeddings based on the input LIDAR data 305. In the illustrated example, the encoder 310 can be implemented as a model that operates on a point cloud (e.g., LIDAR data) to perform efficient convolutions. In some aspects, the encoder 310 includes a PointPillar network.

[0057] In addition, the deep learning model 315 can generally correspond to any suitable architecture, such as a neural network. In some aspects, the deep learning model 315 is used to reduce the dimensionality of the extracted features generated by the encoder 310. In some aspects, the deep learning model 315 is PointNet.

[0058] Generally speaking, feature extraction 308 can be used to generate or extract features for each data point in the LIDAR data 305. That is, for each point cloud (or other data structure used to represent the LIDAR data 305), a corresponding set of features can be generated by feature extraction 308. These features can then be provided (independently or as a sequence) as input to the machine learning model 320. In some aspects, if the model operates on a sequence of data, a GRU network or other recurrent architecture (e.g., RNN, LSTM, etc.) can be used.

[0059] In some aspects, as discussed above, to process sequential data, the machine learning model 320 can include a set of nodes (e.g., one node for each timestamp in the input data), where the input to a given node can include the features at a given timestamp in the data as well as the output of the previous node. In the illustrated example, the machine learning model 320 outputs data or features to a classifier 325. As discussed above, the classifier 325 can be implemented using a variety of architectures such as an MLP. In some aspects, the classifier 325 is an MLP that includes a single linear layer.

[0060] As shown, the classifier 325 outputs a predicted beam 330. As discussed above, this predicted beam 330 can correspond to the beam that is predicted or expected to provide the best available communication to the UE, such as the best robustness, the highest throughput, etc.

[0061] Figure 4 An example architecture 400 for using radar data to provide improved beam selection is depicted. In some aspects, the architecture 400 can be implemented by a machine learning system such asFigure 1 is used by the machine learning system 125. In some aspects, in addition to Figure 2 architecture 200 and / or Figure 3 architecture 300, architecture 300 can also be used. In other aspects, architecture 300 can be used to replace Figure 2 architecture 200 and / or Figure 3 architecture 300.

[0062] Specifically, in the illustrated example, architecture 400 is configured to evaluate radar data 405 (e.g., Figure 1 radar data 117B) to generate a predicted beam 435 (e.g., an optimal beam selected or predicted for communication in terms of robustness, throughput, etc.). In some aspects, some or all of the architectures in architecture 400 can be used as part of a fusion model rather than as an independent model. For example, as discussed in more detail below, the feature extraction performed in preprocessing 410 and feature extractions 415A - 415C can be used to provide radar feature extraction in Figure 5 feature extraction 510B and / or Figure 6 feature extraction 610B.

[0063] In the illustrated example, the radar data 405 is first processed at preprocessing 410. In some aspects, at preprocessing 410, multiple different outputs are generated or extracted for each data point / sample in the radar data 405. For example, preprocessing 410 can generate or extract a range - velocity map (which can be represented as V in some aspects), a range - angle map (which can be represented as R in some aspects), and / or a radar cube (which can be represented as X in some aspects) for each timestamp or data point.

[0064] Generally speaking, preprocessing 410 can be used to generate or extract outputs for each data point in the radar data 405. For example, for each image (or other data structure used to represent radar data), a corresponding range - velocity map, range - angle map, and radar cube can be generated by preprocessing 410.

[0065] In the illustrated example, each of these preprocessed data outputs is provided to the corresponding specific feature extractions 415A - 415C. Generally speaking, each feature extraction 415A - 415C (collectively referred to as feature extraction 415) can include a variety of operations and techniques for extracting or generating features based on the radar data 405. For example, in some aspects, the feature extraction 415 for the radar data 405 can include using a constant false alarm rate (CFAR) adaptive algorithm to generate a feature output.

[0066] As another example, one or more of the feature extractions in feature extraction 415 may correspond to machine learning models, such as neural networks trained to generate or extract features from corresponding radar outputs (e.g., where feature extraction 415A extracts features from a range-velocity map, feature extraction 415B extracts features for a radar cube, and feature extraction 415C extracts features for a range-angle map).

[0067] As shown, for each example in the radar data 405 (e.g., each timestamp or data point), a corresponding set of learned features 420 can thus be compiled (e.g., by combining, concatenating, or reshaping the outputs of each feature extraction 415 of the input radar data 405). By doing so for each piece of radar data in the radar data 405, a set or sequence of learned features 420 is generated.

[0068] The learned features 420 can then be provided (as a sequence and / or independently) as input to the machine learning model 425. In some aspects, if a sequence of features is used, any architecture that receives a sequence of input data points—such as a GRU network, RNN, LSTM architecture, etc.—can be used as the machine learning model 425.

[0069] In some aspects, as discussed above, to process sequential data, the machine learning model 425 can include a set of nodes (e.g., one node for each timestamp in the input data), where the input to a given node can include the features at a given timestamp in the data and the output of the previous node. In the illustrated example, the machine learning model 425 outputs data or features to the classifier 430. In some aspects, the classifier 430 includes an MLP that includes a linear layer, followed by a dropout layer, followed by a non-linear layer (e.g., ReLU), and then a final linear layer.

[0070] As shown, the classifier 430 outputs a predicted beam 435. As discussed above, this predicted beam 435 can correspond to the beam predicted or expected to provide the best available communication to the UE, such as the best robustness, highest throughput, etc.

[0071] In some aspects, fusion techniques can be used to combine some or all of the architectures 200, 300, and / or 400 to generate improved beam predictions. Figure 5 An example architecture 500 for using fusion to provide improved beam selection is depicted. In some aspects, the architecture 500 can be used by a machine learning system such as Figure 1 machine learning system 125.

[0072] As discussed above, in some architectures, data from different modalities can be fused to provide improved beam selection. Although Figure 3 and Figure 4For clarity of concept, single-modal feature extraction and beam selection are depicted, but in some aspects, some or all of the modalities may be fused in a single architecture. For example, architecture 500 is configured to fuse an image modality (represented by image data 505A, which may correspond to Figure 1 image data 117A), a radar modality (represented by radar data 505B, which may correspond to Figure 1 radar data 117B), a LIDAR modality (represented by LIDAR data 505C, which may correspond to Figure 1 LIDAR data 117C), and a position modality (represented by position data 505D, which may correspond to Figure 1 position data 117D).

[0073] Although the illustrated example depicts the use of four different modalities, in various aspects, the architecture may use fewer modalities or may use additional modalities not depicted (e.g., by adding more feature extraction components for new modalities). Additionally, in some aspects, the machine learning system may selectively enable or disable various modalities. That is, in some aspects, the machine learning system may dynamically determine whether to use data from each modality to generate prediction beams (e.g., whether to use a subset of modalities at one or more times).

[0074] In the illustrated example, each modality of the input data 505 undergoes modality-specific feature extraction 510. For example, the image data 505A undergoes image feature extraction 510A (which may correspond to Figure 2 feature extraction 215), the radar data 505B undergoes radar feature extraction 510B (which may correspond to Figure 4 preprocessing 410 and / or feature extraction 415A-415C), the LIDAR data 505C undergoes LIDAR feature extraction 510C (which may correspond to Figure 3 feature extraction 308), and the position data 505D undergoes feature extraction 510D (which may correspond to Figure 2 preprocessing 220).

[0075] In the illustrated architecture 500, image features (output by image feature extraction 510A), radar features (output by radar feature extraction 510B), LIDAR features (output by LIDAR feature extraction 510C), and location features (output by location feature extraction 510D) are provided to a fusion component 515. In some aspects, the fusion component 515 is a machine learning model trained to fuse features from each modality to generate a joint or aggregated set of features. In some aspects, the fusion component 515 uses one or more attention-based mechanisms to fuse the features. In some aspects, the fusion component 515 uses operations such as concatenation, summation, stacking, averaging, etc. to aggregate and fuse the features.

[0076] Advantageously, if the fusion component 515 is a trained model, the machine learning system can thereby learn (during training) an optimal or at least improved way to fuse the extracted features, such as using an attention-based mechanism. In some aspects, as discussed above, the fusion component 515 can selectively or dynamically select which features or modalities to fuse depending on various criteria or implementation details. For example, the fusion component 515 can determine which features are available (e.g., for fusing features from any modality with available data), and / or can evaluate the features themselves to determine whether to fuse them (e.g., determine whether to include features from a given modality based on whether the features meet a defined criterion such as maximum sparsity).

[0077] Thus, in the illustrated example, the fusion component 515 can thereby be used to fuse features from any number of modalities. In some aspects, the architecture 500 can be modified (e.g., adding or removing feature extractions 510 to add or remove modalities) during or prior to training, thereby allowing the fusion model to be trained for any specific combination of modalities. In some aspects, unused modalities can remain in the architecture. For example, during training, the machine learning system can avoid providing input data for any unused or unwanted modalities, and the system can learn to effectively bypass these features when fusing modalities.

[0078] In some aspects, as discussed above, the fusion component 515 can fuse features relative to each timestamp or data point. That is, the machine learning system can fuse the corresponding features from each modality for each given timestamp to generate a set of fused features for the given timestamp. In some aspects, these fused features can be evaluated independently for each timestamp (generating a corresponding prediction beam for each timestamp), as discussed above. In other aspects, the fused features can be evaluated as a series or sequence (e.g., evaluating a window of five sets of fused features) to generate a prediction beam.

[0079] As shown, once the fused feature has been generated, the fused feature is provided as an input to a machine learning model 520. For example, the machine learning model 520 may correspond to or include Figure 2 the machine learning model 225 of Figure 3 the machine learning model 320 of Figure 4 the machine learning model 425 of, etc. In some aspects, as discussed above, the machine learning model 520 processes the fused feature as a time series. For example, the machine learning model 520 may include a GRU network, an RNN, an LSTM, etc.

[0080] In the illustrated example, the machine learning model 520 outputs features or other data to a classifier 525 (e.g., an MLP). For example, in some aspects, the classifier 525 includes an MLP that includes a linear layer, followed by a dropout layer, followed by a non-linear layer (e.g., ReLU), and then a final linear layer.

[0081] As shown, the classifier 525 outputs a predicted beam 530. As discussed above, the predicted beam 530 may correspond to a beam that is predicted or expected to provide the best available communication to the UE, such as the best robustness, the highest throughput, etc. Advantageously, by fusing multiple modalities, the machine learning system is generally able to generate more accurate beam predictions (e.g., more accurately select a beam that will improve or result in good quality communication with the UE).

[0082] Figure 6 An example architecture 600 for using sequential fusion to provide improved beam selection is depicted. In some aspects, the architecture 600 may be used by a machine learning system such as Figure 1 the machine learning system 125 of

[0083] As discussed above, in some architectures, data from different modalities may be fused to provide improved beam selection. In the illustrated example, the architecture 600 is configured to use a sequential fusion process to fuse modalities. Specifically, the architecture 600 includes an image modality (represented by image data 605A, which may correspond to Figure 1 the image data 117A of Figure 1 the radar data 117B of Figure 1 the LIDAR data 117C of Figure 1 the location data 117D of

[0084] Although the illustrated example depicts the use of four different modalities, in all respects, the architecture can use fewer modalities or can use additional modalities not depicted (e.g., by adding more encoder-decoder fusion models 615, as discussed in more detail below). Additionally, although the illustrated example depicts a specific sequence of modalities (e.g., where image data and radar data are processed first, then LIDAR data, and finally location data), the specific ordering used can vary depending on the particular implementation.

[0085] In the illustrated example, each modality of the input data undergoes modality-specific feature extraction. For example, image data 605A undergoes image feature extraction 610A (which may correspond to Figure 2 feature extraction 215), radar data 605B undergoes radar feature extraction 610B (which may correspond to Figure 4 preprocessing 410 and / or feature extraction 415A - 415C), LIDAR data 605C undergoes LIDAR feature extraction 610C (which may correspond to Figure 3 feature extraction 308), and location data 605D undergoes feature extraction 610D (which may correspond to Figure 2 preprocessing 220).

[0086] In the illustrated architecture 600, the image features (output by image feature extraction 610A) and the radar features (output by radar feature extraction 610B) are provided to a first encoder-decoder fusion model 615A. Generally speaking, the encoder-decoder fusion model 615A is an attention-based machine learning model. For example, the encoder-decoder fusion model 615A can be implemented using one or more transformer blocks (e.g., vision transformers). In some aspects, to provide the features, the system may reshape the image features and the radar features from a single timestamp such that these features are in the appropriate format for the encoder-decoder fusion model 615A. The encoder-decoder fusion model 615A can then process these features to generate fused features for the image data 605A and the radar data 605B.

[0087] Advantageously, the encoder-decoder fusion model 615A can thus learn (during training) how to use an attention-based mechanism to fuse the extracted features. As shown, the fused features are then output to a second encoder-decoder fusion model 615B, which also receives the LIDAR features (generated by feature extraction 610C).

[0088] Generally speaking, as discussed above, the encoder-decoder fusion model 615B is also an attention-based machine learning model. For example, the encoder-decoder fusion model 615B can be implemented using one or more transformer blocks (e.g., vision transformers). In some aspects, as discussed above, the system can reshape the fused image features, radar features, and LIDAR features from the corresponding timestamp such that these features are in a proper format for the encoder-decoder fusion model 615B. The encoder-decoder fusion model 615B can then generate a new set of fused features based on a first set of intermediate fused features (generated by the encoder-decoder fusion model 615A based on the image data 605A and the radar data 605B) and the LIDAR data 605C.

[0089] As shown, the fused features are then output to a third encoder-decoder fusion model 615C, which also receives position features (generated by the feature extraction 610D).

[0090] Generally speaking, as discussed above, the encoder-decoder fusion model 615C can also include an attention-based machine learning model, which can be implemented using one or more transformer blocks (e.g., vision transformers). In some aspects, the system can similarly reshape the fused image features, radar features, LIDAR features, and position features from the corresponding timestamp such that these features are in a proper format for the encoder-decoder fusion model 615C. The encoder-decoder fusion model 615C can then generate fused features for all input data modalities at a given timestamp.

[0091] In some aspects, the encoder-decoder fusion model 615 can thus be used to fuse features from any number of modalities (e.g., where encoder-decoder fusion models can be added or removed depending on the modalities used). That is, during or before training, the architecture 600 can be modified (e.g., adding or removing encoder-decoder fusion models to add or remove modalities) to allow the fusion model to be trained for any specific combination of modalities. Alternatively, in some aspects, unused encoder-decoder fusion models and / or modalities can remain in the architecture. During training, since no input data is provided for the missing modality (or modalities), the model can learn to effectively bypass the corresponding encoder-decoder fusion models (e.g., where the data passes through these models unchanged).

[0092] As shown, once the fused features have been generated, the fused features are provided as input to the machine learning model 620. For example, the machine learning model 620 can correspond to or include Figure 2 the machine learning model 225, Figure 3The machine learning model 320, Figure 4 The machine learning model 425, etc. In some aspects, as discussed above, the machine learning model 620 processes the fused features as a time series. For example, the machine learning model 620 may include a GRU network, RNN, LSTM, etc.

[0093] In the illustrated example, the machine learning model 620 outputs features or other data to a classifier 625 (e.g., MLP). For example, in some aspects, the classifier 625 includes an MLP that includes a linear layer, followed by a dropout layer, followed by a non-linear layer (e.g., ReLU), and then a final linear layer.

[0094] As shown, the classifier 625 outputs a predicted beam 630. As discussed above, the predicted beam 630 may correspond to the beam predicted or expected to provide the best available communication to the UE, such as the best robustness, highest throughput, etc. Advantageously, by fusing multiple modalities, the machine learning system is generally able to generate more accurate beam predictions (e.g., more accurately select the beam that will improve or result in good quality communication with the UE).

[0095] Additionally, by using an attention-based model to fuse features from each modality, the architecture 600 is able to achieve high robustness and accuracy (e.g., reliably select or recommend the optimal beam for communication).

[0096] Example Workflow for Pretraining and Scenario Adaptation Using Simulated Data

[0097] Figure 7 Depicts an example workflow 700 for pre-training and scenario adaptation using simulated data. In some aspects, the workflow 700 is performed by a machine learning system such as Figure 1 The machine learning system 125. In some aspects, the workflow 700 is performed entirely or partially by a dedicated training system. For example, one part of the workflow 700 (e.g., pre-training using synthetic data in block 705A) may be performed by a dedicated training system, and the scenario or environment adaptation (in block 705B) may be performed by a machine learning system that uses a trained model to generate predicted beams during runtime. As discussed above, the predicted beam may correspond to the beam predicted or expected to provide the best available communication to the UE, such as the best robustness, highest throughput, etc. In some aspects, by leveraging synthetic data to pre-train and adapt the machine learning model, the workflow 700 enables significant improvement in prediction accuracy with reduced training data.

[0098] In some aspects, to perform pre-training (in block 705A), synthetic data can be created based on codebook 720 (referred to as "CB" in some aspects) and simulator 725. As discussed above, codebook 720 generally includes or indicates a set of beams, each beam targeting a specific angular direction relative to base station 735. Generally speaking, simulator 725 can correspond to a model or representation of the received power at various angles of arrival based on a specific codebook entry (e.g., the selected beam) for transmission. That is, simulator 725 can be a received power simulator that can be used to determine, predict, or indicate the predicted received power (at the UE) of a signal transmitted by base station 735 and / or the predicted received power (at base station 735) of a signal transmitted by the UE, according to the codebook entry for transmission, for a specific angle of arrival and / or angle of departure (e.g., based on relative angle information such as the angle of the receiving UE relative to the transmitting base station 735).

[0099] For example, the transmission characteristics of the antennas of base station 735 can be measured in an anechoic chamber, and a profile of codebook 720 (indicating the power profile of each codebook element at various angles of arrival / departure) can be captured or determined. These profiles can then be applied to predict the best beam based on the angle of arrival (e.g., via simulator 725). In some aspects, since direct generalization through a large amount of training data and measurements is not feasible (e.g., it is not practical or even impossible in some cases to collect real-label data in the actual environment), the synthetic enhancement pre-training step depicted can be deployed at block 705A to facilitate or improve model generalization.

[0100] Specifically, as illustrated in block 705A, the physical locations of transmitting base station 735 and / or the UE can be determined (or simulated), as illustrated by GPS data 710A in the depicted figure. This (real or simulated) GPS data 710A can be used to determine or estimate the angle of arrival and / or angle of departure of the line-of-sight (LOS) signal between base station 735 and the UE, as discussed above. In some aspects, GPS data 710A is also used to determine or estimate the range or distance to the UE. Simulator 725 can then be used to predict the best beam based on the angle of arrival or departure and / or based on the range. As shown, this output prediction (from the simulator) can be used as the target or label for training machine learning model 730.

[0101] Generally speaking, machine learning model 730 can include a variety of components and architectures, including various feature extraction or preprocessing operations, fusion of multiple modalities (using learned fusion models and / or static fusion techniques), deep learning models, beam prediction or classification components (e.g., MLP), etc. For example, depending on the specific modalities supported, machine learning model 730 includes one or more of the following: Figure 2feature extraction 215, preprocessing 220, machine learning model 225, and / or classifier 230; Figure 3 feature extraction 308, machine learning model 320, and / or classifier 325; Figure 4 preprocessing 410, feature extraction 415, machine learning model 425, and / or classifier 430; Figure 5 feature extraction 510, fusion component 515, machine learning model 520, and / or classifier 525; and / or Figure 6 feature extraction 610, encoder-decoder fusion model 615, machine learning model 620, and / or classifier 625.

[0102] In the illustrated example, during the pre-training block 705A, the machine learning model 730 receives (as input) data from one or more available modalities (e.g., GPS data 710A and image data 715A in the illustrated example). In some aspects, these modalities may undergo various feature extraction and / or fusion operations as discussed above. Using these inputs, the machine learning model 730 predicts or selects one or more optimal beams. That is, based on the relative position and / or the captured image, the machine learning model 730 predicts which beams will result in optimal communication with the UE. During the pre-training block 705A, the predicted beams can be compared with the beams selected or predicted by the simulator 725 (used as labels as discussed above), and the difference can be used to generate a loss. The machine learning system can then use this loss to refine the machine learning model 730.

[0103] In various aspects, this pre-training step (at block 705A) can be performed during block 705A using any number of examples (e.g., any number of input samples, each input sample including position information (e.g., GPS data 710A) and / or one or more other modalities (e.g., image data 715A)).

[0104] In some aspects, since the simulator 725 predicts the optimal beam based on LOS (e.g., based on GPS data 710A and / or relative position or angle, predicting that the optimal beam will be a LOS beam), the machine learning model 730 can learn the alignment or mapping between the arrival and / or departure angles and the codebook entries (e.g., beams) in the codebook 720. However, in some aspects, scenario- or environment-specific scenarios (such as terrain, objects that can block or reflect RF waves, etc.) can be considered during the adaptation phase (e.g., at block 705B). That is, during the pre-training in block 705A, the machine learning model 730 can learn to predict the LOS beam as the optimal beam, although in a real scenario, other beams may be more suitable (e.g., due to reflections and / or blocking objects).

[0105] In some aspects, during the pre-training step in block 705A, one or more additional modalities (e.g., image data 715A from a camera) can be utilized in conjunction with positioning data (e.g., GPS data 710A) to assist the machine learning model 730 in learning to become invariant to other (non-communication) changes such as UE changes. For example, if the image data 715A is used as an input during training, the machine learning model 730 can learn to generalize with respect to a specific UE type or appearance (e.g., predict LOS beams regardless of whether the UE is a car, laptop, smartphone, or other object with different visual appearances). Other modalities can be similarly used during pre-training to allow the machine learning model 730 to become invariant with respect to such other modalities.

[0106] As discussed above, the complete measurement of beams (used by some conventional methods to train models) creates a large amount of data and takes a long time to collect (e.g., if all beams are to be scanned for each possible UE location). In practice, this data collection may often be impractical or completely impossible for real-world environments. In some aspects, the machine learning system can use the pre-training block 705A to minimize or eliminate this overhead. Additionally, in some aspects, the machine learning system can reduce this overhead by using the adaptation phase illustrated in block 705B to select the information content of the complete measurement based on only one initial beam measurement, as discussed in more detail below. In some aspects, this immediate measurement selection process is referred to as scenario or environment adaptation, where the machine learning system trains or updates the machine learning model 730 based on data from the specific environment in which the base station 735 is deployed.

[0107] Specifically, using the pre-training step in block 705A, the machine learning system can use a threshold-based metric to determine which real-world measurements should be collected for further refinement of the machine learning model 730. In the illustrated example, once pre-training is complete (e.g., when the machine learning model 730 is sufficiently trained, when there is no additional synthetic data remaining for training, when a limited amount of time or resources have been spent on pre-training, etc.), the system can transition to the adaptation phase at block 705B.

[0108] As illustrated in block 705B, position information (e.g., GPS data 710B) and / or other modality data (e.g., image data 715B) are collected in the real-world environment around the base station 735. That is, the system can capture actual images (e.g., Figure 1 image data 117A) and actual position data (e.g., Figure 1 position data 117D) during scenario adaptation in block 705B.

[0109] In the illustrated example, the location information is again used as an input to the simulator 725 (as discussed above) to predict the best beam (e.g., the LOS beam). Additionally, as shown, this beam is indicated to the base station 735, and the actual received power of the predicted beam (as transmitted and / or received by the base station 735) can be determined. In the illustrated example, the predicted received power (predicted by the simulator) can be compared with the actual received power (determined by the base station) at operation 740.

[0110] Specifically, in some aspects, the difference between the simulated or predicted power and the actual or measured power can be determined. In some aspects, if the difference meets one or more criteria (e.g., meets or exceeds a threshold), the system can determine that additional data should be collected (as indicated by the dashed line between operation 740 and the base station 735). If not, the machine learning system can determine that additional data for the UE location is unnecessary.

[0111] That is, if the difference between the predicted power of the best beam and the actual power of the selected beam (roughly) aligns, the system does not need to measure every other beam for the current location of the UE (e.g., the received power when using every other element in the codebook). In the illustrated example, if the difference exceeds the threshold, the machine learning system can initiate or request a full scan measurement (e.g., measure the actual received power at the remaining (unselected) beams / codebook elements). For example, due to obstacles, RF reflection, refraction, and / or absorption, etc., the predicted beam (e.g., the LOS beam) may have an actual received power lower than predicted. In such cases, additional real-world data can be collected to train an improved model due to the specific environment of the base station 735.

[0112] In the illustrated example, during the scenario adaptation phase in block 705B, these actual measurements (for the initially selected beam and / or for additional beams, if the predicted power is substantially different from the received power) can be used as the target output or label to train the machine learning model 730 (which was pre-trained at block 705A). Additionally, the corresponding input modalities (e.g., GPS data 710B and image data 715B) can be similarly used as inputs during this scenario adaptation. In this way, the machine learning model 730 learns and adapts to the specific scenario or environment of the base station 735, and thus learns to generate or predict the best beam based on the specific environment (e.g., in addition to predicting a simple LOS beam).

[0113] As shown, any number of examples (e.g., multiple UE locations in the usage environment) can be used to perform such adaptation. Once the scenario adaptation is complete, the machine learning system can transition to the deployment or runtime phase (illustrated by block 705C). Although the illustrated example suggests a one-way workflow (moving from pre-training in block 705A to scenario adaptation in block 705B and into deployment in block 705C), in some aspects, the machine learning system can move freely between phases depending on the specific implementation (e.g., re-entering adaptation in block 705B periodically or when model performance or communication efficacy degrades).

[0114] In the illustrated example, for the deployment phase in block 705C, the simulator 725 can be discarded (or otherwise not used), and the trained and / or adapted machine learning model 730 can be used to process input modalities (e.g., GPS data 710C and image data 715C) to select or predict the best beam. The prediction can then be provided to the base station 735 and used to drive the beam selection performed by the base station 735, thereby substantially improving communication, as discussed above.

[0115] In some aspects, the specific loss used to refine the various models described herein (such as the depicted machine learning model 730) can vary depending on the specific implementation. For example, for a classification task, the cross-entropy (CE) loss can be considered the standard choice for training deep models. In some aspects, the CE loss used to train the model can be defined using Equation 1 below, where y n is the index of the (predicted) best beam, C is the number of codebooks (e.g., elements in the codebook 720 or the number of beams), and N is the batch size:

[0116]

[0117] In some aspects, while the CE loss focuses on a single value (e.g., a single beam) for only the correct label, it can also be true that the second (or third or fourth) beam (e.g., on one or more reflections) can receive the same or a similar amount of received power. Thus, in some aspects, the machine learning system casts the task as a multi-class estimation problem, such as using binary cross-entropy (BCE) loss. In some aspects, to provide multi-class estimation, the machine learning system can assign or generate weights for each beam, where the highest weighted beam becomes the target label during training. In some aspects, the BCE loss used to train the model is defined using Equation 2 below, where C is the number of codebook entries in the codebook 720, and the number of "simultaneous" classes (e.g., the number of good candidates or beams to be predicted or selected) is a scan parameter:

[0118]

[0119] In some aspects, using BCE, the ground truth beam vector y corresponding to the best beam (also referred to as the label in some aspects) can be defined using various strategies. In some aspects, the machine learning system assigns weights to each beam (e.g., based on the predicted or actual received power when using that beam), such that beams with higher received power are weighted more highly.

[0120] For example, in some aspects, the system can clip the received power profile to a defined power threshold p t (e.g., defining the target label as all beams with the minimum received power or minimum weight), such as using Equation 3 below. In some aspects, the system can select the top B beams (e.g., defining the target label as the B beams with the most received power or highest weight), such as using Equation 4 below:

[0121]

[0122] In this way, the system can use labels and loss formulas that allow the machine learning model 730 to learn using more than one beam as the target for a given input sample to train the model.

[0123] Example Method for Fusion-Based Beam Selection

[0124] Figure 8 is a flowchart depicting an example method 800 for achieving improved beam selection by fusing data modalities. In some aspects, method 800 is performed by a machine learning system such as Figure 1 the machine learning system 125.

[0125] In some aspects, method 800 uses Figure 5 the architecture 500, Figure 6 the architecture 600, and / or Figure 7 the machine learning model 730 to provide additional details of beam selection. Generally speaking, method 800 can be used during training (e.g., during the forward pass of data through the model, where the backward pass is then used to update the parameters of each component of the architecture) and during inference (e.g., when input data is processed to select one or more beams during runtime).

[0126] At block 805, the machine learning system accesses input data. As used herein, "accessing" data generally includes receiving, retrieving, collecting, generating, determining, measuring, requesting, obtaining, or otherwise gaining access to the data. As discussed above, this input data can include data for any number of modalities, such as images (e.g., Figure 1 the image data 117A), radar data (e.g., Figure 1radar data 117B), LIDAR data (e.g., Figure 1 LIDAR data 117C), position data (e.g., Figure 1 position data 117D), etc. In some aspects, as discussed above, the input data includes a time series of data (e.g., a sequence of data for each modality).

[0127] At block 810, the machine learning system selects two or more modalities from the modalities to be fused. In some aspects, if non-sequential fusion is used (e.g., as discussed above with reference to Figure 5 ), the machine learning system may select all available modalities at block 810. In some aspects, if sequential fusion is used (as discussed above with reference to Figure 6 ), the machine learning system may select two modalities. As discussed above, the specific order for selecting modalities for sequential fusion may vary depending on the particular implementation, and in some aspects, any two modalities may be selected for the first fusion.

[0128] At block 820, the machine learning system performs modality-specific feature extraction on each of the selected modalities. For example, as discussed above, the machine learning system may use Figure 5 modality-specific feature extraction 510 and / or Figure 6 modality-specific feature extraction 610 for each modality to generate corresponding features.

[0129] At block 825, the machine learning system then generates fused features based on the extracted features for the selected modalities. For example, as discussed above, the machine learning system may use a fusion component (e.g., Figure 5 fusion component 515) and / or an attention-based architecture such as a transformer model (e.g., Figure 6 encoder-decoder fusion model 615) to fuse the extracted features.

[0130] At block 830, the machine learning system determines whether there is at least one additional modality reflected in the input data accessed at block 805 (e.g., if sequential fusion is used). If not, method 800 proceeds to block 855, where the machine learning system processes the fused features to generate one or more predicted beams. For example, if a time series is used, the machine learning system may use a GRU network and / or a classifier to process the fused features and generate predicted beams (e.g., Figure 5 predicted beam 530).

[0131] Returning to block 830, if the machine learning system determines that there is at least one additional modality, method 800 proceeds to block 835, where the machine learning system selects one remaining modality from the remaining modalities.

[0132] At block 840, the machine learning system performs modality-specific feature extraction on the selected modality. For example, as discussed above, the machine learning system may use the modality-specific feature extraction 510 of Figure 5 and / or the modality-specific feature extraction 610 of Figure 6 for the specifically selected modality to generate corresponding features.

[0133] At block 845, the machine learning system then generates a fused feature based on the extracted features for the selected modality and the previously generated fused features for the previous modality (e.g., generated at block 825, or generated at block 845 during previous iterations of blocks 835 - 750). For example, as discussed above, the machine learning system may use an attention-based architecture, such as a transformer model (e.g., Figure 6 the encoder-decoder fusion model 615) to fuse the extracted features with the fused features from the previous fusion model (e.g., features that have been fused by earlier layers).

[0134] At block 850, the machine learning system determines whether there is at least one additional modality reflected in the input data accessed at block 805. If so, method 800 returns to block 835 to select the next modality for processing and fusion. If there are no remaining additional modalities, method 800 proceeds to block 855, where the machine learning system processes the fused features to generate one or more predicted beams.

[0135] For example, as discussed above, the machine learning system may use one or more machine learning models and / or classifiers to select one or more beams based on the data. In some aspects, if sequences or time series are used, the machine learning system may use a recurrent model such as a GRU network to process the fused features and generate predicted beams.

[0136] In this way, the machine learning system can generate and fuse modality-specific features to drive improved beam selection. As discussed above, the predicted beam(s) may correspond to the beam(s) predicted or expected to provide the best available communication to the UE, such as the best robustness, highest throughput, etc. Advantageously, by using an attention-based model to fuse features from each modality, the machine learning system can achieve high robustness and accuracy (e.g., reliably select or recommend the optimal beam for communication).

[0137] Although not depicted in the illustrated example, method 800 may additionally include actually facilitating communication with the UE based on the predicted beam. For example, facilitating communication may include indicating or providing the predicted beam to the base station, instructing the base station to use the indicated beam, actually using the beam (e.g., if the machine learning system operates as a component of the base station itself), etc.

[0138] Additionally, although not depicted in the illustrated example, method 800 may then return to block 805 to access new input data to generate an updated set of predicted beams. For example, method 800 may be performed continuously, periodically (e.g., every five seconds), etc.

[0139] Furthermore, although not depicted in the illustrated example, in some aspects, method 800 may be performed separately (sequentially or in parallel) for each UE wirelessly connected to the base station. That is, for each respective associated UE, the machine learning system may access the corresponding input data for that UE to generate predicted beams for communicating with the respective UE. In some aspects, some of the input data (such as image data and radar data) may be shared / global across multiple UEs, while some of the input data (such as location data) is specific to the respective UE.

[0140] Example Method for Pretraining and Scenario Adaptation

[0141] Figure 9 is a flowchart depicting an example method 900 for pre-training and scenario adaptation. In some aspects, method 900 is performed by a machine learning system such as Figure 1 machine learning system 125. In some aspects, method 900 is performed fully or partially by a dedicated training system. For example, one part of method 900 (e.g., the pre-training in blocks 905, 910, 915, and 920) may be performed by a dedicated training system, and the scenario or environment adaptation (in blocks 925, 930, 935, 940, 945, 950, and 955) may be performed by a machine learning system that uses the trained model to generate predicted beams during runtime. In some aspects, method 900 provides additional details for Figure 7 workflow 700.

[0142] At block 905, the machine learning system accesses location information (e.g., Figure 1 location data 117D of Figure 7 and / or

[0143] At block 910, the machine learning system uses a simulator (e.g., the simulator 725 of Figure 7 ) to generate one or more predicted beams based on location information. For example, as discussed above, the simulator can include a mapping between angles of arrival and codebook entries or beams such that the simulator can be used to identify one or more beams corresponding to the location information.

[0144] At block 915, the machine learning system updates one or more parameters of a machine learning model (e.g., the machine learning model 730) based on the predicted beams identified using the simulator. Generally, the specific techniques used to update the model parameters can vary depending on the particular implementation. For example, in the case of a neural network, the machine learning system can generate an output of the model (e.g., one or more selected beams and / or predicted received power of one or more beams) based on inputs (e.g., based on location information and / or one or more other modalities such as image data, radar data, LIDAR data, etc.). The model output can then be compared to the beams and / or received power predicted by the simulator to generate a loss (e.g., using Equation 1 and / or Equation 2 above). The loss can then be used to refine the model parameters (e.g., using backpropagation).

[0145] In some aspects, updating the model parameters can include updating one or more parameters related to beam selection rather than feature extraction. That is, the machine learning system can use a pre-trained feature extractor (trained by the machine learning system or by another system) for each data modality and can avoid modifying these feature extractors during method 900. In other aspects, the machine learning system can also optionally update one or more parameters of the feature extractor during method 900.

[0146] In some aspects, as discussed above, this pre-training operation can be used to train the model to select LOS beams based on location information. Additionally, as discussed above, the optional use of other modalities during this pre-training can make the model invariant to aspects that do not affect communication performance such as the appearance or radar cross-section of the UE.

[0147] At block 920, the machine learning system determines whether one or more pre-training criteria are met. Generally, the specific termination criteria can vary depending on the particular implementation. For example, determining whether the pre-training criteria are met can include determining whether there are additional samples or examples remaining for training, determining whether the machine learning model exhibits a minimum or desired accuracy with respect to beam prediction, determining whether the model accuracy continues to improve (or has stopped), determining whether a limited amount of time or computational resources have been spent during pre-training, etc.

[0148] If, at block 920, the machine learning system determines that the pre-training termination criterion is not met, method 900 returns to block 905 to continue pre-training. Although the illustrated example depicts a stochastic training operation (e.g., where a stochastic gradient descent based on independent data samples is used to update the model) for the sake of conceptual clarity, in some aspects, the machine learning system may use a batch training operation (e.g., using a batch gradient descent based on a set of data samples to refine the model). If, at block 920, the machine learning system determines that the pre-training termination criterion is met, method 900 proceeds to block 925 to begin scenario or environment adaptation.

[0149] At block 925, the machine learning system accesses location information (e.g., Figure 1 location data 117D of Figure 7 and / or GPS data 710B of

[0150] At block 930, the machine learning system uses a simulator (e.g., Figure 7 simulator 725 of

[0151] At block 935, the machine learning system determines the actual power information of the predicted beam. For example, as discussed above, the machine learning system may direct, request, or otherwise cause the base station to communicate with the UE using the indicated beam, thereby measuring the actual received power.

[0152] At block 940, the machine learning system determines whether one or more threshold criteria are met with respect to the actual received power. For example, in some aspects, the machine learning system may determine whether a predicted beam results in a satisfactory or acceptable received power (or the highest received power by testing one or more adjacent beams). In some aspects, the machine learning system may determine the difference between the predicted power (generated by the simulator) and the actual power (determined in the environment). If the difference is less than the threshold, method 900 may proceed to block 950. That is, if the actual received power is similar to the predicted received power, the machine learning system may determine or infer that the location of the UE (determined at block 925) results in a clear LOS to the UE, or otherwise causes the LOS beam to produce an actual received power that closely matches the predicted power. For these locations, the machine learning system may determine to forego further data collection (e.g., scanning the codebook), thereby substantially reducing the time, computational cost, and power consumption for performing such data collection for the UE's location.

[0153] If at block 940 the machine learning system determines that the threshold criteria are not met (e.g., the difference between the actual received power and the predicted received power is greater than the threshold), method 900 proceeds to block 945. That is, if the actual received power is not similar to the predicted received power, the machine learning system may determine or infer that the location of the UE (determined at block 925) does not result in a clear LOS to the UE (e.g., due to obstacles, reflections, etc.), or otherwise causes the LOS beam to produce an actual received power that does not closely match the predicted power. For these locations, the machine learning system may determine that additional data collection (e.g., scanning the codebook) will be beneficial to model accuracy.

[0154] At block 945, the machine learning system determines the power information for one or more additional beams. For example, as discussed above, the machine learning system may direct, request, or otherwise cause the base station to sweep through the codebook, thereby communicating with the UE using each alternative beam and measuring the actual received power from each beam.

[0155] At block 950, the machine learning system updates the model parameters based on the power information (determined at block 935 and / or at block 945). That is, if the machine learning system determines not to scan the codebook (e.g., the threshold criteria are met at block 940), the machine learning system may update the model parameters based on the actual received power of the selected beam (selected at block 930). If the machine learning system determines to scan the codebook (e.g., the threshold criteria are not met at block 940), the machine learning system may update the model parameters based on all the power information determined at blocks 935 and 945.

[0156] As discussed above, the specific techniques for updating model parameters generally can vary depending on the particular implementation. For example, in the case of a neural network, a machine learning system can generate one or more outputs of the model (e.g., one or more selected beams and / or the predicted received power of one or more beams) based on inputs (e.g., based on location information and / or one or more other modalities such as image data, radar data, LIDAR data, etc.). The model output can then be compared to the beams and / or received power predicted by the simulator to generate a loss (e.g., using Equation 1 and / or Equation 2 above). The loss can then be used to refine the model parameters (e.g., using backpropagation).

[0157] In some aspects, as discussed above, updating the model parameters can include updating one or more parameters related to beam selection rather than feature extraction. That is, the machine learning system can use a pre-trained feature extractor (trained by the machine learning system or by another system) for each data modality and can avoid modifying these feature extractors during Method 900. In other aspects, the machine learning system can also optionally update one or more parameters of the feature extractor during Method 900.

[0158] At block 955, the machine learning system determines whether one or more adaptation termination criteria are met. Generally, the specific termination criteria can vary depending on the particular implementation. For example, determining whether the adaptation termination criteria are met can include determining whether there are additional samples or examples remaining for training, determining whether the machine learning model exhibits a minimum or desired accuracy with respect to beam prediction, determining whether the model accuracy continues to improve (or has stopped), determining whether a limited amount of time or computational resources have been spent during environment adaptation, etc.

[0159] If at block 955 the machine learning system determines that the adaptation termination criteria are not met, Method 900 returns to block 925 to continue with environment adaptation. Although the illustrated example depicts a stochastic training operation (e.g., where a stochastic gradient descent based on independent data samples is used to update the model) for clarity of concept, in some aspects, the machine learning system can use a batch training operation (e.g., using a batch gradient descent based on a collection of data samples to refine the model).

[0160] If at block 955 the machine learning system determines that the adaptation termination criteria are met, Method 900 continues to block 960. At block 960, the machine learning system deploys the model (or otherwise enters the runtime or deployment phase). For example, as discussed above with reference to Figure 7 block 705C, the machine learning system can begin using the trained model to process input data (e.g., location information, image data, etc.) to generate or select an optimal beam for communication.

[0161] Example Method for Wireless Communication Configuration

[0162] Figure 10 is a flowchart depicting an example method 1000 for implementing an improved wireless communication configuration using machine learning. In some aspects, method 1000 is performed by a machine learning system such as Figure 1 machine learning system 125.

[0163] At block 1005, multiple data samples corresponding to multiple data modalities are accessed. In some aspects, the multiple data modalities include at least one of the following: (i) image data, (ii) radar data, (iii) LIDAR data, or (iv) relative positioning data.

[0164] At block 1010, multiple features are generated by performing feature extraction for each respective data sample among the multiple data samples, at least partially based on the respective modality of the respective data sample. In some aspects, performing feature extraction includes, for a first data sample among the multiple data samples: determining a first modality of the first data sample from the multiple data modalities; selecting a training feature extraction model based on the first modality; and generating a first set of features by processing the first data sample using the training feature extraction model.

[0165] At block 1015, one or more attention-based models are used to fuse the multiple features.

[0166] At block 1020, a wireless communication configuration is generated based on processing the fused multiple features using a machine learning model.

[0167] In some aspects, the multiple data samples include a sequence of data samples for each respective data modality among the multiple data modalities. In some aspects, the fused multiple features include a sequence of fused features. In some aspects, the machine learning model includes a time-series-based machine learning model that processes the sequence of fused features to generate a wireless communication configuration.

[0168] In some aspects, the wireless communication configuration includes a selection of beams for performing wireless communication with one or more wireless devices. In some aspects, method 1000 further includes facilitating wireless communication with one or more wireless devices using the selected beams.

[0169] In some aspects, the machine learning model is trained using a pre-training operation. The pre-training operation may include: generating a first plurality of predicted beams based on a received power simulator and first relative angle information; and training the machine learning model based on the first plurality of predicted beams and the first relative angle information.

[0170] In some aspects, a machine learning model is refined using adaptation operations. The adaptation operations can involve: generating a second plurality of predicted beams based on a received power simulator and second relative angle information; measuring actual received power information based on the second plurality of predicted beams; and training the machine learning model based on the actual received power information and the second relative angle information. In some aspects, the adaptation operations further include: in response to determining that the actual received power information differs from the predicted received power information by more than a threshold; measuring the actual received power information of at least one additional beam; and training the machine learning model based on the actual received power information of the at least one additional beam and the second relative angle information.

[0171] In some aspects, training the machine learning model includes: generating a plurality of weights for the first plurality of predicted beams based on the received power of each predicted beam in the first plurality of predicted beams; generating a binary cross-entropy loss based on the plurality of weights; and updating one or more parameters of the machine learning model based on the binary cross-entropy loss.

[0172] Example Environment for Fusion-Based Beam Selection

[0173] In some aspects, reference Figures 1 to 10 The workflows, techniques, and methods described can be implemented on one or more devices or systems. Figure 11 An example processing system 1100 is depicted, which is configured to perform various aspects of the present disclosure, including, for example, with respect to Figures 1 to 10 the techniques and methods described. In some aspects, the processing system 1100 can train, implement, or provide a machine learning model for feature fusion, such as Figure 2 architecture 200 of Figure 3 architecture 300 of Figure 4 architecture 400 of Figure 5 architecture 500 of Figure 6 and / or Figure 7 workflow 700 of Figure 8 method 800 of Figure 9 method 900 of Figure 10 and / or method 1000 of

[0174] The processing system 1100 includes a central processing unit (CPU) 1102, which can be a multi-core CPU in some examples. Instructions executed at the CPU 1102 can be loaded, for example, from a program memory associated with the CPU 1102, or from a partition of the memory 1124.

[0175] The processing system 1100 also includes additional processing components customized for specific functions, such as a graphics processing unit (GPU) 1104, a digital signal processor (DSP) 1106, a neural processing unit (NPU) 1108, a multimedia processing unit 1110, and a wireless connection component 1112.

[0176] An NPU such as NPU 1108 is generally a dedicated circuit configured to implement the control and arithmetic logic for executing machine learning algorithms such as algorithms for processing artificial neural networks (ANNs), deep neural networks (DNNs), random forests (RFs), etc. An NPU may sometimes alternatively be referred to as a neural signal processor (NSP), a tensor processing unit (TPU), a neural network processor (NNP), an intelligent processing unit (IPU), a vision processing unit (VPU), or a graphics processing unit.

[0177] An NPU such as NPU 1108 is configured to accelerate the execution of common machine learning tasks such as image classification, machine translation, object detection, and various other prediction models. In some examples, multiple NPUs may be instantiated on a single chip (such as a system-on-chip (SoC)), while in other examples, an NPU may be part of a dedicated neural network accelerator.

[0178] An NPU may be optimized for training or inference, or in some cases be configured to balance the performance between the two. For an NPU capable of performing both training and inference, these two tasks can generally still be performed independently.

[0179] An NPU designed to accelerate training is generally configured to accelerate the optimization of new models, which is a highly computationally intensive operation involving inputting an existing dataset (usually labeled or tagged), iterating over the dataset, and then adjusting model parameters (such as weights and biases) to improve model performance. Generally speaking, optimizing based on incorrect predictions involves backpropagating through the layers of the model and determining gradients to reduce prediction errors.

[0180] An NPU designed to accelerate inference is generally configured to operate on a complete model. Such an NPU can thus be configured to input a new data slice and quickly process this new data through the already trained model to generate a model output (e.g., an inference).

[0181] In some specific implementations, NPU 1108 is part of one or more of CPU 1102, GPU 1104, and / or DSP 1106.

[0182] In some examples, the wireless connection component 1112 may include sub-components for, for example, third-generation (3G) connections, fourth-generation (4G) connections (e.g., 4G Long Term Evolution (LTE)), fifth-generation connections (e.g., 5G or New Radio (NR)), sixth-generation connections (e.g., 6G), Wi-Fi connections, Bluetooth connections, and other wireless data transmission standards. The wireless connection component 1112 is further connected to one or more antennas 1114.

[0183] The processing system 1100 may also include one or more sensor processing units 1116 associated with sensors in any way, one or more image signal processors (ISPs) 1118 associated with image sensors in any way, and / or a navigation component 1120, which may include a satellite-based positioning system component (e.g., for GPS or GLONASS) and an inertial positioning system component.

[0184] The processing system 1100 may also include one or more input and / or output devices 1122, such as a screen, a touch-sensitive surface (including a touch-sensitive display), physical buttons, speakers, microphones, etc.

[0185] In some examples, one or more of the processors in the processor of the processing system 1100 may be based on the ARM or RISC-V instruction set.

[0186] The processing system 1100 also includes a memory 1124, which represents one or more static and / or dynamic memories, such as dynamic random access memory, flash-based static memory, etc. In this example, the memory 1124 includes computer-executable components that can be executed by one or more of the aforementioned processors of the processing system 1100.

[0187] Specifically, in this example, the memory 1124 includes a feature extraction component 1124A, a fusion component 1124B, a prediction component 1124C, and a training component 1124D. Although the illustrated components (and other components not depicted) are depicted as discrete components for conceptual clarity, the illustrated components (and other components not depicted) in various aspects may be implemented jointly or separately. Figure 11 In the illustrated example, the memory 1124 also includes a set of model parameters 1124E. The model parameters 1124E generally may correspond to learnable or trainable parameters of one or more machine learning models, such as for extracting features from various modalities, fusing modality-specific features, classifying based on features, or outputting beam predictions.

[0188]

[0189] ​Although depicted as residing in memory 1124 for clarity of concepts, in some aspects, some or all of the model parameters 1124E may reside in any other suitable location.

[0190] The processing system 1100 also includes a feature extraction circuit 1126, a fusion circuit 1127, a prediction circuit 1128, and a training circuit 1129. The depicted circuits and other non-depicted circuits may be configured to perform various aspects of the techniques described herein.

[0191] In some aspects, the feature extraction component 1124A and the feature extraction circuit 1126 (which may correspond to Figure 2 feature extraction 215 and / or preprocessing 220 of Figure 3 feature extraction 308 of Figure 4 preprocessing 410 and / or feature extraction 415 of Figure 5 feature extraction 510 of Figure 6 feature extraction 610 of Figure 7 and / or a part of the machine learning model 730 of

[0192] can be used to provide modality-specific feature extraction, as discussed above. For example, the feature extraction component 1124A and the feature extraction circuit 1126 can implement the operation of one or more feature extraction blocks to extract or generate features for input data samples based on the specific modality of the samples. Figure 5 Figure 6 The fusion component 1124B and the fusion circuit 1127 (which may correspond to Figure 7 the fusion component 515 of the encoder-decoder fusion model 615 of

[0193] and / or a part of the machine learning model 730 of Figure 2 can be used to fuse modality-specific features, such as using an attention-based mechanism, as discussed above. For example, the fusion component 1124B and the fusion circuit 1127 can be used to generate fused or aggregated features based on features from independent modalities. Figure 3 Figure 4 Figure 5 Figure 6 Figure 7 ​​​​(which is part of the machine learning model 730) can be used to generate beam predictions based on unimodal or multimodal features, as discussed above. For example, the prediction component 1124C and the prediction circuit 1128 can be used to generate or select one or more beams based on the fused input features.

[0194] The training component 1124D and the training circuit 1129 can be used to pre-train, train, refine, adapt, or otherwise update the machine learning model, as discussed above. For example, the training component 1124D and the training circuit 1129 can be used to train the feature extraction component, perform pre-training of the model (e.g., at Figure 7 the box 705A), perform scenario or environment adaptation of the model (e.g., at Figure 7 the box 705B), etc.

[0195] Although depicted as separate components and circuits for clarity in Figure 11 the feature extraction circuit 1126, the fusion circuit 1127, the prediction circuit 1128, and the training circuit 1129 can be implemented jointly or separately in other processing devices of the processing system 1100, such as within the CPU 1102, GPU 1104, DSP 1106, NPU 1108, etc.

[0196] Generally speaking, the processing system 1100 and / or its components can be configured to execute the methods described herein.

[0197] It should be noted that in other aspects, such as when the processing system 1100 is a server computer, etc., elements of the processing system 1100 can be omitted. For example, in other aspects, the multimedia processing unit 1110, the wireless connection component 1112, the sensor processing unit 1116, the ISP 1118, and / or the navigation component 1120 can be omitted. In addition, elements of the processing system 1100 can be distributed among multiple devices.

[0198] Example Clauses

[0199] In addition to the various aspects described above, specific combinations of aspects are also within the scope of the present disclosure, and details of some specific combinations are as follows:

[0200] Clause 1: A method, the method comprising: accessing a plurality of data samples corresponding to a plurality of data modalities; generating a plurality of features by performing feature extraction for each respective data sample of the plurality of data samples at least partially based on the respective modality of the respective data sample; using one or more attention-based models to fuse the plurality of features; and generating a wireless communication configuration based on processing the fused plurality of features using a machine learning model.

[0201] Clause 2: The method according to Clause 1, wherein the plurality of data modalities includes at least one of the following: (i) image data, (ii) radio detection and ranging (radar) data, (iii) light detection and ranging (LIDAR) data, or (iv) relative positioning data.

[0202] Clause 3: The method according to Clause 1 or 2, wherein performing the feature extraction includes, for a first data sample among the plurality of data samples: determining a first modality of the first data sample from the plurality of data modalities; selecting a training feature extraction model based on the first modality; and generating a first set of features by processing the first data sample using the training feature extraction model.

[0203] Clause 4: The method according to any one of Clauses 1 to 3, wherein: the plurality of data samples includes a sequence of data samples for each respective data modality among the plurality of data modalities; the plurality of fused features includes a sequence of fused features; and the machine learning model includes a time-series based machine learning model that processes the sequence of fused features to generate the wireless communication configuration.

[0204] Clause 5: The method according to any one of Clauses 1 to 4, wherein the wireless communication configuration includes a selection of beams for performing wireless communication with one or more wireless devices.

[0205] Clause 6: The method according to Clause 5, the method further comprising: using the selected beams to facilitate wireless communication with the one or more wireless devices.

[0206] Clause 7: The method according to any one of Clauses 1 to 6, wherein the machine learning model is trained using a pre-training operation, the pre-training operation including: generating a first plurality of predicted beams based on a received power simulator and first relative angle information; and training the machine learning model based on the first plurality of predicted beams and the first relative angle information.

[0207] Clause 8: The method according to Clause 7, wherein the machine learning model is refined using an adaptive operation, the adaptive operation including: generating a second plurality of predicted beams based on the received power simulator and second relative angle information; measuring actual received power information based on the second plurality of predicted beams; and training the machine learning model based on the actual received power information and the second relative angle information.

[0208] Clause 9: The method according to Clause 8, wherein the adaptive operation further comprises: measuring the actual received power information of at least one additional beam in response to determining that the difference between the actual received power information and the predicted received power information exceeds a threshold; and training the machine learning model based on the actual received power information of the at least one additional beam and the second relative angle information.

[0209] Clause 10: The method according to any one of Clauses 7 to 9, wherein training the machine learning model comprises: generating a plurality of weights for the first plurality of predicted beams based on the received power of each predicted beam in the first plurality of predicted beams; generating a binary cross-entropy loss based on the plurality of weights; and updating one or more parameters of the machine learning model based on the binary cross-entropy loss.

[0210] Clause 11: A processing system, the processing system comprising: a memory including computer-executable instructions; and one or more processors configured to execute the computer-executable instructions and cause the processing system to perform the method according to any one of Clauses 1 to 10.

[0211] Clause 12: A processing system, the processing system comprising: components for performing the method according to any one of Clauses 1 to 10.

[0212] Clause 13: A non-transitory computer-readable medium, the non-transitory computer-readable medium comprising computer-executable instructions that, when executed by one or more processors of a processing system, cause the processing system to perform the method according to any one of Clauses 1 to 10.

[0213] Clause 14: A computer program product embodied on a computer-readable storage medium, the computer program product comprising: code for performing the method according to any one of Clauses 1 to 10.

[0214] Additional Notes

[0215] The foregoing description is provided to enable any person skilled in the art to practice the various aspects described herein. The examples discussed herein do not limit the scope, applicability, or aspects set forth in the claims. Various modifications to these aspects will be apparent to those skilled in the art, and the general principles defined herein may be applied to other aspects. For example, the functions and arrangements of the elements discussed may be changed without departing from the scope of the disclosure. Various examples may omit, substitute, or add various procedures or components as appropriate. For example, the methods described may be performed in an order different from that described, and various steps may be added, omitted, or combined. In addition, features described with respect to some examples may be combined in some other examples. For example, any number of the aspects set forth herein may be used to implement a device or practice a method. Additionally, the scope of the disclosure is intended to cover such devices or methods that practice using other structures, functionality, or a combination of structure and functionality that supplement or replace the various aspects of the disclosure set forth herein. It should be understood that any aspect of the disclosure herein may be embodied by one or more elements of the invention.

[0216] As used herein, the term "exemplary" means "serving as an example, instance, or illustration." Any aspect described herein as "exemplary" is not necessarily to be construed as preferred or advantageous over other aspects.

[0217] As used herein, the phrase "at least one of" in reference to a list of items refers to any combination of those items (including a single member). For example, "at least one of a, b, or c" is intended to cover a, b, c, a - b, a - c, b - c, and a - b - c, as well as any combination with multiple of the same element (e.g., a - a, a - a - a, a - a - b, a - a - c, a - b - b, a - c - c, b - b, b - b - b, b - b - c, c - c, and c - c - c, or any other ordering of a, b, and c).

[0218] As used herein, the term "determine" encompasses a variety of actions. For example, "determine" may include calculating, computing, processing, deriving, investigating, looking up (e.g., looking up in a table, database, or another data structure), ascertaining, and the like. Additionally, "determine" may include receiving (e.g., receiving information), accessing (e.g., accessing data in a memory), and the like. Further, "determine" may include parsing, selecting, picking, establishing, and the like.

[0219] The methods disclosed herein include one or more steps or acts for implementing the methods. The method steps and / or acts may be interchanged with one another without departing from the scope of the claims. In other words, unless a specific order of the steps or acts is specified, the order and / or use of the specific steps and / or acts may be modified without departing from the scope of the claims. Additionally, the various operations of the methods described above may be performed by any suitable component capable of performing the corresponding functions. The component may include various hardware and / or software components and / or modules, including but not limited to circuits, application specific integrated circuits (ASICs), or processors. Generally, where there are operations illustrated in the figures, those operations may have corresponding components plus function components with similar numbers.

[0220] The following claims are not intended to be limited to the aspects shown herein but should be accorded the full scope consistent with the claim language. In the claims, the recitation of an element in the singular is not intended to mean "one and only one" unless specifically stated, but rather "one or more." The term "some," unless specifically stated otherwise, means one or more. Any claim element is not to be construed under the provisions of 35 U.S.C. § 112(f) unless the element is expressly recited using the phrase "means for" or, in the case of a method claim, the phrase "step for." All structural and functional equivalents of the elements of the various aspects described throughout this disclosure that are known or later become known to those of ordinary skill in the art are expressly incorporated herein by reference and are intended to be covered by the claims. Additionally, nothing disclosed herein is intended to be dedicated to the public, whether or not such disclosure is explicitly recited in the claims.

Claims

1. A processor-implemented method, the processor-implemented method comprises: accessing a plurality of data samples corresponding to a plurality of data modalities; generating a plurality of features by performing feature extraction for each respective data sample among the plurality of data samples, at least partially based on the respective modality of the respective data sample; using one or more attention-based models to fuse the plurality of features; and generating a wireless communication configuration based on processing the fused plurality of features using a machine learning model.

2. The processor-implemented method according to claim 1, wherein the plurality of data modalities comprises at least one of the following: (i) image data, (ii) radio detection and ranging (radar) data, (iii) light detection and ranging (LIDAR) data, or (iv) relative positioning data.

3. The processor-implemented method according to claim 1, wherein performing the feature extraction comprises, for a first data sample among the plurality of data samples: determining a first modality of the first data sample from the plurality of data modalities; selecting a training feature extraction model based on the first modality; and generating a first set of features by processing the first data sample using the training feature extraction model.

4. The processor-implemented method according to claim 1, wherein: the plurality of data samples comprises a sequence of data samples for each respective data modality among the plurality of data modalities; the fused plurality of features comprises a sequence of fused features; and the machine learning model comprises a time-series-based machine learning model that processes the sequence of fused features to generate the wireless communication configuration.

5. The processor-implemented method according to claim 1, wherein the wireless communication configuration comprises a selection of beams for performing wireless communication with one or more wireless devices.

6. The processor-implemented method according to claim 5, the processor-implemented method further comprises: using the selected beams to facilitate wireless communication with the one or more wireless devices.

7. The processor-implemented method according to claim 1, wherein the machine learning model is trained using a pre-training operation, the pre-training operation comprises: generating a first plurality of predicted beams based on a received power simulator and first relative angle information; and training the machine learning model based on the first plurality of predicted beams and the first relative angle information.

8. The processor-implemented method according to claim 7, wherein the machine learning model is refined using an adaptation operation, the adaptation operation comprises: generating a second plurality of predicted beams based on the received power simulator and second relative angle information; measuring actual received power information based on the second plurality of predicted beams; and training the machine learning model based on the actual received power information and the second relative angle information.

9. The processor-implemented method according to claim 8, wherein the adaptation operation further comprises: In response to determining that the actual received power information differs from the predicted received power information by more than a threshold, measure the actual received power information of at least one additional beam; and train the machine learning model based on the actual received power information of the at least one additional beam and the second relative angle information.

10. The processor-implemented method according to claim 7, wherein training the machine learning model comprises: generating a plurality of weights for the first plurality of predicted beams based on the received power of each predicted beam in the first plurality of predicted beams; generating a binary cross-entropy loss based on the plurality of weights; and updating one or more parameters of the machine learning model based on the binary cross-entropy loss.

11. A processing system, the processing system comprising: a memory including computer-executable instructions; and one or more processors configured to execute the computer-executable instructions to cause the processing system to: access a plurality of data samples corresponding to a plurality of data modalities; perform feature extraction at least in part based on the respective modalities of each corresponding data sample among the plurality of data samples to generate a plurality of features for the plurality of data samples; fuse the plurality of features using one or more attention-based models; and generate a wireless communication configuration based on processing the fused plurality of features using a machine learning model.

12. The processing system according to claim 11, wherein the plurality of data modalities includes at least one of the following: (i) image data, (ii) radio detection and ranging (radar) data, (iii) light detection and ranging (LIDAR) data, or (iv) relative positioning data.

13. The processing system according to claim 11, wherein in order to perform the feature extraction for a first data sample among the plurality of data samples, the one or more processors are configured to execute the computer-executable instructions to cause the processing system to: determine a first modality of the first data sample from the plurality of data modalities; select a training feature extraction model based on the first modality; and generate a first set of features by processing the first data sample using the training feature extraction model.

14. The processing system according to claim 11, wherein: the plurality of data samples includes a sequence of data samples for each corresponding data modality among the plurality of data modalities; the fused plurality of features includes a sequence of fused features; and the machine learning model includes a time-series based machine learning model that processes the sequence of fused features to generate the wireless communication configuration.

15. The processing system according to claim 11, wherein the wireless communication configuration includes a selection of beams for performing wireless communication with one or more wireless devices.

16. The processing system according to claim 15, wherein the one or more processors are further configured to execute the computer-executable instructions to cause the processing system to facilitate wireless communication with the one or more wireless devices using the selected beam.

17. The processing system according to claim 11, wherein the machine learning model is trained using a pre-training operation, and wherein, to perform the pre-training operation, the one or more processors are configured to execute the computer-executable instructions to cause the processing system to: generate a first plurality of predicted beams based on a received power simulator and first relative angle information; and train the machine learning model based on the first plurality of predicted beams and the first relative angle information.

18. The processing system according to claim 17, wherein the machine learning model is refined using an adaptation operation, and wherein, to perform the adaptation operation, the one or more processors are configured to execute the computer-executable instructions to cause the processing system to: generate a second plurality of predicted beams based on the received power simulator and second relative angle information; measure actual received power information based on the second plurality of predicted beams; and train the machine learning model based on the actual received power information and the second relative angle information.

19. The processing system according to claim 18, wherein, to perform the adaptation operation, the one or more processors are further configured to execute the computer-executable instructions to cause the processing system to: measure the actual received power information of at least one additional beam in response to determining that the actual received power information differs from predicted received power information by more than a threshold; and train the machine learning model based on the actual received power information of the at least one additional beam and the second relative angle information.

20. The processing system according to claim 17, wherein, to train the machine learning model, the one or more processors are configured to execute the computer-executable instructions to cause the processing system to: generate a plurality of weights for the first plurality of predicted beams based on the received power of each predicted beam in the first plurality of predicted beams; generate a binary cross-entropy loss based on the plurality of weights; and update one or more parameters of the machine learning model based on the binary cross-entropy loss.

21. A non-transitory computer-readable medium comprising computer-executable instructions that, when executed by one or more processors of a processing system, cause the processing system to: access a plurality of data samples corresponding to a plurality of data modalities; perform feature extraction at least in part based on the respective modalities of each corresponding data sample in the plurality of data samples to generate a plurality of features for the plurality of data samples; fuse the plurality of features using one or more attention-based models; and generate a wireless communication configuration based on processing the fused plurality of features using a machine learning model.

22. The non-transitory computer-readable medium according to claim 21, wherein the plurality of data modalities includes at least one of the following: (i) image data, (ii) radio detection and ranging (radar) data, (iii) light detection and ranging (LIDAR) data, or (iv) relative positioning data.

23. The non-transitory computer-readable medium according to claim 21, wherein in order to perform the feature extraction for a first data sample among the plurality of data samples, the one or more processors are configured to execute the computer-executable instructions to cause the processing system to: determine a first modality of the first data sample from the plurality of data modalities; select a training feature extraction model based on the first modality; and generate a first set of features by processing the first data sample using the training feature extraction model.

24. The non-transitory computer-readable medium according to claim 21, wherein: the plurality of data samples includes a sequence of data samples for each respective data modality among the plurality of data modalities; the plurality of fused features includes a sequence of fused features; and the machine learning model includes a time-series-based machine learning model that processes the sequence of fused features to generate the wireless communication configuration.

25. The non-transitory computer-readable medium according to claim 21, wherein the wireless communication configuration includes a selection of a beam for performing wireless communication with one or more wireless devices.

26. The non-transitory computer-readable medium according to claim 25, wherein the computer-executable instructions, when executed by the one or more processors of the processing system, further cause the processing system to facilitate wireless communication with the one or more wireless devices using the selected beam.

27. The non-transitory computer-readable medium according to claim 21, wherein the machine learning model is trained using a pre-training operation, and wherein, in order to perform the pre-training operation, the one or more processors are configured to execute the computer-executable instructions to cause the processing system to: generate a first plurality of predicted beams based on a received power simulator and first relative angle information; and train the machine learning model based on the first plurality of predicted beams and the first relative angle information.

28. The non-transitory computer-readable medium according to claim 27, wherein the machine learning model is refined using an adaptive operation, and wherein, in order to perform the adaptive operation, the one or more processors are configured to execute the computer-executable instructions to cause the processing system to: generate a second plurality of predicted beams based on the received power simulator and second relative angle information; measure actual received power information based on the second plurality of predicted beams; and train the machine learning model based on the actual received power information and the second relative angle information.

29. The non-transitory computer-readable medium according to claim 28, wherein, to perform the adaptive operation, the one or more processors are further configured to execute the computer-executable instructions to cause the processing system to: Measure the actual received power information of at least one additional beam in response to determining that the difference between the actual received power information and the predicted received power information exceeds a threshold; and Train the machine learning model based on the actual received power information of the at least one additional beam and the second relative angle information.

30. A processing system, the processing system comprising: means for accessing a plurality of data samples corresponding to a plurality of data modalities; means for generating a plurality of features by performing feature extraction for each respective data sample of the plurality of data samples based at least in part on the respective modality of the respective data sample; means for fusing the plurality of features using one or more attention-based models; and means for generating a wireless communication configuration based on processing the fused plurality of features using a machine learning model.

Citation Information

Cited By

  • Method and device for determining optimal beam and electronic equipment

    CN120601931A