Vision-aided positioning model management and training data acquisition
By integrating vision-aided techniques for managing RF-based AI/ML models, the approach addresses limitations in indoor positioning accuracy and training data acquisition, improving the efficiency and effectiveness of AI/ML-based localization systems.
Patent Information
- Application Number
- PCT/US2025/027608
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-06-26
- Filing Date
- 2025-05-02
- Publication Date
- 2026-01-02
AI Technical Summary
Existing wireless communication systems, particularly 5G NR, face challenges in improving positioning accuracy and efficiency, especially in indoor environments where vision-based positioning is limited, and there is a need for effective management of RF-based AI/ML models for enhanced localization.
The integration of vision-aided techniques for managing RF-based AI/ML models, utilizing opportunistic use of handheld cameras to acquire ground truth locations, and incorporating temporal and incremental memory matching schemes to enhance training data labeling and model management.
This approach improves the accuracy and efficiency of AI/ML-based positioning systems by leveraging vision-based components to provide highly accurate localization and optimize training data acquisition, thereby enhancing the overall performance of RF-based AI/ML models.
Smart Images

Figure US2025027608_02012026_PF_FP_ABST
Abstract
Description
VISION-AIDED POSITIONING MODEL MANAGEMENT AND TRAINING DATA ACQUISITIONCROSS-REFERENCE TO RELATED APPLICATION
[0001] This application claims the benefit of Greece Application Serial No. 20240100466 entitled “VISION-AIDED POSITIONING MODEL MANAGEMENT AND TRAINING DATA ACQUISITION” and filed on June 26, 2024, which is expressly incorporated by reference herein in its entirety.TECHNICAL FIELD
[0002] The present disclosure relates generally to positioning systems, and more particularly, to positioning systems involving vision-aided positioning.INTRODUCTION
[0003] Wireless communication systems are widely deployed to provide various telecommunication services such as telephony, video, data, messaging, and broadcasts. Typical wireless communication systems may employ multiple-access technologies capable of supporting communication with multiple users by sharing available system resources. Examples of such multiple-access technologies include code division multiple access (CDMA) systems, time division multiple access (TDMA) systems, frequency division multiple access (FDMA) systems, orthogonal frequency division multiple access (OFDMA) systems, single-carrier frequency division multiple access (SC-FDMA) systems, and time division synchronous code division multiple access (TD-SCDMA) systems.
[0004] These multiple access technologies have been adopted in various telecommunication standards to provide a common protocol that enables different wireless devices to communicate on a municipal, national, regional, and even global level. An example telecommunication standard is 5G New Radio (NR). 5G NR is part of a continuous mobile broadband evolution promulgated by Third Generation Partnership Project (3 GPP) to meet new requirements associated with latency, reliability, security, scalability (e.g., with Internet of Things (IoT)), and other requirements. 5G NR includes services associated with enhanced mobile broadband (eMBB), massive machine type communications (mMTC), and ultra-reliable low latencycommunications (URLLC). Some aspects of 5G NR may be based on the 4G Long Term Evolution (LTE) standard. There exists a need for further improvements in 5G NR technology. These improvements may also be applicable to other multi-access technologies and the telecommunication standards that employ these technologies.BRIEF SUMMARY
[0005] The following presents a simplified summary of one or more aspects in order to provide a basic understanding of such aspects. This summary is not an extensive overview of all contemplated aspects. This summary neither identifies key or critical elements of all aspects nor delineates the scope of any or all aspects. Its sole purpose is to present some concepts of one or more aspects in a simplified form as a prelude to the more detailed description that is presented later.
[0006] In an aspect of the disclosure, a method, a computer-readable medium, and an apparatus are provided. The apparatus performs positioningfora user equipment (UE) based on a machine learning (ML) model. The apparatus detects that at least one condition associated with the UE or an environment surrounding of the UE changes or meets a threshold. The apparatus estimates a firstposition oftheUEbased on visual information associated with the UE. The apparatus selects a set of radio frequency (RF) measurements for the UE based on the estimated first position of the UE. The apparatus updates the ML model based on the first position of the UE and the set of RF measurements.
[0007] To the accomplishment of the foregoing and related ends, the one or more aspects may include the features hereinafter fully described and particularly pointed out in the claims. The following description and the drawings set forth in detail certain illustrative features of the one or more aspects. These features are indicative, however, of but a few of the various ways in which the principles of various aspects may be employed.BRIEF DESCRIPTION OF THE DRAWINGS
[0008] FIG. 1 is a diagram illustrating an example of a wireless communications system and an access network.
[0009] FIG. 2A is a diagram illustrating an example of a first frame, in accordance with various aspects of the present disclosure.
[0010] FIG. 2B is a diagram illustrating an example of downlink (DL) channels within a subframe, in accordance with various aspects of the present disclosure.
[0011] FIG. 2C is a diagram illustrating an example of a second frame, in accordance with various aspects of the present disclosure.
[0012] FIG. 2D is a diagram illustrating an example of uplink (UL) channels within a subframe, in accordance with various aspects of the present disclosure.
[0013] FIG. 3 is a diagram illustrating an example of a base station and user equipment (UE) in an access network.
[0014] FIG. 4 is a diagram illustrating an example of a UE positioning based on reference signal measurements.
[0015] FIG. 5 is a diagram illustrating an example of visual-based positioning in accordance with various aspects of the present disclosure.
[0016] FIG. 6 is a diagram illustrating an example scenario of using a wireless infrastructure and camera(s) for locating targets in accordance with various aspects of the present disclosure.
[0017] FIG. 7 is a diagram illustrating an example of artificial intelligence (Al) or machine learning (ML) (AI / ML) positioning models specializing on different parts / areas of a space in accordance with various aspects of the present disclosure.
[0018] FIG. 8 is a communication flow illustrating an example of a server-based AI / ML model selection in accordance with various aspects of the present disclosure.
[0019] FIG. 9 is a communication flow illustrating an example of a UE-based AI / ML model selection in accordance with various aspects of the present disclosure.
[0020] FIG. 10 is a diagram illustrating an example of initiating an AI / ML model selection, training, and / or fine-tuning based on triggering condition(s) in accordance with various aspects of the present disclosure.
[0021] FIG. 11 is a communication flow illustrating an example of a training data acquisition framework in accordance with various aspects of the present disclosure.
[0022] FIG. 12 is a diagram illustrating an example of a memory -based matching of ground truth locations with radio frequency (RF) measurements in accordance with various aspects of the present disclosure.
[0023] FIG. 13 is a diagram illustrating an example of an association matrix that is used for the memory-based matching in accordance with various aspects of the present disclosure.
[0024] FIG. 14 is a communication flow illustrating an example of an incremental memory- based matching scheme in accordance with various aspects of the present disclosure.
[0025] FIG. 15 is a flowchart of a method of wireless communication.
[0026] FIG. 16 is a flowchart of a method of wireless communication.
[0027] FIG. 17 is a diagram illustrating an example of a hardware implementation for an example network entity.DETAILED DESCRIPTION
[0028] Aspects presented herein may improve the overall performance of artificial intelligence (Al) or machine learning (ML) (AI / ML) model managements. Aspects presented herein provide various vision-based techniques for managing RF-based AI / ML models, where vision is used to address the key model management functions. While vision-based positioning (if available) may provide highly accurate positioning, aspects presented herein may leverage existing indoor deployments including the opportunistic usage of handheld cameras to acquire highly accurate ground truth (GT) locations for the training data. For example, a store / warehousemay have several targets distributed around the space and each target is associated with a radio frequency (RF) device / UE which is connected to an RF infrastructure. An RF- based AI / ML positioning engine fuses RF measurements and vision input using AI / ML-based positioning models to localize the targets. As the availability of visionbased components for directpositioningmay be limited due to various factors, aspects presented herein propose vision-based techniques for managing RF-based AI / ML models where vision is used to address key model management functions for AI / ML positioning models. In addition to the AI / ML-based positioning engine driven by RF input and information, the vision based component may be used opportunistically and on-demand for labeling the training data for AI / ML models. Other aspects include temporal and incremental memory matching schemes.
[0029] The detailed description set forth below in connection with the drawings describes various configurations and does not represent the only configurations in which the concepts described herein may be practiced. The detailed description includes specific details for the purpose of providing a thorough understanding of various concepts. However, these concepts may be practiced without these specific details. In someinstances, well known structures and components are shown in block diagram form in order to avoid obscuring such concepts.
[0030] Several aspects of telecommunication systems are presented with ref erenceto various apparatus and methods. These apparatus and methods are described in the following detailed description and illustrated in the accompanying drawings by various blocks, components, circuits, processes, algorithms, etc. (collectively referred to as “elements”). These elements may be implemented using electronic hardware, computer software, or any combination thereof. Whether such elements are implemented as hardware or software depends upon the particular application and design constraints imposed on the overall system.
[0031] By way of example, an element, or any portion of an element, or any combination of elements may be implemented as a “processing system” that includes one or more processors. When multiple processors are implemented, the multiple processors may perform the functions individually or in combination. Examplesof processors include microprocessors, microcontrollers, graphics processing units (GPUs), central processing units (CPUs), application processors, digital signal processors (DSPs), reduced instruction set computing (RISC) processors, systems on a chip (SoC), baseband processors, field programmable gate arrays (FPGAs), programmable logic devices (PLDs), state machines, gated logic, discrete hardware circuits, and other suitable hardware configured to perform the various functionality described throughout this disclosure. One or more processors in the processing system may execute software. Software, whether referred to as software, firmware, middleware, microcode, hardware description language, or otherwise, shall be construed broadly to mean instructions, instruction sets, code, code segments, program code, programs, subprograms, software components, applications, software applications, software packages, routines, subroutines, objects, executables, threads of execution, procedures, functions, or any combination thereof.
[0032] Accordingly, in one or more example aspects, implementations, and / or use cases, the functions described may be implemented in hardware, software, or any combination thereof. If implemented in software, the functions may be stored on or encoded as one or more instructions or code on a computer-readable medium. Computer-readable media includes computer storage media. Storage media may be any available media that can be accessed by a computer. By way of example, such computer-readablemedia can include a random-access memory (RAM), a read-only memory (ROM), an electrically erasable programmable ROM (EEPROM), optical disk storage, magnetic disk storage, other magnetic storage devices, combinations of the types of computer- readable media, or any other medium that can be used to store computer executable code in the form of instructions or data structures that can be accessed by a computer.
[0033] While aspects, implementations, and / or use cases are describedin this application by illustration to some examples, additional or different aspects, implementations and / or use cases may come about in many different arrangements and scenarios. Aspects, implementations, and / oruse cases described herein may be implemented across many differingplatform types, devices, systems, shapes, sizes, and packaging arrangements. For example, aspects, implementations, and / or use cases may come about via integrated chip implementations and other non-module-component based devices (e.g., end-user devices, vehicles, communication devices, computing devices, industrial equipment, retail / purchasing devices, medical devices, artificial intelligence (Al)-enabled devices, etc.). While some examples may or may not be specifically directed to use cases or applications, a wide assortment of applicability of described examples may occur. Aspects, implementations, and / oruse cases may range a spectrum from chip-level or modular components to non-modular, non-chip- level implementations and further to aggregate, distributed, or original equipment manufacturer (OEM) devices or systems incorporating one or more techniques herein. In some practical settings, devices incorporating described aspects and features may also include additional components and features for implementation and practice of claimed and described aspect. For example, transmission and reception of wireless signals necessarily includes a number of components for analog and digital purposes (e.g., hardware components including antenna, RF-chains, power amplifiers, modulators, buffer, processor(s), interleaver, adders / summers, etc.). Techniques described herein may be practiced in a wide variety of devices, chip-level components, systems, distributed arrangements, aggregated or disaggregated components, end-user devices, etc. of varying sizes, shapes, and constitution.
[0034] Deployment of communication systems, such as 5GNR systems, may be arranged in multiple manners with various components or constituent parts. In a 5G NR system, or network, a network node, a network entity, a mobility element of a network, a radio access network (RAN) node, a core network node, a network element, or a networkequipment, such as a base station (BS), or one or more units (or one or more components) performing base station functionality, may be implemented in an aggregated or disaggregated architecture. For example, a BS (such as a Node B (NB), evolved NB (eNB), NR BS, 5GNB, access point (AP), a transmission reception point (TRP), or a cell, etc.) may be implemented as an aggregated base station (also known as a standalone BS or a monolithic BS) or a disaggregated base station.
[0035] An aggregated base station may be configured to utilize a radio protocol stack that is physically or logically integrated within a single RAN node. A disaggregated base station may be configured to utilize a protocol stack that is physically or logically distributed among two or more units (such as one or more central or centralized units (CUs), one or more distributed units (DUs), or one or more radio units (RUs)). In some aspects, a CU may be implemented within a RAN node, and one or more DUs may be co-located with the CU, or alternatively, may be geographically or virtually distributed throughout one or multiple other RAN nodes. The DUs may be implemented to communicate with one or more RUs. Each of the CU, DU and RU can be implemented as virtual units, i.e., a virtual central unit (VCU), a virtual distributed unit (VDU), or a virtual radio unit (VRU).
[0036] Base station operation or network design may consider aggregation characteristics of base station functionality. For example, disaggregated base stations may be utilized in an integrated access backhaul (IAB) network, an open radio access network (O- RAN (such as the network configuration sponsored by the O-RAN Alliance)), or a virtualized radio access network (vRAN, also known as a cloud radio access network (C-RAN)). Disaggregation may include distributing functionality across two or more units at various physical locations, as well as distributing functionality for at least one unit virtually, which can enable flexibility in network design. The various units of the disaggregated base station, or disaggregated RAN architecture, can be configured for wired or wireless communication with at least one other unit.
[0037] FIG. 1 is a diagram 100 illustrating an example of a wireless communications system and an access network. The illustrated wireless communications system includes a disaggregated base station architecture. The disaggregated base station architecture may include one or more CUs 110 that can communicate directly with a core network 120 via a backhaul link, or indirectly with the core network 120 through one or more disaggregated base station units (such as a Near-Real Time (Near-RT) RANIntelligent Controller (RIC) 125 via an E2 link, or a Non-Real Time (Non-RT) RIC 115 associated with a Service Management and Orchestration (SMO) Framework 105, or both). A CU 110 may communicate with one or more DUs 130 via respective midhaul links, such as an Fl interface. The DUs 130 may communicate with one or more RUs 140 via respective fronthaul links. The RUs 140 may communicate with respective UEs 104 via one or more radio frequency (RF) access links. In some implementations, the UE 104 may be simultaneously served by multiple RUs 140.
[0038] Each of the units, i.e., the CUs 110, the DUs 130, the RUs 140, as well as the Near- RT RICs 125, the Non-RT RICs 115, and the SMO Framework 105, may include one or more interfaces or be coupled to one or more interfaces configured to receive or to transmit signals, data, or information (collectively, signals) via a wired or wireless transmission medium. Each of the units, or an associated processor or controller providing instructions to the communication interfaces of the units, can be configured to communicate with one or more of the other units via the transmission medium. For example, the units can include a wired interface configured to receive or to transmit signals over a wired transmission medium to one or more of the other units. Additionally, the units can include a wireless interface, which may include a receiver, a transmitter, or a transceiver (such as an RF transceiver), configured to receive or to transmit signals, or both, over a wireless transmission medium to one or more of the other units.
[0039] In some aspects, the CU 110 may host one or more higher layer control functions. Such control functions can include radio resource control (RRC), packet data convergence protocol (PDCP), service data adaptation protocol (SDAP), or the like. Each control function can be implemented with an interface configured to communicate signals with other control functions hosted by the CU 110. The CU 110 may be configured to handle user plane functionality (i.e., Central Unit - User Plane (CU-UP)), control plane functionality (i.e., Central Unit - Control Plane (CU-CP)), or a combination thereof. In some implementations, the CU 110 can be logically split into one or more CU-UP units and one or more CU-CP units. The CU-UP unit can communicate bidirectionally with the CU-CP unit via an interface, such as an El interface when implemented in an O-RAN configuration. The CU 110 can be implemented to communicate with the DU 130, as necessary, for network control and signaling.
[0040] The DU 130 may correspond to a logical unit that includes one or more base station functions to control the operation of one or more RUs 140. In some aspects, the DU 130 may host one or more of a radio link control (RLC) layer, a medium access control (MAC) layer, and one or more high physical (PHY) layers (such as modules for forward error correction (FEC) encoding and decoding, scrambling, modulation, demodulation, or the like) depending, at least in part, on a functional split, such as those defined by 3 GPP. In some aspects, the DU 130 may further host one or more low PHY layers. Each layer (or module) can be implemented with an interface configured to communicate signals with other layers (and modules) hosted by the DU 130, or with the control functions hosted by the CU 110.
[0041] Lower-layer functionality can be implemented by one or more RUs 140. In some deployments, an RU 140, controlled by a DU 130, may correspond to a logical node that hosts RF processing functions, or low-PHY layer functions (such as performing fast Fourier transform (FFT), inverse FFT (iFFT), digital beamforming, physical random access channel (PRACH) extraction and filtering, or the like), or both, based at least in part on the functional split, such as a lower layer functional split. In such an architecture, the RU(s) 140 can be implemented to handle over the air (OTA) communication with one or more UEs 104. In some implementations, real-time and non-real-time aspects of control and user plane communication with the RU(s) 140 can be controlled by the corresponding DU 130. In some scenarios, this configuration can enable the DU(s) 130 and the CU 110 to be implemented in a cloud-based RAN architecture, such as a vRAN architecture.
[0042] The SMO Framework 105 may be configured to support RAN deployment and provisioning of non-virtualizedandvirtualizednetwork elements. Fornon-virtualized network elements, the SMO Framework 105 may be configured to support the deployment of dedicated physical resources for RAN coverage requirements that may be managed via an operations and maintenance interface (such as an 01 interface). For virtualized network elements, the SMO Framework 105 may be configured to interact with a cloud computing platform (such as an open cloud (O-Cloud) 190) to perform network element life cycle management (such as to instantiate virtualized network elements) via a cloud computing platform interface (such as an 02 interface). Such virtualized network elements can include, but are not limited to, CUs 110, DUs 130, RUs 140 andNear-RTRICs 125. In some implementations, the SMO Framework105 can communicate with a hardware aspect of a 4G RAN, such as an open eNB (O- eNB) 111, via an 01 interface. Additionally, in some implementations, the SMO Framework 105 can communicate directly with one or more RUs 140 via an 01 interface. The SMO Framework 105 also may include aNon-RT RIC 115 configured to support functionality of the SMO Framework 105.
[0043] The Non-RT RIC 115 may be configured to include a logical function that enables non-real-time control and optimization of RAN elements and resources, artificial intelligence (Al) / machine learning (ML) (AI / ML) workflows including model training and updates, or policy-based guidance of applications / features in the Near- RT RIC 125. The Non-RT RIC 115 may be coupled to or communicate with (such as via an Al interface) the Near-RT RIC 125. The Near-RT RIC 125 may be configured to include a logical function that enables near-real-time control and optimization of RAN elements and resources via dataset collection and actions over an interface (such as via an E2 interface) connecting one or more CUs 110, one or more DUs 130, or both, as well as an O-eNB, with the Near-RT RIC 125.
[0044] In some implementations, to generate AI / ML models to be deployed in the Near-RT RIC 125, the Non-RT RIC 115 may receive parameters or external enrichment information from external servers. Such information may be utilized by the Near-RT RIC 125 and may be received at the SMO Framework 105 or the Non-RT RIC 115 from non-network data sources or from network functions. In some examples, the Non-RT RIC 115 or the Near-RT RIC 125 may be configured to tune RANbehavior or performance. For example, the Non-RT RIC 115 may monitor long-term trends and patterns for performanceand employ AI / ML models to perform corrective actions through the SMO Framework 105 (such as reconfiguration via 01) or via creation of RAN management policies (such as Al policies).
[0045] At least one of the CU 110, the DU 130, and the RU 140 maybe referred to as a base station 102. Accordingly, a base station 102 may include one or more of the CU 110, the DU 130, and the RU 140 (each component indicated with dotted lines to signify that each component may or may not be included in the base station 102). The base station 102 provides an access point to the core network 120 for a UE 104. The base station 102 may include macrocells (high power cellular base station) and / or small cells (low power cellular base station). The small cells include femtocells, picocells, and microcells. Anetwork thatincludes both small cell and macrocells may be knownas a heterogeneous network. A heterogeneous network may also include Home Evolved Node Bs (eNBs) (HeNBs), which may provide service to a restricted group known as a closed subscriber group (CSG). The communication links between the RUs 140 and the UEs 104 may include uplink (UL) (also referred to as reverse link) transmissions from a UE 104 to an RU 140 and / or downlink (DL) (also referred to as forward link) transmissions from an RU 140 to a UE 104. The communication links may use multiple-input and multiple-output (MIMO) antenna technology, including spatial multiplexing, beamforming, and / or transmit diversity. The communication links may be through one or more carriers. The base station 102 / UEs 104 may use spectrum up to FMHz (e.g., 5, 10, 15, 20, 100, 400, etc. MHz) bandwidth per carrier allocated in a carrier aggregation of up to a total of Ex MHz (x component carriers) used for transmission in each direction. The carriers may or may not be adjacent to each other. Allocation of carriers may be asymmetric with respecttoDL andUL (e.g., more or fewer carriers may be allocated for DL than for UL). The component carriers may include a primary component carrier and one or more secondary component carriers. A primary component carrier may be referred to as a primary cell (PCell) and a secondary component carrier may be referred to as a secondary cell (SCell).
[0046] Certain UEs 104 may communicate with each other using device-to-device (D2D) communication link 158. The D2D communication link 158 may use the DL / UL wireless wide area network (WWAN) spectrum. The D2D communication link 158 may use one or more sidelink channels, such as a physical sidelink broadcast channel (PSBCH), a physical sidelink discovery channel (PSDCH), a physical sidelink shared channel (PSSCH), and a physical sidelink control channel (PSCCH). D2D communication may be through a variety of wireless D2D communications systems, such as for example, Bluetooth™ (Bluetooth is a trademark of the Bluetooth Special Interest Group (SIG)), Wi-Fi™ (Wi-Fi is a trademark of the Wi-Fi Alliance) based on the Institute of Electrical and Electronics Engineers (IEEE) 802.11 standard, LTE, or NR.
[0047] The wireless communications system may further include a Wi-Fi AP 150 in communication with UEs 104 (also referred to as Wi-Fi stations (STAs)) via communication link 154, e.g., in a 5 GHz unlicensed frequency spectrum orthe like. When communicating in an unlicensed frequency spectrum, the UEs 104 / AP 150may perform a clear channel assessment (CCA) prior to communicating in order to determine whether the channel is available.
[0048] The electromagnetic spectrum is often subdivided, based on frequency / wavelength, into various classes, bands, channels, etc. In 5GNR, two initial operating bands have been identified as frequency range designations FR1 (410 MHz - 7.125 GHz) and FR2 (24.25 GHz - 52.6 GHz). Although a portion of FR1 is greater than 6 GHz, FR1 is often referred to (interchangeably) as a “sub-6 GHz” band in various documents and articles. A similar nomenclature issue sometimes occurs with regard to FR2, which is often referred to (interchangeably) as a “millimeter wave” bandin documents and articles, despite being different from the extremely high frequency (EHF) band (30 GHz - 300 GHz) which is identified by the International Telecommunications Union (ITU) as a “millimeter wave” band.
[0049] The frequencies between FR1 andFR2 are often referred to as mid-band frequencies. Recent 5G NR studies have identified an operating band for these mid-band frequencies as frequency range designation FR3 (7.125 GHz - 24.25 GHz). Frequency bands falling within FR3 may inherit FR1 characteristics and / or FR2 characteristics, and thus may effectively extend features of FR1 and / or FR2 into midband frequencies. In addition, higher frequency bands are currently being explored to extend 5GNRoperationbeyond 52.6GHz. For example, three higher op erating bands have been identified as frequency range designations FR2-2 (52.6 GHz - 71 GHz), FR4 (71 GHz- 114.25 GHz), andFR5 (114.25 GHz- 300 GHz). Each of these hi^ier frequency bands falls within the EHF band.
[0050] With the above aspects in mind, unless specifically stated otherwise, the term “sub-6 GHz” or the like if used herein may broadly represent frequencies that may be less than 6 GHz, may be within FR1 , or may include mid-band frequencies. Further, unless specifically stated otherwise, the term “millimeter wave” or the like if used herein may broadly represent frequencies that may include mid-band frequencies, may be within FR2, FR4, FR2-2, and / or FR5, or may be within the EHF band.
[0051] The base station 102 and the UE 104 may each include a plurality of antennas, such as antenna elements, antenna panels, and / or antenna arrays to facilitate beamforming The base station 102 may transmit a beamformed signal 182 to the UE 104 in one or more transmit directions. The UE 104 may receive the beamformed signal from the base station 102 in one or more receive directions. The UE 104 may also transmit abeamform ed signal 184 to the base station 102 in one or more transmit directions. The base station 102 may receive the beamformed signal from the UE 104 in one or more receive directions. The base station 102 / UE 104 may perform beam training to determine the best receive and transmit directions for each of the base station 102 / UE 104. The transmit and receive directions for the base station 102 may or may not be the same. The transmit and receive directions for the UE 104 may or may not be the same.
[0052] The base station 102 may include and / or be referred to as a gNB, Node B, eNB, an access point, a base transceiver station, a radio base station, a radio transceiver, a transceiver function, a basic service set (BSS), an extended service set (ESS), a TRP, network node, network entity, network equipment, or some other suitable terminology. The base station 102 can be implemented as an integrated access and backhaul (IAB) node, a relay node, a sidelink node, an aggregated (monolithic) base station with a baseband unit (BBU) (including a CU and a DU) and an RU, or as a disaggregated base station including one or more of a CU, a DU, and / or an RU. The set of base stations, which may include disaggregated base stations and / or aggregated base stations, may be referred to as next generation (NG) RAN (NG-RAN).
[0053] The core network 120 may include an Access and Mobility Management Function (AMF) 161, a Session Management Function (SMF) 162, a User Plane Function (UPF) 163, a Unified Data Management (UDM) 164, one or more location servers 168, and other functional entities. The AMF 161 is the control node that processes the signaling between the UEs 104 and the core network 120. The AMF 161 supports registration management, connection management, mobility management, and other functions. The SMF 162 supports session management and other functions. The UPF 163 supports packet routing, packet forwarding, and other functions. The UDM 164 supports the generation of authentication and key agreement (AKA) credentials, user identification handling, access authorization, and subscription management. The one or more location servers 168 are illustrated as including a Gateway Mobile Location Center (GMLC) 165 and a Location Management Function (LMF) 166. However, generally, the one or more location servers 168 may include one or more location / positioning servers, which may include one or more of the GMLC 165, the LMF 166, a position determination entity (PDE), a serving mobile location center (SMLC), a mobile positioning center (MPC), or the like. The GMLC 165 and theLMF 166 support UE location services. The GMLC 165 provides an interface for clients / applications (e.g., emergency services) for accessing UE positioning information. The LMF 166 receives measurements and assistance information from the NG-RAN and the UE 104 via the AMF 161 to compute the position of the UE 104. The NG-RAN may utilize one or more positioning methods in order to determine the position of the UE 104. Positioningthe UE 104 may involve signal measurements, a position estimate, and an optional velocity computation based on the measurements. The signal measurements may be made by the UE 104 and / or the base station 102 serving the UE 104. The signals measured may be based on one or more of a satellite positioning system (SPS) 170 (e.g., one or more of a Global Navigation Satellite System (GNSS), global position system (GPS), non-terrestrial network (NTN), or other satellite position / location system), LTE signals, wireless local area network (WLAN) signals, Bluetooth signals, a terrestrial beacon system (TBS), sensor-based information (e.g., barometric pressure sensor, motion sensor), NR enhanced cell ID (NRE-CID) methods, NRsignals (e.g., multi-round trip time (Multi-RTT), DL angle- of-departure (DL-AoD), DL time difference of arrival (DL-TDOA), UL time difference of arrival (UL-TDOA), and UL angle-of-arrival (UL-AoA) positioning), and / or other systems / signals / sensors.
[0054] Examples of UEs 104 include a cellular phone, a smartphone, a session initiation protocol (SIP) phone, a laptop, a personal digital assistant (PDA), a satellite radio, a global positioning system, a multimedia device, a video device, a digital audio player (e.g., MP3 player), a camera, a game console, a tablet, a smart device, a wearable device, a vehicle, an electric meter, a gas pump, a large or small kitchen appliance, a healthcare device, an implant, a sensor / actuator, a display, or any other similar functioning device. Some of the UEs 104 may be referred to as loT devices (e.g, parking meter, gas pump, toaster, vehicles, heart monitor, etc.). TheUE 104 may also be referred to as a station, a mobile station, a subscriber station, a mobile unit, a subscriber unit, a wireless unit, a remote unit, a mobile device, a wireless device, a wireless communications device, a remote device, a mobile subscriber station, an access terminal, a mobile terminal, a wireless terminal, a remote terminal, a handset, a user agent, a mobile client, a client, or some other suitable terminology. In some scenarios, the term UE may also apply to one or more companion devices such as ina device constellation arrangement. One or more of these devices may collectively access the network and / or individually access the network.
[0055] Referring again to FIG. l, in certain aspects, the one ormore location servers 168 may have an artificial intelligence (AI) / machine learning (ML)(AI / ML) model management component 197 may be configured to perform positioning for a user equipment (UE) based on a ML model; detect that at least one condition associated with the UE or an environment surrounding of the UE changes or meets a threshold; estimate a firstposition of the UEbased on visual information associated with the UE; select a set of radio frequency (RF) measurements for the UEbased on the estimated first position of the UE; and update the ML model based on the first position of the UE and the set of RF measurements. In certain aspects, the UE 104 may have a positioning component 198 that may be configured to perform AI / ML-based positioning. In certain aspects, the base station 102 may have a positioning configuration component 199 that may be configured to provide AI / ML-based positioning related parameters / configurations to the UE 104.
[0056] FIG. 2 A is a diagram 200 illustrating an example of a first subframe within a 5GNR frame structure. FIG. 2B is a diagram 230 illustrating an example of DL channels within a 5G NR subframe. FIG. 2C is a diagram 250 illustrating an example of a second subframe within a 5G NR frame structure. FIG. 2D is a diagram 280 illustrating an example of UL channels within a 5 G NR subframe. The 5 G NR frame structure may be frequency division duplexed (FDD) in which for a particular set of subcarriers (carrier system bandwidth), subframes within the set of subcarriers are dedicated for either DL orUL, or may be time division duplexed (TDD) in which for a particular set of subcarriers (carrier system bandwidth), subframes within the set of subcarriers are dedicated for both DL andUL. In the examples provided by FIGs. 2A, 2C, the 5G NR frame structure is assumed to be TDD, with subframe 4 being configured with slot format 28 (with mostly DL), where D is DL, U is UL, and F is flexible for use between DL / UL, and subframe 3 being configured with slot format 1 (with all UL). While subframes 3, 4 are shown with slot formats 1, 28, respectively, any particular subframe may be configured with any of the various available slot formats 0-61 . Slot formats 0, 1 are all DL, UL, respectively. Other slot formats 2-61 include a mix of DL, UL, and flexible symbols. UEs are configured with the slot format (dynamically through DL control information (DCI), or semi-statically / statically through radio resource control (RRC) signaling) through a received slot format indicator (SFI). Note that the description infra applies also to a 5G NR frame structure that is TDD.
[0057] FIGs. 2 A-2D illustrate a frame structure, and the aspects of the present disclosure may be applicable to other wireless communication technologies, which may have a different frame structure and / or different channels. A frame (10 ms) may be divided into 10 equally sized subframes (1 ms). Each subframe may include one or more time slots. Subframes may also include mini-slots, which may include 7, 4, or 2 symbols. Each slot may include 14 or 12 symbols, depending on whether the cyclic prefix (CP) is normal or extended. For normal CP, each slot may include 14 symbols, and for extended CP, each slot may include 12 symbols. The symbols on DL may be CP orthogonal frequency division multiplexing (OFDM) (CP-OFDM) symbols. The symbols on UL may be CP-OFDM symbols (for high throughput scenarios) or discrete Fourier transform (DFT) spread OFDM (DFT-s-OFDM) symbols (for power limited scenarios; limited to a single stream transmission). The number of slots within a subframe is based on the CP and the numerology. The numerology defines the subcarrier spacing (SCS) (see Table 1). The symbol length / duration may scale with 1 / SCS.Table 1: Numerology, SCS, and CP
[0058] For normal CP (14 symbols / slot), different numerologies p 0 to 4 allow for 1, 2, 4, 8, and 16 slots, respectively, per subframe. For extended CP, the numerology 2 allows for 4 slots per subframe. Accordingly, for normal CP and numerology p, there are 14 symbols / slot and 2.Llsi ots / sub frame. The subcarrier spacing may be equal to 2^ *15 kHz , where is the numerology 0 to 4. As such, the numerology p=0 has a subcarrier spacing of 15 kHz and the numerology p=4 has a subcarrier spacing of 240 kHz. The symbol length / durationis inversely related to the subcarrier spacing. FIGs. 2A-2D provide an example of normal CP with 14 symbols per slot and numerology p=2 with 4 slots per subframe. The slot duration is 0.25 ms, the subcarrier spacing is 60 kHz, and the symbol duration is approximately 16.67 ps. Within a set of frames, there may be one or more different bandwidth parts (BWPs) (see FIG. 2B) that are frequency division multiplexed. Each BWP may have a particular numerology and CP (normal or extended).
[0059] A resource grid may be used to represent the frame structure. Each time slot includes a resource block (RB) (also referred to as physical RBs (PRBs)) that extends 12 consecutive subcarriers. The resource grid is divided into multiple resource elements (REs). The number of bits carried by each RE depends on the modulation scheme.
[0060] As illustrated in FIG. 2 A, some of the REs carry reference (pilot) signals (RS) for the UE. The RS may include demodulation RS (DM-RS) (indicated as Rfor one particular configuration, but other DM-RS configurations are possible) and channel state information reference signals (CSI-RS) for channel estimation attheUE. The RS may also include beam measurement RS (BRS), beam refinement RS (BRRS), and phase tracking RS (PT-RS).
[0061] FIG. 2B illustrates an example of various DL channels within a subframe of a frame. The physical downlink control channel (PDCCH) carries DCI within one or more control channel elements (CCEs) (e.g., 1, 2, 4, 8, or 16 CCEs), each CCE including six RE groups (REGs), each REG including 12 consecutive REs in an OFDM symbol of an RB. A PDCCH within one BWP may be referred to as a control resource set (CORESET). A UE is configured to monitor PDCCH candidates in a PDCCH search space (e.g., common search space, UE-specific search space) during PDCCH monitoring occasions on the CORESET, where the PDCCH candidates have different DCI formats and different aggregation levels. Additional BWPs may be located at greater and / or lower frequencies across the channel bandwidth. A primary synchronization signal (PSS) may be within symbol 2 of particular subframes of a frame. The PSS is used by a UE 104 to determine subframe / symbol timing and a physical layer identity. A secondary synchronization signal (SSS) may be within symbol 4 of particular subframes of a frame. The SSS is used by a UE to determine aphysical layer cell identity group number and radio frame timing. Based on the physical layer identity and the physical layer cell identity group number, the UE can determine a physical cell identifier (PCI). Based on the PCI, the UE can determine the locations of the DM-RS. The physical broadcast channel (PBCH), which carries a master information block (MIB), may be logically grouped with the PSS and SSS to form a synchronization signal (SS) / PBCH block (also referred to as SS block (SSB)). The MIB provides a number of RBs in the system bandwidth and a system frame number (SFN). The physical downlink shared channel (PDSCH) carries user data, broadcast system information not transmitted through the PBCH such as system information blocks (SIBs), and paging messages.
[0062] As illustrated in FIG. 2C, some of the REs carry DM-RS (indicated as R for one particular configuration, but other DM-RS configurations are possible) for channel estimation at the base station. The UE may transmit DM-RS for the physical uplink control channel (PUCCH) and DM-RS for the physical uplink shared channel (PUSCH). The PUSCH DM-RS may be transmitted in the first one or two symbols of the PUSCH. The PUCCH DM-RS may be transmitted in different configurations depending on whether short or long PUCCHs are transmitted and depending on the particular PUCCH format used. The UE may transmit sounding reference signals (SRS). The SRS may be transmitted in the last symbol of a subframe. The SRS may have a comb structure, and a UE may transmit SRS on one of the combs. The SRS may be used by a base station for channel quality estimation to enable frequencydependent scheduling on the UL.
[0063] FIG. 2D illustrates an example of various UL channels within a subframe of a frame. The PUCCH may be located as indicated in one configuration. The PUCCH carries uplink control information (UCI), such as scheduling requests, a channel quality indicator (CQI), a precoding matrix indicator (PMI), a rank indicator (RI), and hybrid automatic repeat request (HARQ) acknowledgment (ACK) (HARQ-ACK) feedback (i.e., one or more HARQ ACK bits indicating one or more ACK and / or negative ACK (NACK)). The PUSCH carries data, and may additionally be used to carry a buffer status report (BSR), a power headroom report (PHR), and / or UCI.
[0064] FIG. 3 is a block diagram of a base station 310 in communication with a UE 350 in an access network. In the DL, Internet protocol (IP) packets may be provided to a controller / processor 375. The controller / processor 375 implements layer 3 and layer2 functionality. Layer 3 includes a radio resource control (RRC) layer, and layer 2 includes a service data adaptation protocol (SDAP) layer, a packet data convergence protocol (PDCP) layer, a radio link control (RLC) layer, and a medium access control (MAC) layer. The controller / processor 375 provides RRC layer functionality associated with broadcasting of system information (e.g., MIB, SIBs), RRC connection control (e.g., RRC connection paging, RRC connection establishment, RRC connection modification, and RRC connection release), inter radio access technology (RAT) mobility, and measurement configuration for UE measurement reporting; PDCP layer functionality associated with header compression / decompression, security (ciphering, deciphering, integrity protection, integrity verification), and handover support functions; RLC layer functionality associated with the transfer of upper layer packet data units (PDUs), error correction through ARQ, concatenation, segmentation, and reassembly of RLC service data units (SDUs), re-segmentation of RLC data PDUs, and reordering of RLC data PDUs; and MAC layer functionality associated with mapping between logical channels and transport channels, multiplexing of MAC SDUs onto transport blocks (TBs), demultiplexing of MAC SDUs from TBs, scheduling information reporting, error correction through HARQ, priority handling, and logical channel prioritization.
[0065] The transmit (TX) processor 316 and the receive (RX) processor 370 implement layer1 functionality associated with various signal processing functions. Layer 1, which includes a physical (PHY) layer, may include error detection on the transport channels, forward error correction (FEC) coding / decoding of the transport channels, interleaving, rate matching, mapping onto physical channels, modulation / demodulation of physical channels, andMIMO antenna processing The TX processor 316 handles mapping to signal constellations based on various modulation schemes (e.g., binary phase-shift keying (BPSK), quadrature phase-shift keying (QPSK), M-phase-shift keying (M-PSK), M-quadrature amplitude modulation (M-QAM)). The coded and modulated symbols may then be split into parallel streams. Each stream may then be mapped to an OFDM subcarrier, multiplexed with a reference signal (e.g., pilot) in the time and / or frequency domain, and then combined together using an Inverse Fast Fourier Transform (IFFT) to produce a physical channel carryingatime domain OFDMsymbol stream. The OFDM stream is spatially precoded to produce multiple spatial streams. Channel estimates from a channelestimator 374 may be used to determine the coding and modulation scheme, as well as for spatial processing. The channel estimate may be derived from a reference signal and / or channel condition feedback transmitted by the UE 350. Each spatial stream may then be provided to a different antenna 320 via a separate transmitter 318Tx. Each transmitter 318Tx may modulate a radio frequency (RF) carrier with a respective spatial stream for transmission.
[0066] At the UE 350, each receiver 354Rx receives a signal through its respective antenna 352. Each receiver 354Rx recovers information modulated onto an RF carrier and provides the information to the receive (RX) processor 356. The TX processor 368 and the RX processor 356 implement layer 1 functionality associated with various signal processing functions. The RX processor 356 may perform spatial processing on the information to recover any spatial streams destined for the UE 350. If multiple spatial streams are destined for the UE 350, they may be combined by the RX processor 356 into a single OFDM symbol stream. The RX processor 356 then converts the OFDM symbol stream from the time-domain to the frequency domain using a Fast Fourier Transform (FFT). The frequency domain signal includes a separate OFDM symbol stream for each subcarrier of the OFDM signal. The symbols on each subcarrier, and the reference signal, are recovered and demodulated by determining the most likely signal constellation points transmitted by the base station 310. These soft decisions may b e based on channel estimates computed by the channel estimator 358. The soft decisions are then decoded and deinterleaved to recover the data and control signals that were originally transmitted by the base station 310 on the physical channel. The data and control signals are then provided to the controller / processor 359, which implements layer 3 and layer 2 functionality.
[0067] The controller / processor 359 can be associated with at least one memory 360 that stores program codes and data. The at least one memory 360 may be referred to as a computer-readable medium. In the UL, the controller / processor 359 provides demultiplexing between transport and logical channels, packet reassembly, deciphering, header decompression, and control signal processing to recover IP packets. The controller / processor 359 is also responsible for error detection using an ACK and / or NACK protocol to support HARQ operations.
[0068] Similar to the functionality described in connection with the DL transmission by the base station 310, the controller / processor 359 provides RRC layer functionalityassociated with system information (e.g., MIB, SIBs) acquisition, RRC connections, and measurement reporting; PDCP layer functionality associated with header compression / decompression, and security (ciphering, deciphering, integrity protection, integrity verification); RLC layer functionality associated with the transfer of upper layer PDUs, error correction through ARQ, concatenation, segmentation, and reassembly of RLC SDUs, re-segmentation of RLC data PDUs, and reordering of RLC data PDUs; and MAC layer functionality associated with mapping between logical channels and transport channels, multiplexing of MAC SDUs onto TBs, demultiplexing of MAC SDUs from TBs, scheduling information reporting, error correction through HARQ, priority handling, and logical channel prioritization.
[0069] Channel estimates derived by a channel estimator 358 from a reference signal or feedback transmitted by the base station 310 may be used by the TX processor 368 to select the appropriate coding and modulation schemes, and to facilitate spatial processing. The spatial streams generated by the TX processor 368 may be provided to different antenna 352 via separate transmitters 354Tx. Each transmitter 354 Tx may modulate an RF carrier with a respective spatial stream for transmission.
[0070] The UL transmission is processed at the base station 310 in a manner similar to that described in connection with the receiver fun ction attheUE 350. Each receiver 318Rx receives a signal through its respective antenna 320. Each receiver 318Rx recovers information modulated onto an RF carrier and provides the information to a RX processor 370.
[0071] The controller / processor 375 can be associated with at least one memory 376 that stores program codes and data. The at least one memory 376 may be referred to as a computer-readable medium. In the UL, the controller / processor 375 provides demultiplexing between transport and logical channels, packet reassembly, deciphering, header decompression, control signal processing to recover IP packets. The controller / processor 375 is also responsible for error detection using an ACK and / or NACK protocol to support HARQ operations.
[0072] At least one of the TX processor 368, the RX processor 356, and the controller / processor 359 may be configured to perform aspects in connection with the positioning component 198 of FIG. 1.
[0073] At least one of the TX processor 316, the RX processor 370, and the controller / processor 375 may be configured to perform aspects in connection with the positioning configuration component 199 of FIG. 1.
[0074] FIG. 4 is a diagram 400 illustrating an example of aUEpositioningbased on reference signal measurements (which may also be referred to as “network-based positioning”) in accordance with variousaspects ofthe present disclosure. The UE404 may transmit UL SRS 412 at time TSRS_TX and receive DL positioning reference signals (PRS) (DL PRS) 410 at time TPRS RX- The TRP 406 may receive the UL SRS 412 at time TSRS RX and transmit the DL PRS 410 at time TpRSTX- The UE 404 may receive the DL PRS 410 before transmitting the UL SRS 412, or may transmit the UL SRS 412 before receiving the DL PRS 410. In both cases, a positioning server (e.g., location servers) 168) or the UE 404 may determine the RTT 414 based on ||TSRS _RX - TPRSTX| - |TSRS TX - TPRS _RX||. Accordingly, multi-RTT positioning may make use of the UE Rx-Tx time difference measurements (i.e., |TSRS TX - TPRS _RX|) and DL PRS reference signal received power (RSRP) (DL PRS-RSRP) of downlink signals received from multiple TRPs 402, 406 and measured by the UE 404, and the measured TRP Rx-Tx time difference measurements (i.e., |TSRS_RX - TPRSTX|) and UL SRS-RSRP at multiple TRPs 402, 406 of uplink signals transmitted from UE 404. The UE 404 measures the UE Rx-Tx time difference measurements (and / or DL PRS-RSRP of the received signals) using assistance data received from the positioning server, and the TRPs 402, 406 measure the gNB Rx-Tx time difference measurements (and / or UL SRS-RSRP of the received signals) using assistance data received from the positioning server. The measurements may be used atthe positioning server or the UE 404 to determine the RTT, which is used to estimate the location of theUE 404. Other methods are possible for determining the RTT, such as for example using DL-TDOA and / or UL-TDOA measurements.
[0075] PRSs may be defined for network-based positioning (e.g., NR positioning) to enable UEs to detect and measure more neighbor transmission and reception points (TRPs), where multiple configurations are supported to enable a variety of deployments (e.g, indoor, outdoor, sub-6, mmW, etc.). To support PRS beam operation, beam sweeping may also be configured for PRS. The UL positioning reference signal may be based on sounding reference signals (SRSs) with enhancements / adjustments for positioning purposes. In some examples, UL-PRS may be referred to as “SRS for positioning”and a new Information Element (IE) may be configured for SRS for positioning in RRC signaling.
[0076] DL PRS-RSRP may be defined as the linear average over the power contributions (in[W]) of the resource elements of the antenna port(s) that carry DL PRS reference signals configured for RSRP measurements within the considered measurement frequency bandwidth. In some examples, for FR1, the ref erencepoint for the DL PRS- RSRP may be the antenna connector of the UE. For FR2, DL PRS-RSRP may be measured based on the combined signal from antenna elements corresponding to a given receiver branch. For FR1 and FR2, if receiver diversity is in use by the UE, the reported DL PRS-RSRP value may not be lower than the corresponding DL PRS- RSRP of any of the individual receiver branches. Similarly, UL SRS-RSRP may be defined as linear average of the power contributions (in [W]) of the resource elements carrying sounding reference signals (SRS). UL SRS-RSRP may be measured over the configured resource elements within the considered measurement frequency bandwidth in the configured measurement time occasions. In some examples, for FR1 , the reference point for the UL SRS-RSRP may be the antenna connector of the base station (e.g., gNB). For FR2, UL SRS-RSRP may be measured based on the combined signal from antenna elements correspondingto a given receiver branch. For FR1 and FR2, if receiver diversity is in use by the base station, the reported UL SRS- RSRP value may not be lower than the corresponding UL SRS-RSRP of any of the individual receiver branches.
[0077] PRS-path RSRP (PRS-RSRPP) may be defined as the power of the linear average of the channel response at the i-th path delay of the resource elements that carry DL PRS signal configured for the measurement, where DL PRS-RSRPP for the 1 st path delay is the power contribution corresponding to the first detected path in time. In some examples, PRS path Phase measurement may refer to the phase associated with an i- th path of the channel derived using a PRS resource.
[0078] DL-AoD positioning may make use of the measured DL PRS-RSRP of downlink signals received from multiple TRPs 402, 406 at the UE 404. The UE 404 measures the DL PRS-RSRP of the received signals using assistance data received from the positioning server, and the resulting measurements are used along with the azimuth angle of departure (A-AoD), the zenith angle of departure (Z-AoD), and otherconfiguration information to locate the UE 404 in relation to the neighboring TRPs 402, 406.
[0079] DL-TDOA positioning may make use of the DL reference signal time difference (RSTD) (and / or DL PRS-RSRP) of downlink signals received from multiple TRPs 402, 406 at the UE 404. The UE 404 measures the DL RSTD (and / or DL PRS-RSRP) of the received signals using assistance data received from the positioning server, and the resulting measurements are used along with other configuration information to locate the UE 404 in relation to the neighboring TRPs 402, 406.
[0080] UL-TDOA positioning may make use of the UL relative time of arrival (RTOA) (and / or UL SRS-RSRP) at multiple TRPs 402, 406 of uplink signals transmitted from UE 404. The TRPs 402, 406 measure the UL-RTOA (and / or UL SRS-RSRP) of the received signals using assistance data received from the positioning server, and the resulting measurements are used along with other configuration information to estimate the location of the UE 404.
[0081] UL-AoApositioningmay make use of the measured azimuth angle of arrival (A-AoA) and zenith angle of arrival (Z-AoA) at multiple TRPs 402, 406 of uplink signals transmitted from the UE 404. The TRPs 402, 406 measure the A-AoA and the Z-AoA of the received signals using assistance data received from the positioning server, and the resulting measurements are used along with other configuration information to estimate the location of the UE 404. For purposes of the present disclosure, a positioning operation in which measurements are provided by a UE to a base station / positioning entity / serverto be used in the computation of the UE’s position may be described as “UE-assisted,” “UE-assisted positioning,” and / or “UE-assisted position calculation,” while a positioning operation in which a UE measures and computes its own position may be described as“UE-based ,” “UE-based positioning,” and / or “UE-based position calculation.”
[0082] Additional positioning methods may be used for estimating the location of the UE 404, such as for example, UE-side UL-AoD and / or DL-AoA. Note that data / measurements from various technologies may be combined in various ways to increase accuracy, to determine and / or to enhance certainty, to supplement / complement measurements, and / or to substitute / provide for missing information.
[0083] Note that the terms “positioning reference signal” and “PRS” generally refer to specific reference signals that are used for positioning in NR and LTE systems. However, as used herein, the terms “positioning reference signal” and “PRS” may also refer to any type of reference signal that can be used for positioning, such as but not limited to, PRS as defined in LTE and NR, TRS, PTRS, CRS, CSLRS, DMRS, PSS, SSS, SSB, SRS, UL-PRS, etc. In addition, the terms “positioning reference signal” and “PRS” may refer to downlink or uplink positioning reference signals, unless otherwise indicated by the context. To further distinguish the type of PRS, a downlink positioning reference signal may be referred to as a “DL PRS,” and an uplink positioning reference signal (e.g., an SRS-for-positioning, PTRS) may be referred to as an “UL-PRS.” In addition, for signals that may be transmitted in both the uplink and downlink (e.g., DMRS, PTRS), the signals may be prepended with “UL” or “DL” to distinguish the direction. For example, “UL-DMRS” may be differentiated from “DL-DMRS.” In addition, the term “location” and “position” may be used interchangeably throughout the specification, which may referto a particular geographical or a relative place.
[0084] In addition to Global Navigation Satellite Systems (GNSS)-based positioning and the network -based positioning described in connection with FIG. 4, various visual-based positioning mechanisms have also been developed over the last few years to improve the overall performance and accuracy of the positioning. Vision-aided positioning which may also be referred to as “vision-based positioning,” “vision positioning” “visual positioning,” “visual-based positioning,” “camera-based positioning,” and / or “camera-based visual positioning,” etc., is a positioning mechanism / mode that is capable of using images / videos captured by at least one camera to determine the location of atarget (e.g., aUE, one or more objects that are in thefield-of-view(FOV) of the at least one camera, etc.) or in some examples the location of the camera itself (e.g., the camera identifies its location based on the surrounding images / videos). For example, images captured by a camera in a warehouse may be used for calculating / estimating the location of inventories in the warehouse, and images captured by a vehicle may be used for calculating / estimating the location of the vehicle, etc.
[0085] FIG. 5 is a diagram 500 illustrating an example of vision-aided positioning in accordance with various aspects of the present disclosure. Vision-aided positioningmay offer high-accuracy position estimation, where coordinates (e.g., two- dimensional (2D) / three-dimensional (3D) coordinates, latitude and longitude coordinates, etc.) of points of interests (e.g., positions of UEs, objects in a warehouse, RF devices used for tracking / identifying inventories, etc.) may be computed from the corresponding pixels of a 2D image using (the inverse of) a projective transformation that maps world points onto the 2D image. For example, as shown at 506, a projective transformation may enable coordinates of an object 504 to be computed from the corresponding pixel of the object 504 on a 2D image 502, where the projective transformation may map the world points of the object 504 onto the 2D image 502. In some scenarios, the (accuracy and reliability of) projective transformation may depend on the camera pose (e.g., the position and the orientation of a camera) and a set of internal camera parameters (e.g., focal length, resolution, image size, etc.). Thus, knowing the (inverse of) projective transformation may be a key for reliable / accurate localization of points / objects of interest in the FOV of a camera.
[0086] For purposes of the present disclosure, a radio frequency (RF) device may referto any device with the capability of transmitting RF signals, such thatthe device may be used for positioning. Examples of RF devices may include an access point (AP), a mobile / smart phone, a user equipment (UE), a tracking device, a network node, etc. An RF-based location estimate may referto an estimation of an RF device’s location using at least one RF-related technology, such as Bluetooth®, Wi-Fi®, RF sensing and / or ultrawide band (UWB) sensing, etc. A location estimate may also referto an estimated / approximated location of an RF device within a threshold distance / radius, and / or the accuracy of the positioning is within certain degree threshold (e.g., within certain degree of accuracy), etc.
[0087] In some implementations, at least one artificial intelligence (AI) / machine learning (ML)(AI / ML) model may be used for vision-aided positioning, such as performing the projective transformation, adjusting parameters of the camera(s), and / or identifying the location of targets, etc. The at least one AI / ML model may be configured / implemented at a network entity / node (e.g., a server, a base station, a cloud, etc.) for assisting the network entity / node with the vision-aided positioning and / or implemented at a wireless device (e.g., a UE, a smart / mobile phone, a tablet, an Intemet-of-Things (loT) device, etc.) for assisting the wireless device with thevision-aided positioning. In most scenarios, using an AI / ML model may significantly improve vision-aided positioning latency, accuracy / reliability, and / or efficiency.
[0088] FIG. 6 is a diagram 600 illustrating an example scenario of using a wireless infrastructure and camera(s) for locating targets in accordance with various aspects of the present disclosure. Consider a setup as shown by the diagram 600, where there may be several targets 602 that are distributed / scattered around a space 604. For example, the space 604 and the targets 602 may be a store with merchandises, a warehouse with storage boxes, an industrial / factory / construction complex with equipments, a museum with collections, etc.
[0089] As shown at 606, each of the targets 602 may be associated with an RF device 608 (e.g., a device capable of transmitting / receivingRF signals, which may include a UE, an RF tag, a tracker, etc.) that is connected to at least one RF infrastructure 610, such as a Wi-Fi® infrastructure, an electronic shelf label (ESL)-Bluetooth® low energy (BLE) (ESL-BLE) infrastructure, an ultra wideband (UWB) infrastructure, a radio frequency identification (RFID) infrastructure, or a combination thereof, etc. The at least one RF infrastructure 610 may enable a user (or a server) to locate a target 602 using RF-based positioning (e.g., determine the position of a target based on measuring RF signals from the target). For examples, as shown at 612, a customer may use an application (e.g., a mobile application) to location a merchandise, an employee may use an RF reader to locate assets, etc. In some examples, the at least one RF infrastructure 610 may also enable different equipments or entities of a complex to locate each other, such as between robots and automated guided vehicles (AGVs), smart carts and so on.
[0090] As shown at 614, in some implementations, a positioning engine 616 (e.g., an RF- based AI / ML positioning engine) may be configured / usedto fuse RF measurements using AI / ML-based positioning models aiming to localize the targets 602. For purposes of the present disclosure, a positioning engine (PE) may refer to a system, an algorithm, or a technology that is configured to determine the location of an object or an entity within a specific space based on various input data. For example, as shown at 617, after the at least one RF infrastructure 610 (e.g., an access point (AP)) measures RF signals from a target 602 (or from the RF device 608 associated with this target 602), the at least one RF infrastructure 610 may provide the RF measurements to one or more RF-based AI / ML positioning models. Then, the one ormoreRF-based AI / ML positioning models may be able to estimate the location of this target 602 based on the RF measurements. Examples of RF measurements may include angle-of-arrival (AoA), time difference of arrival (TDOA), channel impulse response (CIR), and / or radio frequency (RF) fingerprinting, etc.
[0091] In addition to locating targets 602 based on RF measurements (which may be referred to as the “RF component” of the positioning engine 616), the positioning engine 616 may also include a vision componentfor locating the targets 602, or to improve the accuracy / performance of the RF-based positioning. For example, one or more stationary cameras 618, such as closed-circuit television (CCTV) camera system for monitoring and surveillance, over-the-top (OTT) cameras, and / or a set of non- stationary cameras 620 (e.g., a handheld (HH) camera 620 such as a UE or a smart / mobile phone with camera(s), automated guided vehicles (AGVs), smart carts etc.) may be used for obtaining images / videos related to the targets 602 and / or for performing vision-aided positioning as described in connection with FIG. 5. In other words, these stationary and / or non-stationary cameras may provide visual input to the positioning engine 616 for performing vision-based positioning.
[0092] While the positioning engine 616 may use one or more AI / ML positioning models for locating the targets 602, different AI / ML positioning models may have different purposes and / or functions, and may be suitable for different scenarios or settings. For example, a first AI / ML positioningmodel may be trained to locate items in a first area of a warehouse, and a second AI / ML positioningmodel may be trained to locate items in a second area of the warehouse. If the first AI / ML positioning model is used for locating items in the second area of the warehouse, it may not work or the positioning accuracy may be affected / reduced. As such, the positioning engine 616 may be configured to manage the AI / ML positioningmodels it used for positioningfrom time to time. In another example, an AI / ML positioning model may be able to locate a large quantity of items with less accuracy, where as another AI / ML positioning model may be able to locate a smaller quantity of items with higher accuracy, etc.
[0093] In some examples, as shown at 622, management of AI / ML positioning models may include: (1) AI / ML model selection (e.g., selecting the right site-specific AI / ML positioning model(s) to use in given site-specific conditions, (2) change detection (e.g., detecting / identifying condition(s) of a cite or site-specific condition(s) havechanged), and / or (3) training data acquisition (e.g., acquiring data fortraining / fine- tuning the AI / ML positioning model(s) to new, emerging site-specific conditions).
[0094] Aspects presented herein may improve the overall performance of AI / ML model managements. Aspects presented herein provide various vision-based techniques for managing RF -based AI / ML models, where vision (e.g., visual information) is used to address the key model management functions (e.g., as described in connection with 622 of FIG. 6). While vision-based positioning (if available) may provide highly accurate positioning, aspects presented herein may leverage existing indoor deployments including the opportunistic usage of handheld cameras to acquire highly accurate ground truth (GT) locations for the training data.
[0095] FIG. 7 is a diagram 700 illustrating an example of AI / ML positioning models specializing on different parts / areas of a space in accordance with various aspects of the present disclosure. In general, the positioning engine 616 may have the capability to fuse both the RF input as well as the visual input directly for positioning of the targets 602. Nevertheless, the availability of the vision / vision-based component for direct positioning may be limited by a number of factors such as: (1) limited positioning engine access to the vision component (e.g., the stationary cameras 618, the monitoring and surveillance system, etc.) due to privacy and / or management concerns, (2) unavailable feed from cameras (e.g., OTT-cameras) due to occlusion or due to being inactive / off (e.g., shelf cameras, part of ESL deployments), and / or (3) battery and privacy-related constraints for the HH camera(s) (e.g., the set of non- stationary cameras 620), etc.
[0096] In one aspect of the present disclosure, an AI / ML-based positioning engine may be configured to be (primarily) driven by an RF component (e.g., RF input and information), and a vision / vision-based component of the system (e.g., the AI / ML- based positioning engine) may be utilized opportunistically and / or on-demand for labeling the training data for the AI / ML models associated with the AI / ML-based positioning engine. For example, as shown by the diagram 700, for a typical structure of an AI / ML positioning engine (e.g., the positioning engine 616), a single AI / ML model is likely insufficient for performing positioning of the targets 602 for the entire space 604 (e.g., an indoor warehouse) due to the self-similar nature of spaces (e.g., similar shelf / item / package arrangements in indoor spaces) especially with small / medium range RF technologies. Hence, as shown at 702, modern deploymentsare expected to maintain a modular AI / ML positioning system consisting of multiple smaller AI / ML models, where each of them may be specialized or trained on different parts / areas of the space 604. In addition, different conditions may specify / necessitate different AI / ML models even over the same part / area of the space 604. For instance, rush hours might demand different AI / ML models then hours when the occupancy is sparse. As such, the system of specialized AI / ML models is expected to be adequately managed, which may include identifying the relevant AI / ML model for a given target at any given time for any given conditions, and / or for detecting substantial changes in the conditions / structure of the environment such that necessitate AI / ML model fine-tuning or retraining.
[0097] In another aspect of the present disclosure, a positioning AI / ML model may be configured to be associated with a set of model descriptors, where the set of model descriptors may be used for determining the site and / or the site-specific conditions of when the AI / ML model is to be used. Example of relevant pieces of information for the set of model descriptors may include area coordinates in the local coordinate system, area layout descriptors and maps (e.g., number, dimensions, locations of objects, shelfs, etc.), materials, textures and related descriptors (e.g., visual features), and / or area occupancy information (numb er of customers currently at the store), etc. For purpose of the present disclosure, “model selection” and / or AI / ML model selection” may refer to a mechanism / process of choosing (e.g., by a positioning engine, a UE, a server, etc.) a most relevant AI / ML model for a given target by matching the current surrounding conditions of the target with the closest model descriptors.
[0098] In one example, the following mechanisms may be used for acquiring the information that reflects a target’s surrounding conditions: (1) RF-assisted descriptor matching and (2) vision-assisted descriptor matching. For the RF-assisted descriptor matching useful information related to a target’ s surrounding may be obtained by subjecting the RF measurement reports to a positioning algorithm to obtain a coarse, preliminary position estimate for the target. Such position estimate may narrow down the pool / candidates of AI / ML models to just a handful of relevant AI / ML models purely based on the model location-dependence.
[0099] For the vision-assisted descriptor matching, to find the most relevant of the remainingAI / ML model(s), visual cues may be used. In some examples, visual cues may helpdisambiguating the actual area in which the target is, and / or identifying the level of occupancy or other relevant parameters. In addition, visual cues may also be used to detect changes in the environment (discuss below). In some examples, visual cues may be acquired (opportunistically) from image feeds provided by a set of stationary cameras (e.g., the one or more stationary cameras 618, CCTV-cameras, etc.), where the decision of which camera(s) to request / obtain images / videos from may be informed / determined by the RF component (e.g., as describedin connection with 622 of FIG. 6). In some examples, visual cues may be acquired (opportunistically) from image / video feeds provided by a set of non-stationary cameras (e.g., the set of non- stationary cameras 620). This may include image feeds from HH-camera(s) which may come from the target itself or other nearby devices such as store assets (e.g., smart carts etc.).
[0100] Depending on implementations, the AI / ML model selection procedure may be serverbased (e.g., the selection of AI / ML model(s) to be used by a positioning engine is determined at / by a server), UE-based (e.g., the selection of AI / ML model(s) to be used by a positioning engine is determined at / by a UE), or a combination of both. In some examples, with server-based AI / ML model selection procedure, virtually all functions may be configured to be performed by the server, whereas in a UE-based AI / ML model selection procedure, aUEmay be configuredto selectthe AI / ML model from a pool of AI / ML models provided by the server using local RF / visual information as well as assistance information from the server.
[0101] FIG. 8 is a communication flow 800 illustrating an example of a server-based AI / ML model selection in accordance with various aspects of the present disclosure. The numberings associated with the communication flow 800 do not specify a particular temporal order and are merely used as references for the communication flow 800. For purposes of the present disclosure, the UEs / target RF devices on targets, the server / wireless infrastructure that communicates with the UEs / target RF devices, and the cameras (e.g., stationary and non-stationary cameras) may collectively be referred to as a “system.” In some examples, the system may be configured to identify locations of one or more targets / assets based on using AI / ML model(s).
[0102] At 810, a UE 802 (e.g., an RF device associated with or attached to a target (“targetRF device”), the RF device 608, etc.) may be configured to transmitRF signals and / or RF measurements / reports (e.g., by measuring signals transmitted from a wirelessinfrastructure node such as an AP, a Wi-Fi router, a Bluetooth device, etc.) to a server 804 (e.g., a central server for managing targets, an RF infrastructure, etc.). As shown in FIG. 8, UE 802 may include local camera(s) 832.
[0103] At 812, based on measuring RF signals from the UE 802 and / or based on the RF measurements / reports provided by the UE 802, the server 804 may estimate the position of the UE 802 (e.g., using at least one RF-based positioning mechanism). In some configurations, the estimation may be a coarse, preliminary position estimation of the UE 802.
[0104] At 814, based on the estimated position of the UE 802, the server 804 may request one or more stationary cameras 806 (e.g., CCTV-camera(s)) to provide visual feed, such as images / videos captured by the field-of-view (FOV) of the one or more stationary cameras 806. At 816, in response to the request, the one or more stationary cameras 806 may provide their visual information (e.g., captured images / videos) to the server 804.
[0105] At 818, using the visual information provided by the one or more stationary cameras 806, the server 804 may perform model descriptor matching, where the server 804 may select a set of model descriptors that matches the visual information. As discussed above, examples of model descriptors may include area coordinates in the local coordinate system, area layout descriptors and maps (e.g., number, dimensions, locations of objects, shelfs, etc.), materials, textures and related descriptors (e.g., visual features), and / or area occupancy information (number of customers currently at the store), etc. For example, if the visual information show that the UE 802 is located at a specific section of a warehouse, is made of or packaged with a specific material or shape, and has a specific dimension, the server 804 may select a set of model descriptors that the available visual information.
[0106] In some implementations, as shown at 820, in addition or as alternative to requesting visual feed from the one or more stationary cameras 806, the server 804 may request the UE 802 (if it is equipped with at least one camera) and / or at least one non- stationary camera 808 (e.g., a camera deviceheld by auser, a camera on a robot / AGV, etc.) to provide their visual feed. For example, if the server 804 determines that additional visual feed is specified for the model descriptor matching or the AI / ML model selection (e.g., the visual feed provided by the one or more stationary cameras 806 in insufficient or obstructed) and / or if the server 804 determines that visual feedfrom the UE 802 or the at least one non-stationary camera 808 may provide additional useful information, the server 804 may requestthe UE 802 and / orthe at least one non- stationary camera 808 to provide their visual feed. At 822, in response to the request, the UE 802 and / orthe atleast one non-stationary camera 808 may provide their visual information (e.g., captured images / videos) to the server 804.
[0107] Similarly, at 828, usingthe visual information providedby the one or more stationary cameras 806, the server 804 may perform (additional) model descriptor matching or in some implementations, verify the model descriptor matching performed based on visual information from the one or more stationary cameras 806 (e.g., to confirm whether some of the model descriptors are correctly selected).
[0108] At 830, based on the model descriptor matching performed at 818 and / or 828 (or based on the selected model descriptors), the server 804may select at least one AI / ML positioning model that is to be used for estimating the position of a target (e.g., the UE 802, the targets 602, target(s) in proximity to the UE 802, etc.). Then, the server 804 may perform the positioning of the target(s) using the selected AI / ML positioning model(s).
[0109] FIG. 9 is a communication flow 900 illustrating an example of a UE-based AI / ML model selection in accordance with various aspects of the present disclosure. The numberings associated with the communication flow 900 do not specify a particular temporal order and are merely used as references for the communication flow 900.
[0110] At 910, a UE 902 (e.g., a target RF device, the RF device 608, etc.) that has the capability to perform AI / ML-based positioning of a target (which can be the UE 902 itself) may be configured to transmit a request to a server 904 (e.g., a central server for managing targets, an RF infrastructure, etc.) for requesting a list of AI / ML models in which the UE 902 may use for the positioning. In some implementations, the UE 902 may also be configured to transmit RF signals and / or RF measurements / reports (e.g., by measuring signals transmitted from a wireless infrastructure node such as an AP, a Wi-Fi router, a Bluetooth device, etc.) to the server 904. As shown in FIG. 9, UE 902 may include local camera(s) 932.
[0111] At 912, based on the request (and the RF signals / measurements / reports), the server 904 may send a list of candidate AI / ML models (and also their corresponding / associated model descriptors) to the UE 902.
[0112] At 914, the UE 902 may estimate its position (e.g., using at least one RF-based positioning mechanism). In some configurations, the estimation may be a coarse, preliminary position estimation of the UE 902.
[0113] At 916, based on its estimated position, the UE 902 may request the server 904 to provide visual feed related to its estimated position. As such, this visual feed request may also include notifying the server 904 of the estimated position of the UE 902.
[0114] At 918, in response to the visual feed request from the UE 902, the server 904 may request one or more stationary cameras 906 (e.g., CCTV-camera(s)) and / or at least one non-stationary camera 908 (e.g., a camera device held by a user, a camera on a robot / AGV, etc.) to provide their visual feed. The server 904 may select the one or more stationary cameras 906 and / or the at least one non-stationary camera 908 based on the estimated position of the UE 902. At 920, in response to the request from the server 904, the one or more stationary cameras 906 and / or the at least one non- stationary camera 908 may provide their visual information (e.g., captured images / videos) to the server 904. Then, at 922, the server 904 may transmit / forward the visual information (received from the one or more stationary cameras 906 and / or the at least one non-stationary camera 908) to the UE 902. In some implementations, the server 904 may also configure the one or more stationary cameras 906 and / or the at least one non-stationary camera 908 to provide visual information directly to the UE 902.
[0115] At 924, using the visual information provided by the server 904 (or directly from the one or more stationary cameras 906 and / or the at least one non-stationary camera 908), the UE 902 may perform model descriptor matching, where the UE 902 may select a set of model descriptors that matches the visual information. As discussed above, examples of model descriptors may include area coordinates in the local coordinate system, area layout descriptors and maps (e.g., number, dimensions, locations of objects, shelfs, etc.), materials, textures and related descriptors (e.g., visual features), and / or area occupancy information (number of customers currently at the store), etc. For example, if the visual information show that the UE 902 is located at a specific section of a warehouse, is made of or packaged with a specific material or shape, and has a specific dimension, theUE 902 may select a set of model descriptors that the available visual information.
[0116] At 926, based on the model descriptor matching performed at 924 (or based on the selected model descriptors), the UE 902 may select at least one AI / ML positioning model that is to be used for estimating the position of a target (e.g., the UE 902, the targets 602, target(s) in proximity to the UE 902, etc.). Then, theUE 902 may perform the positioning of the target(s) using the selected AI / ML positioning model(s).
[0117] In some scenarios, a system (e.g., a UE, a server, a wireless infrastructure, etc.) may be configured to initialize a set of AI / ML positioning models, either through initial training (e.g., dedicated offline survey at the time of the system deployment) and / or by initializing the weights of the AI / ML models using pre-trained third-party models. However, conditions in given areas may change significantly, to the point that may render the AI / ML models ineffective or useless in which case the system may be demanded to identify the new conditions and change / update / modify the AI / ML model(s), or to trigger a retraining / fine-tuning procedure for the AI / ML model(s) used.
[0118] The following are some example events that may trigger an AI / ML model change or an AI / ML model retraining or fine-tuning. In one example, an AI / ML model change / retraining / fine-tuning may be triggered when a target (or targets) leaves the area over which the AI / ML model was trained. This may be one of most common trigger evens, where a new AI / ML model from a pool of available trained AI / ML models is to be selected using the inputs and the procedure discussed in connection with FIGs. 8 and 9. In another example, an AI / ML model change / retraining / fine- tuning may be triggered when there are changes in the wireless / RF infrastructure. The AI / ML models may be trained for a given fixed RF infrastructure, and any change in the RF infrastructure should be correspondingly reflected in the AI / ML models. Example changes in an RF infrastructure may include changes in the AP / transmitter locations, transmit parameters (e.g., frequency, power levels, bandwidth), etc. In another example, an AI / ML model change / retraining / fine-tuning may be triggered when there are changes in the environment (surrounding the targets). The environment may affect the RF propagation channel, and any change should be reflected in the AI / ML models as well. For example, changes in the layout such as aisle / shelf re-organization, systematic inventory replacement, or any other structural change that may deem the previously trained AI / ML model(s) outdated.
[0119] In another aspect of the present disclosure, the system described herein (e.g., the system for identifying locations of one or more targets / assets based on using AI / ML model(s), or a positioning engine) may be configured with a set of metrics and / or criteria for detecting trigger events that specify an AI / ML model change or an AI / ML model retraining or fine-tuning. In one implementation, the system may include a positioning confidence that is configured to be associated with the output of the AI / ML models, such as a type of metric that is capable of quantifying the quality of the inference (e.g., output of an AI / ML model). For example, a trigger event may occur when the positioning confidence drops below a specified threshold andthe drop persists for a specified duration or a time limit.
[0120] In another implementation, the trigger events may be based on visual cues and indicators from images / videos. For example, images / videos provided by the stationary camera(s) (e.g., CCTV-cameras) and / or non-stationary camera(s) (e.g., HH-Camera(s) of an RF device, a UE, or a target) may be processed, analyzed, and compared against a database to help identify any major relevant changes (discuss below). This visual information may be configured to be received periodically or on- demand (e.g., in response to a request from a UE / server).
[0121] In another implementation, the trigger events may be based on RF-based positioning information obtained using a positioning method (e.g., an RF-based positioning method). For example, the UE / server may be configured to estimate the RF position of targets periodically to detect whether any of the targets has changed its position significantly. In some examples, in the case that the computed metrics are inconclusive, the system / server may ask for additional input from the target RF device / UE and / or the visual component of the system (e.g., the stationary / non- stationary camera(s)).
[0122] FIG. 10 is a diagram 1000 illustrating an example of initiating an AI / ML model selection, training, and / or fine-tuning (e.g., retraining) based on triggering condition(s) in accordance with various aspects of the present disclosure. In one example, as shown at 1002, the system described herein or an entity of the system (e.g., a UE, a server, ora wireless infrastructure, etc. depending on where this function is implemented) may be configured to detect change(s) / condition(s) that triggers the system or an entity of the system (hereafter “system / entity”) to change / select at leastone new AI / ML model, and / or to train / fine-tune at least one new AI / ML model based on various types of information as discussed above.
[0123] For example, as shown at 1004, types of information that may be used by the system / entity for detecting changes / conditions may include: (1) the positioning confidence and / or the position estimates (of the targets) obtained from a UE (e.g., the UE 802, 902), (2) visual information from one or more camera(s) (e.g., the one or more stationary cameras 806, 906, the at least one non-stationary camera 808, 908, etc.), and / or (3) additional information from other sources, such as maps (e.g., map data showing locations of targets, warehouse layout, equipment arrangements, etc.), database for digital twins, etc.
[0124] In some examples, the system / entity may request the visual information (opportunistically) from the vision component of the system (e.g., the stationary camera(s), the monitoring and surveillance system, CCTV-cameras, etc.). Visual information from the stationary camera(s) may typically pertain to detections of potential targets such as bounding boxes. Additionally, the system / entity may request the visual information from non-stationary camera(s) such as the target itself, store assets (e.g., smart carts, robots, automated guided vehicles), or other targets in the vicinity (e.g., active collaborative targets (ACTs)) from their devices. In some examples, non-stationary camera(s) may provide various pieces of useful visual information. For instance, if the visual feed is coming from a target RF device (e.g, the target RF device includes at least one camera), the information may be contextual, represented in surrounding visual features (such as scale-invariant feature transform (SIFT)) that encode the location of the target implicitly. If the visual feed is coming from other non-stationary camera(s) (e.g., neighboring collaborative devices), the visual information may be represented by both potential target detections as well as contextual features.
[0125] In some implementations, as shown at 1006, based on detecting change(s) / condition(s), the system / entity may also be configured to determine whether there is sufficient information to initiate the AI / ML model selection, training or fine-tuning. For example, the system / entity may generate / compute a change indication metrics to determine whether there is sufficient information to initiate the AI / ML model selection, training, or fine-tuning. As shown at 1008 and 1010, if the system / entity determines there is sufficient information, the system / entity may initiateat least one of an AI / ML model selection (e.g., as described in connection with FIGs. 8 and 9), or an AI / ML model training / fine-tuning (discussed below). On the other hand, as shown at 1012, if the system / entity determines there is insufficient information, the system / entity may be configured to request additional information from the one or more entities of the system, such as additional RF measurements and / or position estimates from UEs / target RF devices, additional visual information from cameras, and / or information from additional source(s), etc.
[0126] As discussed above (and shown at 1010 of FIG. 10), in some scenarios, the AI / ML positioning models (e.g., the RF-based AI / ML positioning models) used by the system / entity may be specified to be trained or retrained / fine-tuned if the (triggering) condition(s) have changed sufficiently suchthatthe previously trained AI / ML models (e.g., the AI / ML model(s) currently used) become less effective or not useful anymore. Note while the initial training of the AI / ML model(s) is not discussed in details here, aspects presented herein may also apply to the initial training of the AI / ML model(s), such as for gathering data for the initial AI / ML model training. As such, aspects presented herein may focus on the processes and protocols to acquire sufficient amount of fresh data for retraining / fine-tuning the AI / ML RF-based positioning models.
[0127] In general, the amount of data specified for updating an AI / ML model after change of conditions may be application dependent (e.g., different applications may demand different conditions). Nevertheless, aspects presented herein is to outline the protocol and information exchange framework specified to acquire enough fresh data that can be used to update the RF-based positioning models as timely as possible. Therefore, aspects presented herein may adhere to following design principles related to the process of collecting training data for AI / ML model updates: (1) accuracy, (2) speed and timeliness, and (3) privacy.
[0128] In terms of accuracy, training data may be specified to be as accurate as possible, which means that the ground truth locations (discussed below in connection with training data formatting) may be expected to have high precisions (e.g., the precision is above a precision threshold). In addition, the training data may be expected to be correctly labeled as mismatches in labeling may be very costly and lead to performance degradations.
[0129] In terms of speed and timeliness, after a change of the condition(s) has been detected and the AI / ML model is determined to be retrained, this retraining process may be specified to be performed as soon as possible. Therefore, the framework described herein may be configured to support quick acquisition of sufficient amounts of fresh data to make sure that the acquired data is as accurate as possible.
[0130] In terms of privacy, as the framework may be configured to use camera(s) and visual data opportunistically for collecting training data, aspects presented herein may account for privacy concerns especially if people (e.g., store customers) are used as visual training targets.
[0131] In one aspect of the present disclosure, an example format of the training data (e.g., for training / fine-tuning AI / ML model(s)) that is to be used / acquired may include at least the ground truth location and the RF measurement(s) (e.g., {Ground Truth Location, RF measurement(s)}). A ground truth location may refer to the actual / precise location of a target, which may be represented by a set of coordinates, a place label, etc. The RF measurement(s) may refer to the RF measurements associated with the target (e.g., with the UE / RF device on the target), such as described in connection with FIGs. 8 and 9. For example, RF measurements may be the signal pattem / strength of signals transmitted by the UE / RF device on the target when the target is at the ground truth location. Depending on implementations, additional information may also be included in the training data such as RF-related information (e.g., infrastructure AP locations, types of wireless technology, frequency / bandwidth, etc.) as well as contextual information such as three- dimensional (3D) maps, digital twins, layout changes indications, etc. For purposes of the present disclosure, a digital twin may refer to a digital representation of a physical object, person, or process, contextualized in a digital version of its environment.
[0132] The AI / ML model(s) may be configured to be trained by a minimum amount of training data that includes at least the pair of information (the ground truth location and the RF measurement(s)). As such, the system / entity may also specify effective ways to acquire ground truth locations and match the ground truth locations with the corresponding RF measurements.
[0133] In one example, to acquire the ground truth locations of targets, the system / entity may be configured to use visual positioning of dedicated RF devices carried / attachedto adedicated target (which may be referred to as a “training target”) from a set of calibrated cameras (e.g., a set of calibrated monitoring and surveillance cameras) to obtain very accurate ground truth location of the training target. In another example, the system / entity may be configured to use RF-based position estimates (e.g., non- AI / ML-based and / or AI / ML-based positioning mechanism(s)) as featuresto match a set of training targets that are visually detected with the corresponding RF measurementsusedastheinputforupdatingtheRF-based AI / ML positioningmodels. In another example, the system / entity may be configured to use a memory -based matching procedure with flexible / incremental sliding window to ensure highly accurate matching between the RF measurements and the corresponding ground truth locations (i.e., visual position estimates).
[0134] In some implementations, the training targets may include infrastructure-based targets and / or active collaborative targets. Infrastructure-based targets may refer to assets of the deployment such as AGVs, robots, smart carts, or even employees that are part of the deployment and are able to be readily used to produce data for AI / ML training when specified. Active collaborative targets may refer to extraneous RF devices typically associated with people such as customers who are expected to explicitly consent to participate in training data acquisition while the session is on if specified.
[0135] In addition, each training target may be associated with an identification / identifier (ID) and stored with the training data. For example, the system / entity may be configured to store, at minimum, the medium access control (MAC) ID (or other unique ID) of an RF device / UE associated with the target (and the corresponding training data may also be linked to / labelled with this ID). The system / entity may also be configured to store additional information for the infrastructure-based targets and the active collaborative targets depending on the relevant cases.
[0136] For example, as infrastructure-based targets may be any asset from the deployment, privacy issues may not be a concern (even if the assets are employees as they may understand the usage of the data and consent in being tracked thoroughly if needed). In such cases, matching the visual detections and the visual position estimates (e.g., the ground truth locations) with the RF measurements may be done easily and very accurately based on visual features of the training targets. Hence, in relation to the MAC ID of the RF device / UE, the system / entity may also be configured to store avariety of visual features and related information (e.g., shape, color, type facial recognition information, etc.) of the infrastructure-based targets.
[0137] In another example, as active collaborative targets (ACTs) may typically correspond to consenting users such as customers in a store, privacy issues may exist. Therefore, the system / entity may be implemented with a design that does not rely on any visual feature information pertaining to the active collaborative targets to do the matching besides the typical visual detections for visual positioning (such as bounding boxes). For example, the system / entity may be configured to use RF-based information, namely the RF-based position estimates to match the visual detections with the RF measurements. Furthermore, as the infrastructure-based targets may not be available in a sufficient number in some scenarios, it is possible that training data for AI / ML model updating may be acquired more often opportunistically, using active collaborative targets.
[0138] FIG. 11 is a communication flow 1100 illustrating an example of a training data acquisition framework in accordance with various aspects of the present disclosure. The numberings associated with the communication flow 1100 do not specify a particular temporal order and are merely used as references for the communication flow 1100.
[0139] As shown at 1110, after a training / fine-tuning trigger event is detected, such as described in connection with 1010 of FIG. 10, a server 1104 (e.g., the server 804, 904) may be configured to request RF measurement(s) from a set of training targets 1102 (e.g., may be UE 802, 902, a set of RF devices / UEs, etc.) and to request visual information from a set of cameras 1106 (e.g., may be a set of stationary cameras, a set of non-stationary cameras, or a combination of both). In one example, the server 1104 may select the set of training targets 1102 from a set of infrastructure-based targets if they are available, or from a list of active collaborative targets with focus on the area affectedby the trigger event (e.g., certain part of the store, certain time of the day, etc.). Similarly, the server 1104 may select the set of cameras 1106 based on their locations, FOVs, resolutions, etc. Also, the server 1104 may have the capability to activate additional cameras if available, and keep them inactive if not used.
[0140] Based on the request from the server 1104, as shown at 1112, the set of training targets1102 may provide their (fresh) RF measurements to the server 1104, such as described in connection with 810 of FIG. 8, and the set of cameras 1106 may also provide theirvisual information to the server 1104. The RF measurements from the set of training targets 1102 may serve a variety of purposes. For example, the RF measurements may be used (e.g., by the server 1104) as the relevant and fresh RF input for updating the AI / ML models. In another example, as shown at 1114, the server 1104 may use the RF measurements to perform matching with the corresponding visual position estimates (e.g., ground truth locations). In other words, the server 1104 may match the RF measurements related to the set of training targets 1102 to their corresponding visual features visual-based position estimates. For matching, RF-based position estimates (discuss below) may be obtained using either RF-based positioning or even the AI / ML-based algorithm if the positioning confidence is within a defined threshold. In some implementations, the RF-based positioning may be performed at the set of training target(s) and send over to the server 1104 with the raw RF measurements.
[0141] The visual information may be used (e.g., by the server 1104) for obtaining visual features associated with the set of training targets 1102, and / or for performing visual positioning for the set of training targets 1102 (e.g., to obtain a set of visual-based position estimates). Depending on implementations, the set of cameras 1106 may send the server 1104 either frames from the cameras or processed detections (e.g., bounding boxes) including relevant camera calibration information (e.g., camera matrices used for visual positioning). Then, the server 1104 may use the RF-based position estimates to match the set of training targets 1102 with the corresponding visual detections (e.g., to match the RF-based position estimates to the visual features and / or visual-based position estimates).
[0142] In some examples, as shown at 1116, the server 1104 may be configured to perform the above steps (e.g., 1110, 1112, and 1114) incrementally by using a memory-based approach for matching (discussed below). The memory-based approach may be used to ensure that the matching of ground truth locations with RF measurements is highly accurate which may be important for training purposes (discussed below).
[0143] At 1118, after the RF measurements are matched to their corresponding training targets (e.g., their ground truth locations obtained via visual positioning), the server 1104 may obtain a set of training data that includes the pair of information (e.g, ground truth locations and RF measurements). Then, at 1120, the server 1104 may train / fine-tune (e.g., update) one or more AI / ML models based on the set of trainingdata. In some examples, as shown at 1122, after an AI / ML model is trained, the server 1104 may save it with the corresponding model descriptors pertaining to the conditions under which the AI / ML model is valid (e.g., as described in connection with FIGs. 8 and 9).
[0144] FIG. 12 is a diagram 1200 illustrating an example of a memory-based matching of ground truth locations with RF measurements in accordance with various aspects of the present disclosure. Have a high accuracy matching rate (e.g., as close to 100% matching rate as possible) may be a key for accurate / correct ground truth location labeling, i.e., pairing the ground truth locations with the corresponding RF measurements for AI / ML model training and updating. However, when trying to match visual detections / position estimates with the RF-based position estimates, a single snapshot matching is likely to be insufficient due to the relatively high RF positioning error when typical RF-based positioning is used. Accordingly, in another aspect of the present disclosure, the server 1104 (or the system / entity) may be implemented with a temporal memory-based matching mechanism that is capable of performing an accurate matching of ground truth locations with RF measurements.
[0145] As shown by the diagram 1200, in each memory period 1204 of a given duration, each training target 1102 may be configured to spawn an RF tracklet and a visual tracklet (if the training target is not occluded). For purposes of the present disclosure, an RF tracklet may refer to a sequence of timestamped RF-based position estimates belonging to the same training target, and a visual tracklet may refer to a sequence of timestamped vision-based position estimates belonging to the same training target.
[0146] Then, as shown at 1210, the server 1104 may be configured to compute pairwise matching costs between the RF tracklet and the visual trackets (which may be referred to as a “pair of RF and visual tracklets”) provided by the training target 1102. In each pair of RF and visual tracklets, multiple timestamped RF and vision-based position estimates may be available, and the pairwise association cost between the RF tracklet and visual tracklet may be computed using a function of the pairwise distances between corresponding timestamped RF and vision-based position estimates. Example pairwise cost functions may include sum, mean, etc.
[0147] FIG. 13 is adiagram 1300 illustrating an example of an association matrix that is used for the memory -based matching in accordance with various aspects of the present disclosure. In one example, as shown at 1302, the server 1104 may be configured toform an association matrix using the pairwise association costs for all RF and visual tracklet pairs, where the association matrix may then be used as the input to the matching (e.g., using a Hungarian algorithm or other technique(s)). After the matching is done, the training data set may be constructed as shown at 1304.
[0148] FIG. 14 is a communication flow 1400 illustrating an example of an incremental memory -based matching scheme in accordance with various aspects of the present disclosure. The numberings associated with the communication flow 1400 do not specify a particular temporal order and are merely used as references for the communication flow 1400.
[0149] For the incremental memory -based matching discussed in connection with FIGs. 11 to 13, the duration of the memory period (e.g., the memory period 1204, which may be referred to as a “memory window” hereafter) may be a key parameter that determines the accuracy of the matching (e.g., the matching between ground truth locations and RF measurements). For example, if the memory window is too short, the matching performance may be close to a single-snapshot and might not be sufficient (e.g., the server is just able to obtain one or very few RF measurements and ground truth locations within this short memory window). On the other hand, while longer memory windows may generally be better, they might lead to occlusion issues and visual tracklet fragmentation resulting in multiple visual tracklets per detected training target in each memory period. For example, as more people (e.g., customers and employees)may movebetweenthe camera / RF-infrastructure and a training target during a longer period of time, the RF tracklets and / or visual tracklets spawned by the training target may be in fragments (e.g., with interruptions).
[0150] Accordingly, in another aspect of the present disclosure, the server 1104 may further be implemented with an incremental approach where the server 1104 is configured to gradually increase the duration of the memory window until certain criteria for termination is satisfied. Such criteria may include a set of metrics that is capable of inferring the accuracy of the matching achieved at the current memory window. For example, the termination metrics that may be used by the server 1104 may include the value of the association cost achieved at the result of the matching (e.g., an optimal value, adjusted for the duration of the memory window), where lower values of such metric may indicate that the matching is more accurate. In another example, the termination metrics may include a mean number of visual tracklets formed perdetected training target for a given memory duration due to tracklet fragmentation, where a larger number of such fragments may indicate that the matching performance is degraded / jeopardized. In another example, the termination metrics may be associated with the latency induced by increasing the memory window to ensure timeliness of the AI / ML model training.
[0151] For example, at 1410, the server 1104 may be configured to set a memory window (e.g., a duration of X seconds / minutes, etc.), where this memory window may be increased by small increments (e.g., quanta of memory). For instance, such increment may correspond to several additional RF-based and vision-based position estimates (e.g., measurements).
[0152] As shown at 1412, for each increment (or each decision to increase the memory window by the duration of the increment), the server 1104 may notify / request the set of training targets 1102 (e.g., the active collaborative targets) to send further / additional RF measurement(s) until further notice. Similarly, the server 1104 may notify / request the set of cameras 1106 to send further / additional visual information until further notice. At 1414, in response to the notification / request, the set of training targets 1102 and / or the set of cameras 1106 may send their corresponding RF measurements and / or visual information to the server 1104.
[0153] At 1416 and 1418, when the counter related to the increment expires or the set of training targets 1102 has provided / sent sufficient numb er of RF measurements, the server 1104 may send an acknowledgement (ACK) to the set of training targets 1102 (and the set of training targets 1102 may stop sending the RF measurements based on the acknowledgement). Similarly, when the counter related to the increment expires or the set of cameras 1106 has provided / sent sufficient visual information, the server 1104 may send an acknowledgement to the set of cameras 1106 (and the set of cameras 1106 may stop sending the visual information based on the acknowledgement). If the counter has not expired and the server 1104 has not obtained sufficient RF measurements / visual information, then as shown at 1420, the server 1104 may continue to requestRF measurements / visual information and / or receive RF measurements / visual information from the set of training targets 1102 and / or the set of cameras 1106.
[0154] Based on the reception of the new batch of RF measurements from the set of training targets 1102 as well as the visual information from the set of cameras 1106, the server1104 may perform the matching as describedin connection with FIGs. 12 and 13. For example, at 1422, the server 1104 may perform RF-based positioning and visualbased positioning for each of the set of training targets 1102 to obtain a set of RF- based positioning estimates and a set of visual-based positioning estimates for each of the set of training targets 1102. At 1424, the server 1104 may form a set of RF tracklets and a set of visual tracklets (i.e., a set of RF and visual tracklet pairs) for each of the set of training targets 1102. At 1426, the server 1104 may be configured to form an association matrix usingthe pairwise association costs for all RF and visual tracklet pairs. At 1428, the server 1104 may perform the matching using the association matrix as the input to the matching (e.g., using a Hungarian algorithm or other technique(s)).
[0155] At 1430, after performing the matching, the server 1104 may determine whether criteria forthe termination have been satisfied (e.g., whether a termination metric used has crossed a defined threshold). If the termination has not been satisfied (e.g., the termination metric has not crossed the defined threshold, as shown at 1432, the server 1104 may be configured to set another memory window (e.g., to a longer period of time), and repeat the steps described above. However, if the termination has been satisfied (e.g., the termination metric has crossed the defined threshold, as shown at 1434, the server 1104 may proceed to the training pairs formation, such as described in connection with 1118 of FIG. 11 .
[0156] Aspects presented herein may improve the overall performance of AI / ML model managements. Aspects presented herein provide various vision-based techniques for managing RF-based AI / ML models, where vision is used to address the key model management functions. While vision-based positioning (if available) may provide highly accurate positioning, aspects presented herein may leverage existing indoor deployments including the opportunistic usage of handheld cameras to acquire highly accurate GT locations for the training data. For example, a store / warehouse may have several targets distributed around the space and each target is associated with an RF device / UE which is connected to an RF infrastructure. An RF-based ML positioning engine fuses RF measurements and vision input using ML-based positioning models to localize the targets. As the availability of vision-based component for direct positioning may be limited due to various factors, aspects presented herein propose vision-based techniques for managing RF-based ML models where vision is used toaddress key model management functions for ML positioning models. In addition to theML-based positioning engine driven by RF input and information, the visionbased component is used opportunistically and on-demand for labeling the training data for ML models. Other aspects include temporal and incremental memory matching schemes.
[0157] FIG. 15 is a flowchart 1500 of a method of wireless communication. The methodmay be performed by a server (e.g., the one or more location servers 168; the atleast one RF infrastructure 610; the positioning engine 616; the server 804, 904, 1104; network entity 1760). The method may enable the server to use vision-based techniques for managing RF-based AI / ML positioning models, thereby improving the overall performance of the RF-based AI / ML positioning.
[0158] At 1504, the server may perform positioning for a UE based on a ML model, such as described in connection with FIGs. 8 to 14. For example, as described in connection with 812 of FIG. 8, based on measuringRF signals from the UE 802 and / or based on the RF measurements / reports provided by the UE 802, the server 804 may estimate the position of the UE 802 (e.g., using at least one RF-based positioning mechanism). The positioning of the UE may be performed by, e.g., the AI / ML model management component 197, the network processor(s) 1712, and / or the network interface 1780 of the network entity 1760 in FIG. 17.
[0159] At 1506, the server may detect that at least one condition associated with the UE or an environment surrounding oftheUE changes ormeets athreshold, such as described in connection with FIGs. 8 to 14. For example, as described in connection with 1002 ofFIG. 10, the system described herein or an entity ofthe system (e.g., a UE, a server, or a wireless infrastructure, etc. depending on where this function is implemented) may be configured to detect change(s) / condition(s) that specifies the system or the entity of the system (hereafter “system / entity”) to change / select at least one new AI / ML model, and / or to train / fine-tune at least one new AI / ML model based on various types of information as discussed above. The detection of the at least one condition may be performed by, e.g., the AI / ML model management component 197, the network processor(s) 1712, and / or the network interface 1780 of the network entity 1760 in FIG. 17.
[0160] In one example, the at least one condition includes at least one of: a position of the UE is outside of a designated area over which the ML model is trained, a change inRF infrastructure associated with the UE, or a change in the environment surrounding the UE.
[0161] In another example, the detection of the at least one condition is based on at least one of: a positioning confidence, a set of visual cues and indicators, or RF-based positioning information.
[0162] At 1512, the server may estimate a first position of the UEbased on visual information associated with the UE, such as described in connection with FIGs. 8 to 14. For example, as described in connection with 1112 of FIG. 11, the visual information may be used (e.g., by the server 1104) for obtaining visual features associated with the set of training targets 1102, and / or for performingvisualpositioningfor the set of training targets 1102 (e.g., to obtain a set of visual-based position estimates). The estimation of the first position of the UE may be performed by, e.g., the AI / ML model management component 197, the network processor(s) 1712, and / or the network interface 1780 of the network entity 1760 in FIG. 17.
[0163] In one example, the visual information corresponds to a set of visual features, a set of frames, or a set of processed detections.
[0164] At 1514, the server may select a set of RF measurements for the UE based on the estimated first position of the UE, such as described in connection with FIGs. 8 to 14. For example, as described in connection with 1114 of FIG. 14, the server 1104 may use the RF measurements to perform matching with the corresponding visual position estimates (e.g., ground truth locations). In other words, the server 1104 may match the RF measurements related to the set of training targets 1102 to their corresponding visual features visual-based position estimates. The selection of the set of RF measurements may be performed by, e.g., the AI / ML model management component 197, the network processor(s) 1712, and / or the network interface 1780 of the network entity 1760 in FIG. 17. In some implementations, to selectthe set of RF measurements for the UE based on the estimated first position of the UE, the server may be configured to estimate or receive a second position of the UE based on the set of RF measurements, and match the UE to the set of RF measurements based on mapping the first position to the second position. In some implementations, to match the UE to the set of RF measurements based on mappingthe firstposition to the secondposition, the server may be configured to compare a first sequence of time-stamped RF-based position estimates associated with the UE with a second sequence of time-stampedvision-based position estimates associated with the UE over a defined duration, and map the first position to the second position based on the comparison. In some implementation, the first position corresponds to a ground truth location of the UE, and the second position corresponds to a coarse position of the UE.
[0165] At 1516, the server may update the ML model based on the first position of the UE and the set of RF measurements, such as described in connection with FIGs. 8 to 14. For example, as described in connection with 1118 and 1120 of FIG. 11 , after the RF measurements are matched to their corresponding training targets (e.g., their ground truth locations obtained via visual positioning), the server 1104 may obtain a set of training data that includes the pair of information (e.g., ground truth locations and RF measurements). Then, the server 1104 may train / fine-tune (e.g., update) one or more AI / ML models based on the set of training data. The positioning of the UE may be performed by, e.g., the AI / ML model management component 197, the network processor(s) 1712, and / orthe network interface 1780 of the network entity 1760 in FIG. 17. In some implementations, to update the ML model comprises at least one of: retrain the ML model, fine-tune the ML model, replace the ML model with a second ML model, or switch the ML model to the second ML model.
[0166] In one example, the server may transmit, to the UE based on the detection that the at least one condition changes or meets the threshold, a request to perform the set of RF measurements, and receive, from the UE based on the request, the set of RF measurements, such as described in connection with FIGs. 8 to 14. For example, as described in connection with 1110 and 1112 of FIG. 11, after a training / fine-tuning trigger event is detected, such as described in connection with 1010 of FIG. 10, a server 1104 (e.g., the server 804, 904) may be configured to request RF measurement(s) from a set of training targets 1102 (e.g., may be UE 802, 902, a set of RF devices / UEs, etc.) and to request visual information from a set of cameras 1106 (e.g., may be a set of stationary cameras, a set of non-stationary cameras, or a combination of both). Based on the request from the server 1104, as shown at 1112, the set of training targets 1102 may provide their (fresh) RF measurements to the server 1104, such as describedin connection with 810 ofFIG. 8, and the set of cameras 1106 may also provide their visual information to the server 1104. The transmission of the request and / orthe reception of the set of RF measurements may be performedby, e.g., the AI / ML model management component 197, the network processors) 1712, and / or the network interface 1780 of the network entity 1760 in FIG. 17.
[0167] In another example, the server may transmit, to at least one camera based on the detection that the at least one condition changes or meets the threshold, a request to provide the visual information associated with the UE, and receive, from the at least one camera based on the request, the visual information, such as described in connection with FIGs. 8 to 14. For example, as described in connection with 1110 and 1112 of FIG. 11, after a training / fine-tuning trigger event is detected, such as described in connection with 1010 of FIG. 10, a server 1104 (e.g., the server 804, 904) may be configured to request RF measurement(s) from a set of training targets 1102 (e.g., may be UE 802, 902, a set of RF devices / UEs, etc.) and to request visual information from a set of cameras 1106 (e.g., may be a set of stationary cameras, a set of non-stationary cameras, or a combination of both). Based on the request from the server 1104, as shown at 1112, the set of training targets 1102 may provide their (fresh) RF measurements to the server 1104, such as describedin connection with 810 of FIG. 8, and the set of cameras 1106 may also provide their visual information to the server 1104. The transmission of the request and / or the reception of the visual information may be performed by, e.g., the AI / ML model management component 197, the network processor(s) 1712, and / or the network interface 1780 of the network entity 1760 in FIG. 17.
[0168] In another example, the server may request, based on an estimated position of the UE, a vision feed from at least one of the UE, a stationary camera, or a non-stationary camera, receive the vision feed from at least one of the UE, the stationary camera, or the non-stationary camera, identify, based on the received vision feed, at least one model descriptor from a set of model descriptors, and select, based on the identified at least one model descriptor, the ML model for the positioning of the UE, such as described in connection with FIGs. 8 to 14. For example, at 814, based on the estimated position of the UE 802, the server 804 may request one or more stationary cameras 806 (e.g., CCTV-camera(s)) to provide visual feed, such as images / videos captured by the field-of-view (FOV) of the one or more stationary cameras 806. At 816, in response to the request, the one or more stationary cameras 806 may provide theirvisual information (e.g., capturedimages / videos)tothe server 804. At 818, using the visual information provided by the one or more stationary cameras 806, the server804 may perform model descriptor matching, where the server 804 may select a set of model descriptors that matches the visual information. At 830, based on the model descriptor matching performed at 818 and / or 828 (or based on the selected model descriptors), the server 804 may select at least one AI / ML positioning model that is to be used for estimating the position of a target (e.g., the UE 802, the targets 602, target(s) in proximity to the UE 802, etc.). Then, the server 804 may perform the positioning of the target(s) using the selected AI / ML positioning model(s). The transmission of the request, the reception of the vision feed, the identification of the at least one model descriptor, and / orthe selection of theMLmodel may be performed by, e.g., the AI / ML model management component 197, the network processors) 1712, and / or the network interface 1780 of the network entity 1760 in FIG. 17.
[0169] In some implementations, to identify, based on the received vision feed, the at least one model descriptor from the set of model descriptors, the server may be configured to determine a site and a set of site-specific conditions based on the received vision feed, and select the at least one model descriptor from the set of model descriptors based on the determined site and the determined set of site-specific conditions. In some implementations, the set of model descriptors includes one or more of: a set of area coordinates in a local coordinate system, a set of area layout descriptors, a set of maps, a set of materials, a set of textures, or area occupancy information. In some implementations, to select, based on the identified at least one model descriptor, the ML model for the positioning of the U, the server maybe configured to select a most relevant ML model from a set of ML models that matches or approximately matches the identified at least one model descriptor. In some implementations, the identification of the at least one model descriptor from the set of model descriptors is based on at least one of RF-assisted descriptor matching or vision-assisted descriptor matching. In some implementations, the stationary camera is a closed-circuit television (CCTV) camera, and the non-stationary camera is a handheld (HH) camera.
[0170] FIG. 16 is a flowchart 1600 of a method of wireless communication. The methodmay be performed by a server (e.g., the one or more location servers 168; the at least one RF infrastructure 610; the positioning engine 616; the server 804, 904, 1104; network entity 1760). The method may enable the server to use vision-based techniques for managing RF-based AI / ML positioning models, thereby improving the overall performance of the RF-based AI / ML positioning.
[0171] At 1604, the server may perform positioning for a UE based on a ML model, such as described in connection with FIGs. 8 to 14. For example, as described in connection with 812 of FIG. 8, based on measuringRF signals from the UE 802 and / or based on the RF measurements / reports provided by the UE 802, the server 804 may estimate the position of the UE 802 (e.g., using at least one RF-based positioning mechanism). The positioning of the UE may be performed by, e.g., the AI / ML model management component 197, the network processor(s) 1712, and / or the network interface 1780 of the network entity 1760 in FIG. 17.
[0172] At 1606, the server may detect that at least one condition associated with the UE or an environment surrounding oftheUE changes or meets a threshold, such as described in connection with FIGs. 8 to 14. For example, as described in connection with 1002 of FIG. 10, the system described herein oran entity of the system (e.g., a UE, a server, or a wireless infrastructure, etc. depending on where this function is implemented) may be configured to detect change(s) / condition(s) that specifies the system or the entity of the system (hereafter “system / entity”) to change / select at least one new AI / ML model, and / or to train / fine-tune at least one new AI / ML model based on various types of information as discussed above. The detection of the at least one condition may be performed by, e.g., the AI / ML model management component 197, the network processor(s) 1712, and / or the network interface 1780 of the network entity 1760 in FIG. 17.
[0173] In one example, the at least one condition includes at least one of: a position of the UE is outside of a designated area over which the ML model is trained, a change in RF infrastructure associated with the UE, or a change in the environment surrounding the UE.
[0174] In another example, the detection of the at least one condition is based on at least one of: a positioning confidence, a set of visual cues and indicators, or RF-based positioning information.
[0175] At 1612, the server may estimate a firstposition of theUEbased on visual information associated with the UE, such as described in connection with FIGs. 8 to 14. For example, as described in connection with 1112 of FIG. 11, the visual information may be used (e.g., by the server 1104) for obtaining visual features associated with the set of training targets 1102, and / orforperformingvisualpositioningforthe set of training targets 1102 (e.g., to obtain a set of visual-based position estimates). The estimationof the first position of the UE may be performed by, e.g., the AI / ML model management component 197, the network processor(s) 1712, and / or the network interface 1780 of the network entity 1760 in FIG. 17.
[0176] In one example, the visual information corresponds to a set of visual features, a set of frames, or a set of processed detections.
[0177] At 1614, the server may select a set of RF measurements for the UE based on the estimated first position of the UE, such as described in connection with FIGs. 8 to 14. For example, as described in connection with 1114 of FIG. 14, the server 1104 may use the RF measurements to perform matching with the corresponding visual position estimates (e.g., ground truth locations). In other words, the server 1104 may match the RF measurements related to the set of training targets 1102 to their corresponding visual features visual-based position estimates. The selection of the set of RF measurements may be performed by, e.g., the AI / ML model management component 197, the network processor(s) 1712, and / or the network interface 1780 of the network entity 1760 in FIG. 17. In some implementations, to selectthe setofRF measurements for the UE based on the estimated first position of the UE, the server may be configured to estimate or receive a second position of the UE based on the set of RF measurements, and match the UE to the set of RF measurements based on mapping the first position to the second position. In some implementations, to match the UE to the set of RF measurements based on mappingthe firstposition to the secondposition, the server may be configured to compare a first sequence of time-stamped RF -based position estimates associated with the UE with a second sequence of time-stamped vision-based position estimates associated with the UE over a defined duration, and map the first position to the second position based on the comparison. In some implementation, the firstposition corresponds to a ground truth location of the UE, and the second position corresponds to a coarse position of the UE.
[0178] At 1616, the server may update the ML model based on the first position of the UE and the set of RF measurements, such as described in connection with FIGs. 8 to 14. For example, as described in connection with 1118 and 1120 of FIG. 11 , after the RF measurements are matched to their corresponding training targets (e.g., their ground truth locations obtained via visual positioning), the server 1104 may obtain a set of training data that includes the pair of information (e.g., ground truth locations and RF measurements). Then, the server 1104 may train / fine-tune (e.g., update) one or moreAI / ML models based on the set of training data. The positioning of the UE may be performed by, e.g., the AI / ML model management component 197, the network processor(s) 1712, and / orthe network interface 1780 of the network entity 1760 in FIG. 17. In some implementations, to update the ML model comprises at least one of: retrain the ML model, fine-tune the ML model, replace the ML model with a second ML model, or switch the ML model to the second ML model.
[0179] In one example, as shown at 1608, the server may transmit, to the UE based on the detection that the at least one condition changes or meets the threshold, a request to perform the set of RF measurements, and receive, from the UE based on the request, the set of RF measurements, such as described in connection with FIGs. 8 to 14. For example, as described in connection with 1110 and 1112 of FIG. 11, after a training / fine-tuning trigger event is detected, such as described in connection with 1010 of FIG. 10, a server 1104 (e.g. , the server 804, 904) maybe configuredto request RF measurement(s) from a set of training targets 1102 (e.g., may be UE 802, 902, a set of RF devices / UEs, etc.) and to request visual information from a set of cameras 1106 (e.g., may be a set of stationary cameras, a set of non-stationary cameras, or a combination of both). Based on the request from the server 1104, as shown at 1112, the set of training targets 1102 may provide their (fresh) RF measurements to the server 1104, such as describedin connection with 810 ofFIG. 8, and the set of cameras 1106 may also provide their visual information to the server 1104. The transmission of the request and / orthe reception of the set of RF measurements may be performed by, e.g., the AI / ML model management component 197, the network processors) 1712, and / or the network interface 1780 of the network entity 1760 in FIG. 17.
[0180] In another example, as shown at 1610, the server may transmit, to at least one camera based on the detection that the at least one condition changes or meets the threshold, a request to provide the visual information associated with the UE, and receive, from the at least one camera based on the request, the visual information, such as described in connection with FIGs. 8 to 14. For example, as described in connection with 1110 and 1112 of FIG. 11, after a training / fine-tuning trigger event is detected, such as described in connection with 1010 of FIG. 10, a server 1104 (e.g., the server 804, 904) may be configured to request RF measurement(s) from a set of training targets 1102 (e.g., may be UE 802, 902, a set of RF devices / UEs, etc.) and to request visual information from a set of cameras 1106 (e.g., may be a set of stationary cameras, a setof non-stationary cameras, or a combination of both). Based on the request from the server 1104, as shown at 1112, the set of training targets 1102 may provide their (fresh) RF measurements to the server 1104, such as describedin connection with 810 of FIG. 8, and the set of cameras 1106 may also provide their visual information to the server 1104. The transmission of the request and / or the reception of the visual information may be performed by, e.g., the AI / ML model management component 197, the network processor(s) 1712, and / or the network interface 1780 of the network entity 1760 in FIG. 17.
[0181] In another example, as shown at 1602, the server may request, based on an estimated position of the UE, a vision feed from at least one of the UE, a stationary camera, or a non-stationary camera, receive the vision feed from at least one of the UE, the stationary camera, or the non-stationary camera, identify , based on the received vision feed, at least one model descriptor from a set of model descriptors, and select, based on the identified at least one model descriptor, the ML model for the positioning of the UE, such as described in connection with FIGs. 8 to 14. For example, at 814, based on the estimated position of the UE 802, the server 804 may request one or more stationary cameras 806 (e.g., CCTV-camera(s)) to provide visual feed, such as images / videos captured by the field-of-view (FOV) of the one or more stationary cameras 806. At 816, in response to the request, the one or more stationary cameras 806 may provide their visual information (e.g., captured images / videos) to the server 804. At 818, using the visual information provided by the one or more stationary cameras 806, the server 804may perform model descriptor matching, where the server 804 may select a set of model descriptors that matches the visual information. At 830, based on the model descriptor matching performed at 818 and / or 828 (or based on the selected model descriptors), the server 804 may select at least one AI / ML positioning model that is to be used for estimating the position of a target (e.g., the UE 802, the targets 602, target(s) in proximity to the UE 802, etc.). Then, the server 804 may perform the positioning of the target(s) using the selected AI / ML positioning model(s). The transmission of the request, the reception of the vision feed, the identification of the at least one model descriptor, and / or the selection of the ML model may be performed by, e.g., the AI / ML model management component 197, the network processor(s) 1712, and / or the network interface 1780 of the network entity 1760 in FIG. 17.
[0182] In some implementations, to identify, based on the received vision feed, the at least one model descriptor from the set of model descriptors, the server may be configured to determine a site and a set of site-specific conditions based on the received vision feed, and select the at least one model descriptor from the set of model descriptors based on the determined site and the determined set of site-specific conditions. In some implementations, the set of model descriptors includes one or more of: a set of area coordinates in a local coordinate system, a set of area layout descriptors, a set of maps, a set of materials, a set of textures, or area occupancy information. In some implementations, to select, based on the identified at least one model descriptor, the ML model for the positioning of the U, the server maybe configured to select a most relevant ML model from a set of ML models that matches or approximately matches the identified at least one model descriptor. In some implementations, the identification of the at least one model descriptor from the set of model descriptors is based on at least one of RF-assisted descriptor matching or vision-assisted descriptor matching. In some implementations, the stationary camera is a CCTV camera, and the non-stationary camera is a HH camera.
[0183] FIG. 17 is a diagram 1700 illustrating an example of a hardware implementation for a network entity 1760 (e.g., a server). In one example, the network entity 1760 may be within the core network 120. The network entity 1760 may include at least one network processor 1712. The network processor(s) 1712 may include on-chip memory 1712'. In some aspects, the network entity 1760 may further include additional memory modules 1714. The network entity 1760 communicates via the network interface 1780 directly (e.g., backhaul link) orindirectly (e.g., through a RIC) with the CU 1702. The on-chip memory 1712' and the additional memory modules 1714 may each be considered a computer-readable medium / memory. Each computer-readable medium / memory may be non-transitory. The network processor(s) 1712 is responsible for general processing, including the execution of software stored on the computer-readable medium / memory. The software, when executed by the corresponding processor(s) causes the processor(s) to perform the various functions described supra. The computer-readable medium / memory may also be used for storing data that is manipulated by the processor(s) when executing software.
[0184] As discussed supra, the AI / ML model management component 197 may be configured to perform positioning for a UE based on a ML model. The AI / ML model management component 197 may also be configured to detect that at least one condition associated with the UE or an environment surrounding of the UE changes or meets a threshold. The AI / ML model management component 197 may also be configured to estimate a first position of the UE based on visual information associated with the UE. The AI / ML model management component 197 may also be configured to select a set of RF measurements for the UE based on the estimated first position of the UE. The AI / ML model management component 197 may also be configured to update the ML model based on the first position of the UE and the set of RF measurements. The AI / ML model management component 197 may be within the network processor(s) 1712. The AI / ML model management component 197 may be one or more hardware components specifically configured to carry out the stated processes / algorithm, implemented by one or more processors configured to perform the stated processes / algorithm, stored within a computer-readable medium for implementation by one or more processors, or some combination thereof. When multiple processors are implemented, the multiple processors may perform the stated processes / algorithm individually or in combination. The network entity 1760 may include a variety of components configured for various functions. In one configuration, the network entity 1760 may include means f or performing positioning for a UE based on a ML model. The network entity 1760 may further include means for detecting that at least one condition associated with the UE or an environment surrounding of the UE changes or meets a threshold. The network entity 1760 may further include means for estimating a first position of the UE based on visual information associated with the UE. The network entity 1760 may further include means for selecting a set of RF measurements for the UE based on the estimated first position of the UE. The network entity 1760 may further include means for updating the ML model based on the first position of the UE and the set of RF measurements.
[0185] In one configuration, the at least one condition includes at least one of : a position of the UE is outside of a designated area over which the ML model is trained, a change in RF infrastructure associated with the UE, or a change in the environment surrounding the UE.
[0186] In another configuration, the detection of the at least one condition is based on at least one of: a positioning confidence, a set of visual cues and indicators, or RF-based positioning information.
[0187] In another configuration, the visual information corresponds to a set of visual features, a set of frames, or a set of processed detections.
[0188] In another configuration, the means for selecting the set of RF measurements for the UE based on the estimated first position of the UE may include configuring the network entity 1760 to estimate or receive a second position of the UEbased on the set of RF measurements, and match the UE to the set of RF measurements based on mapping the first position to the second position. In some implementations, to match the UE to the set of RF measurements based on mapping the first position to the second position, the network entity 1760 may be configured to compare a first sequence of time-stamped RF-based position estimates associated with the UE with a second sequence of time-stamped vision-based position estimates associated with the UE over a defined duration, and map the first position to the second position based on the comparison. In some implementation, the first position corresponds to a ground truth location of the UE, and the second position corresponds to a coarse position of the UE.
[0189] In another configuration, the means forupdatingthe ML model comprises atleast one of: means for retraining the ML model, means for fine-tuning the ML model, means for replacing the ML model with a second ML model, or means for switching the ML model to the second ML model.
[0190] In another configuration, the network entity 1760 may further include means for transmitting, to the UE based on the detection that the at least one condition changes or meets the threshold, a request to perform the set of RF measurements, and means for receiving, from the UE based on the request, the set of RF measurements.
[0191] In another configuration, the network entity 1760 may further include means for transmitting, to at least one camera based on the detection that the at least one condition changes or meets the threshold, a request to provide the visual information associated with the UE, and means for receiving, from the at least one camera based on the request, the visual information.
[0192] In another configuration, the network entity 1760 may further include means for requesting, based on an estimated position of the UE, a vision feed from at least oneof the UE, a stationary camera, or a non-stationary camera, means for receiving the vision feed from at least one of the UE, the stationary camera, or the non-stationary camera, means for identifying, based on the received vision feed, at least one model descriptor from a set of model descriptors, and means for selecting, based on the identified at least one model descriptor, the ML model for the positioning of the UE. In some implementations, the means for identifying, based on the received vision feed, the at least one model descriptor from the set of model descriptors may include configuring the network entity 1760 to determine a site and a set of site-specific conditions based on the received vision feed, and select the at least one model descriptor from the set of model descriptors based on the determined site and the determined set of site-specific conditions. In some implementations, the set of model descriptors includes one or more of: a set of area coordinates in a local coordinate system, a set of area layout descriptors, a set of maps, a set of materials, a set of textures, or area occupancy information. In some implementations, the means for selecting, based on the identified at least one model descriptor, the ML model for the positioning of the U may include configuring the network entity 1760 to select a most relevant ML model from a set of ML models that matches or approximately matches the identified at least one model descriptor. In some implementations, the identification of the at least one model descriptor from the set of model descriptors is based on at least one of RF-assisted descriptor matching or vision-assisted descriptor matching. In some implementations, the stationary camera is a CCTV camera, and the non-stationary camera is a HH camera.
[0193] The means may be the AI / ML model management component 197 of the network entity 1760 configured to perform the functions recited by the means.
[0194] It is understood that the specific order or hierarchy of blocks in the processes / flowcharts disclosed is an illustration of example approaches. Based upon design preferences, it is understood that the specific order or hierarchy of blocks in the processes / flowcharts maybe rearranged. Further, some blocks may be combined or omitted. The accompanying method claims present elements of the various blocks in a sample order, and are not limited to the specific order or hierarchy presented.
[0195] The previous description is provided to enable any person skilled in the art to practice the various aspects described herein. Various modifications to these aspects will be readily apparent to those skilled in the art, and the generic principles defined hereinmay be applied to other aspects. Thus, the claims are not limited to the aspects described herein, but are to be accorded the full scope consistent with the language claims. Reference to an element in the singular does not mean “one and only one” unless specifically so stated, but rather “one or more.” Terms such as “if,” “when,” and “while” do not imply an immediate temporal relationship or reaction. That is, these phrases, e.g., “when,” do notimply an immediate action in response to or during the occurrence of an action, but simply imply that if a condition is met then an action will occur, butwithoutrequiringa specific or immediate time constraint for the action to occur. The word “exemplary” is used herein to mean “serving as an example, instance, or illustration.” Any aspect described herein as “exemplary” is not necessarily to be construed as preferred or advantageous over other aspects. Unless specifically stated otherwise, the term “some” refers to one or more. Combinations such as “at least one of A, B, or C,” “one or more of A, B, or C,” “at least one of A, B, and C,” “one or more of A, B, and C,” and “A, B, C, or any combination thereof’ include any combination of A, B, and / or C, and may include multiples of A, multiples of B, or multiples of C. Specifically, combinations such as “at least one of A, B, or C,” “one or more of A, B, or C,” “at least one of A, B, and C,” “one or more of A, B, and C,” and “A, B, C, or any combination thereof’ may be A only, B only, C only, A and B, A and C, B and C, or A and B and C, where any such combinations may contain one or more memb er or members of A, B, or C. Sets should b e interpreted as a set of elements where the elements number one or more. Accordingly, for a set of X, X would include one or more elements. When at least one processor is configured to perform a set of functions, the at least one processor, individually or in any combination, is configured to perform the set of functions. Accordingly, each processor of the at least one processor may be configured to perform a particular subset of the set of functions, where the subset is the full set, a proper subset of the set, or an empty subset of the set. A processor may be referred to as processor circuitry. A memory / memory module may be referred to as memory circuitry. If a first apparatus receives data from or transmits data to a second apparatus, the data may be received / transmitted directly between the first and second apparatuses, or indirectly between the first and second apparatuses through a set of apparatuses. A device configured to “output” data or “provide” data, such as a transmission, signal, or message, may transmit the data, for example with a transceiver, or may send thedata to a device that transmits the data. A device configured to “obtain” data, such as a transmission, signal, or message, may receive, for example with a transceiver, or may obtain the data from a device that receives the data. Information stored in a memory includes instructions and / or data. All structural and functional equivalents to the elements of the various aspects described throughout this disclosure that are known or later come to be known to those of ordinary skill in the art are expressly incorporated herein by reference and are encompassed by the claims. Moreover, nothing disclosed herein is dedicated to the public regardless of whether such disclosure is explicitly recited in the claims. The words “module,” “mechanism,” “element,” “device,” and the like may notbe a substitute forthe word “means.” As such, no claim element is to be construed as a means plus function unless the element is expressly recited using the phrase “means for.”
[0196] As used herein, the phrase “based on” shall notbe construed as a reference to a closed set of information, one or more conditions, one or more factors, or the like. In other words, the phrase “based on A” (where “A” may be information, a condition, a factor, or the like) shall be construed as “based at least on A” unless specifically recited differently.
[0197] The following aspects are illustrative only and may be combined with other aspects or teachings described herein, without limitation.
[0198] Aspect 1 is a method of wireless communication at a server, comprising: performing positioning for a user equipment (UE) based on a machine learning (ML) model; detecting that at least one condition associated with the UE or an environment surrounding of the UE changes or meets a threshold; estimating a first position of the UE based on visual information associated with the UE; selecting a set of radio frequency (RF) measurements for the UE based on the estimated first position of the UE; and updating the ML model based on the first position of the UE and the set of RF measurements.
[0199] Aspect 2 is the method of aspect 1, wherein selecting the set of RF measurements for the UE based on the estimated first position of the UE comprises: estimating or receiving a second position of the UE based on the set of RF measurements; and matching the UE to the set of RF measurements based on mapping the first position to the second position.
[0200] Aspect 3 is the method of aspect 1 or aspect 2, wherein matching the UE to the set of RF measurements based on mapping the first position to the second position comprises: comparing a first sequence of time-stamped RF-based position estimates associated with the UE with a second sequence of time-stamped vision-based position estimates associated with the UE over a defined duration; and mapping the first position to the second position based on the comparison.
[0201] Aspect 4 is the method of any of aspects 1 to 3, wherein the first position corresponds to a ground truth location of the UE, and wherein the second position corresponds to a coarse position of the UE.
[0202] Aspect 5 is the method of any of aspects 1 to 4, wherein updating the ML model comprises at least one of: retraining the ML model, fine-tuning the ML model, replacing the ML model with a second ML model, or switching the ML model to the second ML model.
[0203] Aspect 6 is the method of any of aspects 1 to 5, further comprising: transmitting, to the UE based on the detection that the at least one condition changes or meets the threshold, a request to perform the set of RF measurements; and receiving, from the UE based on the request, the set of RF measurements.
[0204] Aspect 7 is the method of any of aspects 1 to 6, further comprising: transmitting to at least one camera based on the detection that the at least one condition changes or meets the threshold, a request to provide the visual information associated with the UE; and receiving, from the at least one camera based on the request, the visual information.
[0205] Aspect 8 is the method of any of aspects 1 to 7, wherein the at least one condition includes atleast one of: a position oftheUE is outside of a designated area over which the ML model is trained, a change in RF infrastructure associated with the UE, or a change in the environment surrounding the UE.
[0206] Aspect 9 is the method of any of aspects 1 to 8, wherein the detection of the at least one condition is based on at least one of: a positioning confidence, a set of visual cues and indicators, or RF-based positioning information.
[0207] Aspect 10 is the method of any of aspects 1 to 9, wherein the visual information corresponds to a set of visual features, a set of frames, or a set of processed detections.
[0208] Aspect 11 is the method of any of aspects 1 to 10, further comprising: requesting based on an estimated position of the UE, a vision feed from at least one of the UE, astationary camera, or a non-stationary camera; receiving the vision feed from at least one of the UE, the stationary camera, or the non-stationary camera; identifying, based on the received vision feed, at least one model descriptor from a set of model descriptors; and selecting, based on the identified at least one model descriptor, the ML model for the positioning of the UE.
[0209] Aspect 12 is the method of any of aspects 1 to 11, wherein identifying, based on the received vision feed, the at least one model descriptor from the set of model descriptors comprises: determining a site and a set of site-specific conditions based on the received vision feed; and selecting the at least one model descriptor from the set of model descriptors based on the determined site and the determined set of sitespecific conditions.
[0210] Aspect 13 is the method of any of aspects 1 to 12, wherein the set of model descriptors includes one or more of: a set of area coordinates in a local coordinate system, a set of area layout descriptors, a set of maps, a set of materials, a set of textures, or area occupancy information.
[0211] Aspect 14 is the method of any of aspects 1 to 13, wherein selecting, based on the identified at least one model descriptor, the ML model for the positioning of the UE comprises: selectinga mostrelevantML model from a set of ML models that matches or approximately matches the identified at least one model descriptor.
[0212] Aspect 15 is the method of any of aspects 1 to 14, wherein the identification of the at least one model descriptor from the set of model descriptors is based on at least one of RF-assisted descriptor matching or vision-assisted descriptor matching.
[0213] Aspect 16 is the method of any of aspects 1 to 15, wherein the stationary camera is a closed-circuit television (CCTV) camera, and wherein the non-stationary camera is a handheld (HH) camera.
[0214] Aspect 17 is an apparatus for wireless communication at a server, including: atleast one memory; and atleast one processor coupled to the atleast one memory and, based at least in part on stored information that is stored in the at least one memory, the at least one processor, individually or in any combination, is configured to implement any of aspects 1 to 16.
[0215] Aspect 18 is the apparatus of aspect 17, further including at least one transceiver coupled to the at least one processor, wherein the at least one processor, individuallyor in any combination, is further configured to: transmit an indication of the updated the ML model via the at least one transceiver.
[0216] Aspect 19 is an apparatus for wireless communication at a server including means for implementing any of aspects 1 to 16.
[0217] Aspect 20 is a computer-readable medium (e.g., a non-transitory computer-readable medium) storing computer executable code, where the code when executed by a processor causes the processor to implement any of aspects 1 to 16.
Claims
CLAIMSWHAT IS CLAIMED IS:1 . An apparatus for wireless communication at a server, comprising: at least one memory; and at least one processor coupled to the at least one memory, the at least one processor, individually or in any combination, is configured to: perform positioning for a user equipment (UE) based on a machine learning (ML) model; detect that at least one condition associated with the UE or an environment surrounding of the UE changes or meets a threshold; estimate a first position of the UE based on visual information associated with the UE; select a set of radio frequency (RF) measurements for the UE based on the estimated first position of the UE; and update the ML model based on the first position of the UE and the set of RF measurements.
2. The apparatus of claim 1, wherein to select the set of RF measurements for the UE based on the estimated firstposition of the UE, the at least one processor, individually or in any combination, is configured to: estimate or receive a second position of the UE based on the set of RF measurements; and match the UE to the set of RF measurements based on mapping the first position to the second position.
3. The apparatus of claim 2, wherein to match the UE to the set of RF measurements based on mapping the first position to the second position, the at least one processor, individually or in any combination, is configured to: compare a first sequence of time-stamped RF-based position estimates associated with the UE with a second sequence of time-stamped vision-based position estimates associated with the UE over a defined duration; and map the first position to the second position based on comparison of the first sequence with the second sequence.
4. The apparatus of claim 2, wherein the first position corresponds to a ground truth location of the UE, and wherein the second position corresponds to a coarse position of the UE.
5. The apparatus of claim 1, wherein to update the ML model, the at least one processor, individually or in any combination, is configured to at least one of: retrain the ML model, fine-tune the ML model, replace the ML model with a second ML model, or switch the ML model to the second ML model.
6. The apparatus of claim 1, wherein the at least one processor, individually or in any combination, is further configured to: transmit, to the UE based on the detection that the at least one condition changes or meets the threshold, a request to perform the set of RF measurements; and receive, from the UE based on the request, the set of RF measurements.
7. The apparatus of claim 1, wherein the at least one processor, individually or in any combination, is further configured to: transmit, to at least one camera based on the detection that the at least one condition changes or meets the threshold, a request to provide the visual information associated with the UE; and receive, from the at least one camera based on the request, the visual information.
8. The apparatus of claim 1 , wherein the at least one condition includes at least one of: a position of the UE is outside of a designated area over which the ML model is trained, a change in RF infrastructure associated with the UE, or a change in the environment surrounding the UE.
9. The apparatus of claim 1, wherein the detection of the at least one condition is based on at least one of: a positioning confidence, a set of visual cues and indicators, orRF-based positioning information.
10. The apparatus of claim 1, wherein the visual information corresponds to a set of visual features, a set of frames, or a set of processed detections.11 . The apparatus of claim 1, wherein the at least one processor, individually or in any combination, is further configured to: request, based on an estimated position of the UE, a vision feed from at least one of the UE, a stationary camera, or a non-stationary camera; receive the vision feed from at least one of the UE, the stationary camera, or the non-stationary camera; identify, based on the received vision feed, at least one model descriptor from a set of model descriptors; and select, based on the identified at least one model descriptor, the ML model for the positioning of the UE.
12. The apparatus of claim 11, wherein to identify, based on the received vision feed, the at least one model descriptor from the set of model descriptors, the at least one processor, individually or in any combination, is configured to: determine a site and a set of site-specific conditions based on the received vision feed; and select the at least one model descriptor from the set of model descriptors based on the determined site and the determined set of site-specific conditions.
13. The apparatus of claim 11 , wherein the set of model descriptors includes one or more of: a set of area coordinates in a local coordinate system, a set of area layout descriptors, a set of maps,a set of materials, a set of textures, or area occupancy information.
14. The apparatus of claim 11, wherein to select, based on the identified atleast one model descriptor, the ML model for the positioning of the UE, the at least one processor, individually or in any combination, is configured to: select a most relevant ML model from a set of ML models that matches or approximately matches the identified at least one model descriptor.
15. The apparatus of claim 11, wherein identification of the at least one model descriptor from the set of model descriptors is based on at least one of RF-assisted descriptor matching or vision-assisted descriptor matching, wherein the stationary camera is a closed-circuit television (CCTV) camera, and wherein the non-stationary camera is a handheld (HH) camera.
16. The apparatus of claim 1, further comprising at least one transceiver coupled to the at least one processor, wherein the at least one processor, individually or in any combination, is further configured to: transmit an indication of the updated the ML model via the at least one transceiver.
17. A method of wireless communication at a server, comprising: performing positioning for a user equipment (UE) based on a machine learning (ML) model; detecting that at least one condition associated with the UE or an environment surrounding of the UE changes or meets a threshold; estimating a first position of the UE based on visual information associated with the UE; selecting a set of radio frequency (RF) measurements for the UE based on the estimated first position of the UE; and updating the ML model based on the first position of the UE and the set of RF measurements.
18. The method of claim 17, wherein selectingthe set of RF measurements fortheUE based on the estimated first position of the UE comprises: estimating or receiving a second position of the UE based on the set of RF measurements; and matching the UE to the set of RF measurements based on mapping the first position to the second position.
19. The method of claim 17, further comprising: requesting, based on an estimated position of the UE, a vision feed from at least one of the UE, a stationary camera, or a non-stationary camera; receiving the vision feed from at least one of the UE, the stationary camera, or the non-stationary camera; identifying, based on the received vision feed, at least one model descriptor from a set of model descriptors; and selecting, based on the identified at least one model descriptor, the ML model for the positioning of the UE.
20. A computer-readable medium storing computer executable code, the code when executed by at least one processor causes the at least one processor to: perform positioning for a user equipment (UE) based on a machine learning (ML) model; detect that at least one condition associated with the UE or an environment surrounding of the UE changes or meets a threshold; estimate a first position of the UE based on visual information associated with the UE; select a set of radio frequency (RF) measurements for the UE based on the estimated first position of the UE; and update the ML model based on the first position of the UE and the set of RF measurements.
Citation Information
Patent Citations
Location tracking
US20190333245A1
Positioning model training based on radio frequency fingerprint positioning (RFFP) measurements corresponding to position displacements
US20240133995A1
Training machine learning positioning models in a wireless communications network
WO2024027939A1
Machine learning (ML)-based measurements for uplink positioning
WO2024091340A1