Collaborative RF-assisted camera calibration for visual localization
By using radio frequency assisted camera calibration technology, the device positioning within the camera's field of view is improved by utilizing the RF position estimation of network devices. This solves the problems of insufficient accuracy and reliability in visual positioning and enables more efficient device positioning in network scenarios.
Patent Information
- Application Number
- CN202480033570.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-05-26
- Filing Date
- 2024-04-09
- Publication Date
- 2025-12-16
AI Technical Summary
Existing wireless communication systems lack accuracy and reliability in visual positioning, especially in network scenarios where it is difficult to effectively utilize the device position estimation within the camera's field of view.
By using radio frequency-assisted camera calibration technology, the location estimation of devices within the camera's field of view is improved by using RF-based position estimation of network devices as scene features for camera calibration and combining network assets for scene representation, thus achieving reliable localization of other devices, objects, and points of interest.
It improves the accuracy and reliability of visual positioning, provides privacy advantages, and avoids the limitations of traditional camera calibration methods.
Smart Images

Figure CN121152979A_ABST
Abstract
Description
[0001] Cross-references to related applications
[0002] This application claims the benefit of Greek patent application No. 20230100423, filed on May 26, 2023, entitled “COOPERATIVE RF-ASSISTED CAMERACALIBRATION FOR VISUAL POSITIONING”, the entire contents of which are expressly incorporated herein by reference. Technical Field
[0003] This disclosure relates generally to communication systems, and more specifically to wireless communication relating to visual positioning. Background Technology
[0004] Wireless communication systems are widely deployed to provide a variety of telecommunications services, such as telephone, video, data, messaging, and broadcasting. Typical wireless communication systems may employ multiple access technologies capable of supporting communication with multiple users by sharing available system resources. Examples of such multiple access technologies include Code Division Multiple Access (CDMA) systems, Time Division Multiple Access (TDMA) systems, Frequency Division Multiple Access (FDMA) systems, Orthogonal Frequency Division Multiple Access (OFDMA) systems, Single Carrier Frequency Division Multiple Access (SC-FDMA) systems, and Time Division Synchronous Code Division Multiple Access (TD-SCDMA) systems.
[0005] These multiple access technologies have been adopted in various telecommunications standards to provide a common protocol that enables different wireless devices to communicate at the city, national, regional, and even global levels. An example telecommunications standard is 5G New Radio (NR). 5G NR is part of the Continuous Evolution of Mobile Broadband (CEM) program issued by the 3rd Generation Partnership Project (3GPP) to meet new requirements associated with latency, reliability, security, scalability (e.g., with the Internet of Things (IoT),) and other requirements. 5G NR includes services associated with enhanced mobile broadband (eMBB), massive machine-type communications (mMTC), and ultra-reliable low-latency communications (URLLC). Some aspects of 5G NR can be based on the 4G Long Term Evolution (LTE) standard. Further improvements to 5G NR technology are needed. Furthermore, these improvements can also be applied to other multiple access technologies and telecommunications standards that adopt these technologies. Summary of the Invention
[0006] The following is a simplified overview of one or more aspects to provide a basic understanding of them. This overview is not a comprehensive summary of all hypothetical aspects. It neither identifies key or important elements of all aspects nor describes the scope of any or all aspects. Its sole purpose is to present some concepts of one or more aspects in a simplified form as a prelude to the more detailed descriptions that follow.
[0007] In one aspect of this disclosure, a method, computer-readable medium, and apparatus are provided. The apparatus receives from a network entity an indication of a training point set and at least one image of a region captured by at least one camera of at least one second network node, wherein the first network node is within the field of view (FOV) of the at least one camera of the at least one second network node, and wherein the indication for each training point in the training point set includes a pixel position and a world position in the at least one image. The apparatus estimates a projection transformation between the world position of an object and the pixel position of the object in the at least one image based on the training point set and the known or estimated position of the first network node.
[0008] In one aspect of this disclosure, a method, computer-readable medium, and apparatus are provided. The apparatus selects at least one second network node for a first network node including a known or estimated location, wherein the first network node is within the field of view (FOV) of the at least one second network node. The apparatus sends a request to the at least one second network node to capture at least one image of a region using at least one camera, wherein the at least one image includes the first network node. The apparatus receives the at least one image of the region from the at least one second network node based on the request.
[0009] To achieve the foregoing and related objectives, one or more aspects may include the features fully described below and specifically pointed out in the claims. The following description and drawings set forth some exemplary features of one or more aspects in detail. However, these features indicate only a few of the various ways in which the principles of the various aspects may be employed. Attached Figure Description
[0010] Figure 1 This is a diagram illustrating an example of a wireless communication system and an access network.
[0011] Figure 2A is an illustration of an example of the first frame according to various aspects of this disclosure.
[0012] Figure 2B is a diagram illustrating examples of downlink (DL) channels within a subframe according to various aspects of this disclosure.
[0013] Figure 2C is an illustration of an example of a second frame according to various aspects of this disclosure.
[0014] Figure 2D is a diagram illustrating examples of uplink (UL) channels within a subframe according to various aspects of this disclosure.
[0015] Figure 3 This is a diagram illustrating examples of base stations and user equipment (UEs) in an access network.
[0016] Figure 4 This is a diagram illustrating an example of UE positioning based on reference signal measurements.
[0017] Figure 5 This is an illustration illustrating examples of vision-based positioning according to various aspects of this disclosure.
[0018] Figure 6 This is an illustration of an example scenario in which a camera performs camera calibration using radio frequency (RF) features according to various aspects of this disclosure.
[0019] Figure 7 This is a diagram illustrating example projection transformations according to various aspects of this disclosure.
[0020] Figure 8 This is an illustration of an example of a fusion engine architecture associated with visual positioning according to various aspects of this disclosure.
[0021] Figure 9 This is a communication flow illustrating an example process of a UE-assisted visual positioning initiated by an RF-enabled cooperative network device (CpRFDev) according to various aspects of this disclosure.
[0022] Figure 10 This is a communication flow that illustrates an example process of CpRFDev initiating UE-based visual positioning according to various aspects of this disclosure.
[0023] Figure 11 This is a communication flow illustrating an example process of a network device (CamDev) with camera and RF capabilities initiating UE-assisted visual positioning according to various aspects of this disclosure.
[0024] Figure 12 This is a communication flow that illustrates an example process of CamDev initiating UE-based visual positioning according to various aspects of this disclosure.
[0025] Figure 13 This is a flowchart of a wireless communication method.
[0026] Figure 14 This is a flowchart of a wireless communication method.
[0027] Figure 15 These are illustrations illustrating specific hardware implementations used for example devices and / or network entities.
[0028] Figure 16 This is a flowchart of a wireless communication method.
[0029] Figure 17 This is a flowchart of a wireless communication method.
[0030] Figure 18 This is a diagram illustrating an example of a hardware implementation used for an example network entity. Detailed Implementation
[0031] The aspects presented in this paper improve the accuracy and reliability of vision-based localization, where devices (e.g., UEs, network nodes, etc.) can be configured to use radio frequency (RF)-based features for camera calibration (for performing visual localization). For example, in a network scenario, RF-based position estimates of various network devices (e.g., UEs, access points (APs), base stations / transmitter-receiver points (TRPs), etc.) within a camera's field of view (FOV) can be used as scene features for camera calibration. The aspects presented in this paper also improve RF-based position estimates of connected devices within a camera's FOV and enable the camera (or devices associated with the camera) to reliably estimate the positions of other devices, objects, and / or points of interest. The aspects presented in this paper can rely entirely on network assets for scene representation (e.g., network devices, RF localization engines, algorithms, servers, etc.), which can provide advantages over other camera calibration methods in terms of privacy.
[0032] The detailed descriptions following, illustrated with reference to the accompanying drawings, describe various configurations and do not represent the only configurations in which the concepts described herein can be practiced. To provide a thorough understanding of the various concepts, the detailed descriptions include specific details. However, these concepts can be practiced without these specific details. In some instances, well-known structures and components are shown in block diagram form to avoid obscuring these concepts.
[0033] Various apparatuses and methods are presented with reference to several aspects of a telecommunications system. These apparatuses and methods are described in detail below and illustrated in the accompanying drawings by various blocks, components, circuits, processes, algorithms, etc. (collectively, “elements”). These elements can be implemented using electronic hardware, computer software, or any combination thereof. Whether such elements are implemented as hardware or software depends on the specific application and the design constraints imposed on the system as a whole.
[0034] As an example, an element, any part of an element, or any combination of elements may be implemented as a “processing system” including one or more processors. Examples of processors include microprocessors, microcontrollers, graphics processing units (GPUs), central processing units (CPUs), application processors, digital signal processors (DSPs), reduced instruction set computing (RISC) processors, system-on-a-chip (SoCs), baseband processors, field-programmable gate arrays (FPGAs), programmable logic devices (PLDs), state machines, gate logic, discrete hardware circuits, and other suitable hardware configured to perform the various functionalities described throughout this disclosure. One or more processors in a processing system may execute software. Whether referred to as software, firmware, middleware, microcode, hardware description language, or other terms, software should be broadly interpreted as instructions, instruction sets, code, code segments, program code, programs, subroutines, software components, applications, software applications, software packages, routines, subroutines, objects, executable files, threads of execution, procedures, functions, or any combination thereof.
[0035] Therefore, in one or more example aspects, specific implementations, and / or use cases, the described functionality may be implemented in hardware, software, or any combination thereof. If implemented in software, the functionality may be stored or encoded as one or more instructions or code on a computer-readable medium. Computer-readable media include computer storage media. Storage media can be any available medium that can be accessed by a computer. By way of example, such computer-readable media may include random access memory (RAM), read-only memory (ROM), electrically erasable programmable ROM (EEPROM), optical disc storage devices, magnetic disk storage devices, other magnetic storage devices, combinations of these types of computer-readable media, or any other medium that can be used to store computer-executable code in the form of instructions or data structures accessible by a computer.
[0036] While aspects, implementations, and / or use cases are described herein by way of example, additional or different aspects, implementations, and / or use cases may arise in many different arrangements and scenarios. The aspects, implementations, and / or use cases described herein can be implemented across many different platform types, devices, systems, shapes, sizes, and package arrangements. For example, aspects, implementations, and / or use cases may arise via integrated chip implementations and other devices based on non-modular components (e.g., end-user devices, vehicles, communication devices, computing devices, industrial equipment, retail / purchasing devices, medical devices, AI-enabled devices, etc.). While some examples may or may not be specific to a use case or application, the described examples may exhibit broad applicability. Aspects, implementations, and / or use cases can range from chip-level or modular components to non-modular, non-chip-level implementations, and further to aggregated, distributed, or original equipment manufacturer (OEM) devices or systems incorporating one or more of the technologies described herein. In some practical settings, devices incorporating the described aspects and features may also include additional components and features for implementing and practicing the claimed and described aspects. For example, the transmission and reception of wireless signals necessarily involve multiple components for analog and digital purposes (e.g., hardware components including antennas, RF chains, power amplifiers, modulators, buffers, processors, interleavers, adders / summers, etc.). The techniques described herein can be practiced in a wide variety of devices, chip-level components, systems, distributed arrangements, aggregated or decomposed components, end-user equipment, etc., of various sizes, shapes, and configurations.
[0037] Communication systems, such as 5G NR systems, can be deployed in various ways with a variety of components or parts. In a 5G NR system or network, network nodes, network entities, network mobility elements, radio access network (RAN) nodes, core network nodes, network elements or network equipment (such as base stations (BS)), or one or more units (or components) performing base station functions can be implemented in aggregated or decomposed architectures. For example, BSs (such as Node B (NB), evolved NB (eNB), NR BS, 5G NB, access point (AP), transmit / receive point (TRP), or cell, etc.) can be implemented as aggregated base stations (also known as standalone BS or monolithic BS) or decomposed base stations.
[0038] Aggregated base stations can be configured to utilize a radio protocol stack that is physically or logically integrated within a single RAN node. Decentralized base stations can be configured to utilize a protocol stack that is physically or logically distributed across two or more units, such as one or more central or centralized units (CUs), one or more distributed units (DUs), or one or more radio units (RUs). In some respects, the CU can be implemented within a RAN node, and one or more DUs can be co-located with the CU, or alternatively, can be geographically or virtually distributed across one or more other RAN nodes. DUs can be implemented to communicate with one or more RUs. Each of the CU, DU, and RU can be implemented as a virtual unit, namely a virtual central unit (VCU), a virtual distributed unit (VDU), or a virtual radio unit (VRU).
[0039] Base station operation or network design can take into account the aggregation characteristics of base station functionality. For example, decomposed base stations can be utilized in Integrated Access Backhaul (IAB) networks, Open Radio Access Networks (O-RAN (such as network configurations advocated by the O-RAN Alliance)), or Virtualized Radio Access Networks (vRAN, also known as Cloud Radio Access Networks (C-RAN)). Decomposition can include distributing functionality across two or more units in various physical locations, as well as virtually distributing the functionality of at least one unit, which allows for flexibility in network design. The various units in a decomposed base station or decomposed RAN architecture can be configured to communicate wirelessly with at least one other unit.
[0040] Figure 1 Figure 100 illustrates an example of a wireless communication system and access network. The illustrated wireless communication system includes a decomposed base station architecture. The decomposed base station architecture may include one or more CUs 110, which may communicate directly with the core network 120 via a backhaul link, or indirectly with the core network 120 via one or more decomposed base station units, such as a near real-time (near-RT) RAN Intelligent Controller (RIC) 125 via an E2 link, or a non-real-time (non-RT) RIC 115 associated with a Service Management and Orchestration (SMO) framework 105, or both. CUs 110 may communicate with one or more DUs 130 via a corresponding midhaul link (such as an F1 interface). DUs 130 may communicate with one or more RUs 140 via a corresponding fronthaul link. RUs 140 may communicate with a corresponding UE 104 via one or more radio frequency (RF) access links. In some implementations, a UE 104 may be served simultaneously by multiple RUs 140.
[0041] Each of the units (i.e., CU 110, DU 130, RU 140, and near-RT RIC 125, non-RT RIC 115, and SMO frame 105) may include or be coupled to one or more interfaces configured to receive or transmit signals, data, or information (collectively, signals) via wired or wireless transmission media. Each of the units, or an associated processor or controller providing instructions to the communication interfaces of these units, may be configured to communicate with one or more other units via transmission media. For example, these units may include wired interfaces configured to receive signals or transmit signals to one or more other units via wired transmission media. Additionally, these units may include wireless interfaces that may include receivers, transmitters, or transceivers (such as RF transceivers) configured to receive signals via wireless transmission media and / or transmit signals to one or more other units.
[0042] In some aspects, the CU 110 can host one or more higher-level control functions. Such control functions may include Radio Resource Control (RRC), Packet Data Convergence Protocol (PDCP), Serving Data Adaptation Protocol (SDAP), etc. Each control function can be implemented using an interface configured to signal to other control functions hosted by the CU 110. The CU 110 can be configured to handle user plane functionality (i.e., Central Unit-User Plane (CU-UP)), control plane functionality (i.e., Central Unit-Control Plane (CU-CP)), or a combination thereof. In some implementations, the CU 110 can be logically divided into one or more CU-UP units and one or more CU-CP units. When implemented in an O-RAN configuration, the CU-UP units can communicate bidirectionally with the CU-CP units via an interface such as an E1 interface. The CU 110 can be implemented to communicate with the DU 130 for network control and signaling, as needed.
[0043] DU 130 may correspond to a logical unit that includes one or more base station functions for controlling the operation of one or more RU 140s. In some aspects, DU 130 may at least partially host one or more of the Radio Link Control (RLC) layer, Media Access Control (MAC) layer, and one or more high physical (PHY) layers (such as modules for forward error correction (FEC) encoding and decoding, scrambling, modulation, demodulation, etc.) according to functional splits (such as those defined by 3GPP). In some aspects, DU 130 may also host one or more low PHY layers. Each layer (or module) may be implemented using an interface configured to communicate signaling with other layers (and modules) hosted by DU 130 or with control functions hosted by CU 110.
[0044] Lower-layer functionality can be implemented by one or more RU 140s. In some deployments, an RU140 controlled by a DU 130 may correspond to a logical node that hosts RF processing functions or low-PHY layer functions (such as performing Fast Fourier Transform (FFT), Inverse FFT (iFFT), digital beamforming, Physical Random Access Channel (PRACH) extraction and filtering, or both, at least in part based on functional decomposition (such as lower-layer functional decomposition). In this architecture, the RU 140 may be implemented to handle over-the-air (OTA) communications with one or more UE 104s. In some specific implementations, the real-time and non-real-time aspects of control plane and user plane communications with the RU 140 may be controlled by the corresponding DU 130. In some scenarios, this configuration enables the implementation of the DU 130 and CU 110 in a cloud-based RAN architecture (such as a vRAN architecture).
[0045] SMO framework 105 can be configured to support RAN deployment and provisioning of both non-virtualized and virtualized network elements. For non-virtualized network elements, SMO framework 105 can be configured to support the deployment of dedicated physical resources for RAN coverage requirements, which can be managed via operation and maintenance interfaces such as the O1 interface. For virtualized network elements, SMO framework 105 can be configured to interact with a cloud computing platform such as Open Cloud (O-Cloud) 190 to perform network element lifecycle management (such as instantiating virtualized network elements) via a cloud computing platform interface such as the O2 interface. Such virtualized network elements may include, but are not limited to, CU 110, DU 130, RU 140, and near-RT RIC 125. In some implementations, SMO framework 105 can communicate with hardware aspects of the 4G RAN, such as Open eNB (O-eNB) 111, via the O1 interface. Additionally, in some implementations, SMO framework 105 can communicate directly with one or more RU 140s via the O1 interface. SMO framework 105 may also include a non-RT RIC 115 configured to support the functionality of SMO framework 105.
[0046] The non-RT RIC 115 can be configured to include logical functions enabling non-real-time control and optimization of RAN elements and resources, including artificial intelligence (AI) / machine learning (ML) workflows for model training and updates, or policy-based guidance for applications / features in the near-RT RIC 125. The non-RT RIC 115 can be coupled to or communicate with the near-RT RIC 125, such as via an A1 interface. The near-RT RIC 125 can be configured to include logical functions enabling near real-time control and optimization of RAN elements and resources via an interface, such as via an E2 interface, connecting one or more CU 110s, one or more DU 130s, or both, and O-eNBs to the near-RT RIC 125.
[0047] In some implementations, to generate AI / ML models to be deployed in the near-RT RIC 125, the non-RT RIC 115 may receive parameters or external enrichment information from an external server. This information can be utilized by the near-RT RIC 125 and can be received from non-network data sources or network functions at the SMO framework 105 or the non-RT RIC 115. In some examples, the non-RT RIC 115 or the near-RT RIC 125 may be configured to tune RAN behavior or performance. For example, the non-RT RIC 115 may monitor long-term trends and patterns in performance and use AI / ML models to perform corrective actions via the SMO framework 105 (such as reconfiguration via O1) or by creating RAN management policies (such as A1 policies).
[0048] At least one of CU 110, DU 130, and RU 140 may be referred to as base station 102. Therefore, base station 102 may include one or more of CU 110, DU 130, and RU 140 (each component is indicated by a dashed line to indicate that each component may or may not be included in base station 102). Base station 102 provides UE 104 with an access point to core network 120. Base station 102 may include macro cells (high-power cellular base stations) and / or small cells (low-power cellular base stations). Small cells include femtocells, picocells, and microcells. A network that includes both small cells and macro cells may be referred to as a heterogeneous network. A heterogeneous network may also include an evolved home node B (eNB) (HeNB), which can provide service to a restricted group referred to as a closed subscriber group (CSG). The communication link between RU 140 and UE 104 may include uplink (UL) transmission (also known as reverse link) from UE 104 to RU 140 and / or downlink (DL) transmission (also known as forward link) transmission from RU 140 to UE 104. The communication link may utilize multiple-input multiple-output (MIMO) antenna techniques, including spatial multiplexing, beamforming, and / or transmit diversity. The communication link may use one or more carriers. For each carrier allocated in a carrier aggregation of up to Yx MHz (x component carriers) for transmission in each direction, base station 102 / UE 104 may use a spectrum with a bandwidth of up to Y MHz (e.g., 5MHz, 10MHz, 15MHz, 20MHz, 100MHz, 400MHz, etc.). Carriers may be adjacent to each other or may not be adjacent to each other. Carrier allocation may be asymmetric for DL and UL (e.g., more or fewer carriers may be allocated to DL compared to UL). Component carriers may include primary component carriers and one or more secondary component carriers. The primary component carrier can be referred to as the primary cell (PCell) and the secondary component carrier can be referred to as the secondary cell (SCell).
[0049] Some UEs 104 can communicate with each other using device-to-device (D2D) communication link 158. D2D communication link 158 can use DL / UL wireless wide area network (WWAN) spectrum. D2D communication link 158 can use one or more sidelink channels, such as Physical Sidelink Broadcast Channel (PSBCH), Physical Sidelink Discovery Channel (PSDCH), Physical Sidelink Shared Channel (PSSCH), and Physical Sidelink Control Channel (PSCCH). D2D communication can be achieved through a variety of wireless D2D communication systems, such as Bluetooth. ® Wi-Fi based on the IEEE 802.11 standard ® LTE or NR.
[0050] The wireless communication system may also include a Wi-Fi AP 150, which communicates with the UE 104 (also referred to as a Wi-Fi station (STA)) via a communication link 154, for example, in an unlicensed spectrum such as 5 GHz. When communicating in unlicensed spectrum, the UE 104 / AP 150 may perform a free channel assessment (CCA) to determine whether the channel is available before communication.
[0051] The electromagnetic spectrum is typically subdivided into various categories, bands, channels, etc., based on frequency / wavelength. In 5G NR, two initial operating bands have been designated as frequency ranges FR1 (410MHz–7.125GHz) and FR2 (24.25GHz–52.6GHz). Although a portion of FR1 is greater than 6GHz, in various documents and articles, FR1 is often (interchangeably) referred to as the “sub-6GHz” band. Similar naming issues sometimes occur with FR2, which is often (interchangeably) referred to as the “millimeter wave” band in documents and articles, although this is distinct from the Extremely High Frequency (EHF) band (30GHz–300GHz) designated as a “millimeter wave” band by the International Telecommunication Union (ITU).
[0052] The frequencies between FR1 and FR2 are generally referred to as mid-band frequencies. Recent 5G NR studies have identified the operating bands used for these mid-band frequencies as the frequency range designation FR3 (7.125GHz–24.25GHz). Bands falling within FR3 can inherit FR1 and / or FR2 characteristics, thus effectively extending the features of FR1 and / or FR2 to mid-band frequencies. Furthermore, higher frequency bands are currently being explored to extend 5G NR operation beyond 52.6GHz. For example, three higher operating frequency bands have been identified as the frequency range designations FR2-2 (52.6GHz–71GHz), FR4 (71GHz–114.25GHz), and FR5 (114.25GHz–300GHz). Each of these higher frequency bands falls within the EHF band.
[0053] In view of the above, unless otherwise specified, the term "below 6 GHz" as used herein can broadly refer to frequencies less than 6 GHz, within FR1, or including intermediate frequency band frequencies. Furthermore, unless otherwise specified, the term "millimeter wave" as used herein can broadly refer to frequencies that can include intermediate frequency band frequencies, within FR2, FR4, FR2-2 and / or FR5, or within the EHF band.
[0054] Base station 102 and UE 104 may each include multiple antennas (such as antenna elements, antenna panels, and / or antenna arrays) to facilitate beamforming. Base station 102 may transmit beamformed signals 182 to UE 104 in one or more transmit directions. UE 104 may receive beamformed signals from base station 102 in one or more receive directions. UE 104 may also transmit beamformed signals 184 to base station 102 in one or more transmit directions. Base station 102 may receive beamformed signals from UE 104 in one or more receive directions. Base station 102 / UE 104 may perform beamforming training to determine the optimal receive and transmit directions for each of base station 102 / UE 104. The transmit and receive directions of base station 102 may be the same or different. The transmit and receive directions of UE 104 may be the same or different.
[0055] Base station 102 may include and / or be referred to as gNB, Node B, eNB, access point, base transceiver, radio base station, radio transceiver, transceiver function, basic service set (BSS), extended service set (ESS), TRP, network node, network entity, network equipment, or some other suitable terminology. Base station 102 may be implemented as an integrated access and backhaul (IAB) node, relay node, sidelink node, aggregated (monolithic) base station with baseband units (BBU) (including CU and DU) and RU, or may be implemented as a decomposed base station including one or more of CU, DU, and / or RU. A collection of base stations that may include decomposed base stations and / or aggregated base stations may be referred to as Next Generation (NG) RAN (NG-RAN).
[0056] The core network 120 may include Access and Mobility Management Function (AMF) 161, Session Management Function (SMF) 162, User Plane Function (UPF) 163, Unified Data Management (UDM) 164, one or more location servers 168, and other functional entities. AMF 161 is the control node that handles signaling between UE 104 and the core network 120. AMF 161 supports registration management, connection management, mobility management, and other functions. SMF 162 supports session management and other functions. UPF 163 supports packet routing, packet forwarding, and other functions. UDM 164 supports authentication and key agreement (AKA) credential generation, user identity processing, access authorization, and subscription management. One or more location servers 168 are exemplified as including a Gateway Mobile Location Center (GMLC) 165 and a Location Management Function (LMF) 166. However, generally, one or more location servers 168 may include one or more location / positioning servers, which may include one or more of GMLC 165, LMF 166, Position Determination Entity (PDE), Serving Mobile Location Center (SMLC), Mobile Location Center (MPC), etc. GMLC 165 and LMF 166 support UE location services. GMLC 165 provides an interface for clients / applications (e.g., emergency services) to access UE location information. LMF 166 receives measurement and auxiliary information from NG-RAN and UE 104 via AMF 161 to calculate the location of UE 104. NG-RAN may use one or more positioning methods to determine the location of UE 104. Positioning UE 104 may involve signal measurement, location estimation, and optional speed calculation based on these measurements. Signal measurement may be performed by UE 104 and / or base station 102 serving UE 104. The measured signals may be based on a satellite positioning system (SPS) 170 (e.g., one or more of the Global Navigation Satellite System (GNSS), Global Positioning System (GPS), Non-Terrestrial Network (NTN) or other satellite positioning / location systems), LTE signals, wireless local area network (WLAN) signals, Bluetooth, etc. ® One or more of the following: signal, terrestrial beacon system (TBS), sensor-based information (e.g., barometric pressure sensor, motion sensor), NR enhanced cell ID (NR E-CID) method, NR signal (e.g., multiple round trip time (multiple RTT), DL departure angle (DL-AoD), DL time difference of arrival (DL-TDOA), UL time difference of arrival (UL-TDOA) and UL angle of arrival (UL-AoA) positioning) and / or other systems / signals / sensors.
[0057] Examples of UE 104 include cellular phones, smartphones, Session Initiation Protocol (SIP) phones, laptops, personal digital assistants (PDAs), satellite radios, GPS devices, multimedia devices, video devices, digital audio players (e.g., MP3 players), cameras, game consoles, tablet devices, smart devices, wearable devices, vehicles, electricity meters, air pumps, large or small kitchen appliances, healthcare devices, implants, sensors / actuators, displays, or any other similarly functional device. Some UEs in UE 104 may be referred to as IoT devices (e.g., parking meters, air pumps, toasters, vehicles, heart monitors, etc.). UE 104 may also be referred to as a station, mobile station, subscriber station, mobile unit, subscriber unit, wireless unit, remote unit, mobile device, wireless device, wireless communication device, remote device, mobile subscriber station, access terminal, mobile terminal, wireless terminal, remote terminal, mobile phone, user agent, mobile client, client, or some other suitable terminology. In some scenarios, the term UE may also be applied to one or more companion devices, such as in a device constellation arrangement. One or more of these devices may access the network together and / or individually.
[0058] Refer again Figure 1 In some aspects, UE 104 may include a visual positioning component 198 (and / or base station 102 may include a visual positioning component 199), which may be configured to: receive from a network entity an indication of a training point set and at least one image of a region captured by at least one camera of at least one second network node, wherein the first network node is in the field of view (FOV) of the at least one camera of the at least one second network node, wherein the indication includes a pixel position and a world position in the at least one image for each training point in the training point set; and estimate a projection transformation between the world position of an object and the pixel position of the object in the at least one image based on the training point set and the known or estimated position of the first network node.
[0059] In some aspects, one or more location servers 168 may include a visual positioning coordination component 197, which may be configured to: select at least one second network node for a first network node including a known or estimated location, wherein the first network node is in the field of view (FOV) of the at least one second network node; send a request to the at least one second network node to capture at least one image of an area using at least one camera, wherein the at least one image includes the first network node; and receive the at least one image of the area from the at least one second network node based on the request.
[0060] Figure 2A is a diagram 200 illustrating an example of a first subframe within a 5G NR frame structure. Figure 2B is a diagram 230 illustrating an example of a DL channel within a 5G NR subframe. Figure 2C is a diagram 250 illustrating an example of a second subframe within a 5G NR frame structure. Figure 2D is a diagram 280 illustrating an example of a UL channel within a 5G NR subframe. The 5G NR frame structure can be Frequency Division Duplex (FDD) (where subframes within a specific set of subcarriers (carrier system bandwidth) are dedicated to either DL or UL) or Time Division Duplex (TDD) (where subframes within a specific set of subcarriers (carrier system bandwidth) are dedicated to both DL and UL). In the examples provided in Figures 2A and 2C, the 5G NR frame structure is assumed to be TDD, where subframe 4 is configured with slot format 28 (most of which are DL), where D is DL, U is UL, and F is flexible and can be used between DL / UL, and subframe 3 is configured with slot format 1 (all of which are UL). Although subframes 3 and 4 are shown as having slot formats 1 and 28 respectively, any particular subframe can be configured with any of the various available slot formats 0-61. Slot formats 0 and 1 are both DL and UL, respectively. Other slot formats 2-61 include a mixture of DL, UL, and flexible symbols. The slot format is configured for the UE via the received Slot Format Indicator (SFI) (dynamically configured via DL Control Information (DCI) or semi-statically / statically configured via Radio Resource Control (RRC) signaling). Note that the following description also applies to the 5G NR frame structure as TDD.
[0061] Figures 2A to 2D illustrate the frame structure, and aspects of this disclosure are applicable to other wireless communication technologies that may have different frame structures and / or different channels. A frame (10 ms) can be divided into 10 equal-sized subframes (1 ms). Each subframe may include one or more time slots. Subframes may also include micro-time slots, which may include 7, 4, or 2 symbols. Each time slot may include 14 or 12 symbols, depending on whether the cyclic prefix (CP) is normal or extended. For normal CP, each time slot may include 14 symbols, and for extended CP, each time slot may include 12 symbols. Symbols on the DL may be CP Orthogonal Frequency Division Multiplexing (OFDM) (CP-OFDM) symbols. Symbols on the UL may be CP-OFDM symbols (for high-throughput scenarios) or Discrete Fourier Transform (DFT) Extended OFDM (DFT-s-OFDM) symbols (for power-constrained scenarios; limited to single-stream transmission). The number of time slots within a subframe is based on the CP and parameter set. The parameter set defines the subcarrier spacing (SCS) (see Table 1). The symbol length / duration can be scaled with 1 / SCS.
[0062]
[0063] Table 1: Parameter Set, SCS, and CP
[0064] For a normal CP (14 symbols / slot), different parameter sets µ 0 through 4 allow 1, 2, 4, 8, and 16 slots per subframe, respectively. For an extended CP, parameter set 2 allows 4 slots per subframe. Therefore, for a normal CP and parameter set µ, there are 14 symbols / slot and 2... µ One time slot / subframe. Subcarrier spacing can be equal to ,in The parameter sets are 0 through 4. Therefore, the subcarrier spacing is 15 kHz for parameter set µ=0 and 240 kHz for parameter set µ=4. Symbol length / duration is negatively correlated with subcarrier spacing. Figures 2A through 2D provide examples of a normal CP with 14 symbols per slot and a parameter set µ=2 with 4 slots per subframe. The slot duration is 0.25 ms, the subcarrier spacing is 60 kHz, and the symbol duration is approximately 16.67 μs. Within a frame set, there may be one or more different bandwidth portions (BWPs) of frequency division multiplexing (see Figure 2B). Each BWP can have a specific parameter set and CP (normal or extended).
[0065] A resource grid can be used to represent the frame structure. Each time slot consists of a resource block (RB) extending for 12 consecutive subcarriers (also known as a physical RB (PRB)). The resource grid is divided into multiple resource elements (REs). The number of bits carried by each RE depends on the modulation scheme.
[0066] As illustrated in Figure 2A, some REs carry reference (pilot) signals (RS) for the UE. RS may include demodulation RS (DM-RS) (indicated as R for a particular configuration, but other DM-RS configurations are possible) and channel state information reference signals (CSI-RS) for channel estimation at the UE. RS may also include beam measurement RS (BRS), beam refinement RS (BRRS), and phase tracking RS (PT-RS).
[0067] Figure 2B illustrates examples of various DL channels within a subframe of a frame. The Physical Downlink Control Channel (PDCCH) carries the DCI within one or more Control Channel Elements (CCEs) (e.g., 1, 2, 4, 8, or 16 CCEs), each CCE comprising six RE Groups (REGs), each REG comprising 12 consecutive REs in the OFDM symbol of the RB. A PDCCH within a BWP can be referred to as a Control Resource Set (CORESET). The UE is configured to monitor PDCCH candidates in a PDCCH search space (e.g., a common search space, a UE-specific search space) during PDCCH monitoring timing on the CORESET, where the PDCCH candidates have different DCI formats and different aggregation levels. Additional BWPs may be located at higher and / or lower frequencies on the channel bandwidth. The Primary Synchronization Signal (PSS) may be located within symbol 2 of a specific subframe of the frame. The PSS is used by the UE 104 to determine subframe / symbol timing and physical layer identification. The Secondary Synchronization Signal (SSS) may be located within symbol 4 of a specific subframe of the frame. The SSS is used by the UE to determine the Physical Layer Cell Identifier Group Number and radio frame timing. Based on the Physical Layer Identifier and the Physical Layer Cell Identifier Group Number, the UE can determine the Physical Cell Identifier (PCI). Based on the PCI, the UE can determine the location of the DM-RS. The Physical Broadcast Channel (PBCH), carrying the Master Information Block (MIB), can be logically grouped with the PSS and SSS to form a Synchronization Signal (SS) / PBCH block (also known as an SS block (SSB)). The MIB provides the System Frame Number (SFN) and the number of Restricted Blocks (RBs) in the system bandwidth. The Physical Downlink Shared Channel (PDSCH) carries user data, broadcast system information not transmitted via the PBCH (such as System Information Blocks (SIBs)), and paging messages.
[0068] As illustrated in Figure 2C, some REs in the REs carry DM-RS (indicated as R for one specific configuration, but other DM-RS configurations are possible) for channel estimation at the base station. The UE can transmit DM-RS for the Physical Uplink Control Channel (PUCCH) and DM-RS for the Physical Uplink Shared Channel (PUSCH). The PUSCH DM-RS can be transmitted in the first or second symbol of the PUSCH. Depending on whether a short or long PUCCH is transmitted and depending on the specific PUCCH format used, the PUCCH DM-RS can be transmitted in different configurations. The UE can transmit a Sounding Reference Signal (SRS). The SRS can be transmitted in the last symbol of a subframe. The SRS can have a comb structure, and the UE can transmit the SRS on one of the comb teeth. The SRS can be used by the base station for channel quality estimation to enable frequency-dependent scheduling of the UL.
[0069] Figure 2D illustrates examples of various UL channels within a subframe of a frame. The PUCCH can be located as indicated in one configuration. The PUCCH carries uplink control information (UCI), such as scheduling requests, channel quality indicators (CQI), pre-decoding matrix indicators (PMI), rank indicators (RI), and hybrid automatic repeat request (HARQ) acknowledgment (ACK) (HARQ-ACK) feedback (i.e., one or more HARQ ACK bits indicating one or more ACKs and / or negative ACKs (NACKs)). The PUCCH carries data and may additionally be used to carry buffer status reports (BSR), power clearance reports (PHR), and / or UCI.
[0070] Figure 3 This is a block diagram illustrating communication between base station 310 and UE 350 in the access network. In the DL, Internet Protocol (IP) packets can be provided to controller / processor 375. Controller / processor 375 implements Layer 3 and Layer 2 functionality. Layer 3 includes the Radio Resource Control (RRC) layer, and Layer 2 includes the Service Data Adaptation Protocol (SDAP) layer, Packet Data Convergence Protocol (PDCP) layer, Radio Link Control (RLC) layer, and Media Access Control (MAC) layer. The controller / processor 375 provides RRC layer functionality associated with broadcasting system information (e.g., MIB, SIB), RRC connection control (e.g., RRC connection paging, RRC connection establishment, RRC connection modification, and RRC connection release), inter-Radio Access Technology (RAT) mobility, and measurement configuration for UE measurement reporting; PDCP layer functionality associated with header compression / decompression, security (encryption, decryption, integrity protection, integrity verification), and handover support functions; RLC layer functionality associated with the delivery of upper-layer packet data units (PDUs), error correction via ARQ, concatenation, segmentation, and reassembly of RLC service data units (SDUs), resegmentation of RLC data PDUs, and reordering of RLC data PDUs; and MAC layer functionality associated with mapping between logical channels and transport channels, multiplexing of MAC SDUs onto transport blocks (TBs), demultiplexing of MAC SDUs from TBs, scheduling information reporting, error correction via HARQ, priority handling, and logical channel priority ordering.
[0071] Transmit (TX) processor 316 and receive (RX) processor 370 implement Layer 1 functionality associated with various signal processing functions. Layer 1 (which includes the physical (PHY) layer) may include error detection on the transport channel, forward error correction (FEC) decoding / decoding of the transport channel, interleaving, rate matching, mapping to the physical channel, modulation / demodulation of the physical channel, and MIMO antenna processing. TX processor 316 processes the mapping to the signal constellation based on various modulation schemes (e.g., binary phase shift keying (BPSK), quadrature phase shift keying (QPSK), M-order phase shift keying (M-PSK), M-order quadrature amplitude modulation (M-QAM)). The decoded and modulated symbols can then be divided into parallel streams. Each stream can then be mapped to OFDM subcarriers, multiplexed with a reference signal (e.g., a pilot) in the time and / or frequency domains, and subsequently combined using inverse fast Fourier transform (IFFT) to produce a physical channel carrying a stream of time-domain OFDM symbols. The OFDM stream undergoes spatial pre-decoding to generate multiple spatial streams. Channel estimates from channel estimator 374 can be used to determine the decoding and modulation scheme, as well as for spatial processing. The channel estimates can be derived from reference signals transmitted by UE 350 and / or channel condition feedback. Each spatial stream can then be provided to different antennas 320 via a separate transmitter 318Tx. Each transmitter 318Tx can utilize the corresponding spatial stream to modulate a radio frequency (RF) carrier for transmission.
[0072] At UE 350, each receiver 354Rx receives signals via its corresponding antenna 352. Each receiver 354Rx recovers the information modulated onto the RF carrier and provides that information to the receive (RX) processor 356. The TX processor 368 and RX processor 356 implement Layer 1 functionality associated with various signal processing functions. The RX processor 356 can perform spatial processing on the information to recover any spatial stream destined for UE 350. If multiple spatial streams are destined for UE 350, the RX processor 356 can combine them into a single OFDM symbol stream. The RX processor 356 then uses a Fast Fourier Transform (FFT) to transform the OFDM symbol stream from the time domain to the frequency domain. The frequency domain signal consists of a separate OFDM symbol stream for each subcarrier of the OFDM signal. The symbols on each subcarrier, along with the reference signal, are recovered and demodulated by determining the most probable signal constellation points transmitted by base station 310. These soft decisions can be based on a channel estimate calculated by channel estimator 358. Subsequently, the soft decision is decoded and deinterleaved to recover the data and control signals originally transmitted by base station 310 on the physical channel. The data and control signals are then provided to controller / processor 359, which implements layer 3 and layer 2 functionality.
[0073] The controller / processor 359 may be associated with a memory 360 that stores program code and data. The memory 360 may be referred to as a computer-readable medium. In the UL, the controller / processor 359 provides demultiplexing, packet reassembly, decryption, header decompression, and control signal processing between transport and logical channels to recover IP packets. The controller / processor 359 is also responsible for error detection using ACK and / or NACK protocols to support HARQ operation.
[0074] Similar to the functionality described in conjunction with DL transmission performed by base station 310, controller / processor 359 provides RRC layer functionality associated with system information (e.g., MIB, SIB) acquisition, RRC connectivity, and measurement reporting; PDCP layer functionality associated with header compression / decompression and security (encryption, decryption, integrity protection, integrity verification); RLC layer functionality associated with upper-layer PDU delivery, error correction via ARQ, concatenation, segmentation and reassembly of RLC SDUs, resegmentation of RLC data PDUs, and reordering of RLC data PDUs; and MAC layer functionality associated with mapping between logical channels and transport channels, multiplexing of MAC SDUs onto TBs, demultiplexing of MAC SDUs from TBs, scheduling information reporting, error correction via HARQ, priority handling, and logical channel priority ordering.
[0075] The TX processor 368 can use the channel estimate derived from the reference signal or feedback transmitted by the channel estimator 358 from the base station 310 to select an appropriate decoding and modulation scheme and facilitate spatial processing. The spatial stream generated by the TX processor 368 can be provided to different antennas 352 via individual transmitters 354Tx. Each transmitter 354Tx can use the corresponding spatial stream to modulate an RF carrier for transmission.
[0076] UL transmission is processed at base station 310 in a manner similar to that described in conjunction with the receiver function at UE 350. Each receiver 318Rx receives signals via its corresponding antenna 320. Each receiver 318Rx recovers the information modulated onto the RF carrier and provides that information to RX processor 370.
[0077] The controller / processor 375 may be associated with a memory 376 that stores program code and data. The memory 376 may be referred to as a computer-readable medium. In UL, the controller / processor 375 provides demultiplexing, packet reassembly, decryption, header decompression, and control signal processing to recover IP packets between transport and logical channels. The controller / processor 375 is also responsible for error detection using ACK and / or NACK protocols to support HARQ operation.
[0078] At least one of the TX processor 368, RX processor 356, and controller / processor 359 can be configured to perform and Figure 1 The visual positioning component 198 combines various aspects.
[0079] At least one of the TX processor 316, RX processor 370, and controller / processor 375 can be configured to perform and Figure 1 The visual positioning component 199 combines various aspects.
[0080] Figure 4 Figure 400 illustrates an example of UE positioning (which may also be referred to as "network-based positioning") based on reference signal measurements according to various aspects of this disclosure. UE 404 can [operate at time T]. SRS_TX Send UL-SRS 412 and at time T PRS_RX Receives the DL positioning reference signal (PRS) (DL-PRS) 410. TRP 406 can be used at time T. SRS_RX Receive UL-SRS 412 and at time T PRS_TX Send DL-PRS 410. UE 404 may receive DL-PRS 410 before sending UL-SRS 412, or may send UL-SRS 412 before receiving DL-PRS 410. In both cases, the location server (e.g., location server 168) or UE 404 may base its response on ||T SRS_RX – T PRS_TX | – |T SRS_TX – T PRS_RX || to determine RTT 414. Therefore, multi-RTT positioning can utilize the UE Rx-Tx time difference measurement (i.e., |T) of downlink signals received from multiple TRPs 402, 406 and measured by UE 404. SRS_TX – T PRS_RX |) and DL-PRS reference signal received power (RSRP) (DL-PRS-RSRP), and the measured TRP Rx-Tx time difference measurement (i.e., |T) of the uplink signal transmitted from UE404 at multiple TRPs 402, 406. SRS_RX – T PRS_TX|) and UL-SRS-RSRP. UE 404 uses auxiliary data received from the location server to measure the UE Rx-Tx time difference (and / or the DL-PRS-RSRP of the received signal), and TRPs 402 and 406 use auxiliary data received from the location server to measure the gNB Rx-Tx time difference (and / or the UL-SRS-RSRP of the received signal). These measurements can be used at the location server or at UE 404 to determine the RTT, which is used to estimate the location of UE 404. Other methods for determining the RTT are possible, such as, for example, using DL-TDOA and / or UL-TDOA measurements.
[0081] PRS can be defined for network-based positioning (e.g., NR positioning) to enable the UE to detect and measure more neighboring transmit and receive points (TRPs), supporting various configurations for diverse deployments (e.g., indoor, outdoor, sub-6, mmW, etc.). Beam scanning can also be configured for PRS to support PRS beam operation. The UL positioning reference signal can be based on an enhanced / adjusted probe reference signal (SRS) for positioning purposes. In some examples, the UL-PRS may be referred to as "SRS for Positioning," and new information elements (IEs) can be configured for the SRS for positioning in RRC signaling.
[0082] DL PRS-RSRP can be defined as the linear average of the power contribution (in [W]) of a resource element carrying a DL PRS reference signal configured for RSRP measurement at an antenna port within the considered measurement frequency bandwidth. In some examples, for FR1, the reference point for DL PRS-RSRP can be the UE's antenna connector. For FR2, DL PRS-RSRP can be measured based on a combined signal from an antenna element corresponding to a given receiver branch. For FR1 and FR2, if the UE uses receiver diversity, the reported DL PRS-RSRP value can be no less than the corresponding DL PRS-RSRP of any individual receiver branch within the individual receiver branch. Similarly, UL SRS-RSRP can be defined as the linear average of the power contribution (in [W]) of a resource element carrying a sounding reference signal (SRS). UL SRS-RSRP can be measured by configured resource elements within the considered measurement frequency bandwidth at a configured measurement time. In some examples, for FR1, the reference point for UL SRS-RSRP can be the antenna connector of a base station (e.g., gNB). For FR2, the UL SRS-RSRP can be measured based on the combined signal from the antenna element corresponding to a given receiver branch. For FR1 and FR2, if the base station uses receiver diversity, the reported UL SRS-RSRP value may not be lower than the corresponding UL SRS-RSRP of any individual receiver branch within the individual receiver branch.
[0083] PRS-Path RSRP (PRS-RSRPP) can be defined as the power of the linear average of the channel response at the i-th path delay of a resource element carrying a DL PRS signal configured for measurement, where the DL PRS-RSRPP at the first path delay is the power contribution corresponding to the first detected path in time. In some examples, the PRS path phase measurement may refer to the phase associated with the i-th path of the channel derived using the PRS resource.
[0084] DL-AoD positioning utilizes the measured DL-PRS-RSRP of downlink signals received at UE 404 from multiple TRPs 402, 406. UE 404 uses auxiliary data received from the positioning server to measure the DL-PRS-RSRP of the received signals, and the resulting measurement, along with the azimuth departure (A-AoD), zenith departure (Z-AoD), and other configuration information, is used to position UE 404 relative to adjacent TRPs 402, 406.
[0085] DL-TDOA positioning utilizes the DL Reference Signal Time Difference (RSTD) (and / or DL-PRS-RSRP) of downlink signals received at UE 404 from multiple TRPs 402, 406. UE 404 uses auxiliary data received from the positioning server to measure the DL RSTD (and / or DL-PRS-RSRP) of the received signals, and the resulting measurement, along with other configuration information, is used to position UE 404 relative to adjacent TRPs 402, 406.
[0086] UL-TDOA positioning utilizes the UL relative time of arrival (RTOA) (and / or UL-SRS-RSRP) of the uplink signal transmitted from UE 404 at multiple TRPs 402, 406. TRPs 402, 406 use auxiliary data received from the positioning server to measure the UL-RTOA (and / or UL-SRS-RSRP) of the received signal, and the resulting measurements, along with other configuration information, are used to estimate the location of UE 404.
[0087] UL-AoA positioning utilizes the measured azimuth (A-AoA) and zenith (Z-AoA) of the uplink signal transmitted from UE 404 at multiple TRPs 402, 406. TRPs 402, 406 use auxiliary data received from a positioning server to measure the A-AoA and Z-AoA of the received signal, and the resulting measurements, along with other configuration information, are used to estimate the position of UE 404. For the purposes of this disclosure, a positioning operation in which the UE provides measurements to a base station / positioning entity / server for calculating the UE's position can be described as "UE-assisted," "UE-assisted positioning," and / or "UE-assisted position calculation," while a positioning operation in which the UE measures and calculates its own position can be described as "UE-based," "UE-based positioning," and / or "UE-based position calculation."
[0088] Additional positioning methods can be used to estimate the location of UE 404, such as, for example, UE-side UL-AoD and / or DL-AoA. It should be noted that data / measurements from various technologies can be combined in various ways to increase accuracy, determine and / or enhance certainty, supplement / improve measurements, and / or replace / provide missing information.
[0089] It should be noted that the terms "location reference signal" and "PRS" generally refer to specific reference signals used for positioning in NR and LTE systems. However, as used herein, the terms "location reference signal" and "PRS" can also refer to any type of reference signal that can be used for positioning, such as, but not limited to: PRS, tracking reference signal (TRS), PTRS, cell-specific reference signal (CRS), CSI-RS, DMRS, PSS, SSS, SSB, SRS, UL-PRS, etc., as defined in LTE and NR. Furthermore, the terms "location reference signal" and "PRS" can refer to downlink or uplink positioning reference signals, unless otherwise indicated by the context. To further distinguish the types of PRS, downlink positioning reference signals may be referred to as "DL PRS," and uplink positioning reference signals (e.g., SRS, PTRS used for positioning) may be referred to as "UL-PRS." Additionally, for signals that can be transmitted in both uplink and downlink (e.g., DMRS, PTRS), these signals may be prefixed with "UL" or "DL" to distinguish direction. For example, “UL-DMRS” can be distinguished from “DL-DMRS”.
[0090] In addition to positioning based on Global Navigation Satellite System (GNSS) and network-based positioning (e.g., combining...), Figure 4 In addition to those described, various vision-based localization methods have been developed to provide alternative / additional localization mechanisms / modes. Vision-based localization (which may also be referred to as "visual localization," "vision-based localization," "camera-based localization," and / or "camera-based visual localization") is a localization mechanism / mode that uses images captured by at least one camera to determine the location of a target (e.g., a UE, one or more objects within the field of view (FOV) of that at least one camera). For example, images captured by cameras in a warehouse can be used to calculate / estimate the location of inventory in the warehouse.
[0091] Figure 5Figure 500 illustrates examples of vision-based localization according to various aspects of this disclosure. Vision-based localization can provide highly accurate position estimation, where the coordinates (e.g., two-dimensional (2D) / three-dimensional (3D) coordinates, latitude and longitude coordinates, etc.) of a point of interest (e.g., the location of a UE, an object in a warehouse, etc.) can be calculated from the corresponding pixels of a 2D image using a projection transformation (inverse transformation) that maps world points onto the image. For example, as shown at 506, a projection transformation makes it possible to calculate the coordinates of an object 504 from the corresponding pixels of the object 504 in a 2D image 502, where the projection transformation maps the world point of the object 504 onto the 2D image 502. In some scenarios, the accuracy and reliability of the projection transformation can depend on the camera pose (e.g., the camera's position and orientation) and a set of internal camera parameters (e.g., focal length, resolution, image size, etc.). Therefore, knowing the projection transformation (inverse transformation) can be key to reliably / accurately locating a point of interest / object of interest within the camera's field of view (FOV).
[0092] In one example, the projection transformation can be learned by a device (e.g., UE, TRP, network node, etc.) based on a process called camera calibration. During camera calibration, training data can be provided to the device, which pairs 2D world points / 3D world points of one or more objects with known coordinates with the pixel locations of their corresponding image projections. For example, scene features can be extracted from an image and then matched with features obtained from a scene representation (e.g., a 3D model). In some scenarios, the camera calibration process may be unreliable if the extracted features are not descriptive (e.g., in an indoor environment, under low light, at night, etc.) or if a scene representation does not exist (e.g., a 3D model is not available for localization or it is impossible to access an external server for 3D model retrieval, etc.). Furthermore, in some use cases, the projection transformation may change frequently due to mobility or occasional internal camera adjustments (examples include mobile smartphones, virtual reality (VR) glasses, unmanned aerial vehicles (UAVs), autonomous vehicles, etc.).
[0093] The aspects presented in this paper can improve the accuracy and reliability of vision-based positioning, where devices (e.g., UEs, network nodes, etc.) can be configured to use radio frequency (RF) based features for camera calibration. For example, Figure 6Figure 600 illustrates an example scenario of a camera performing camera calibration using RF-based features according to various aspects of this disclosure. As shown at 604, in a network scenario, RF-based location estimates of various network devices (e.g., UEs, access points (APs), base stations / TRPs, etc.) within the FOV of camera 602 can be used as scene features for camera calibration. The aspects presented herein also improve RF-based location estimates of connected devices within the FOV of a camera and enable the camera (or devices associated with the camera) to reliably estimate the locations of other devices, objects, and / or points of interest. The aspects presented herein can rely entirely on network assets for scene representation (e.g., network devices, RF localization engines, algorithms, servers, etc.), which can provide advantages over other camera calibration methods in terms of privacy.
[0094] For the purposes of this disclosure, RF device may refer to any device capable of transmitting RF signals, making the device usable for location estimation. RF-based location estimation may refer to the use of at least one RF-related technology (such as Bluetooth). ® Wi-Fi ® Location estimation refers to the estimation of the location of an RF device using RF sensing and / or ultra-wideband (UWB) sensors. Location estimation can also refer to the estimated / approximate location of an RF device, the accuracy of the RF device's positioning within a threshold distance / radius, and / or within a specific threshold range (e.g., within a specific level of accuracy).
[0095] Figure 7 This is a diagram 700 illustrating example projection transformations according to various aspects of this disclosure. In one example, parameters that can be used to determine the projection of world points of one or more objects onto the image plane 702 may include extrinsic parameters (such as camera pose relative to the world coordinate system (WCS) (e.g., camera position and orientation)) and intrinsic parameters (such as camera focal length, size, resolution, skew and / or distortion, etc.).
[0096] Projective transformation can be a linear operation in homogeneous coordinates, with a size of 3 4 calibration matrix express:
[0097] and
[0098] in, It can indicate the world point of an object (e.g., the object's WCS), and This can indicate the pixel point of the object on the image plane 702, such as shown at 704. This is valid across the scale in homogeneous coordinates. If the world point... The pixel coordinates are determined by This indicates the use of homogeneous vectors. We can obtain the following equation:
[0099] , ,
[0100] in, It is a vector, and and It can indicate the first and second elements of the vector.
[0101] In some scenarios, camera calibration can be performed by minimizing geometric errors. For example, given a dataset of matching training points ( ):
[0102]
[0103] Camera calibration matrix The solution can be found as follows:
[0104]
[0105] in, This indicates the Euclidean distance between the non-homogeneous pixel coordinates and the projection of the world point onto the image plane 702. Minimization can be iterative, where... Suitable initial values can be derived by direct linear transformation (DLT) (which can be a standard procedure). The training data can first be normalized (e.g., to zero mean and fixed standard deviation), where normalization can be performed separately for pixel and world coordinates. In some examples, Random Sample Consensus (RANSAC), an iterative method for estimating a mathematical model from a dataset containing outliers, can be used to... Select to generate A subset of training points for the well-state / optimal / most suitable estimate.
[0106] As shown at point 706, once Given that the projection of world coordinates onto the ground plane (e.g., This can be calculated from the relevant pixels on the image plane 702. For example, the columns corresponding to the coordinates can be removed to obtain an invertible matrix. :
[0107]
[0108]
[0109]
[0110] Then, pixels The coordinates of the ground plane can be obtained based on the following equation:
[0111] .
[0112] In some examples, if the z-coordinate of an object is not on the ground plane, the following technique can be used to determine the z-coordinate of the object: detect the object (e.g., a bounding box) on the image, and then search for pixels on the image where the vertical projection of the bounding box intersects with the detected ground plane.
[0113] In one aspect of this disclosure, world points in the training dataset used for camera calibration can be derived from the 3D location of network devices within the camera's field of view (FOV). For example, in the case of UEs, VR glasses, vehicles, UAVs, and other mobile nodes, world points can be derived from estimated locations obtained via RF (e.g., via wireless communication). In the case of base station / TRPs (e.g., NBs), APs, or other fixed infrastructure nodes, world points can be derived from their actual locations. Furthermore, in addition to location servers (e.g., LMFs), four types of nodes can be used in conjunction with the camera calibration proposed herein.
[0114] The first type of node is a network device with a camera and RF capabilities (e.g., the ability to perform wireless communication), which may be referred to as "CamDev" for the purposes of this disclosure. CamDev can be an RF-enabled device with a camera, where images obtained from CamDev can be used to locate world points / objects of interest within the FOV of CamDev's camera. In most scenarios, the projection transformation of CamDev's camera is unknown; for example, CamDev could be a mobile UE, VR glasses, UAV, etc.
[0115] The second type of node is a cooperative network device with RF capabilities (e.g., the ability to perform wireless communication) participating in CamDev camera calibration; for the purposes of this disclosure, this network device may be referred to as "CpRFDev". CpRFDev may be an RF-enabled node whose world location (or an estimate of its world location) is known to a location server. Furthermore, CpRFDev is within the FOV of CamDev, and their image pixels and world location are cooperatively used for camera calibration.
[0116] The third type of node is a non-cooperative network device that does not participate in CamDev camera calibration. For the purposes of this disclosure, this network device may be referred to as "NonCpRFDev". NonCpRFDev can be any RF-enabled network device that does not participate in camera calibration. For example, NonCpRFDev can be a node defined as capturing a consensus concept among RF-enabled network devices. For instance, some devices / UEs (e.g., CpRFDev) may not have the ability or willingness to participate in the camera calibration process, as this may specify the analysis and fusion of their measurement results.
[0117] The fourth type of node is an RF-unenabled object (e.g., one that does not have the ability to perform wireless communication), which may be referred to as "NonRFobj" for the purposes of this disclosure. For example, NonRFobj can be any point of interest in the FOV of CamDev (e.g., a wall, obstacle, object, object that does not have RF communication capabilities, etc.).
[0118] The aspects presented in this paper enable the camera to be calibrated using data from the CpRFDev set and measured pixels from the CamDev image. This allows for improved position estimation of CpRFDev and / or the generation of position estimates for NonCpRFDev and NonRFobj with higher accuracy, potentially as a result of pixel information fusion. For example, at higher planes, during fusion, RF-based information can be fused with relevant pixels from the image, introducing the UE's position information from the pixels. Therefore, the UE can use the calibrated camera to obtain improved position estimates.
[0119] Figure 8 Figure 800 illustrates an example of a fusion engine architecture associated with visual positioning according to various aspects of this disclosure. In one example, as shown at 820, the fusion engine 810 may receive images / pixels captured by a camera of CamDev 802 (e.g., a UE with a camera) and the locations of one or more CpRFDev 804 (e.g., base stations / TRPs, APs, etc.) in the FOV of CamDev 802.
[0120] As shown at 830, the fusion engine 810 may include three main functional blocks. The first functional block 812 may be associated with data preparation and processing (e.g., as in combination). Figure 7(As described). In one example, the first functional block 812 may be responsible for process initiation (e.g., initialization of a visual localization process). Depending on the node type initiating the process (discussed in detail below), the fusion engine 810 may determine the relevant set of CamDev (e.g., CamDev 802) and / or CpRFDev (e.g., CpRFDev 804), and the fusion engine 810 may allocate resources (e.g., time and / or frequency resources) specified for information exchange (between different nodes), such as requests, responses, training data, and / or images. In some scenarios, the size of the UE (e.g., CpRFDev) may be small compared to network entities (e.g., base stations). Therefore, in order to detect the UE, the algorithm run by the fusion engine 810 may be configured to detect the object carrying the UE. For example, this could be a pedestrian or an asset in a warehouse. When generating a bounding box, a reference pixel may be selected (e.g., in the case of a pedestrian, the reference pixel could be near the pedestrian's hand or head). In another example, the first functional block 812 may also be responsible for pixel location extraction, where the fusion engine 810 may extract the pixel locations of CpRFDev 804, for example, using an object detection mechanism. In some examples, to determine the correspondence between multiple UEs detected in an image and their locations, certain RANSAC-based matching algorithms may be used for this purpose. Matching may be a process prior to camera calibration, and the algorithm may be configured to use off-the-shelf methods. In another example, the first functional block 812 may also be responsible for matching, where the fusion engine 810 may match the world location / coordinates (e.g., WCS coordinates) of CpRFDev 804 in the FOV of CamDev 802 with its corresponding image pixel locations. In another example, the first functional block 812 may also be responsible for data normalization (e.g., standardization, whitening, etc.), where the fusion engine 810 may normalize the world location and the image separately in each domain. In one aspect, the purpose of data normalization is to make the range of data variation equal in the world domain and the image domain, so that the estimation / calibration problem is adaptable. Standardization is commonly used for camera calibration. The Fusion Engine 810 may also have two domains, the first of which is a world coordinate system domain, in which position is measured in world coordinates (e.g., in meters), and the second of which is an image coordinate system domain, in which position is measured in pixel position on the image.
[0121] The second functional block 814 can be associated with camera calibration, such as combining... Figure 7As described. In one example, the second functional block 814 may be responsible for outlier rejection, where the fusion engine 810 may perform training point selection, such as using RANSAC. In another example, the second functional block 814 may also be responsible for projection transformation (calibration matrix) estimation. As shown at 820, during the calibration phase, a process (e.g., a visual positioning process) is initiated, a communication channel is established, and specified data is exchanged. Additionally, data may be prepared / preprocessed, and the camera may be calibrated.
[0122] The third functional block 816 can be associated with positioning, such as combining... Figure 7 As described. In one example, the third functional block 816 enables the fusion engine 810 to compute location updates for cooperative CpRFDev 804 and / or location estimates for NonCpRFDev 806 based on image pixel locations. As shown at 840, during the localization phase, world locations can be estimated from pixels and distributed to relevant devices / applications. The location information from CpRFDev can be used as training data updates. In some examples, the estimated location can be more accurate than the location in the training data, such as in the depth direction. This may be a result / consequence of fusing location information from relevant pixel locations on the image.
[0123] In one respect, the fusion engine 810 can be implemented in either a centralized or distributed mode. In the centralized mode (which may be referred to as UE-assisted visual localization), all functions (e.g., those performed by function blocks 812, 814, and 816) except for pixel extraction in privacy-conscious implementations can be implemented on a server (e.g., a location server, LMF, learning management system (LMS), etc.). The centralized mode is suitable for lightweight devices (e.g., UEs with reduced capabilities) because their processing power may be limited. In the distributed mode (which may be referred to as UE-based visual localization), all functions except for some performed by the first function block 812 (such as process initiation and matching) can be implemented on the UE side.
[0124] Combination Figure 8 The described visual localization process can be initiated by different types of nodes, such as CamDev (e.g., CamDev 802), CpRFDev (e.g., CpRFDev 804) and / or NonCpRFDev (e.g., NonCpRFDev 806), etc.
[0125] If the visual localization process is initiated by either CpRFDev or NonCpRFDev (hereinafter collectively referred to as "RFDev"), the RFDev (e.g., either CpRFDev or NonCpRFDev) may transmit a message requesting a position estimate to a server (e.g., a location server, LMF, etc.). NonCpRFDev can explicitly indicate through a dedicated field that it will not participate in camera calibration. For example, the default value of this dedicated field can be set to 1 for CpRFDev and 0 for NonCpRFDev. The remainder of the visual localization process remains the same for both CpRFDev and NonCpRFDev.
[0126] Using an approximate RF-based estimate of the CpRFDev's location, the server can locate the active CamDev in the area. For example, such information can be stored and retrieved from a database (e.g., UAV surveillance could be an example use case). Alternatively, an approximate location estimate can be used to determine the CamDev based on its proximity to the CpRFDev (e.g., extended reality (XR) / VR or smartphone applications could be an example use case). In some examples, because the server is configured to consider the CamDev's FOV when finding the appropriate CamDev, various techniques can be used to determine the CamDev's FOV. In one example, to determine the FOV of a CamDev (e.g., a mobile device such as a UE or XR / VR glasses), additional orientation information associated with the CamDev (e.g., from the associated IMU) can be used. By combining the CamDev's orientation information with the camera's relative position on the CamDev and coarse location information available via RF, an approximate FOV of the CamDev can be determined.
[0127] On the other hand, if the visual positioning process is initiated by a CamDev, the CamDev can send a message requesting its calibration matrix to a server (e.g., a location server, LMF, etc.). The server can then determine one or more nearby CpRFDevs based on the area covered by the CamDev. Similarly, the server can determine the one or more CpRFDevs based on information stored in a database (e.g., UAV surveillance could be an example use case). Alternatively, the server can also determine the one or more CpRFDevs based on their proximity to the CamDev, such as using approximate location estimation (e.g., extended reality (XR) / VR or smartphone applications could be an example use case).
[0128] Such as combination Figure 8Depending on where the fusion engine functions are performed (e.g., functions performed in association with function blocks 812, 814, and 816), visual positioning can be UE-assisted or UE-based. For UE-assisted visual positioning, after the visual positioning process is initiated, the server may send a ping request (e.g., notification, message, etc.) to the CamDev, which in turn may transmit an image or a list of relevant pixel locations from the CpRFDev to the server. In response, the server may then perform the remaining functions, such as processing the data, calibrating the camera using the RF-based location estimation from the CpRFDev, generating vision-based location estimates, and transmitting them to the initiating device. For UE-based visual positioning, after the visual positioning process is initiated, the server sends a ping request (e.g., notification, message, etc.) to the CamDev, which in turn may transmit an image or a list of relevant pixel locations to the server. In response, the server may then transmit a list of matching pairs between pixels and the CpRFDev to the initiating device, which may perform the remaining functions.
[0129] In another aspect of this disclosure, the aspects presented herein may also consider privacy concerns when CamDev is designated to share visual information with a server. In one example (or according to the first option), CamDev may be configured to transmit the entire image to the server, and the server may perform the function of extracting relevant pixels. Such a configuration may have the advantage of potentially more reliable matching due to the rich information present in the image. In another example (or according to the second option), CamDev may be configured to transmit only the location of relevant pixels to the server, where CamDev may perform the function of extracting relevant pixels. Such a configuration may have the advantage of increased privacy (e.g., the server has no access to the actual visual context that CamDev is observing). In some specific implementations, the matching function may be configured to be performed on the server, regardless of which option is adopted, because (1) matching may be computationally demanding, and (2) the server has access to more environmental information that can assist in matching relevant objects / pixels on the image with the RF-based locations of CpRFDev.
[0130] Figure 9 This is a communication flow 900 illustrating an example process of initiating UE-assisted visual positioning using CpRFDev according to various aspects of this disclosure. The numbers associated with communication flow 900 do not specify a particular time sequence and are used only as a reference to communication flow 900.
[0131] At 920, a CpRFDev 904 configured to perform or participate in a UE-assisted visual positioning session (e.g., for determining / estimating the location of one or more objects in an area) can send a location estimation request to server 906. (As in combination) Figure 8 The CpRFDev 904 discussed here can be a cooperative network device with RF capabilities participating in CamDev camera calibration (e.g., a UE, base station / TRP, or network node capable of performing wireless communication, etc.). Server 906 can be a location server or LMF. In one example, under UE-assisted visual localization, most functions (e.g., by...) are performed except for pixel extraction in a privacy-conscious implementation. Figure 8 The functions performed by function blocks 812, 814 and 816 can be implemented on server 906.
[0132] The location (or estimated location) of CpRFDev 904 may be known to server 906. If server 906 does not know the location of CpRFDev 904, server 906 may query CpRFDev 904 (or a database), and CpRFDev 904 may indicate its location to server 906 (or the database may provide the location to server 906). In some examples, if CpRFDev 904 is a mobile device (e.g., a UE), the location of CpRFDev 904 may be an approximate location (e.g., obtained by CpRFDev 904 via GNSS-based positioning, network-based positioning, etc.). If CpRFDev 904 is a fixed device (e.g., a base station, TRP, roadside unit (RSU), etc.), the location of CpRFDev 904 may be a true (e.g., precise) location.
[0133] In one example, at 922, in response to a location estimation request from CpRFDev 904, server 906 may send a query to database 908 to find a list of CamDevs surrounding CpRFDev 904. For example, database 908 may maintain a set of CamDevs with known locations, and database 908 may select the list of CamDevs surrounding CpRFDev 904 based on the location of the CamDev or its distance from CpRFDev 904. Then, at 924, database 908 may respond to the query by sending a list of CamDevs surrounding CpRFDev 904 (and the areas they cover) to server 906.
[0134] At 926, server 906 can determine the relevant CamDev 902 for participating in a UE-assisted visual positioning session. For example, server 906 can select a CamDev where CpRFDev 904 (and the one or more objects and / or areas to be detected) is within the FOV of CamDev 902. In some examples, to obtain (or be notified) the FOV of CamDev, the server can determine an approximate FOV based on reports from other devices. For example, in many mobile applications (such as smartphones and VR / XR), devices can report measurements from multiple sensors (such as RF modules and IMUs) that can be used to determine the direction the camera is looking. (As in combination...) Figure 8 As described, CamDev 902 can be a network device with at least a camera and RF capabilities (e.g., a UE, base station / TRP, or network node capable of performing wireless communications). For example, CamDev 902 can be an RF-enabled device with a camera, where images obtained from CamDev 902 can be used to locate world points (e.g., WCS coordinates, latitude and longitude coordinates, 2D / 3D coordinates, etc.) of one or more objects within the FOV of the camera of CamDev 902.
[0135] At 928, after server 906 determines that CamDev 902 will participate in UE-assisted visual positioning, server 906 may send an image request to CamDev 902 to request CamDev 902 to take an image (e.g., capture the FOV of CamDev 902).
[0136] At 930, based on the image request, CamDev 902 may capture one or more images (e.g., based on their FOV). In some implementations, CamDev 902 may also be configured to extract relevant pixels from the captured images. Relevant pixels may refer to pixels / features that may be useful / descriptive for visual localization, such as pixels of one or more objects captured by CamDev 902 and pixels of CpRFDev 904. In some examples, relevant pixels may be obtained by removing irrelevant pixels, such as non-descriptive pixels / features, such as pixels / features that are unlikely to be useful for visual localization, such as background, natural objects (e.g., sun, sky, clouds, etc.), reflections, blurred objects, etc. Then, at 932, CamDev 902 may transmit the captured images and / or the extracted relevant pixels (if performed by CamDev 902) to server 906.
[0137] At 934, based on the image from CamDev 902 and / or the extracted relevant pixels, server 906 can perform data preparation and processing on the image and / or the extracted relevant pixels, such as combining... Figure 8 The first functional block 812 is described. For example, data preparation and processing may include:
[0138] (1) The visual positioning process is initiated.
[0139] (2) Determine the relevant CamDev and / or CpRFDev for visual positioning (e.g., if not performed in other steps).
[0140] (3) Allocate the resources specified for information exchange between different nodes (e.g., between CamDev, CpRFDev and the server) (e.g., if not performed in other steps).
[0141] (4) Pixel location extraction (e.g., extracting pixel locations of CpRFDev 904, such as using an object detection mechanism).
[0142] (5) Matching (e.g., matching the world position (e.g., WCS / 2D / 3D coordinates) of CpRFDev 904 in the FOV of CamDev 902 with its corresponding image pixel position).
[0143] (6) Data normalization (e.g., standardization, whitening, etc.), or
[0144] (7) Their combination.
[0145] At position 936, server 906 can perform camera calibration based on data preparation and processing, such as combining... Figure 8 As described in the second functional block 814. For example, camera calibration may include outlier rejection (e.g., performing training point selection, such as using RANSAC) and / or projection transformation (calibration matrix) estimation, such as in combination with... Figures 5 to 7 As described.
[0146] At 938, after performing camera calibration, server 906 can perform visual localization (e.g., UE-assisted visual localization) on one or more objects in the captured image (or in the extracted relevant pixels), such as combining... Figure 8The third functional block 816 is described. For example, visual localization may include calculating a location update for CpRFDev 904 (and / or a location estimate for NonCpRFDev) based on image pixel locations, estimating the world location of one or more objects from pixels, and / or distributing the estimated locations of the one or more objects to relevant devices / applications. For example, at 940, server 906 may send the results of visual localization to CpRFDev 904, such as location estimates of one or more objects detected within the FOV of CamDev 902 (e.g., estimated locations of CpRFDev, NonCpRFDev, NonRFobj, etc.).
[0147] Combination Figure 9 The described UE-assisted visual positioning is applicable to CpRFDev with limited (or low) processing capabilities, such as UEs with reduced capabilities and lightweight devices.
[0148] Figure 10 This is a communication flow 1000 illustrating an example process for initiating UE-based visual positioning using CpRFDev according to various aspects of this disclosure. The numbers associated with the communication flow 1000 do not specify a particular time sequence and are used only as a reference for the communication flow 1000.
[0149] At 1020, CpRFDev 1004, configured to perform or participate in a UE-based visual positioning session (e.g., for determining / estimating the location of one or more objects in an area), can send a location estimation request to server 1006. (As in combination) Figure 8 The CpRFDev 1004 discussed here can be a cooperative network device with RF capabilities participating in CamDev camera calibration (e.g., a UE, base station / TRP, or network node capable of performing wireless communication, etc.). Server 1006 can be a location server or LMF. In one example, under UE-based visual positioning, most functions associated with visual positioning can be implemented on (e.g., performed by) CpRFDev 1004. For example, besides those performed by… Figure 8 Apart from certain functions performed by the first function block 812 (e.g., initiating and matching), the remaining functions performed by function blocks 812, 814, and 816 can be implemented on CpRFDev 1004.
[0150] The location (or estimated location) of CpRFDev 1004 may be known to server 1006. If server 1006 does not know the location of CpRFDev 1004, server 1006 may query CpRFDev 1004 (or a database), and CpRFDev 1004 may indicate its location to server 1006 (or the database may provide the location of CpRFDev 1004 to server 1006). In some examples, if CpRFDev 1004 is a mobile device (e.g., a UE), the location of CpRFDev 1004 may be an approximate location (e.g., obtained by CpRFDev 1004 via GNSS-based positioning, network-based positioning, etc.). If CpRFDev 1004 is a fixed device (e.g., a base station, TRP, roadside unit (RSU), etc.), the location of CpRFDev 1004 may be a true (e.g., precise) location.
[0151] In one example, at 1022, in response to a location estimation request from CpRFDev 1004, server 1006 may send a query to database 1008 to find a list of CamDevs surrounding CpRFDev 1004. For example, database 1008 may maintain a set of CamDevs with known locations, and database 1008 may select the list of CamDevs surrounding CpRFDev 1004 based on the location of the CamDev or its distance to CpRFDev 1004. Then, at 1024, database 1008 may send the list of CamDevs surrounding CpRFDev 1004 (and the areas they cover) to server 1006 in response to the query.
[0152] At 1026, server 1006 can determine the relevant CamDev 1002 for participating in the UE-based visual positioning session. For example, server 1006 can select CamDev where CpRFDev 1004 (and the one or more objects and / or regions to be detected) is in the FOV of CamDev 1002. (As in combination...) Figure 8 As described, CamDev 1002 can be a network device with at least a camera and RF capabilities (e.g., a UE, base station / TRP, or network node capable of performing wireless communications). For example, CamDev 1002 can be an RF-enabled device with a camera, wherein images obtained from CamDev 1002 can be used to locate world points (e.g., WCS coordinates, latitude and longitude coordinates, 2D / 3D coordinates, etc.) of one or more objects within the FOV of the camera of CamDev 1002.
[0153] At 1028, after server 1006 determines that CamDev 1002 is to participate in UE-based visual positioning, server 1006 may send an image request to CamDev 1002 to request CamDev 1002 to take an image (e.g., capture the FOV of CamDev 1002).
[0154] At 1030, based on the image request, CamDev 1002 may capture one or more images (e.g., based on their FOV). In some implementations, CamDev 1002 may also be configured to extract relevant pixels from the captured images. Relevant pixels may refer to pixels / features that may be useful / descriptive for visual localization, such as pixels of one or more objects captured by CamDev 1002 and pixels of CpRFDev 1004. In some examples, relevant pixels may be obtained by removing irrelevant pixels, such as non-descriptive pixels / features, such as pixels / features that are unlikely to be useful for visual localization, such as background, natural objects (e.g., sun, sky, clouds, etc.), reflections, blurred objects, etc. Then, at 1032, CamDev 1002 may transmit the captured images and / or the extracted relevant pixels (if performed by CamDev 1002) to server 1006.
[0155] At 1034, if CamDev 1002 is not configured to perform the extraction of relevant pixels from the captured image, server 1006 can be configured to perform pixel extraction instead. Server 1006 can then perform a matching of the positions of one or more objects in the captured image (or in the extracted relevant pixels) with their pixel positions. For example, server 1006 can match the world position (e.g., WCS / 2D / 3D coordinates) of CpRFDev 1004 (or multiple CpRFDevs) in the FOV of CamDev 1002 with its corresponding image pixel position. If CamDev 1002 is configured to perform the extraction of relevant pixels from the captured image, server 1006 can simply perform the matching. Then, at 1036, server 1006 can send a set of training points (from the combined...) to CpRFDev 1004. Figure 7 (obtained from the described matching) and the captured image (or the relevant pixels extracted).
[0156] At position 1038, based on training points from server 1006 and the captured image (or extracted relevant pixels), CpRFDev 1004 can perform data preparation and processing on the image and / or extracted relevant pixels, such as combining... Figure 8 The first functional block 812 is described. For example, data preparation and processing may include:
[0157] (1) Determine the relevant CamDev and / or CpRFDev for visual positioning (e.g., if not performed in other steps).
[0158] (2) Allocate the resources specified for information exchange between different nodes (e.g., between CamDev, CpRFDev and the server) (e.g., if not performed in other steps).
[0159] (3) Pixel location extraction (e.g., extracting pixel locations of CpRFDev 1004, such as using an object detection mechanism).
[0160] (4) Data normalization (e.g., standardization, whitening, etc.), or
[0161] (5) Their combination.
[0162] At 1040, the CpRFDev 1004 can perform camera calibration based on data preparation and processing, such as combining... Figure 8 As described in the second functional block 814. For example, camera calibration may include outlier rejection (e.g., performing training point selection, such as using RANSAC) and / or projection transformation (calibration matrix) estimation, such as in combination with... Figures 5 to 7 As described.
[0163] At 1042, after performing camera calibration, CpRFDev 1004 can perform visual localization (e.g., UE-based visual localization) on one or more objects in the captured image (or in the extracted relevant pixels), such as combining... Figure 8 The third functional block 816 is described. For example, visual localization may include calculating a location update (and / or a location estimate of NonCpRFDev) of CpRFDev 1004 based on image pixel locations, estimating the world location of one or more objects from pixels, and / or distributing the estimated locations of the one or more objects to relevant devices / applications. For example, at 1044, CpRFDev 1004 may send the results of visual localization to server 1006 (if specified), such as location estimates of one or more objects detected within the FOV of CamDev 1002 (e.g., estimated locations of CpRFDev, NonCpRFDev, NonRFobj, etc.).
[0164] Combination Figure 10 The described UE-based visual localization is applicable to CpRFDev with high processing capabilities, such as TRPs, base stations, or high-performance UEs.
[0165] Figure 11This is a communication flow 1100 illustrating an example process of CamDev initiating UE-assisted visual positioning according to various aspects of this disclosure. The numbers associated with communication flow 1100 do not specify a particular time sequence and are used only as a reference for communication flow 1100.
[0166] At 1120, the CamDev 1102, configured to perform or participate in a UE-assisted visual positioning session (e.g., for determining / estimating the position of one or more objects in the FOV of the CamDev 1102), can send a camera calibration request to server 1106. (As in combination) Figure 8 As described, CamDev 1102 can be a network device with at least a camera and RF capabilities (e.g., a UE, base station / TRP, or network node capable of performing wireless communication). For example, CamDev 1102 can be an RF-enabled device with a camera, where images obtained from CamDev 1102 can be used to locate world points (e.g., WCS coordinates, latitude and longitude coordinates, 2D / 3D coordinates, etc.) of one or more objects within the FOV of CamDev 1102's camera. Server 1106 can be a location server or LMF. In one example, under UE-assisted visual localization, most functions (e.g., by...) are performed except for pixel extraction in a privacy-conscious implementation. Figure 8 The functions performed by function blocks 812, 814 and 816 can be implemented on server 1106.
[0167] In one example, at 1122, in response to a camera calibration request from CamDev 1102, server 1106 may send a query to database 1108 to find the area covered by CamDev 1102. For example, database 1108 may maintain a set of areas with different CpRFDev distributions, and database 1108 may select a list of areas around CamDev 1102 based on the location of the areas or their distance from CamDev 1102. Then, at 1124, database 1108 may, in response to the query, send a list of areas around CamDev 1102 (e.g., regions of interest) (and the associated CpRFDev in each area) to server 1106.
[0168] At 1126, server 1106 can determine the set of relevant CpRFDev 1104 used to participate in the UE-assisted visual positioning session. For example, server 1106 can select the set of CpRFDev in the FOV of CamDev 1102. (As in combination) Figure 8The set of relevant CpRFDev 1104 discussed can be cooperative network devices with RF capabilities (e.g., UEs, base stations / TRPs, or network nodes capable of performing wireless communication) that can participate in CamDev 1102 camera calibration.
[0169] The location (or estimated location) of the set of relevant CpRFDev 1104 can be known to server 1106. If server 1106 does not know the location of CpRFDev, server 1106 can query CpRFDev, and CpRFDev can indicate its location to server 1106. In some examples, if CpRFDev is a mobile device (e.g., UE), the location of CpRFDev can be an approximate location (e.g., obtained via GNSS-based positioning, network-based positioning, etc.). If CpRFDev is a fixed device (e.g., base station, TRP, roadside unit (RSU), etc.), the location of CpRFDev can be a true (e.g., precise) location.
[0170] At 1128, server 1106 may send an image request to CamDev 1102 to request CamDev 1102 to take an image (e.g., capture the FOV of CamDev 1102).
[0171] At 1130, based on the image request, CamDev 1102 may capture one or more images (e.g., based on their FOV). In some implementations, CamDev 1102 may also be configured to extract relevant pixels from the captured images. Relevant pixels may refer to pixels / features that may be useful / descriptive for visual localization, such as pixels of one or more objects captured by CamDev 1102 and the set of relevant CpRFDev 1104. In some examples, relevant pixels may be obtained by removing irrelevant pixels, such as non-descriptive pixels / features, such as pixels / features that are unlikely to be useful for visual localization, such as background, natural objects (e.g., sun, sky, clouds, etc.), reflections, blurred objects, etc. Then, at 1132, CamDev 1102 may transmit the captured images and / or the extracted relevant pixels (if performed by CamDev 1102) to server 1106.
[0172] At 1134, based on the image from CamDev 1102 and / or the extracted relevant pixels, server 1106 can perform data preparation and processing on the image and / or the extracted relevant pixels, such as combining... Figure 8 The first functional block 812 is described. For example, data preparation and processing may include:
[0173] (1) The visual positioning process is initiated.
[0174] (2) Determine the relevant CamDev and / or CpRFDev for visual positioning (e.g., if not performed in other steps).
[0175] (3) Allocate the resources specified for information exchange between different nodes (e.g., between CamDev, CpRFDev and the server) (e.g., if not performed in other steps).
[0176] (4) Pixel location extraction (e.g., extracting pixel locations from a set of related CpRFDev 1104 data, such as using an object detection mechanism).
[0177] (5) Matching (e.g., matching the world location (e.g., WCS / 2D / 3D coordinates) of the set of related CpRFDev 1104 in the FOV of CamDev 1102 with its corresponding image pixel location).
[0178] (6) Data normalization (e.g., standardization, whitening, etc.), or
[0179] (7) Their combination.
[0180] At position 1136, server 1106 can perform camera calibration based on data preparation and processing, such as combining... Figure 8 As described in the second functional block 814. For example, camera calibration may include outlier rejection (e.g., performing training point selection, such as using RANSAC) and / or projection transformation (calibration matrix) estimation, such as in combination with... Figures 5 to 7 As described. Then, at 1138, server 1106 can send the calibration matrix (associated with or obtained from camera calibration) to CamDev 1102.
[0181] At 1140, based on the calibration matrix from server 1106, CamDev 1102 can perform visual localization (e.g., UE-assisted visual localization) on one or more objects in the captured image (or in the extracted relevant pixels), such as combining... Figure 8The third functional block 816 is described. For example, visual localization may include calculating location updates (and / or location estimates of NonCpRFDev) of a set of related CpRFDev 1104 based on image pixel locations, estimating the world location of one or more objects from pixels, and / or distributing the estimated locations of the one or more objects to related devices / applications. For example, at 1142, CamDev 1102 may send the results of visual localization to server 1106, such as location estimates of one or more objects detected within the FOV of CamDev 1102 (e.g., estimated locations of CpRFDev, NonCpRFDev, NonRFobj, etc.). In some examples, at 1144, server 1106 may also send / forward the location estimates of the one or more objects detected within the FOV of CamDev 1102 to a set of related CpRFDev 1104.
[0182] Combination Figure 11 The described UE-assisted visual positioning is applicable to CamDev devices with limited (or low) processing capabilities, such as UEs with reduced capabilities and lightweight devices.
[0183] Figure 12 This is a communication flow 1200 illustrating an example process of CamDev initiating UE-based visual positioning according to various aspects of this disclosure. The numbers associated with communication flow 1200 do not specify a particular time sequence and are used only as a reference for communication flow 1200.
[0184] At 1220, the CamDev 1202, configured to perform or participate in a UE-based visual positioning session (e.g., for determining / estimating the position of one or more objects in the FOV of the CamDev 1202), can send a camera calibration request to server 1206. (As in combination) Figure 8 As described, CamDev 1202 can be a network device with at least a camera and RF capabilities (e.g., a UE, base station / TRP, or network node capable of performing wireless communication, etc.). For example, CamDev 1202 can be an RF-enabled device with a camera, where images obtained from CamDev 1202 can be used to locate world points (e.g., WCS coordinates, latitude and longitude coordinates, 2D / 3D coordinates, etc.) of one or more objects within the FOV of the camera of CamDev 1202. Server 1206 can be a location server or LMF. In one example, under UE-based visual positioning, most functions associated with visual positioning can be implemented (e.g., performed by) CamDev 1202. For example, in addition to those performed by... Figure 8Apart from certain functions performed by the first function block 812 (e.g., initiating and matching), the remaining functions performed by function blocks 812, 814, and 816 can be implemented on CamDev 1202.
[0185] In one example, at 1222, in response to a camera calibration request from CamDev 1202, server 1206 may send a query to database 1208 to find the area covered by CamDev 1202. For example, database 1208 may maintain a set of areas with different CpRFDev distributions, and database 1208 may select a list of areas around CamDev 1202 based on the location of the areas or their distance from CamDev 1202. Then, at 1224, database 1208 may, in response to the query, send a list of areas around CamDev 1202 (e.g., regions of interest) (and the associated CpRFDev in each area) to server 1206.
[0186] At 1226, server 1206 can determine the set of relevant CpRFDev 1204 used to participate in the UE-based visual positioning session. For example, server 1206 can select the set of CpRFDev in the FOV of CamDev 1202. (As in combination) Figure 8 The set of relevant CpRFDev 1204 discussed can be cooperative network devices with RF capabilities (e.g., UEs, base stations / TRPs, or network nodes capable of performing wireless communication) that can participate in CamDev 1202 camera calibration.
[0187] The location (or estimated location) of the set of relevant CpRFDev 1204 can be known to server 1206. If server 1206 does not know the location of CpRFDev, server 1206 can query CpRFDev, and CpRFDev can indicate its location to server 1206. In some examples, if CpRFDev is a mobile device (e.g., UE), the location of CpRFDev can be an approximate location (e.g., obtained via GNSS-based positioning, network-based positioning, etc.). If CpRFDev is a fixed device (e.g., base station, TRP, roadside unit (RSU), etc.), the location of CpRFDev can be a true (e.g., precise) location.
[0188] At 1228, server 1206 may send an image request to CamDev 1202 to request CamDev 1202 to take an image (e.g., capture the FOV of CamDev 1202).
[0189] At 1230, based on the image request, CamDev 1202 may capture one or more images (e.g., based on their FOV). In some implementations, CamDev 1202 may also be configured to extract relevant pixels from the captured images. Relevant pixels may refer to pixels / features that are potentially useful / descriptive for visual localization, such as pixels of one or more objects captured by CamDev 1202 and the set of relevant CpRFDev 1204. In some examples, relevant pixels may be obtained by removing irrelevant pixels, such as non-descriptive pixels / features, such as pixels / features that are unlikely to be useful for visual localization, such as background, natural objects (e.g., sun, sky, clouds, etc.), reflections, blurred objects, etc. Then, at 1232, CamDev 1202 may transmit the captured images and / or the extracted relevant pixels (if performed by CamDev 1202) to server 1206.
[0190] At 1234, if CamDev 1202 is not configured to perform the extraction of relevant pixels from the captured image, server 1206 may be configured to perform pixel extraction instead. Server 1206 may then perform matching of the positions of one or more objects in the captured image (or in the extracted relevant pixels) with their pixel positions. For example, server 1206 may match the world positions (e.g., WCS / 2D / 3D coordinates) of a set (or multiple CpRFDevs) of relevant CpRFDev 1204 in the FOV of CamDev 1202 with their corresponding image pixel positions. If CamDev 1202 is configured to perform the extraction of relevant pixels from the captured image, server 1206 may simply perform the matching. Then, at 1236, server 1206 may send a set of training points (from the combined...) to CamDev 1202... Figure 7 (obtained from the described match).
[0191] At point 1238, based on training points from server 1206 and images (or extracted relevant pixels) captured by CamDev 1202, CamDev 1202 can perform data preparation and processing on the images and / or the extracted relevant pixels, such as combining... Figure 8 The first functional block 812 is described. For example, data preparation and processing may include:
[0192] (1) Determine the relevant CamDev and / or CpRFDev for visual positioning (e.g., if not performed in other steps).
[0193] (2) Allocate the resources specified for information exchange between different nodes (e.g., between CamDev, CpRFDev and the server) (e.g., if not performed in other steps).
[0194] (3) Pixel location extraction (e.g., extracting pixel locations from a set of related CpRFDev 1204 data, such as using an object detection mechanism).
[0195] (4) Data normalization (e.g., standardization, whitening, etc.), or
[0196] (5) Their combination.
[0197] At 1240, the CamDev 1202 can perform camera calibration based on data preparation and processing, such as combining... Figure 8 As described in the second functional block 814. For example, camera calibration may include outlier rejection (e.g., performing training point selection, such as using RANSAC) and / or projection transformation (calibration matrix) estimation, such as in combination with... Figures 5 to 7 As described.
[0198] At 1242, after performing camera calibration, CamDev 1202 can perform visual localization (e.g., UE-based visual localization) on one or more objects in the captured image (or in the extracted relevant pixels), such as combining... Figure 8 The third functional block 816 is described. For example, visual localization may include calculating location updates (and / or location estimates of NonCpRFDev) of a set of related CpRFDev 1204 based on image pixel locations, estimating the world location of one or more objects from pixels, and / or distributing the estimated locations of the one or more objects to related devices / applications. For example, at 1244, CamDev 1202 may send the results of visual localization to server 1206 (if specified), such as location estimates of one or more objects detected within the FOV of CamDev 1202 (e.g., estimated locations of CpRFDev, NonCpRFDev, NonRFobj, etc.). In some examples, at 1246, server 1206 may also send / forward the location estimates of the one or more objects detected within the FOV of CamDev 1202 to a set of related CpRFDev 1204.
[0199] Combination Figure 12 The described UE-based visual localization is applicable to CamDev systems with high processing capabilities, such as TRPs, base stations, or high-performance UEs.
[0200] This paper presents various aspects of vision-based localization that provide high-accuracy position estimation, where the coordinates of a point of interest (such as the location of an object in a UE or warehouse) can be calculated from corresponding pixels using the inverse transformation of a projection transformation that maps world points onto a 2D image. The projection transformation can depend on the camera pose (position + orientation) and a set of camera parameters, and knowing the projection transformation (known as camera calibration) can be key to reliably locating an object of interest within the camera's field of view (FOV). In one aspect, RF-based position estimation of network devices (UE, AP, NB, etc.) within the camera's FOV can be used as scene features for camera calibration. The fusion engine receives images / pixels from the camera devices and the positions of RF devices within the FOV. In various aspects, the fusion engine can be centralized (UE-assisted vision localization) or distributed (UE-based vision localization).
[0201] Figure 13 This is a flowchart 1300 of a wireless communication method. The method may be implemented by a first network node (e.g., UE 104, 404; camera 602; CpRFDev 804, 904, 1004, 1104, 1204); device 1504 (hereinafter referred to as...). Figure 15 The method, as discussed below, enables the first network node (e.g., the UE) to initiate UE-assisted visual localization or UE-based visual localization by utilizing camera calibration performed on a network device with known location and RF capabilities.
[0202] At 1304, the first network node can receive from the network entity an indication of a set of training points and at least one image of a region captured by at least one camera of at least one second network node, wherein the first network node is within the FOV of the at least one camera of the at least one second network node, and wherein the indication for each training point in the set of training points includes the pixel location and world location in the at least one image, such as combining... Figure 10 As described. For example, at 1036, CpRFDev 1004 can receive a set of training points from server 1006 (from the combination) Figure 7 The image set (or extracted relevant pixels) obtained from the described matching and the region. The at least one image was captured by a CamDev 1002, and the CpRFDev 1004 is within the FOV of the CamDev. The reception of the training point set and the at least one image can be, for example, by... Figure 15 The device 1504 uses a visual positioning component 198, a camera 1532, an application processor 1506, a cellular baseband processor 1524, and / or a transceiver 1522 to perform the operation.
[0203] In one example, as shown at 1302, a first network node may send a request to a network entity to estimate the location of one or more objects in a region using vision-based localization, wherein indications of a training point set and at least one image may be received from the network entity based on the request, such as combining... Figure 10 As described. For example, at 1020, CpRFDev 1004, configured to perform or participate in a UE-based visual positioning session (e.g., for determining / estimating the location of one or more objects in an area), can send a location estimation request to server 1006. The sending of the request can be, for example, by Figure 15 The device 1504 uses a visual positioning component 198, a camera 1532, an application processor 1506, a cellular baseband processor 1524, and / or a transceiver 1522 to perform the operation.
[0204] In another example, at 1306, the first network node may extract the pixel position of the first network node from the at least one image; match the world position of the first network node with the pixel position of the first network node; and perform data normalization on the training point set or the at least one image based on the world position and the pixel position of the first network node, such as combining... Figure 10 As described. For example, at 1038, based on training points from server 1006 and the captured image (or extracted relevant pixels), CpRFDev 1004 can perform data preparation and processing on the image and / or the extracted relevant pixels, such as combining... Figure 8 The first functional block 812 is described. For example, data preparation and processing may include: (1) determining relevant CamDev and / or CpRFDev for visual localization, (2) allocating resources specified for information exchange between different nodes (e.g., between CamDev, CpRFDev and the server), (3) pixel location extraction (e.g., extracting pixel locations of CpRFDev 1004, such as using an object detection mechanism), (4) data normalization (e.g., standardization, whitening, etc.), or (5) combinations thereof. The extraction of pixel locations of the first network node, the matching of the world location of the first network node with the pixel locations of the first network node, and / or the data normalization of the training point set or the at least one image may be performed by, for example Figure 15 The device 1504 utilizes a visual positioning component 198, a camera 1532, an application processor 1506, a cellular baseband processor 1524, and / or a transceiver 1522 to perform this function. In some implementations, the pixel position of the first network node is extracted using object detection.
[0205] At 1308, the first network node can estimate the projection transformation between the world position of the object and the pixel position of the object in the at least one image based on the training point set and the known or estimated position of the first network node, such as combining... Figure 10 As described. For example, at 1040, CpRFDev 1004 can perform camera calibration (e.g., projection transformation) based on data preparation and processing, such as combining... Figure 8 As described in the second functional block 814. For example, camera calibration may include outlier rejection (e.g., performing training point selection, such as using RANSAC) and / or projection transformation (calibration matrix) estimation, such as in combination with... Figures 5 to 7 As described. Projection transformation can be, for example, Figure 15 The device 1504 comprises a visual positioning component 198, a camera 1532, an application processor 1506, a cellular baseband processor 1524, and / or a transceiver 1522. In some embodiments, to estimate the projection transformation, the first network node may select a subset of training points from the training point set and generate a camera calibration matrix based on that subset. In some embodiments, the subset of training points may be selected from the training point set based on Random Sample Consensus (RANSAC).
[0206] At 1310, the first network node can estimate the location of one or more objects in the region based on the at least one image and projection transformation, such as by combining... Figure 10 As described. For example, at 1042, after performing camera calibration, CpRFDev 1004 can perform visual localization (e.g., UE-based visual localization) on one or more objects in the captured image (or in the extracted relevant pixels), such as combining Figure 8 The third functional block 816 describes the estimation of the location of the one or more objects in the region, for example... Figure 15 The device 1504 utilizes a visual positioning component 198, a camera 1532, an application processor 1506, a cellular baseband processor 1524, and / or a transceiver 1522 to perform this function. In some embodiments, to estimate the position of the one or more objects in the region based on the at least one image and projection transformation, the first network node may calculate the image pixel position of the one or more objects based on the at least one image. In some embodiments, the one or more objects in the region are within the FOV of the at least one camera of the at least one second network node.
[0207] In one example, the first network node could be a base station or a TRP with a known location.
[0208] In another example, the first network node may be a UE that includes an estimated location. In some implementations, the estimated location of the first network node may be obtained based on network-based positioning or GNSS-based positioning.
[0209] In another example, the one or more objects may be non-radio frequency (non-RF) objects, or the one or more objects may not include RF capabilities.
[0210] In another example, vision-based localization is based on the user's visual location.
[0211] At point 1312, the first network node can send an indication to the network entity of the estimated location of one or more objects in the region, such as combining... Figure 10 As described. For example, at 1044, CpRFDev 1004 can send the results of visual positioning to server 1006 (if specified), such as the estimated positions of one or more objects detected within the FOV of CamDev 1002 (e.g., estimated positions of CpRFDev, NonCpRFDev, NonRFobj, etc.). The transmission of indications of the estimated positions of these one or more objects in the area can be, for example, by... Figure 15 The device 1504 uses a visual positioning component 198, a camera 1532, an application processor 1506, a cellular baseband processor 1524, and / or a transceiver 1522 to perform the operation.
[0212] Figure 14 This is a flowchart 1400 of a wireless communication method. The method may be implemented by a first network node (e.g., UE 104, 404; camera 602; CpRFDev 804, 904, 1004, 1104, 1204); device 1504 (hereinafter referred to as...). Figure 15 The method, as discussed below, enables the first network node (e.g., the UE) to initiate UE-assisted visual localization or UE-based visual localization by utilizing camera calibration performed on a network device with known location and RF capabilities.
[0213] At 1404, the first network node can receive from the network entity an indication of a training point set and at least one image of a region captured by at least one camera of at least one second network node, wherein the first network node is within the FOV of the at least one camera of the at least one second network node, and wherein the indication for each training point in the training point set includes the pixel location and world location in the at least one image, such as combining... Figure 10 As described. For example, at 1036, CpRFDev 1004 can receive a set of training points from server 1006 (from the combination) Figure 7The image set (or extracted relevant pixels) obtained from the described matching and the region. The at least one image was captured by a CamDev 1002, and the CpRFDev 1004 is within the FOV of the CamDev. The reception of the training point set and the at least one image can be, for example, by... Figure 15 The device 1504 uses a visual positioning component 198, a camera 1532, an application processor 1506, a cellular baseband processor 1524, and / or a transceiver 1522 to perform the operation.
[0214] In one example, a first network node may send a request to a network entity to estimate the location of one or more objects in a region using vision-based localization, wherein indications of a training point set and at least one image may be received from the network entity based on the request, such as combining... Figure 10 As described. For example, at 1020, CpRFDev 1004, configured to perform or participate in a UE-based visual positioning session (e.g., for determining / estimating the location of one or more objects in an area), can send a location estimation request to server 1006. The sending of the request can be, for example, by Figure 15 The device 1504 uses a visual positioning component 198, a camera 1532, an application processor 1506, a cellular baseband processor 1524, and / or a transceiver 1522 to perform the operation.
[0215] In another example, the first network node may extract its pixel position from the at least one image; match the world position of the first network node with the pixel position of the first network node; and perform data normalization on the training point set or the at least one image based on the world position and the pixel position of the first network node, such as combining... Figure 10 As described. For example, at 1038, based on training points from server 1006 and the captured image (or extracted relevant pixels), CpRFDev 1004 can perform data preparation and processing on the image and / or the extracted relevant pixels, such as combining... Figure 8The first functional block 812 is described. For example, data preparation and processing may include: (1) determining relevant CamDev and / or CpRFDev for visual localization, (2) allocating resources specified for information exchange between different nodes (e.g., between CamDev, CpRFDev and the server), (3) pixel location extraction (e.g., extracting pixel locations of CpRFDev 1004, such as using an object detection mechanism), (4) data normalization (e.g., standardization, whitening, etc.), or (5) combinations thereof. The extraction of pixel locations of the first network node, the matching of the world location of the first network node with the pixel locations of the first network node, and / or the data normalization of the training point set or the at least one image may be performed by, for example Figure 15 The device 1504 utilizes a visual positioning component 198, a camera 1532, an application processor 1506, a cellular baseband processor 1524, and / or a transceiver 1522 to perform this function. In some implementations, the pixel position of the first network node is extracted using object detection.
[0216] At 1408, the first network node can estimate the projection transformation between the world position of the object and the pixel position of the object in the at least one image based on the training point set and the known or estimated position of the first network node, such as combining... Figure 10 As described. For example, at 1040, CpRFDev 1004 can perform camera calibration based on data preparation and processing, such as combining... Figure 8 As described in the second functional block 814. For example, camera calibration may include outlier rejection (e.g., performing training point selection, such as using RANSAC) and / or projection transformation (calibration matrix) estimation, such as in combination with... Figures 5 to 7 As described. Projection transformation can be, for example, Figure 15 The device 1504 comprises a visual positioning component 198, a camera 1532, an application processor 1506, a cellular baseband processor 1524, and / or a transceiver 1522. In some embodiments, to estimate the projection transformation, the first network node may select a subset of training points from the training point set and generate a camera calibration matrix based on that subset. In some embodiments, the subset of training points may be selected from the training point set based on Random Sample Consensus (RANSAC).
[0217] In one example, the first network node can estimate the location of one or more objects in the region based on the at least one image and projection transformation, such as by combining... Figure 10 As described. For example, at 1042, after performing camera calibration, CpRFDev 1004 can perform visual localization (e.g., UE-based visual localization) on one or more objects in the captured image (or in the extracted relevant pixels), such as combining Figure 8 The third functional block 816 describes the estimation of the location of the one or more objects in the region, for example... Figure 15 The device 1504 utilizes a visual positioning component 198, a camera 1532, an application processor 1506, a cellular baseband processor 1524, and / or a transceiver 1522 to perform this function. In some embodiments, to estimate the position of the one or more objects in the region based on the at least one image and projection transformation, the first network node may calculate the image pixel position of the one or more objects based on the at least one image. In some embodiments, the one or more objects in the region are within the FOV of the at least one camera of the at least one second network node.
[0218] In another example, the first network node could be a base station or a TRP with a known location.
[0219] In another example, the first network node may be a UE that includes an estimated location. In some implementations, the estimated location of the first network node may be obtained based on network-based positioning or GNSS-based positioning.
[0220] In another example, the one or more objects may be non-radio frequency (non-RF) objects, or the one or more objects may not include RF capabilities.
[0221] In another example, vision-based localization is based on the user's visual location.
[0222] In another example, the first network node may send an indication to the network entity of the estimated location of one or more objects in the region, such as combining... Figure 10 As described. For example, at 1044, CpRFDev 1004 can send the results of visual positioning to server 1006 (if specified), such as the estimated positions of one or more objects detected within the FOV of CamDev 1002 (e.g., estimated positions of CpRFDev, NonCpRFDev, NonRFobj, etc.). The transmission of indications of the estimated positions of these one or more objects in the area can be, for example, by... Figure 15 The device 1504 uses a visual positioning component 198, a camera 1532, an application processor 1506, a cellular baseband processor 1524, and / or a transceiver 1522 to perform the operation.
[0223] Figure 15Figure 1500 illustrates an example of a hardware implementation for device 1504. Device 1504 may be a UE, a component of a UE, or implement UE functionality. In some aspects, device 1504 may include a cellular baseband processor 1524 (also referred to as a modem) coupled to one or more transceivers 1522 (e.g., cellular RF transceivers). Cellular baseband processor 1524 may include on-chip memory 1524'. In some aspects, device 1504 may also include one or more Subscriber Identity Module (SIM) cards 1520 and an application processor 1506 coupled to a Secure Digital Card (SD) card 1508 and a screen 1510. Application processor 1506 may include on-chip memory 1506'. In some aspects, device 1504 may also include a Bluetooth module 1512, a WLAN module 1514, an SPS module 1516 (e.g., a GNSS module), an ultra-wideband (UWB) module 1536, one or more sensor modules 1518 (e.g., an atmospheric pressure sensor / altimeter; motion sensors such as an inertial measurement unit (IMU), a gyroscope, and / or an accelerometer; light detection and ranging (LIDAR), radio-assisted detection and ranging (RADAR), sound navigation and ranging (SONAR), a magnetometer, audio, and / or other technologies for positioning), an additional memory module 1526, a power source 1530, and / or a camera 1532. Bluetooth module 1512, WLAN module 1514, UWB module 1536, and SPS module 1516 may include an on-chip transceiver (TRX) (or in some cases, only a receiver (RX)). Bluetooth module 1512, WLAN module 1514, UWB module 1536, and SPS module 1516 may include their own dedicated antennas and / or communicate using antenna 1580. Cellular baseband processor 1524 communicates with UE 104 and / or RUs associated with the same network entity 1502 via transceiver 1522 through one or more antennas 1580. Cellular baseband processor 1524 and application processor 1506 may each include computer-readable media / memory 1524', 1506', respectively. Additional memory module 1526 may also be considered as computer-readable media / memory. Each computer-readable media / memory 1524', 1506', 1526 may be non-transitory. Cellular baseband processor 1524 and application processor 1506 are each responsible for general processing, including executing software stored on the computer-readable media / memory. When executed by the cellular baseband processor 1524 / application processor 1506, the software causes the cellular baseband processor 1524 / application processor 1506 to perform the various functions described above. The computer-readable medium / memory can also be used to store data manipulated by the cellular baseband processor 1524 / application processor 1506 during software execution.Cellular baseband processor 1524 / application processor 1506 may be a component of UE 350 and may include memory 360 and / or at least one of TX processor 368, RX processor 356, and controller / processor 359. In one configuration, device 1504 may be a processor chip (modem and / or application) and may only include cellular baseband processor 1524 and / or application processor 1506, while in another configuration, device 1504 may be the entire UE (see, for example). Figure 3 The UE 350 includes an additional module of the device 1504.
[0224] As discussed above, the visual positioning component 198 may be configured to receive from a network entity an indication of a training point set and at least one image of a region captured by at least one camera of at least one second network node, wherein the first network node is within the FOV of the at least one camera of the at least one second network node, and wherein the indication for each training point in the training point set includes a pixel position and a world position in the at least one image. The visual positioning component 198 may also be configured to estimate a projection transformation between the world position of an object and the pixel position of the object in the at least one image based on the training point set and the known or estimated position of the first network node. The visual positioning component 198 may be within a cellular baseband processor 1524, an application processor 1506, or both. The visual positioning component 198 may be one or more hardware components specifically configured to execute the process / algorithm, implemented by one or more processors configured to execute the process / algorithm, stored in a computer-readable medium for implementation by one or more processors, or some combination thereof. As shown, the apparatus 1504 may include a variety of components configured for various functions. In one configuration, apparatus 1504, and particularly cellular baseband processor 1524 and / or application processor 1506, may include components for receiving from a network entity an indication of a training point set and at least one image of a region captured by at least one camera of at least one second network node, wherein the first network node is within the field of view (FOV) of the at least one camera of the at least one second network node, and wherein the indication for each training point in the training point set includes a pixel position and a world position in the at least one image. Apparatus 1504 may also include components for estimating a projection transformation between the world position of an object and the pixel position of the object in the at least one image based on the training point set and the known or estimated position of the first network node.
[0225] In one configuration, the device 1504 may further include: components for extracting pixel positions of the device 1504 from the at least one image; components for matching the world position of the device 1504 with the pixel positions of the device 1504; and components for performing data normalization on a training point set or the at least one image based on the world position and pixel positions of the device 1504. In some specific implementations, the pixel positions of the device 1504 are extracted using object detection.
[0226] In another configuration, the components for estimating the projection transformation may include configuration device 1504 to select a subset of training points from the training point set and to generate a camera calibration matrix based on that subset. In some specific implementations, the subset of training points may be selected from the training point set based on RANSAC.
[0227] In another configuration, the apparatus 1504 may further include: components for sending a request to a network entity to estimate the location of one or more objects in a region using vision-based localization, wherein indications of a training point set and at least one image are received from the network entity based on the request; and components for estimating the location of the one or more objects in the region based on the at least one image and a projection transformation. In some embodiments, the components for estimating the location of the one or more objects in the region based on the at least one image and a projection transformation may include configuring the apparatus 1504 to image the pixel locations of the one or more objects based on the at least one image. In some embodiments, the one or more objects in the region are within the field of view (FOV) of the at least one camera of the at least one second network node.
[0228] In another configuration, device 1504 may be a base station or TRP that includes a known location.
[0229] In another configuration, device 1504 may be a UE that includes an estimated location. In some specific implementations, device 1504 may also include components for obtaining the estimated location of device 1504 based on network-based positioning or GNSS-based positioning.
[0230] In another configuration, the one or more objects may be non-radio frequency (non-RF) objects, or the one or more objects may not include RF capabilities.
[0231] In another configuration, vision-based positioning is based on the UE's visual positioning.
[0232] In another configuration, the device 1504 may also include components for sending to a network entity an indication of the estimated location of the one or more objects in the area.
[0233] The component may be a visual positioning component 198 of device 1504 configured to perform the functions described therein. As described above, device 1504 may include a TX processor 368, an RX processor 356, and a controller / processor 359. Therefore, in one configuration, the component may be the TX processor 368, the RX processor 356, and / or the controller / processor 359 configured to perform the functions described therein.
[0234] Figure 16 This is a flowchart 1600 of a wireless communication method. The method may be implemented by network entities (e.g., one or more location servers 168; servers 906, 1006, 1106, 1206; network entity 1860 (hereinafter referred to as...)). Figure 18 The discussion section discusses how this method can be implemented. This approach enables network entities to coordinate multiple network devices (e.g., CpRFDev, non-CpRFDev, and / or CamDev) to perform visual localization using camera calibration performed by network devices with known locations and RF capabilities.
[0235] At 1604, a network entity may select at least one second network node for a first network node, including a known or estimated location, wherein the first network node is within the FOV of the at least one second network node, such as in combination. Figure 9 As described. For example, at 926, server 906 may determine the relevant CamDev 902 for participating in a UE-assisted visual positioning session. For example, server 906 may select CamDev where CpRFDev 904 (and the one or more objects and / or areas to be detected) is in the FOV of CamDev 902. The selection of this at least one second network node may be, for example, by Figure 18 The visual positioning coordination component 197, network processor 1812, and / or network interface 1880 of the network entity 1860 are used to perform this function.
[0236] In one example, as shown at 1602, a network entity may receive from a first network node a request to estimate the location of one or more objects in a region using vision-based localization, wherein at least one second network node is selected based on the request to estimate the location of the one or more objects in the region, such as combining... Figure 9 As described. For example, at 920, server 906 may receive a location estimation request from CpRFDev 904 configured to perform or participate in a UE-assisted visual positioning session (e.g., for determining / estimating the location of one or more objects in an area). Reception of a request to estimate the location of one or more objects may be achieved by, for example... Figure 18 The visual positioning coordination component 197, network processor 1812, and / or network interface 1880 of the network entity 1860 are used to perform this function.
[0237] In another example, in order to select the at least one second network node, the network entity may receive from the first network node a request to estimate the location of one or more objects in the region using vision-based positioning, wherein the at least one second network node is selected based on the request to estimate the location of the one or more objects in the region; send a query to a database for a list of second network nodes surrounding the first network node; and receive the list of second network nodes surrounding the first network node from the database.
[0238] At point 1606, the network entity may send a request to the at least one second network node to capture at least one image of the area using at least one camera, wherein the at least one image includes the first network node, such as in conjunction with... Figure 9 As described. For example, at 928, after server 906 determines that CamDev 902 is to participate in UE-assisted visual positioning, server 906 may send an image request to CamDev 902 to request CamDev 902 to capture an image (e.g., capture the FOV of CamDev 902). The sending of the request for at least one image of the capture area may be, for example, by Figure 18 The visual positioning coordination component 197, network processor 1812, and / or network interface 1880 of the network entity 1860 are used to perform this function.
[0239] At 1608, the network entity may receive at least one image of the region from the at least one second network node based on the request, such as in combination. Figure 9 As described. For example, at 932, server 906 may receive captured images and / or extracted relevant pixels from CamDev 902 (if performed by CamDev 902). Reception of at least one image of the area may be achieved by, for example... Figure 18 The network entity 1860 is executed by the visual positioning coordination component 197, the network processor 1812, and / or the network interface 1880. In some specific implementations, the network entity may extract a set of relevant pixels of one or more objects from the at least one image.
[0240] In one example, a network entity may extract a set of training points from the at least one image; and send an instruction to a first network node regarding the set of training points and the at least one image, wherein the instruction for each training point in the set of training points includes the pixel location and world location in the at least one image.
[0241] At 1610, the network entity can match the known or estimated location of the first network node with the relevant pixels of the first network node in the at least one image, such as combining... Figure 9As described. For example, at 934, based on the image from CamDev 902 and / or the extracted relevant pixels, server 906 can perform data preparation and processing on the image and / or the extracted relevant pixels, such as combining... Figure 8 The first functional block 812 is described. For example, data preparation and processing may include: (1) initiating a visual localization process, (2) determining the relevant CamDev and / or CpRFDev for visual localization, (3) allocating resources specified for information exchange between different nodes (e.g., between CamDev, CpRFDev and the server), (4) pixel location extraction (e.g., extracting the pixel location of CpRFDev 904, such as using an object detection mechanism), (5) matching (e.g., matching the world location (e.g., WCS / 2D / 3D coordinates) of CpRFDev 904 in the FOV of CamDev 902 with its corresponding image pixel location), (6) data normalization (e.g., standardization, whitening, etc.), or (7) combinations thereof. The matching of the known or estimated location of the first network node with the relevant pixels of the first network node in the at least one image may be performed by, for example Figure 18 The visual positioning coordination component 197, network processor 1812, and / or network interface 1880 of the network entity 1860 are used to perform this function.
[0242] In one example, a network entity may extract the pixel position of a first network node from the at least one image; match the world position of the first network node with the pixel position of the first network node; and perform data normalization on the at least one image based on the world position and the pixel position of the first network node.
[0243] At 1612, the network entity can estimate the projection transformation between the world position of the object and the pixel position of the object in the at least one image based on the known or estimated position of the first network node, such as combining... Figure 9 As described. For example, at 936, server 906 can perform camera calibration (e.g., projection transformation) based on data preparation and processing, such as combining... Figure 8 As described in the second functional block 814. For example, camera calibration may include outlier rejection (e.g., performing training point selection, such as using RANSAC) and / or projection transformation (calibration matrix) estimation, such as in combination with... Figures 5 to 7 As described. Projection transformation can be, for example, Figure 18 The visual positioning coordination component 197, network processor 1812, and / or network interface 1880 of the network entity 1860 are used to perform this function.
[0244] At point 1614, the network entity can estimate the location of one or more objects in the region based on at least one image and projection transformation, such as by combining... Figure 9 As described. For example, at 938, after performing camera calibration, server 906 may perform visual localization (e.g., UE-assisted visual localization) on one or more objects in the captured image (or in the extracted relevant pixels), such as combining... Figure 8 The third functional block 816 describes the estimation of the location of the one or more objects in the region, which can be achieved by, for example... Figure 18 The visual positioning coordination component 197, network processor 1812, and / or network interface 1880 of the network entity 1860 are used to perform this function.
[0245] In one example, in order to estimate the location of one or more objects in a region based on the at least one image and projection transformation, the network entity may calculate the image pixel location of the one or more objects based on the at least one image.
[0246] In another example, a network entity may send an indication to a first network node of the estimated location of one or more objects in the region.
[0247] In another example, the first network node is a base station or TRP with a known location.
[0248] In another example, the first network node is a UE with an estimated location. The network entity can obtain the estimated location of the first network node.
[0249] In another example, the one or more objects are non-RF objects, or the one or more objects do not include RF capabilities.
[0250] Figure 17 This is a flowchart 1700 of a wireless communication method. The method may be implemented by network entities (e.g., one or more location servers 168; servers 906, 1006, 1106, 1206; network entity 1860 (hereinafter referred to as...)). Figure 18 The discussion section discusses how this method can be implemented. This approach enables network entities to coordinate multiple network devices (e.g., CpRFDev, non-CpRFDev, and / or CamDev) to perform visual localization using camera calibration performed by network devices with known locations and RF capabilities.
[0251] At 1704, a network entity may select at least one second network node for a first network node, including a known or estimated location, wherein the first network node is within the FOV of the at least one second network node, such as in combination. Figure 9As described. For example, at 926, server 906 may determine the relevant CamDev 902 for participating in a UE-assisted visual positioning session. For example, server 906 may select CamDev where CpRFDev 904 (and the one or more objects and / or areas to be detected) is in the FOV of CamDev 902. The selection of this at least one second network node may be, for example, by Figure 18 The visual positioning coordination component 197, network processor 1812, and / or network interface 1880 of the network entity 1860 are used to perform this function.
[0252] In one example, a network entity may receive from a first network node a request to estimate the location of one or more objects in a region using vision-based localization, wherein at least one second network node is selected based on the request to estimate the location of the one or more objects in the region, such as combining... Figure 9 As described. For example, at 920, server 906 may receive a location estimation request from CpRFDev 904 configured to perform or participate in a UE-assisted visual positioning session (e.g., for determining / estimating the location of one or more objects in an area). Reception of a request to estimate the location of one or more objects may be achieved by, for example... Figure 18 The visual positioning coordination component 197, network processor 1812, and / or network interface 1880 of the network entity 1860 are used to perform this function.
[0253] In another example, in order to select the at least one second network node, the network entity may receive from the first network node a request to estimate the location of one or more objects in the region using vision-based positioning, wherein the at least one second network node is selected based on the request to estimate the location of the one or more objects in the region; send a query to a database for a list of second network nodes surrounding the first network node; and receive the list of second network nodes surrounding the first network node from the database.
[0254] At point 1606, the network entity may send a request to the at least one second network node to capture at least one image of the area using at least one camera, wherein the at least one image includes the first network node, such as in conjunction with... Figure 9 As described. For example, at 928, after server 906 determines that CamDev 902 is to participate in UE-assisted visual positioning, server 906 may send an image request to CamDev 902 to request CamDev 902 to capture an image (e.g., capture the FOV of CamDev 902). The sending of the request for at least one image of the capture area may be, for example, by Figure 18The visual positioning coordination component 197, network processor 1812, and / or network interface 1880 of the network entity 1860 are used to perform this function.
[0255] At 1608, the network entity may receive at least one image of the region from the at least one second network node based on the request, such as in combination. Figure 9 As described. For example, at 932, server 906 may receive captured images and / or extracted relevant pixels from CamDev 902 (if performed by CamDev 902). Reception of at least one image of the area may be achieved by, for example... Figure 18 The network entity 1860 is executed by the visual positioning coordination component 197, the network processor 1812, and / or the network interface 1880. In some specific implementations, the network entity may extract a set of relevant pixels of one or more objects from the at least one image.
[0256] In one example, a network entity may extract a set of training points from the at least one image; and send an instruction to a first network node regarding the set of training points and the at least one image, wherein the instruction for each training point in the set of training points includes the pixel location and world location in the at least one image.
[0257] In another example, a network entity can match the known or estimated location of a first network node with relevant pixels of that first network node in the at least one image, such as by combining... Figure 9 As described. For example, at 934, based on the image from CamDev 902 and / or the extracted relevant pixels, server 906 can perform data preparation and processing on the image and / or the extracted relevant pixels, such as combining... Figure 8 The first functional block 812 is described. For example, data preparation and processing may include: (1) initiating a visual localization process, (2) determining the relevant CamDev and / or CpRFDev for visual localization, (3) allocating resources specified for information exchange between different nodes (e.g., between CamDev, CpRFDev and the server), (4) pixel location extraction (e.g., extracting the pixel location of CpRFDev 904, such as using an object detection mechanism), (5) matching (e.g., matching the world location (e.g., WCS / 2D / 3D coordinates) of CpRFDev 904 in the FOV of CamDev 902 with its corresponding image pixel location), (6) data normalization (e.g., standardization, whitening, etc.), or (7) combinations thereof. The matching of the known or estimated location of the first network node with the relevant pixels of the first network node in the at least one image may be performed by, for example Figure 18The visual positioning coordination component 197, network processor 1812, and / or network interface 1880 of the network entity 1860 are used to perform this function.
[0258] In another example, the network entity may extract the pixel position of a first network node from the at least one image; match the world position of the first network node with the pixel position of the first network node; and perform data normalization on the at least one image based on the world position and the pixel position of the first network node.
[0259] In another example, a network entity may estimate the projection transformation between the world position of an object and the pixel position of the object in the at least one image based on the known or estimated positions of the at least one image and the first network node, such as by combining... Figure 9 As described. For example, at 936, server 906 can perform camera calibration (e.g., projection transformation) based on data preparation and processing, such as combining... Figure 8 As described in the second functional block 814. For example, camera calibration may include outlier rejection (e.g., performing training point selection, such as using RANSAC) and / or projection transformation (calibration matrix) estimation, such as in combination with... Figures 5 to 7 As described. Projection transformation can be, for example, Figure 18 The visual positioning coordination component 197, network processor 1812, and / or network interface 1880 of the network entity 1860 are used to perform this function.
[0260] In another example, a network entity can estimate the location of one or more objects in a region based on at least one image and projection transformation, such as by combining... Figure 9 As described. For example, at 938, after performing camera calibration, server 906 may perform visual localization (e.g., UE-assisted visual localization) on one or more objects in the captured image (or in the extracted relevant pixels), such as combining... Figure 8 The third functional block 816 describes the estimation of the location of the one or more objects in the region, which can be achieved by, for example... Figure 18 The visual positioning coordination component 197, network processor 1812, and / or network interface 1880 of the network entity 1860 are used to perform this function.
[0261] In another example, in order to estimate the location of one or more objects in a region based on the at least one image and projection transformation, the network entity may calculate the image pixel location of the one or more objects based on the at least one image.
[0262] In another example, a network entity may send an indication to a first network node of the estimated location of one or more objects in the region.
[0263] In another example, the first network node is a base station or TRP with a known location.
[0264] In another example, the first network node is a UE with an estimated location. The network entity can obtain the estimated location of the first network node.
[0265] In another example, the one or more objects are non-RF objects, or the one or more objects do not include RF capabilities.
[0266] Figure 18 Figure 1800 illustrates an example of a hardware implementation for network entity 1860. In one example, network entity 1860 may be within core network 120. Network entity 1860 may include network processor 1812. Network processor 1812 may include on-chip memory 1812'. In some aspects, network entity 1860 may also include additional memory module 1814. Network entity 1860 communicates with CU 1802 directly (e.g., via a backhaul link) or indirectly (e.g., via RIC) through network interface 1880. On-chip memory 1812' and additional memory module 1814 may each be considered as computer-readable media / memory. Each computer-readable media / memory may be non-transitory. Processor 1812 is responsible for general processing, including executing software stored on the computer-readable media / memory. This software, when executed by the corresponding processor, causes the processor to perform the various functions described above. The computer-readable media / memory may also be used to store data manipulated by the processor while executing the software.
[0267] As discussed above, the visual positioning coordination component 197 can be configured to: select at least one second network node for a first network node including a known or estimated location, wherein the first network node is within the FOV of the at least one second network node. The visual positioning coordination component 197 can also be configured to: send a request to the at least one second network node to capture at least one image of an area using at least one camera, wherein the at least one image includes the first network node. The visual positioning coordination component 197 can also be configured to: receive the at least one image of the area from the at least one second network node based on the request. The visual positioning coordination component 197 can be within processor 1812. The visual positioning coordination component 197 can be one or more hardware components specifically configured to perform the process / algorithm, implemented by one or more processors configured to perform the process / algorithm, stored in a computer-readable medium for implementation by one or more processors, or some combination thereof. Network entity 1860 can include a variety of components configured for various functions. In one configuration, network entity 1860 may include: components for selecting at least one second network node for a first network node including a known or estimated location, wherein the first network node is within the field of view (FOV) of the at least one second network node. Network entity 1860 may also include: components for sending a request to the at least one second network node to capture at least one image of an area using at least one camera, wherein the at least one image includes the first network node. Network entity 1860 may further include: components for receiving the at least one image of the area from the at least one second network node based on the request.
[0268] In one configuration, components for selecting the at least one second network node may include: configuring network entity 1860 to receive from a first network node a request to estimate the location of one or more objects in a region using vision-based positioning, wherein the at least one second network node is selected based on the request to estimate the location of the one or more objects in the region; sending a query to a database for a list of second network nodes surrounding the first network node; receiving from the database the list of second network nodes surrounding the first network node; and determining the at least one second network node from the list of second network nodes.
[0269] In another configuration, network entity 1860 may also include a component for extracting a set of related pixels of one or more objects from the at least one image.
[0270] In another configuration, network entity 1860 may further include: a component for extracting a set of training points from the at least one image; and a component for sending the set of training points and the at least one image to a first network node.
[0271] In another configuration, network entity 1860 may further include a component for matching a known or estimated location of a first network node with a relevant pixel of the first network node in the at least one image.
[0272] In another configuration, network entity 1860 may further include: a component for extracting a set of training points from the at least one image; and a component for sending an instruction to a first network node regarding the set of training points and the at least one image, wherein the instruction for each training point in the set of training points includes a pixel location and a world location in the at least one image.
[0273] In another configuration, network entity 1860 may further include a component for estimating the projection transformation between the world position of an object and the pixel position of the object in the at least one image based on the known or estimated position of the at least one image and the first network node.
[0274] In another configuration, network entity 1860 may further include a component for estimating the location of one or more objects in the region based on the at least one image and projection transformation.
[0275] In another configuration, the component for estimating the location of the one or more objects in the region based on the at least one image and projection transformation may include configuring network entity 1860 to calculate the image pixel location of the one or more objects based on the at least one image.
[0276] In another configuration, network entity 1860 may further include a component for sending an indication to a first network node of the estimated location of the one or more objects in the area.
[0277] In another configuration, the first network node is a base station or TRP with a known location.
[0278] In another configuration, the first network node is a UE that includes an estimated location. Network entity 1860 may also include components for obtaining the estimated location of the first network node.
[0279] In another configuration, the one or more objects are non-RF objects, or the one or more objects do not include RF capabilities.
[0280] The component may be a visual positioning coordination component 197 of network entity 1860 configured to perform the functions described therein.
[0281] It should be understood that the specific order or hierarchy of the boxes in the disclosed process / flowcharts is merely an example of the exemplary method. It should be understood that the specific order or hierarchy of the boxes in the process / flowcharts may be rearranged based on design preferences. Furthermore, some boxes may be combined or omitted. The appended method claims present the elements of various boxes in a sample order, but are not limited to the specific order or hierarchy presented.
[0282] The foregoing description is provided to enable any person skilled in the art to practice the various aspects described herein. Various modifications to these aspects will be apparent to those skilled in the art, and the general principles defined herein may be applied to other aspects. Therefore, the claims are not limited to the aspects described herein but should be given the full scope consistent with the language of the claims. Unless specifically stated otherwise, references to elements in the singular form do not mean “one and only one” but rather “one or more.” Terms such as “if,” “when,” and “simultaneously” do not imply a direct temporal relationship or reaction. That is, these phrases, such as “when…”, do not imply an immediate action in response to the occurrence of an action or during the occurrence of an action, but simply suggest that if a condition is met, then the action will occur, without requiring a specific or immediate time limit for the occurrence of the action. The word “exemplary” is used herein to mean “serving as an example, instance, or illustration.” Any aspect described herein as “exemplary” is not necessarily to be construed as preferred or superior to other aspects. Unless otherwise specifically stated, the term “some” refers to one or more. Combinations such as "at least one of A, B, or C", "one or more of A, B, or C", "at least one of A, B, and C", "one or more of A, B, and C", and "A, B, C, or any combination thereof" include any combination of A, B, and / or C, which may include multiple A, multiple B, or multiple C. Specifically, combinations such as "at least one of A, B, or C", "one or more of A, B, or C", "at least one of A, B, and C", "one or more of A, B, and C", and "A, B, C, or any combination thereof" can be only A, only B, only C, A and B, A and C, B and C, or A and B and C, where any such combination may contain one or more members of A, B, or C. A set should be interpreted as a collection of elements, where the number of elements is one or more. Therefore, for a set of X, X will include one or more elements. If the first device receives data from or sends data to the second device, data can be received / sent directly between the first and second devices, or indirectly between the first and second devices via a set of devices. A device configured to "output" data (such as transmission, signaling, or messaging) can, for example, transmit the data using a transceiver, or can transmit the data to the device that sent the data. A device configured to "receive" data (such as transmission, signaling, or messaging) can, for example, receive the data using a transceiver, or can obtain the data from the device that received the data. All structural and functional equivalents of the elements throughout the various aspects described herein that are known to those skilled in the art or will later be known are expressly incorporated herein by reference and are covered by the claims.Furthermore, nothing disclosed herein is intended to be offered to the public, whether or not such disclosure is explicitly stated in the claims. Terms such as “module,” “mechanism,” “element,” and “device” cannot replace the term “component.” Therefore, no claim element will be interpreted as a functional component unless the element is explicitly stated using the phrase “component for…”.
[0283] As used in this article, the phrase “based on” should not be interpreted as referring to a closed set of information, one or more conditions, one or more factors, etc. In other words, the phrase “based on A” (where “A” can be information, conditions, factors, etc.) should be interpreted as “based on at least A”, unless specifically stated differently.
[0284] The following aspects are merely illustrative and may be combined with other aspects or teachings described herein without limitation.
[0285] Aspect 1 is a method for wireless communication at a first network node, the method comprising: receiving from a network entity an indication of a training point set and at least one image of a region captured by at least one camera of at least one second network node, wherein the first network node is within the field of view (FOV) of the at least one camera of the at least one second network node, wherein the indication for each training point in the training point set includes a pixel position and a world position in the at least one image; and estimating a projection transformation between the world position of an object and the pixel position of the object in the at least one image based on the training point set and a known or estimated position of the first network node.
[0286] Aspect 2 is the method according to aspect 1, the method further comprising: sending a request to the network entity to estimate the location of one or more objects in the region using vision-based localization, wherein the indication of the training point set and the at least one image is received from the network entity based on the request; and estimating the location of the one or more objects in the region based on the at least one image and the projection transformation.
[0287] Aspect 3 is the method according to aspect 2, wherein estimating the position of the one or more objects in the region based on the at least one image and the projection transformation comprises: calculating the image pixel position of the one or more objects based on the at least one image.
[0288] Aspect 4 is the method according to any one of aspects 2 to 3, the method further comprising: sending to the network entity an indication of the estimated location of the one or more objects in the region.
[0289] Aspect 5 is the method according to any one of Aspects 2 to 4, wherein the one or more objects in the region are in the FOV of the at least one camera of the at least one second network node.
[0290] Aspect 6 is a method according to any one of Aspects 1 to 5, the method further comprising: extracting pixel positions of the first network node from the at least one image; matching the world position of the first network node with the pixel positions of the first network node; and performing data normalization on the training point set or the at least one image based on the world position of the first network node and the pixel positions of the first network node.
[0291] Aspect 7 is the method according to aspect 6, wherein the pixel position of the first network node is extracted using object detection.
[0292] Aspect 8 is the method according to any one of Aspects 1 to 7, wherein estimating the projection transformation comprises: selecting a subset of training points from the set of training points; and generating a camera calibration matrix based on the subset of training points.
[0293] Aspect 9 is the method according to aspect 8, wherein the subset of training points is selected from the set of training points based on RANSAC.
[0294] Aspect 10 is the method according to any one of aspects 1 to 9, wherein the first network node is a base station or TRP including the known location.
[0295] Aspect 11 is a method according to any one of Aspects 1 to 10, wherein the first network node is a UE including the estimated location, and the method further includes: obtaining the estimated location of the first network node based on network-based positioning or GNSS-based positioning.
[0296] Aspect 12 is the method according to any one of aspects 1 to 11, wherein the one or more objects are non-radio frequency (non-RF) objects, or the one or more objects do not include RF capabilities.
[0297] Aspect 13 is the method according to any one of aspects 1 to 12, wherein the vision-based positioning is based on the visual positioning of the UE.
[0298] Aspect 14 is an apparatus for wireless communication at a first network node, the apparatus comprising: a memory; and at least one processor coupled to the memory and based at least in part on information stored in the memory, the at least one processor being configured to implement any one of aspects 1 to 13.
[0299] Aspect 15 is the apparatus according to aspect 14, the apparatus further comprising at least one of a transceiver or an antenna coupled to the at least one processor.
[0300] Aspect 16 is an apparatus for wireless communication, the apparatus comprising: components for implementing any one of aspects 1 to 13.
[0301] Aspect 17 is a computer-readable medium (e.g., a non-transitory computer-readable medium) storing computer-executable code, wherein the code, when executed by a processor, causes the processor to implement any one of aspects 1 to 13.
[0302] Aspect 18 is a method for wireless communication at a network entity, the method comprising: selecting at least one second network node for a first network node including a known or estimated location, wherein the first network node is within the field of view (FOV) of the at least one second network node; sending a request to the at least one second network node to capture at least one image of a region using at least one camera, wherein the at least one image includes the first network node; and receiving the at least one image of the region from the at least one second network node based on the request.
[0303] Aspect 19 is the method according to aspect 18, wherein selecting the at least one second network node comprises: receiving from the first network node a request to estimate the location of one or more objects in the region using vision-based localization, wherein the at least one second network node is selected based on the request to estimate the location of the one or more objects in the region; sending a query to a database for a list of second network nodes surrounding the first network node; receiving from the database the list of second network nodes surrounding the first network node; and determining the at least one second network node from the list of second network nodes.
[0304] Aspect 20 is the method according to aspect 18 or 19, the method further comprising: extracting a set of relevant pixels of one or more objects from the at least one image.
[0305] Aspect 21 is a method according to any one of aspects 18 to 20, the method further comprising: matching the known location or the estimated location of the first network node with a relevant pixel of the first network node in the at least one image.
[0306] Aspect 22 is a method according to any one of aspects 18 to 21, the method further comprising: extracting a set of training points from the at least one image; and sending an indication to the first network node of the set of training points and the at least one image, wherein the indication for each training point in the set of training points includes a pixel location and a world location in the at least one image.
[0307] Aspect 23 is a method according to any one of aspects 18 to 22, the method further comprising: estimating a projection transformation between the world position of an object and the pixel position of the object in the at least one image based on the known position or the estimated position of the first network node and the at least one image.
[0308] Aspect 24 is the method according to aspect 23, the method further comprising: estimating the position of one or more objects in the region based on the at least one image and the projection transformation.
[0309] Aspect 25 is the method according to aspect 24, wherein estimating the position of the one or more objects in the region based on the at least one image and the projection transformation comprises: calculating the image pixel position of the one or more objects based on the at least one image.
[0310] Aspect 26 is the method according to any one of aspects 24 to 25, the method further comprising: sending to the first network node an indication of the estimated location of the one or more objects in the region.
[0311] Aspect 27 is a method according to any one of aspects 18 to 26, the method further comprising: extracting pixel positions of the first network node from the at least one image; matching the world position of the first network node with the pixel positions of the first network node; and performing data normalization on the at least one image based on the world position of the first network node and the pixel positions of the first network node.
[0312] Aspect 28 is the method according to any one of aspects 18 to 27, wherein the first network node is a base station or TRP including the known location.
[0313] Aspect 29 is a method according to any one of aspects 18 to 28, wherein the first network node is a UE including the estimated location, and the method further includes: obtaining the estimated location of the first network node.
[0314] Aspect 30 is the method according to any one of aspects 18 to 29, wherein the one or more objects are non-radio frequency (non-RF) objects, or the one or more objects do not include RF capabilities.
[0315] Aspect 31 is an apparatus for wireless communication at a network entity, the apparatus comprising: a memory; and at least one processor coupled to the memory and based at least in part on information stored in the memory, the at least one processor being configured to implement any one of aspects 18 to 30.
[0316] Aspect 32 is the apparatus according to aspect 31, the apparatus further comprising at least one of a transceiver or an antenna coupled to the at least one processor.
[0317] Aspect 33 is an apparatus for wireless communication, the apparatus comprising: components for implementing any one of aspects 18 to 30.
[0318] Aspect 34 is a computer-readable medium (e.g., a non-transitory computer-readable medium) that stores computer-executable code, wherein the code, when executed by a processor, causes the processor to implement any one of aspects 18 to 30.
Claims
1. An apparatus for wireless communication at a first network node, the apparatus comprising: Memory; and At least one processor, the at least one processor being coupled to the memory, and the at least one processor being configured to: Send a request to a network entity to estimate the location of one or more objects within a defined area using vision-based localization; The first network entity receives a set of training points and a set of images of the defined region captured by at least one camera of at least one second network node, wherein the first network node is within the field of view (FOV) of the at least one camera of the at least one second network node, and wherein the set of training points is extracted from the set of images. as well as Camera calibration is performed on the at least one camera of the at least one second network node based on the training point set and the known or estimated location of the first network node.
2. The apparatus of claim 1, wherein the at least one processor is further configured to: The positions of the one or more objects in the defined region are estimated based on the image set and the camera calibration.
3. The apparatus of claim 2, wherein, in order to estimate the position of the one or more objects in the defined region based on the image set and the camera calibration, the at least one processor is configured to: The image pixel positions of the one or more objects are calculated based on the image set.
4. The apparatus of claim 2, wherein the at least one processor is further configured to: Send the estimated location of the one or more objects in the defined area to the network entity.
5. The apparatus of claim 2, wherein the one or more objects in the defined region are in the FOV of the at least one camera of the at least one second network node.
6. The apparatus of claim 1, wherein the at least one processor is further configured to: Extract the pixel position of the first network node from the image set; Match the world position of the first network node with the pixel position of the first network node; and Data normalization is performed on the training point set or the image set based on the world location and pixel location of the first network node.
7. The apparatus of claim 6, wherein the at least one processor is configured to use object detection to extract the pixel position of the first network node.
8. The apparatus of claim 1, wherein, in order to perform the camera calibration on the at least one camera of the at least one second network node, the at least one processor is configured to: Select a subset of training points from the set of training points; and A camera calibration matrix is generated based on the subset of training points.
9. The apparatus of claim 8, wherein the at least one processor is configured to select the subset of training points from the set of training points based on Random Sampling Consensus (RANSAC).
10. The apparatus of claim 1, wherein the first network node is a base station or transmit / receive point (TRP) including the known location.
11. The apparatus of claim 1, wherein the first network node is a user equipment (UE) including the estimated location, and wherein the at least one processor is further configured to: The estimated position of the first network node is obtained based on network-based positioning or Global Navigation Satellite System (GNSS)-based positioning.
12. The apparatus of claim 1, wherein the one or more objects are non-radio frequency (non-RF) objects, or the one or more objects do not include RF capabilities.
13. The apparatus of claim 1, wherein the vision-based positioning is based on the visual positioning of the UE.
14. The apparatus of claim 1, further comprising a transceiver coupled to the at least one processor, wherein, in order to send the request to estimate the location of the one or more objects in the defined region, the at least one processor is configured to send, via the transceiver, the request to estimate the location of the one or more objects in the defined region using the vision-based localization, and wherein, in order to receive the training point set and the image set of the defined region, the at least one processor is configured to receive, via the transceiver, the training point set and the image set of the defined region captured by the at least one camera of the at least one second network node.
15. A method for wireless communication at a first network node, the method comprising: Send a request to a network entity to estimate the location of one or more objects within a defined area using vision-based localization; The first network entity receives a set of training points and a set of images of the defined region captured by at least one camera of at least one second network node, wherein the first network node is within the field of view (FOV) of the at least one camera of the at least one second network node, and wherein the set of training points is extracted from the set of images. as well as Camera calibration is performed on the at least one camera of the at least one second network node based on the training point set and the known or estimated location of the first network node.
16. An apparatus for wireless communication at a network entity, the apparatus comprising: Memory; and At least one processor, the at least one processor being coupled to the memory, and the at least one processor being configured to: Receive a first request from a first network node, including known or estimated locations, to estimate the location of one or more objects in a defined area using vision-based localization; Based on the first request, at least one second network node is selected, wherein the first network node is in the field of view (FOV) of the at least one second network node; Send a second request to the at least one second network node to capture a set of images of the defined area using at least one camera, wherein the set of images includes the first network node; as well as Based on the second request, the image set of the defined region is received from the at least one second network node.
17. The apparatus of claim 16, wherein, in order to select the at least one second network node based on the first request, the at least one processor is configured to: Send a query to the database for a list of second network nodes surrounding the first network node; Receive the list of second network nodes surrounding the first network node from the database; as well as The at least one second network node is determined from the list of second network nodes.
18. The apparatus of claim 16, wherein the at least one processor is further configured to: Extract the relevant pixel set from the image set.
19. The apparatus of claim 16, wherein the at least one processor is further configured to: The known or estimated location of the first network node is matched with the relevant pixels of the first network node in the image set.
20. The apparatus of claim 16, wherein the at least one processor is further configured to: Extract a training point set from the image set; and The training point set and the image set are sent to the first network node.
21. The apparatus of claim 16, wherein the at least one processor is further configured to: Camera calibration is performed on the at least one camera of the at least one second network node based on the image set and the known or estimated location of the first network node.
22. The apparatus of claim 21, wherein the at least one processor is further configured to: The positions of the one or more objects in the defined region are estimated based on the image set and the camera calibration.
23. The apparatus of claim 22, wherein, in order to estimate the position of the one or more objects in the defined region based on the image set and the camera calibration, the at least one processor is configured to: The image pixel positions of the one or more objects are calculated based on the image set.
24. The apparatus of claim 22, wherein the at least one processor is further configured to: Send an indication of the estimated location of the one or more objects in the defined area to the first network node.
25. The apparatus of claim 16, wherein the at least one processor is further configured to: Extract the pixel position of the first network node from the image set; Match the world position of the first network node with the pixel position of the first network node; and Data normalization is performed on the image set based on the world location of the first network node and the pixel location of the first network node.
26. The apparatus of claim 16, wherein the first network node is a base station or transmit / receive point (TRP) including the known location.
27. The apparatus of claim 16, wherein the first network node is a user equipment (UE) including the estimated location, and wherein the at least one processor is further configured to: Obtain the estimated position of the first network node.
28. The apparatus of claim 16, wherein the one or more objects are non-radio frequency (non-RF) objects, or the one or more objects do not include RF capabilities.
29. The apparatus of claim 16, further comprising a transceiver coupled to the at least one processor, wherein, in order to receive the first request for estimating the position of the one or more objects in the defined region, the at least one processor is configured to receive, via the transceiver, the first request for estimating the position of the one or more objects in the defined region using the vision-based localization; wherein, in order to send the second request for capturing the image set of the defined region, the at least one processor is configured to send, via the transceiver, the second request for capturing the image set of the defined region using the at least one camera; and wherein, in order to receive the image set of the defined region, the at least one processor is configured to receive, via the transceiver, the image set of the defined region based on the second request.
30. A method for wireless communication at a network entity, the method comprising: Receive a first request from a first network node, including known or estimated locations, to estimate the location of one or more objects in a defined area using vision-based localization; Based on the first request, at least one second network node is selected, wherein the first network node is in the field of view (FOV) of the at least one second network node; Send a second request to the at least one second network node to capture a set of images of the defined area using at least one camera, wherein the set of images includes the first network node; as well as Based on the second request, the image set of the defined region is received from the at least one second network node.