Generating environmental depth map in wireless communication system
By training the ML model in a wireless communication system, and using network nodes and user equipment to jointly generate an environment depth map, the problem of signal distortion and low accuracy of RF generation depth maps in the prior art is solved, and the generation of high-resolution depth maps is achieved.
Patent Information
- Application Number
- CN202380070130.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-10-31
- Filing Date
- 2023-10-30
- Publication Date
- 2025-05-09
AI Technical Summary
The prior art faces the problem of signal distortion when generating depth maps using radio frequency (RF), and the generated depth map has a low resolution and poor accuracy.
By training machine learning (ML) models in a wireless communication system, network nodes and user equipment (UEs) are used to jointly generate an environment depth map. The specific steps include the network node sending a sub-cell area identifier request to the UE, the UE sensing the environment data and inputting it into the trained ML model, and generating a depth map.
It realizes efficient generation of high-resolution depth maps in wireless communication systems, solves the problems of signal distortion and low accuracy in traditional methods, and improves the resolution and accuracy of depth maps.
Smart Images

Figure CN119968874A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to a wireless communication system, and more particularly, to a method and a user equipment (UE) for generating an environment depth map. Background Art
[0002] Considering the development of mobile communications from generation to generation, technology has been developed mainly for human-oriented services such as voice calls, multimedia services, and data services. With the commercialization of 5G (fifth generation) communication systems, the number of connected devices is expected to grow exponentially. More and more devices will be connected to the communication network. Examples of connected devices can include vehicles, robots, drones, home appliances, displays, smart sensors connected to various infrastructures, construction machinery, and factory equipment. Mobile devices are expected to develop in various forms, such as augmented reality glasses, virtual reality headsets, and holographic devices. In order to provide various services by connecting hundreds of billions of devices in the 6G (sixth generation) era, efforts have been made to develop improved 6G communication systems. For these reasons, 6G communication systems are also called super 5G systems.
[0003] It is expected that the 6G communication system will be commercialized around 2030, with a peak data rate of terabit (1000 Gigabits) per second and wireless delay of less than 100 microseconds. Therefore, the speed will be 50 times that of the 5G communication system, and the wireless delay will be one tenth of that of 5G.
[0004] In order to achieve such high data rates and ultra-low latency, 6G communication systems are being considered for implementation in the terahertz band (e.g., 95 GHz to 3 THz band). It is expected that since the terahertz band has more severe path loss and atmospheric absorption than the millimeter wave band introduced in 5G, technologies that can ensure the signal transmission distance (i.e., coverage) will become more critical. As the main technologies to ensure coverage, it is necessary to develop multi-antenna transmission technologies including radio frequency (RF) components, antennas, new waveforms with better coverage than OFDM, beamforming and massive MIMO, full-dimensional MIMO (FD-MIMO), array antennas, and massive antennas. In addition, new technologies for improving the coverage of terahertz band signals are also under continuous discussion, such as metamaterial-based lenses and antennas, orbital angular momentum (OAM), and reconfigurable smart surfaces (RIS).
[0005] In addition, in order to improve spectrum efficiency and overall network performance, the following technologies have been developed for 6G communication systems: full-duplex technology, which enables uplink (UE transmission) and downlink (Node B transmission) to use the same frequency resources simultaneously; network technology, which utilizes satellites, high altitude platform stations (HAPS), etc. in an integrated manner; improved network structure, which is used to support mobile base stations, etc. and realize network operation optimization and automation, etc.; use of AI in wireless communications, by considering AI from the initial stage of developing 6G technology and internalizing end-to-end AI support functions to improve overall network operations; and next-generation distributed computing technology, which overcomes the limitations of UE computing capabilities through ultra-high performance communication and computing resources (MEC, cloud, etc.) accessible on the network.
[0006] Such research and development of 6G communication systems is expected to bring the next hyper-connected experience to every corner of life. In particular, it is expected that services such as truly immersive XR, high-fidelity mobile holograms, and digital replicas can be provided through 6G communication systems.
[0007] Generally speaking, depth maps are usually generated by depth sensors, including but not limited to stereo cameras or lidar sensors. Depth sensors measure the distance of objects in the scene and reconstruct depth maps of objects and scenes in 3D models. Depth sensors can generate high-resolution and high-precision depth maps, but depth sensors have a limited working range.
[0008] Conventional methods and systems utilize radio frequency (RF) in depth map generation to overcome this limited range problem. RF-based depth map generation uses existing communication infrastructure such as spectrum, equipment, and protocols to simultaneously communicate and sense. However, conventional methods and systems face signal distortion issues when detecting the position, motion, and even orientation of an object. In addition, RF-based depth maps have lower resolution and poorer accuracy than depth maps generated by depth sensors including cameras or lidar sensors.
[0009] Therefore, there is a need to address the above-mentioned shortcomings or other deficiencies, or at least provide a useful alternative method to generate high-resolution depth maps using RF.
[0010] The main purpose of the embodiments of this invention is to provide a method and a network node for generating an environment depth map in a wireless communication system. The method includes generating a depth map of one or more sub-cell areas by inputting input data of one or more sub-cell area identifiers into at least one trained ML model.
[0011] Another object of embodiments herein is to train at least one untrained ML model for at least one sub-cell area, wherein the UE trains the at least one untrained ML model in a training phase.
[0012] Another object of embodiments of the present invention is to receive at least one channel data in a bistatic format, wherein the bistatic format includes data of an RF signal received directly or indirectly from at least one transmitter. The method also includes determining input data by converting the channel data from the bistatic format to a monostatic format. Summary of the invention
[0013] Technical issues
[0014] The present disclosure relates to wireless communication systems, and more particularly, to generating an environment depth map in a wireless communication system.
[0015] Technical Solution
[0016] Accordingly, an embodiment of the present invention discloses a method for generating an environmental depth map. The method includes training, by a network node, at least one untrained ML model for at least one sub-cell area. The method includes sending, by the network node, a request for at least one sub-cell area identifier to a user equipment (UE). In addition, the method includes receiving, by the network node, a response including the at least one sub-cell area identifier and input data corresponding to the at least one sub-cell area identifier; wherein the input data is sensed data of the at least one sub-cell area environment. Further, the method includes generating, by the network node, a depth map of the at least one sub-cell area by inputting the input data of the at least one sub-cell area identifier into at least one trained ML model.
[0017] In an embodiment, the method comprises sending, by the network node, a location request for the at least one sub-cell area identifier of the at least one sub-cell area to the UE. In addition, the method comprises receiving, by the network node, a location response comprising the at least one sub-cell area identifier of the UE. Further, the method comprises sending, by the network node, the at least one untrained ML model corresponding to the at least one sub-cell area identifier to the UE for training. In addition, the method comprises receiving, by the network node, at least one trained model corresponding to the at least one sub-cell area identifier from the UE.
[0018] In an embodiment, training the at least one untrained ML model refers to updating, by the UE, at least one parameter based on the sensing data.
[0019] In an embodiment, the network node (110) stores a plurality of trained ML models for a plurality of sub-cell areas, wherein the network node (110) uses the stored plurality of trained ML models for virtual beam selection without sending actual beams in a physical environment.
[0020] Accordingly, an embodiment of the present invention discloses a method for generating an environmental depth map. Further, the method includes the UE requesting at least one trained ML model of at least one sub-cell area by sending at least one sub-cell area identifier to a network node. In addition, the method includes the UE receiving from the network node at least one trained ML model of at least one sub-cell area corresponding to the at least one sub-cell area identifier. Further, the method includes the UE determining input data by sensing the environment of the at least one sub-cell area. In addition, the method includes the UE generating a depth map of the at least one sub-cell area by inputting the input data into at least one trained ML model.
[0021] In an embodiment, the method includes the UE receiving at least one channel data in a dual-base format; wherein the dual-base format includes data of an RF signal received directly or indirectly from at least one transmitter. Further, the method includes the UE determining the input data by converting the channel data from the dual-base format to a single-base format.
[0022] In an embodiment, the UE trains at least one untrained ML model during a training phase.
[0023] In an embodiment, the method includes the UE generating sensing data based on sensing of an environment using the at least one untrained ML model. Further, the method includes the UE using light detection and ranging (LiDAR) data to verify the sensing data. Further, the method includes the UE determining whether the sensing data meets a threshold. Further, when the sensing data does not meet the threshold, the method includes the UE updating parameters of at least one of the convolutional layers and upsampling layers of the at least one untrained ML model. Further, when the sensing data meets the threshold, the method includes the UE treating the at least one untrained ML model as a trained ML model.
[0024] Accordingly, an embodiment of the present invention discloses a network node for generating an environmental depth map. The network node includes a memory, a processor, and a network node depth map controller communicatively connected to the memory and the processor. The network node depth map controller is configured to train at least one untrained ML model for at least one sub-cell area. Further, the network node depth map controller is configured to send a request for at least one sub-cell area identifier to a UE. In addition, the network node depth map controller is configured to receive a response including the at least one sub-cell area identifier and input data corresponding to the at least one sub-cell area identifier; wherein the input data is sensed data of the environment of the at least one sub-cell area. Further, the network node depth map controller is configured to generate a depth map of the at least one sub-cell area by inputting the input data of the at least one sub-cell area identifier into at least one trained ML model.
[0025] Accordingly, an embodiment of the present invention discloses a UE for generating an environmental depth map. The UE includes a memory, a processor, and a UE depth map controller communicatively connected to the memory and the processor. The UE depth map controller is configured to request at least one trained ML model of at least one sub-cell area by sending at least one sub-cell area identifier to a network node. Further, the UE depth map controller is configured to receive from the network node at least one trained ML model of the at least one sub-cell area corresponding to the at least one sub-cell area identifier. In addition, the UE depth map controller is configured to determine input data by sensing the environment of the at least one sub-cell area. Further, the UE depth map controller is configured to generate a depth map of the at least one sub-cell area by inputting the input data into the at least one trained ML model.
[0026] These and other aspects of the embodiments herein will be better understood and appreciated when considered in conjunction with the following description and accompanying drawings. It should be understood that the following description, while indicating preferred embodiments and numerous specific details thereof, is given by way of example and not limitation. Many changes and modifications may be made within the scope of the embodiments herein without departing from the inventive concept of the embodiments herein, and the embodiments herein include all such modifications.
[0027] Advantageous Effects of the Invention
[0028] Advantages and salient features of the invention will become apparent to those skilled in the art from the following detailed description, which, taken in conjunction with the annexed drawings, discloses exemplary embodiments of the invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] These and other features, aspects and advantages of the present method are illustrated in the accompanying drawings, in which the same reference letters represent corresponding parts. From the following description in conjunction with the accompanying drawings, the various embodiments described herein will be better understood, in which:
[0030] Figure 1 A block diagram of a network node and a UE for generating an environment depth map in a wireless communication system according to an embodiment disclosed herein is shown;
[0031] Figure 2 A flowchart of a method for generating an environment depth map according to an embodiment disclosed herein is shown;
[0032] Figure 3 An example of using RF communication signals for inference according to embodiments disclosed herein is shown;
[0033] Figure 4 shows an example of a depth map of the physical world according to embodiments disclosed herein;
[0034] Figure 5 An example of a base station connected to multiple UEs to estimate a depth map for a smaller area according to an embodiment disclosed herein is shown;
[0035] Figure 6 An example of estimating a depth map for a UE for each smaller area having a fixed transmitter and a mobile receiver according to an embodiment disclosed herein is shown;
[0036] Figure 7 A flow chart showing an end-to-end process of generating an environment depth map according to an embodiment disclosed herein;
[0037] Figure 8 A flowchart showing the steps involved in the training phase of an untrained ML model according to an embodiment disclosed herein;
[0038] Fig. 9 A flowchart showing the steps involved in the inference phase of a trained ML model according to embodiments disclosed herein;
[0039] Fig.10 An example of a base station connected to multiple UEs during the training phase of an ML model according to an embodiment disclosed herein is shown;
[0040] Fig.11 A schematic diagram showing conversion of dual-base format data into single-base format data according to an embodiment disclosed herein;
[0041] Fig.12 A schematic diagram showing converting dual-base format data into monostatic format data by preprocessing according to an embodiment disclosed herein;
[0042] Fig.13 A schematic diagram showing an overview of preprocessing according to an embodiment disclosed herein;
[0043] Fig.14 A schematic diagram showing preprocessing accompanying spatial transformation according to an embodiment disclosed herein;
[0044] Fig.15 A schematic diagram showing preprocessing combining spatial transformation and angle of arrival information according to an embodiment disclosed herein is shown;
[0045] Fig.16 A schematic diagram showing pre-processing accompanying line of sight (LOS) search according to an embodiment disclosed herein;
[0046] Fig.17 A schematic diagram showing preprocessing with removal of transmission effects according to an embodiment disclosed herein;
[0047] Fig.18 A schematic diagram showing preprocessing combining removal of transmission influence and arrival angle information according to an embodiment disclosed herein is shown;
[0048] Fig.19 A schematic diagram showing pre-processing accompanied by re-adjustment of power according to an embodiment disclosed herein;
[0049] Fig. 20 A schematic diagram showing a comparison between the LIDAR PCD according to the embodiment disclosed herein and the PCD of the proposed method;
[0050] Fig.21 A schematic diagram showing input and output for ML model training according to an embodiment disclosed herein;
[0051] Fig. 22 A schematic diagram showing an ML model architecture according to an embodiment disclosed herein; and
[0052] Fig.23 A schematic diagram of a test sample with error analysis according to an embodiment disclosed herein is shown. DETAILED DESCRIPTION
[0053] The embodiments of this document and their various features and beneficial details will be more fully described in conjunction with the non-limiting embodiments shown in the accompanying drawings and described below. In order to avoid unnecessarily obscuring the embodiments of this document, descriptions of well-known components and processing technologies will be omitted. In addition, the various embodiments described herein are not mutually exclusive, because some embodiments can be combined with one or more other embodiments to form new embodiments. The word "or" used in this document, unless otherwise specified, refers to a non-exclusive "or". The examples given herein are only for the convenience of understanding the implementation methods of the embodiments of this document, and further help those skilled in the art to implement the embodiments of this document. Therefore, these examples should not be regarded as limiting the scope of the embodiments of this document.
[0054] According to the convention in the art, the embodiments can be described and illustrated by the modules that perform the described functions. These modules may be referred to as units or modules, etc. in this article, and they are physically implemented by analog or digital circuits (such as logic gates, integrated circuits, microprocessors, microcontrollers, storage circuits, passive electronic components, active electronic components, optical components, hard-wired circuits, etc.), and can be selectively driven by firmware. For example, these circuits may be embodied in one or more semiconductor chips, or embodied on substrate supports such as printed circuit boards. The circuits constituting the modules may be implemented by dedicated hardware, or by processors (such as one or more programmed microprocessors and related circuits), or by dedicated hardware that performs part of the functions of the modules in combination with processors that perform other functions of the modules. Without departing from the scope of the present disclosure, each module in the embodiment may be physically separated into two or more interacting and independent modules. Similarly, without departing from the scope of the present disclosure, the modules in the embodiments may also be physically combined into more complex modules.
[0055] The accompanying drawings help to intuitively understand various technical features, and it should be understood that the embodiments presented herein are not limited by the accompanying drawings. Therefore, the present disclosure should be understood to include any changes, equivalents and alternatives in addition to the contents specifically described in the accompanying drawings. Although the first, second and other terms may be used herein to describe various elements, these elements should not be limited by these terms. These terms are usually only used to distinguish different elements.
[0056] Accordingly, an embodiment of the present invention discloses a method for generating an environmental depth map. The method includes training, by a network node, at least one untrained ML model for at least one sub-cell area. The method also includes sending, by the network node, a request for at least one sub-cell area identifier to a user equipment (UE). In addition, the method includes receiving, by the network node, a response including at least one sub-cell area identifier and input data corresponding to the at least one sub-cell area identifier; wherein the input data is sensed data of the environment of the at least one sub-cell area. Further, the method includes generating, by the network node, a depth map of the at least one sub-cell area by inputting the input data of the at least one sub-cell area identifier into at least one trained ML model.
[0057] Accordingly, an embodiment of the present invention discloses a method for generating an environmental depth map. Further, the method includes the UE requesting at least one trained ML model of at least one sub-cell area by sending at least one sub-cell area identifier to a network node. In addition, the method includes the UE receiving at least one trained ML model of at least one sub-cell area corresponding to the at least one sub-cell area identifier from the network node. Further, the method includes the UE determining input data by sensing the environment of at least one sub-cell area. In addition, the method includes the UE generating a depth map of the at least one sub-cell area by inputting the input data into the at least one trained ML model.
[0058] Accordingly, an embodiment of the present invention discloses a network node for generating an environmental depth map. The network node includes a memory, a processor, and a network node depth map controller communicatively connected to the memory and the processor. The network node depth map controller is configured to train at least one untrained ML model for at least one sub-cell area. Further, the network node depth map controller is configured to send a request for at least one sub-cell area identifier to a UE. In addition, the network node depth map controller is configured to receive a response including at least one sub-cell area identifier and input data corresponding to the at least one sub-cell area identifier; wherein the input data is sensed data of the environment of at least one sub-cell area. Further, the network node depth map controller is configured to generate a depth map of at least one sub-cell area by inputting the input data of at least one sub-cell area identifier into at least one trained ML model.
[0059] Accordingly, an embodiment of the present invention discloses a UE for generating an environmental depth map. The UE includes a memory, a processor, and a UE depth map controller communicatively connected to the memory and the processor. The UE depth map controller is configured to request at least one trained ML model of at least one sub-cell area by sending at least one sub-cell area identifier to a network node. Further, the UE depth map controller is configured to receive at least one trained ML model of at least one sub-cell area corresponding to at least one sub-cell area identifier from the network node. In addition, the UE depth map controller is configured to determine input data by sensing the environment of at least one sub-cell area. Further, the UE depth map controller is configured to generate a depth map of at least one sub-cell area by inputting the input data into at least one trained ML model.
[0060] In traditional systems and methods for AR application scenarios such as video calling, traditional depth map sensors such as LiDAR, 3D cameras, etc. need to estimate the depth map first, and then transmit that high-throughput information to the base station / router in a low-latency manner, possibly through a millimeter wave link, to obtain a high-quality calling experience. However, the proposed system utilizes the millimeter wave link itself used to communicate with the base station / router to sense the depth map information, making the depth map perception and communication process more efficient than using other sensors such as LiDAR and cameras.
[0061] Conventional systems and methods generate depth maps using LIDAR data, for which specialized sensor hardware is required, which older UEs do not have. Unlike conventional methods and systems, the proposed system performs depth map estimation (DME) using only RF channel data (CIR) that is readily available in any UE with communication capabilities. The proposed system and method is able to generate high-resolution LIDAR point clouds (PCDs) or depth maps from low-resolution RF data.
[0062] Different from traditional methods and systems, the proposed framework demonstrates the possibility of generating depth maps by utilizing environmental information (obstacle information) in commonly received communication signals in next-generation systems.
[0063] Unlike conventional methods and systems, the proposed system estimates the depth map of the environment by resolving the positions of obstacles from RF signals obtained in a bistatic format, where RF signals from TX are reflected from obstacles and then propagate to RX.
[0064] Unlike conventional methods and systems, the proposed system performs a pre-processing process to extract the monostatic power (MSP) seen from the perspective of the receiver at a specific location in the environment from the conventional signal used for communication between the transmitter and the receiver. The pre-processing step handles the deterministic aspects of converting the MIMO channel data in a bistatic format to a monostatic format. This makes the learning of the ML model faster and more understandable.
[0065] Different from conventional methods and systems, the proposed system exchanges information between the base station and the UE during the training and inference phases to deploy the proposed solution.
[0066] Different from traditional methods and systems, the proposed system converts low-resolution, low-dimensional input data (MSP) into high-resolution, feature-rich LiDAR data through ML models, which captures the surfaces of objects, obstacles, people, etc. in the surrounding environment in detail through dense point clouds. The perception capabilities of the proposed system far exceed basic applications such as counting people in a room and locating specific objects, which are traditionally achieved through wireless sensing.
[0067] Unlike conventional methods and systems, the proposed system utilizes low-resolution RF signals to return high-resolution depth maps, which are traditionally generated using LiDAR point clouds (PCDs) or camera images as input.
[0068] Different from conventional methods and systems, the proposed system utilizes low-resolution RF data to generate LiDAR-like high-resolution point clouds in terms of distance and angle.
[0069] Unlike conventional methods and systems, the proposed system estimates the depth map of the environment by solving for the positions of obstacles using RF signals obtained in a bistatic format, where RF signals from TX are reflected from obstacles and then propagate to RX.
[0070] Unlike conventional methods and systems, the proposed system uses a pre-processing process to extract the MSP from the perspective of the receiver at a specific location in the environment from the conventional signal used for communication between the transmitter and the receiver. The pre-processing step deals with the deterministic aspects of converting the MIMO channel data in a bistatic format to a monostatic format.
[0071] Different from traditional methods and systems, the proposed system extracts spatial information of the surrounding environment by reusing the existing RF hardware on 5G-enabled mobile phones in ISAC format. The proposed system is deployed on the basis of existing hardware with minimal protocol additions.
[0072] Unlike traditional methods and systems, the proposed system leverages the existing smart device ecosystem to enable the creation of digital twins by integrating readily available RF sensor data from multiple sensors. The proposed system is also capable of performing activity recognition such as posture estimation, human activity detection, fall detection, etc. using only wireless RF signals.
[0073] Unlike traditional methods and systems, the proposed system integrates data from multiple sensors (such as cameras, LiDAR, RF, etc.) to create a powerful digital twin, enabling virtual testing, validation, and self-optimization of smart city B5G networks.
[0074] Different from the traditional methods and systems, the proposed system uses the mmWave signals in B5G to realize AR / VR in COTS hardware without the need for dedicated hardware such as LiDAR and 3D cameras.
[0075] Different from traditional methods and systems, the proposed system provides rich point cloud data from widely available RF data that can perceive higher-level activities such as human gestures, pedestrian movements, etc.
[0076] Referring now to the drawings, and in particular to Figure 1 to Figure 2 5, wherein like reference characters indicate corresponding features throughout the various figures, which illustrate preferred embodiments.
[0077] Figure 1 A block diagram of a network node (110) and a UE (120) for generating an environment depth map in a wireless communication system (100) according to an embodiment disclosed herein is shown.
[0078] In an embodiment, the network node (110) includes a memory (111), a processor (113), a communicator (112), a network node deep map controller (114), and an ML memory (115).
[0079] The memory (111) is configured to store image frames representing various stages of an event. The memory (111) includes a non-volatile storage element. Examples of such non-volatile storage elements include a magnetic hard disk, an optical disk, a floppy disk, a flash memory, or an electrically programmable memory (EPROM) or an electrically erasable programmable (EEPROM) memory. In addition, in some examples, the memory (111) is considered to be a non-transitory storage medium. The term "non-transitory" indicates that the storage medium is not embodied in a carrier or propagating signal. However, the term "non-transitory" should not be understood to mean that the memory (111) is non-removable. In some examples, the memory (111) is configured to store a large amount of information. In some examples, the non-transitory storage medium stores data that changes over time (e.g., in a random access memory (RAM) or a cache).
[0080] The processor (113) includes one or more processors. The one or more processors may be general-purpose processors, such as a central processing unit (CPU), an application processor (AP), etc., or may be a unit used only for graphics processing, such as a graphics processing unit (GPU), a visual processing unit (VPU), and / or an ML-specific processor such as a neural processing unit (NPU). The processor (113) includes multiple cores and is configured to execute instructions stored in the memory (111).
[0081] In an embodiment, the communicator (112) includes electronic circuitry specific to a standard supporting wired or wireless communication. The communicator (112) is configured to communicate internally between internal hardware components of the wearable AR device (100) and to communicate with external devices over one or more networks.
[0082] In an embodiment, the network node depth map controller (114) is configured to train at least one untrained ML model for at least one sub-cell area. In addition, the network node depth map controller (114) is configured to send a request for at least one sub-cell area identifier to the UE (120). Further, the network node depth map controller (114) is configured to receive a response including at least one sub-cell area identifier and input data corresponding to the at least one sub-cell area identifier; wherein the input data is sensed data of at least one sub-cell area environment. Further, the network node depth map controller (114) is configured to generate a depth map of at least one sub-cell area by inputting the input data of the at least one sub-cell area identifier into at least one trained ML model.
[0083] In an embodiment, the network node depth map controller (114) is configured to send a location request for at least one sub-cell area identifier of at least one sub-cell area to the UE (120). In addition, the network node depth map controller (114) is configured to receive a location response including at least one sub-cell area identifier of the UE (120), and send at least one untrained ML model corresponding to the at least one sub-cell area identifier to the UE (120) for training. Further, the network node depth map controller (114) is configured to receive at least one trained model corresponding to the at least one sub-cell area identifier from the UE (120).
[0084] In an embodiment, at least one untrained ML model is trained by the UE (120) by updating at least one parameter based on the sensing data.
[0085] In an embodiment, the network node (110) stores a plurality of trained ML models for a plurality of sub-cell areas, wherein the network node (110) uses the stored plurality of trained ML models to virtually select beams in a physical environment without transmitting actual beams.
[0086] The network node depth map controller (114) is implemented by processing circuits such as logic gates, integrated circuits, microprocessors, microcontrollers, memory circuits, passive electronic components, active electronic components, optical components, hard-wired circuits, etc., and can be selectively driven by firmware. For example, these circuits can be embodied in one or more semiconductor chips, or on a substrate support such as a printed circuit board.
[0087] At least one of the multiple modules / components of the network node depth map controller (114) is implemented by an ML model. Functions related to the ML model are executed by the memory (111) and the processor (113). One or more processors (113) control the processing of input data according to predefined operating rules or ML models stored in non-volatile memory and volatile memory. The predefined operating rules or artificial intelligence models are obtained through training or learning.
[0088] Here, acquisition through learning means that a predefined operation rule or ML model with desired characteristics is obtained by applying a learning process to a plurality of learning data. The learning is performed in the device itself that executes ML according to the embodiment and / or is implemented by a separate server / system.
[0089] ML models consist of multiple neural network layers. Each layer has multiple weight values, and layer operations are performed by the judgment of the previous layer and the operation of multiple weights. Examples of neural networks include, but are not limited to, convolutional neural networks (CNNs), deep neural networks (DNNs), recurrent neural networks (RNNs), restricted Boltzmann machines (RBMs), deep belief networks (DBNs), bidirectional recurrent deep neural networks (BRDNNs), generative adversarial networks (GANs), and deep Q networks.
[0090] The learning process is a method of using a plurality of learning data to train a predetermined target device (e.g., a robot) so that the target device can make a decision or prediction. Examples of the learning process include, but are not limited to, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning.
[0091] In an embodiment, the network node (110) includes a memory (121), a processor (123), a communicator (122), a UE depth map controller (124), and an ML memory (115).
[0092] The memory (121) is configured to store image frames representing various stages of an event. The memory (121) includes a non-volatile storage element. Examples of such non-volatile storage elements include a magnetic hard disk, an optical disk, a floppy disk, a flash memory, or an electrically programmable memory (EPROM) or an electrically erasable programmable (EEPROM) memory. In addition, in some examples, the memory (121) is considered to be a non-transitory storage medium. The term "non-transitory" indicates that the storage medium is not embodied in a carrier or propagating signal. However, the term "non-transitory" should not be understood to mean that the memory (121) is non-removable. In some examples, the memory (121) is configured to store a large amount of information. In some examples, the non-transitory storage medium stores data that changes over time (e.g., in a random access memory (RAM) or a cache).
[0093] The processor (123) includes one or more processors. The one or more processors may be general-purpose processors, such as a central processing unit (CPU), an application processor (AP), etc., or may be a unit used only for graphics processing, such as a graphics processing unit (GPU), a visual processing unit (VPU), and / or an AI-specific processor such as a neural processing unit (NPU). The processor (123) includes multiple cores and is configured to execute instructions stored in the memory (121).
[0094] In an embodiment, the communicator (122) includes electronic circuitry specific to a standard supporting wired or wireless communication. The communicator (122) is configured to communicate internally between internal hardware components of the wearable AR device (100) and to communicate with external devices over one or more networks.
[0095] In an embodiment, the UE depth map controller (124) includes an ML model trainer (125), a depth map generator (126), and an input data determiner (127).
[0096] In an embodiment, the depth map generator (126) requests at least one trained ML model for at least one sub-cell area by sending at least one sub-cell area identifier to the network node (110). In addition, the depth map generator (126) receives at least one trained ML model for at least one sub-cell area corresponding to the at least one sub-cell area identifier from the network node (110). Further, the depth map generator (126) determines input data by sensing the environment of the at least one sub-cell area. Further, the depth map generator (126) generates a depth map for the at least one sub-cell area by inputting the input data into the at least one trained ML model.
[0097] In an embodiment, the input data determiner (127) receives at least one channel data in a bistatic format; wherein the bistatic format includes data of an RF signal received directly or indirectly from at least one transmitter. Further, the input data determiner (127) determines the input data by converting the channel data from the bistatic format to a monostatic format.
[0098] In an embodiment, the ML model trainer (125) trains at least one untrained ML model during a training phase.
[0099] In an embodiment, the ML model trainer (125) generates sensory data based on sensing of an environment using at least one untrained ML model. In addition, the ML model trainer (125) uses light detection and ranging (LiDAR) data to verify the sensory data. Further, the ML model trainer (125) determines whether the sensory data meets a threshold. When the sensory data does not meet the threshold, the ML model trainer (125) updates the parameters of at least one of the convolutional layers and upsampling layers of at least one untrained ML model. Wherein, the threshold is defined as a certain percentage of the error between the sensory data prediction value and the LiDar data (for example: 1%-5%). Further, when the sensory data meets the threshold, the ML model trainer (125) treats at least one untrained ML model as a trained ML model.
[0100] The UE depth map controller (124) is implemented by processing circuits such as logic gates, integrated circuits, microprocessors, microcontrollers, memory circuits, passive electronic components, active electronic components, optical components, hard-wired circuits, etc., and can be selectively driven by firmware. For example, these circuits can be embodied in one or more semiconductor chips, or on a substrate support such as a printed circuit board.
[0101] At least one of the multiple modules / components of the UE depth map controller (124) is implemented by an ML model. Functions related to the ML model are executed by the memory (121) and the processor (123). One or more processors (123) control the processing of input data according to predefined operating rules or ML models stored in non-volatile memory and volatile memory. The predefined operating rules or artificial intelligence models are obtained through training or learning.
[0102] Here, acquisition through learning means that a predefined operation rule or ML model with desired characteristics is obtained by applying a learning process to a plurality of learning data. The learning is performed in the device itself that executes ML according to the embodiment and / or is implemented by a separate server / system.
[0103] ML models consist of multiple neural network layers. Each layer has multiple weight values, and layer operations are performed by the judgment of the previous layer and the operation of multiple weights. Examples of neural networks include, but are not limited to, convolutional neural networks (CNNs), deep neural networks (DNNs), recurrent neural networks (RNNs), restricted Boltzmann machines (RBMs), deep belief networks (DBNs), bidirectional recurrent deep neural networks (BRDNNs), generative adversarial networks (GANs), and deep Q networks.
[0104] The learning process is a method of using a plurality of learning data to train a predetermined target device (e.g., a robot) so that the target device can make a decision or prediction. Examples of the learning process include, but are not limited to, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning.
[0105] although Figure 1 The hardware elements of the network node (110) and the UE (120) are shown, but it should be understood that other embodiments are not limited thereto. In other embodiments, the network node (110) and the UE (120) include fewer or more elements. In addition, the labels or names of the elements are only for illustrative purposes and do not limit the scope of the embodiments. One or more components can be combined to perform the same or substantially similar functions.
[0106] Figure 2 A flowchart of a method for generating an environment depth map according to an embodiment disclosed herein is shown.
[0107] At step 201, the network node (110) trains at least one untrained ML model for at least one sub-cell area.
[0108] At step 202, the network node (110) sends a request for at least one sub-cell area identifier to the UE (120).
[0109] At step 203, the network node (110) receives a response including the at least one sub-cell area identifier and input data corresponding to the at least one sub-cell area identifier.
[0110] At step 204, the network node (110) generates a depth map of at least one sub-cell region by inputting the input data of the at least one sub-cell region identifier into at least one trained ML model.
[0111] In an embodiment, the network node (110) may also be referred to as a base station, a network element, an access point, etc.
[0112] The various actions, behaviors, modules, steps, etc. in the method may be performed in the order shown, or in a different order or simultaneously. In addition, in some embodiments, certain actions, behaviors, modules, steps, etc. may be omitted, added, modified, skipped, etc. without departing from the scope of the embodiments.
[0113] Figure 3 An example of using RF communication signals for inference according to embodiments disclosed herein is shown.
[0114] RF communication signals utilize existing communication infrastructure such as spectrum, equipment, and protocols to simultaneously communicate and sense, detecting the location, movement, and direction of objects including but not limited to persons or obstacles by inferring signal distortions.
[0115] Figure 4 An example of a physical world depth map according to embodiments disclosed herein is shown.
[0116] In an embodiment, the proposed system estimates a depth map of the environment (401) at each receiver location using mmWave signals and an ML model.
[0117] In an embodiment, the proposed system trains an ML model with MIMO channel impulse response (402) as input and generated Lidar point cloud (403) as output.
[0118] Figure 5 An example of a base station connected to multiple UEs to estimate a depth map for a smaller area according to embodiments disclosed herein is shown. Figure 5 There is one base station in a cell area in , and there are multiple UEs in the cell area. The cell area is divided into smaller location areas, called sub-cell areas, because it is feasible to determine depth maps for these smaller areas, some of which are indoors and some are outdoors.
[0119] BS B is located in a cell, and the cell is divided into multiple indoor and outdoor sub-areas: A_=1... Each of these areas has multiple UEs: UE_([j,A]_i): j=0...M,i=1...N. ML models are trained for smaller areas for depth map estimation.
[0120] The information exchange at the BS for the depth map estimation model is as follows:
[0121] -BS->UE: request DME information
[0122] -UE->BS: Send: 1. (MSP) [B1]; 2. Position and direction
[0123] -At BS:
[0124] Parameters for each UE (120):
[0125] oDetermine the location area based on the location
[0126] o Using the weights stored in the region, run the ML model on the MSP of the UE (120) [B2]
[0127] o Inferring a depth map and, based on the location and orientation of the UE (120), adjusting the depth map in the GCS for the location area.
[0128] oRecord the estimated depth maps at all locations to get a complete depth map of the cell area
[0129] Figure 6 An example of estimating a depth map for a UE (120) for each small area having a fixed transmitter and a mobile receiver according to embodiments disclosed herein is shown.
[0130] For indoor scenarios, the UE (120) estimates a depth map in the indoor environment. Figure 6 Shown is a room where measurements are taken. The room has a transmitter at a fixed position on the left, and receivers that move around in different areas of the room. As the position and orientation of the receiver changes, the observation of the room changes. The problem is further complicated by the fact that the position and orientation of the transmitter and receiver are unknown.
[0131] Figure 7 A flow chart showing an end-to-end process of generating an environment depth map according to embodiments disclosed herein.
[0132] The UE (120) converts the tapped channel into reflected single-base power (MSP) by preprocessing (701). In addition, the ML model (702) is trained at the UE (120) and inferred at the BS and UE (120) to predict the depth map (703). The estimation (704) of the depth map includes an ML model training phase (705) and an ML model inference phase (706).
[0133] The steps of the ML model training phase (705) are as follows:
[0134] At step 709, the base station sends a location request for at least one sub-cell area identifier of at least one sub-cell area to the UE (120).
[0135] At step 710, the base station receives a location response including at least one sub-cell area identifier of the UE (120).
[0136] In step 711, the base station sends at least one untrained ML model corresponding to the at least one sub-cell area identifier to the UE (120) for training, and in step 712, the UE (120) trains the ML model.
[0137] In step 711, the base station receives at least one trained model corresponding to the at least one sub-cell area identifier from the UE (120).
[0138] At step 714, the base station stores the trained ML model.
[0139] The ML model inference phase (706) and application steps at the base station (707) are as follows:
[0140] At step 715, the base station sends a request for at least one sub-cell area identifier to the UE.
[0141] In step 716, the base station receives a response including at least one sub-cell area identifier and input data corresponding to the at least one sub-cell area identifier, wherein the input data is sensed data of the at least one sub-cell area environment. In an embodiment, the input data is reflected single-base power (MSP).
[0142] In step 717 , the base station generates a depth map of the at least one sub-cell area by inputting the input data of the at least one sub-cell area identifier into at least one trained ML model.
[0143] The ML model inference phase (706) and application steps at the US (708) are as follows:
[0144] At step 718, the UE (120) requests at least one trained ML model for the at least one sub-cell area by sending at least one sub-cell area identifier to the base station.
[0145] At step 719, the UE (120) receives at least one trained ML model of at least one sub-cell area corresponding to the at least one sub-cell area identifier from the base station.
[0146] In step 720, the UE (120) determines the MSP by sensing the environment of the at least one sub-cell area, and generates a depth map of the at least one sub-cell area by inputting the MSP into at least one trained ML model.
[0147] Figure 8 A flowchart showing the steps involved in the training phase of an untrained ML model according to an embodiment disclosed herein is shown. The ML model is trained for a small area for depth map estimation.
[0148] In step 801, the base station asks the UE for a location area. In step 802, the UE sends its location to the base station. In step 803, the BS determines the location area A. i , and sends the untrained ML model to the UE, where BS B is located in a cell and the cell is divided into multiple indoor and outdoor sub-areas: A i , i = 1…N, each of these areas has multiple UEs distributed: UE j,Ai ,j=0…M,i=1…N.
[0149] At step 804, the UE (120) performs preprocessing, and at step 805, the UE (120) trains an ML model to generate a depth map (at step 806).
[0150] At step 807, the UE (120) sends the trained model to the BS, and the BS stores the ML model based on the location area and decides when to end the training phase (at step 808).
[0151] In an embodiment, during ML model training, the UE (120) estimates the channel and pre-processes it to convert it to MSP: reflected monostatic power at the UE (120) (at step 804). The MSP is used as input to the ML model (at step 805), which predicts a depth map of the surrounding environment.
[0152] In an embodiment, the UE sends the trained ML model to the BS considering its own position, and the BS stores the trained ML model. Therefore, in the inference phase, when any UE (120) needs the trained ML model, the BS provides the trained ML model to the UE (120), so that the UE (120) can estimate the depth map without retraining.
[0153] Fig. 9 A flowchart showing the steps involved in the inference phase of a trained ML model according to embodiments disclosed herein.
[0154] In step 901, the UE (120) requests DME information (i.e., for the ML model obtained from the BS based on the UE (120) application). The BS sends the ML model to the UE (120), and the UE (120) uses the model to generate a depth map using the MSP obtained by preprocessing the input channel as input (steps 904 and 903).
[0155] In step 902, the BS determines whether it has a trained model, wherein the BS cloud or database stores the trained ML model and its weight for each sub-region (step 912). The trained ML model generates a depth map (step 906).
[0156] In step 907, the BS requests the UE for a location area and MSP (ie, generated by the UE (120) by preprocessing the channel for BS-based applications). The BS uses the corresponding ML model on the MSP to obtain the corresponding depth map according to the location area.
[0157] In step 909, the BS determines a location area based on the location of the UE. Further, the BS applies the MSP to the trained ML model based on the location area to generate a depth map. In addition, the parameters "DMEtrainingInfoRequest", "DMEInferInfoRequest" are exchanged between the PHY / MAC of the UE and the BS, and involve physical layer-based processing, and require standard support and additional functions. Moreover, the sub-cell area and MSP parameters are processed in the PHY, and these parameters can be transmitted during the depth map estimation window (ie, when the BS enables the depth map estimation flag).
[0158] In an embodiment, the information is exchanged through RRC reconfiguration on the BS side and UE assistance information on the UE side.
[0159] In conventional methods and systems, RRC reconfiguration (BS->UE) is as follows:
[0160]
[0161] In the conventional method and system, the UE assistance information IE (UE->BS) is as follows:
[0162] DMETrainingInfoResponse-IEs::=SEQUENCE{UE-pos
[0163] position value OPTIONAL,
[0164] }
[0165] DMEinferInfoResponse-IEs::=SEQUENCE{UE-pos position value MSP mono-static power OPTIONAL,
[0166] }
[0167] DME-TrainedModel-IEs::=SEQUENCE{MLModelParams Model-params OPTIONAL,
[0168] }
[0169] UE-DMEinferInfoResponse-IEs::=SEQUENCE{UE-pos position valueOPTIONAL,
[0170] }
[0171] In the embodiments, the detailed information of each term is as follows:
[0172] -DMEtrainingInfoRequest: A message sent by the BS to the UE to request the start of DME training and inquire about the location.
[0173] -DMETrainingInfoResponse: A response message sent by the UE to the BS and sends location / area information.
[0174] -DMEInferInfoRequest: A message sent by the BS to the UE to request the MSP and location area for inference at the BS.
[0175] -DMEInferInfoResponse: A response message sent by the UE to the BS to send location / area information and MSP.
[0176] -UE-DMEinferInfoResponse: A message sent by the UE to the BS to request DME information such as the ML model from the BS and to send location area information.
[0177] -DME-MLModel: message sent by BS to UE.
[0178] -After receiving DMETrainingInfoResponse, send the ML model for training.
[0179] -After receiving UE-DMEinferInfoResponse, send the ML model for inference.
[0180] Fig.10 An example of a base station connected to multiple UEs during the training phase of an ML model according to an embodiment disclosed herein is shown.
[0181] In an embodiment, the cell area of the BS is divided into multiple (a total of N areas) smaller areas, such as Fig.10Location areas 1 to 16 are shown. At the beginning of the training phase, the BS provides random weights and the ML model type to the UE. The UE then trains the ML model and provides the weights to the BS. In addition, the BS stores the weights and information about the location areas.
[0182] In an embodiment, when a new UE (120) enters a location area, the BS provides the stored weights, and then the new UE (120) trains the ML model and provides the trained weights to the BS. The BS updates the weights for that location. In each smaller location area including indoors (501) and outdoors (502), the UE trains the ML model and provides the trained weights to the BS.
[0183] The information exchange during the training phase is as follows:
[0184] -BS->UE(120): Request DME training information
[0185] -UE(120)->BS: Location
[0186] The BS determines a location area for the UE (120).
[0187] -BS->UE(120): 1. The stored weights for the location area
[0188] 2. ML Model Types
[0189] The UE (120) uses the pre-processed MSP to train the ML model.
[0190] -UE(120)->BS: training weight for this location area
[0191] The BS stores the weight of the location area.
[0192] The information exchange for the depth map estimation model at the UE (120) is as follows:
[0193] -UE(120)->BS:
[0194] 1. Location
[0195] 2. Request for model weights: DME parameters
[0196] The BS determines the location area.
[0197] - BS->UE (120): transmits the stored weights for the location area.
[0198] The UE (120) uses the MSP (single-base power) obtained through preprocessing ([B1]) to run the ML model [B2] and predict the depth map.
[0199] Fig.11 A schematic diagram showing conversion of dual-base format data into monobase format data according to an embodiment disclosed herein is shown.
[0200] In an embodiment, the ML model is trained using an input RF dataset (step 1101) with training lidar data as ground truth for training (step 1105).
[0201] The input RF data is sent and received at physically separated locations, which is also called bistatic mode (step 1102), while the LiDAR sends and receives laser signals from the same location, which is also called monostatic mode. Therefore, in order to standardize the two types of data during training, the proposed system converts the bistatic RF data into monostatic data (step 1106).
[0202] In an embodiment, in a bistatic format, the RF signal from TX reflects from obstacles and then propagates to RX. The proposed system aims to estimate the depth map of the environment by solving for the positions of these obstacles.
[0203] In an embodiment, the input to the ML model is RF data, i.e., channel impulse response (CIR) (step 1103), and the output is the LiDAR point cloud. Therefore, pre-processing is required to handle the deterministic aspects of converting the MIMO channel data in a bistatic format to a monostatic format, so that the learning of the ML model is faster and easier to understand.
[0204] Fig.12 A schematic diagram showing conversion of dual-base format data into monostatic format data through preprocessing according to an embodiment disclosed herein is shown.
[0205] The pre-processing step 1201 handles the deterministic aspects of converting the MIMO channel data in a bistatic format to a monostatic format. This makes the learning of the ML model faster and more understandable.
[0206] At each RX location (step 1202), the bistatic RF data is converted to a monostatic format (1203). At step 1104, the ML model is trained by fitting it to a LiDAR ground truth of similar structure.
[0207] The steps involved in preprocessing are as follows:
[0208] Input H CIR : TX antenna × RX antenna × delay
[0209] (nH TX ×nV TX )×(nH RX ×nV RX )×Delay
[0210] Channel Impulse Response 5D Complex
[0211] Output P RX :AoA×Delay
[0212] θ AoA ×Φ AoA ×Delay
[0213] Single base power 3D real scene
[0214] The proposed method makes the training part of the ML model simpler and interpretable. This avoids the model learning the existence of MIMO TX and its location and other characteristics by itself through a large amount of training data, as such learning results are not guaranteed and depend on the DL architecture.
[0215] In an embodiment, the proposed system handles deterministic aspects of model fitting such as data transformation outside of the DL model.
[0216] Fig.13 A schematic diagram showing an overview of preprocessing according to embodiments disclosed herein.
[0217] Reference Fig.13 , the figure (P RX The ) shows the pre-processed data for one RX location and the power at the RX for 64 different angles of arrival (AOA) and 100 delays in each direction. The RF data (P RX ) is similar in structure to LiDAR data (both are in single-base format), which makes it easier for ML models to fit RF data to LiDAR ground truth data. Figures 14 to 20 The steps involved are shown in detail.
[0218] Fig.14 A schematic diagram showing pre-processing of adjoint spatial transformation (1301) according to an embodiment disclosed herein.
[0219] In an embodiment, the proposed system pre-processes the input channel impulse response (1305) to perform a spatial transform (1301). The proposed system converts the signal data from antenna space to angle beam space. The spatial transform (1301) helps to obtain the sensing frame from the communication frame. For a delayed matrix H_beam diagram, each entry represents the power of an AOA × angle of departure (AOD) beam pair.
[0220] Fig.15 A schematic diagram showing the combination of spatial transformation (1301) and pre-processing of AoA information according to an embodiment disclosed herein is shown.
[0221] The proposed system converts the antenna space into angle space by performing 2D-FFT on the channel response H_CRI, where the TX antenna dimension is transformed to obtain AoD, and the RX antenna dimension is transformed to obtain AoA.
[0222] -TX Antenna(nH TX ×nV TX )<-2D->FFT AoD(θ AoD ×Φ AoD )
[0223] -RX Antenna(nH RX ×nV RX )<-2D->FFT AoA(θ AoA ×Φ AoA )
[0224] Matrix H beam The reflected power is given at points separated by 0.17 m or 17 cm in distance and by an average of 22.5 degrees in azimuth and elevation.
[0225] -Delay The distance resolution is 0.17 meters.
[0226] -Horizontal antenna (nH) 8 <=> Azimuth resolution is 22.5 degrees (average)
[0227] -Vertical antenna (nV) 8 <=> Elevation resolution is 22.5 degrees (average)
[0228] In an embodiment, the H_beam is converted to a monostatic format (ie, as if both transmission and reception occur at the RX side), where the first step of the conversion is to find the LOS path.
[0229] Fig.16 A schematic diagram showing pre-processing accompanying finding LOS according to embodiments disclosed herein.
[0230] In an embodiment, the proposed system selects a beam with at least one of higher power, lower latency, and narrower beam width to find LOS (1302). In an embodiment, the LOS path propagates directly from TX to RX, and it does not contain information about the environment. Therefore, the proposed system removes these LOS entries from H_beam and stores LOS path parameters separately because they provide position and direction information of RX to be used in subsequent steps.
[0231] Fig.17 A schematic diagram of pre-processing with removal of transmission effects (step 1303 ) according to an embodiment disclosed herein is shown.
[0232] In an embodiment, the purpose of preprocessing is to obtain sensing data in a monostatic format consistent with the output LiDAR format. With the monostatic format, the proposed system intends that the RF signal is transmitted and received at the RX location. Whereas in the bistatic format, the RF signal is transmitted from the TX located at a different location and is received at the RX location. Therefore, in order to convert the bistatic format (step 1701) to the monostatic format, the proposed system (step 1303) converts the bistatic format (step 1702) to the monostatic format. beam Remove the influence of TX (step 1702).
[0233] In the matrix H beam In , multiple entries of the AOD dimension are projected into the delay dimension.
[0234]
[0235] Fig.18 A schematic diagram of combining removal of transmission influence and preprocessing of arrival angle information according to an embodiment disclosed herein is shown.
[0236] In an embodiment, Fig.18 The signal shown starts from TX and propagates r 2 distance after encountering an obstacle, then propagating r before reaching RX 1 Distance. The signal travels from TX to a distance r 1 +r 2 Therefore, the proposed system reduces the delay of this entry in H_beam from r 1 +r 2 Convert to r 1 , making it appear as if it reaches RX directly from the reflection point. Therefore, the proposed system calculates the ratio r using the following formula 2 / r 1 :
[0237]
[0238] In an embodiment, a triangle (1801) is formed by the LOS path and the current reflection, and the angle delta AOD is the difference between the departure angle of the LOS and the departure angle of the current reflection. delta AOA can be defined similarly. Here we use Fig.17 By applying the law of sines, the proposed system obtains the ratio r 2 / r 1 , which is the ratio of the sine of ΔAOA to ΔAOD (step 1802). The proposed system projects (step 1802) the entries of the AOD dimension to the appropriate delay to obtain a power map from the RX perspective.
[0239] Fig.19A schematic diagram of pre-processing with re-adjusting power (1304) according to embodiments disclosed herein is shown.
[0240] In an embodiment, the proposed system removes the effect of TX in both the delay and angle dimensions (step 1303, and also shown as the transition from 1901 to 1902), but the resulting power is still related to the distance propagated in the bistatic format. Further, the proposed system rescales the power according to the distance propagated in the monostatic format (step 1903), so that the reflection appears to propagate the receive path twice.
[0241] The proposed system transforms the reflected power P′ into RX Scale to scale Makes it appear to spread (r 1 ,r 1 ) instead of (r 1 ,r 2 ). The distance of propagation in bistatic format is r 2 +r 1 , the propagation distance in a single-base format is 2r 2 Therefore, the delay and power of the received signal (pre-processed signal) appear to be in monostatic format (same as LiDAR).
[0242] Fig. 20 A schematic diagram comparing LIDAR PCD with the proposed method PCD according to the embodiments disclosed herein is shown.
[0243] Reference Fig. 20 , shows the comparison between LIDAR PCD (2001) and the proposed method PCD (2002) on two samples (2000A) and (2000B). The proposed system returns a high-resolution point cloud similar to LIDAR. Even when the Tx and Rx positions are unknown, the changes in perception can be shown.
[0244] Fig.21A schematic diagram of input and output for ML model training according to an embodiment disclosed herein is shown. For ML model training, input data (2101) and output data (2104) are given for multiple locations in a room, where the input data is the channel impulse response of Rx in a dual-base format, and the output data is the LIDAR point cloud data of the surrounding environment visible to RX. The room perception visible to RX changes as the RX position and direction change. The output data (2106) shows a snapshot of the output data across locations to explain this change. The output LIDAR PCD (2014) is converted into a voxel grid for ML model training and prediction, where the voxel grid is a three-dimensional grid with binary entries for each surrounding environment location. In the voxel grid, the entry "0" indicates that there is no obstacle at the location, while "1" indicates that there is an obstacle. In addition, the larger the size of the voxel grid (2107), the lower the complexity of the ML model and the lower the output resolution, so there is a trade-off.
[0245] Fig. 22 A schematic diagram of an ML model architecture according to an embodiment disclosed herein is shown.
[0246] In an embodiment, the input to the ML model is the MSP in N directions, and for each direction, the MSP is available at M distance points. Here, N=64 and M=100, so the input size is 6400 MSP values. The proposed system applies a fully connected layer to the input, followed by an upsampling layer. Since the input to output size ratio is 3% here and the LIDAR point cloud has 128×128×16 voxels, an upsampling layer is used. Further, the proposed system applies several CNN layers to extract correlations between adjacent reflections. In addition, the proposed system uses a sigmoid activation function to predict 0 or 1 for each voxel of the voxel grid, thereby detailing whether there is an obstacle in that voxel.
[0247] In an embodiment, the custom loss function is a weighted binary cross entropy, where W 1 =10,W 0 = 1. Since the output voxel grid is sparse, training will be biased towards label “1”.
[0248]
[0249] The details of the dataset with test samples (2301) are as follows:
[0250] -Total training samples: 3400 samples
[0251] Region 1: 2400 samples
[0252] Region 3: 1000 samples
[0253] - Validation samples: 350 (from two regions)
[0254] -Test data: 529 samples from region 2
[0255] The optimizer details are as follows:
[0256] -Adam optimizer with a learning rate of 0.0005
[0257] - Decay rate: 0.9 per 10,000 steps
[0258] - epochs = 100, batch size: 32
[0259] Fig.23 A schematic diagram of a test sample with error analysis according to an embodiment disclosed herein is shown.
[0260] In an embodiment, the chamfer distance (2302) is used to measure the difference between point clouds. The chamfer distance (2302) is calculated using the following formula:
[0261]
[0262] The chamfer distance (2302) of test data 1 is as follows:
[0263] -Average chamfer distance = 2m 2 , this is suitable for a room with a size of 16m×16m×4m
[0264] -Average chamfer distance = 1m 2 , for samples closer to Tx.
[0265] The chamfer distance (2302) of test data 2 is as follows:
[0266] -Average chamfer distance = 1.5m 2 , targeted voxel size = 0.25m
[0267] -Average chamfer distance = 2.2m 2 , targeted voxel size = 0.5m
[0268] In an embodiment, the depth map contains 3D information of the environment presented at high resolution from the perspective of the receiving end. The proposed system supports various imaging applications such as AR, VR, simultaneous localization and mapping (SLAM), 3D scanners, and sensing applications such as activity recognition. Traditionally, LiDAR and cameras are used to estimate 3D depth maps. LiDAR emits a high-frequency laser beam to estimate the depth map. The proposed system uses the existing communication infrastructure for sensing, and can perceive the environment with good spatial resolution compared to traditional systems due to the higher carrier frequency and the larger number of antenna elements. The proposed system estimates the depth map using existing communication signals and creates a digital twin using RF signals, which means that an entire room containing hundreds of objects can be sensed. Therefore, the proposed system no longer relies solely on manual features for inference, which has a higher complexity and scale. The proposed system first performs preprocessing and then inputs the processed data into the ML model to match the LiDAR point cloud data (PCD). The proposed system is the first of its kind and is able to generate high-resolution point clouds similar to LiDAR in terms of distance and angle using low-resolution RF data. The proposed system also proposes information exchange between the BS and the UE (120), which is necessary to apply the solution to the next generation wireless system.
[0269] In an embodiment, the depth map contains 3D information of the environment presented in high resolution from the perspective of the receiver. Unlike traditional wireless sensing, which can only sense low-dimensional outputs such as location, creating a digital twin using RF signals means sensing an entire room containing hundreds of objects. The proposed system no longer relies solely on hand-crafted features for inference, a problem with higher complexity and scale. The proposed system obtains a high-resolution depth map from a low-resolution RF signal, which is not an easy task. Therefore, the proposed system uses AI to estimate the depth map, and even for AI, the input is pre-processed, which makes the input format the same as the LIDAR output format, while also making it easy to handle for learning.
[0270] In an embodiment, the ML model architecture converts low-resolution, low-dimensional input data or MSP into high-resolution, feature-rich LiDAR data that captures the surfaces of objects, obstacles, people, etc. in the surrounding environment in detail through dense point clouds. The sensing capabilities of the proposed system far exceed the traditional primary applications of wireless sensing such as counting people in a room and locating specific objects.
[0271] In an embodiment, DME is a basic module for building AR / VR applications on the device and in the cloud. Therefore, finding an effective way to implement DME in a low-power, low-latency, and low-computational complexity manner is critical to extending AR / VR applications to UEs. In addition, in traditional methods and systems, the resolution and size of each sample from the camera, LIDAR are high, and processing these samples on the UE or transmitting them to the BS / cloud requires more computing resources, consuming power and time. However, the proposed method uses low-resolution RF channel data that has already been processed for communication purposes to perform depth map estimation, which has a negligible increase in the load on resource usage.
[0272] Metaverse applications require high throughput on the backhaul because they are used to transmit spatial information of the user and the surrounding environment from camera / LIDAR sensors. This requirement is greatly reduced by using an RF2LiDAR platform that needs to share RF sensing data (i.e., MSP information). The proposed system and method utilizes the existing 3GPP protocol to extract 3D information of the environment through the RF2LiDAR platform and reuses the existing communication infrastructure such as spectrum, base stations, etc. for sensing.
[0273] As shown in Table 1, LIDAR data has much higher resolution than RF data in terms of distance and angle. In addition, LiDAR frequency is higher and guided by a laser beam, while RF sensing is limited by sampling rate and number of antennas. Moreover, LiDAR can give more reflections than RF sensors when the TX and RX positions are unknown. In an embodiment, for training ML models, the input dimension is much lower than the output dimension (input / output = 3%).
[0274]
Table 1
[0275]
[0276] In an embodiment, similar to the global training of SLAM, the field of view in wireless data collected from different locations is integrated based on the proposed one-to-one RF to LiDAR mapping.
[0277] The applications of the proposed system in next generation communication systems are as follows:
[0278] -In next-generation communication systems, or 6G, digital twins can be used to solve problems such as virtual selection of beams without actually transmitting the beams in the physical environment.
[0279] - The proposed system is able to infer from communication signals (i.e., ISAC) and does not require specialized equipment (unlike LiDAR, radar, 3D cameras) when used for sensing purposes in next-generation communication systems or 6G.
[0280] - For commercial use cases such as self-updating 6G networks, smart cities, IoT, AR / VR glasses and applications, metaverse, etc., the proposed system supports various applications such as virtual beam selection (even without measuring physical beams), digital twins of next-generation wireless systems. The proposed system can be used for imaging applications such as AR, VR, simultaneous localization and mapping (SLAM), 3D scanners, and sensing applications such as activity recognition.
[0281] The applications of the proposed system in connecting wireless and visual fields are as follows:
[0282] - The proposed system is able to generate LiDAR-like high-resolution 3D point clouds from this readily available RF data and use this platform as a bridge between the wireless and vision domains.
[0283] -RF2Lidar can then be used to build a variety of complex applications in areas such as wireless communications, wireless sensing and 3D vision.
[0284] -AR / VR applications can be implemented on traditional mobile phones without the need for expensive and specialized hardware such as AR glasses, and 3D sensors such as LiDAR and 3D cameras.
[0285] The applications of the proposed system in terms of low-cost solutions are as follows:
[0286] -RF channel data is readily available in any UE (120) with communication capabilities. The existing RF signal chain and 3GPP protocol are sufficient to extract 3D information of the environment through the RF2LiDAR platform.
[0287] The above description of specific embodiments will fully reveal the overall nature of the embodiments herein, so that others can easily modify and / or adapt these specific embodiments for various applications without departing from the overall concept by applying current knowledge. Therefore, such adaptations and modifications should and are intended to be understood as being within the equivalent meaning and scope of the disclosed embodiments. It should be understood that the wording or terminology used herein is for description rather than limitation. Therefore, although the embodiments herein are described in the form of preferred embodiments, those skilled in the art will recognize that within the scope of the embodiments described herein, these embodiments can still be implemented after modification.
Claims
1. A method for a network node (110) to generate an environment depth map in a wireless communication system, the method comprising: Training at least one untrained machine learning (ML) model for at least one sub-cell area; sending a request for at least one sub-cell area identifier to a user equipment UE (120); receiving a response including the at least one sub-cell area identifier and input data corresponding to the at least one sub-cell area identifier; wherein the input data is sensed data of an environment of the at least one sub-cell area; as well as A depth map of the at least one sub-cell region is generated by inputting the input data of the at least one sub-cell region identifier into at least one trained ML model.
2. The method according to claim 1, wherein: Training at least one ML model for the at least one sub-cell area comprises: sending a location request for the at least one sub-cell area identifier of the at least one sub-cell area to the UE (120); receiving a location response including the at least one sub-cell area identifier of the UE (120); sending the at least one untrained ML model corresponding to the at least one sub-cell area identifier to the UE (120) for training; and At least one trained model corresponding to the at least one sub-cell area identifier is received from the UE (120).
3. The method according to claim 1, wherein: Training the at least one untrained ML model refers to updating at least one parameter by the UE (120) based on the sensing data.
4. The method according to claim 1, wherein: The network node (110) stores a plurality of trained ML models for a plurality of sub-cell areas, wherein the network node (110) uses the stored plurality of trained ML models to perform virtual beam selection without transmitting actual beams in a physical environment.
5. A method for a user equipment UE (120) to generate an environment depth map in a wireless communication system, comprising: Requesting at least one trained machine learning (ML) model for at least one sub-cell area by sending at least one sub-cell area identifier to a network node (110); receiving, from the network node (110), the at least one trained ML model for the at least one sub-cell area corresponding to the at least one sub-cell area identifier; determining input data by sensing the environment of the at least one sub-cell area; as well as A depth map of the at least one sub-cell area is generated by inputting the input data into the at least one trained ML model.
6. The method according to claim 5, wherein: Determining input data by sensing the environment of the at least one sub-cell area includes: receiving at least one channel of data in a bistatic format; wherein the bistatic format includes data of an RF signal received directly or indirectly from at least one transmitter; and The input data is determined by converting the channel data from a bistatic format to a monostatic format.
7. The method according to claim 5, wherein: The UE (120) trains at least one untrained ML model in a training phase.
8. The method according to claim 7, wherein: Training the at least one untrained ML model comprises: generating sensory data based on sensing of an environment using the at least one untrained ML model; Verifying the sensing data using Light Detection and Ranging (LiDAR) data; determining whether the sensing data satisfies a threshold; and Do one of the following: When the sensing data does not satisfy the threshold, updating parameters of at least one of a convolutional layer and an upsampling layer of the at least one untrained ML model; or When the sensory data satisfies the threshold, the at least one untrained ML model is considered as a trained ML model.
9. A network node (110) for generating an environment depth map in a wireless communication system, the network node comprising: Memory; processor; as well as A network node depth map controller (114) in communication with the memory and the processor, configured to: Training at least one untrained machine learning (ML) model for at least one sub-cell area; sending a request for at least one sub-cell area identifier to a user equipment UE (120) (120); receiving a response including the at least one sub-cell area identifier and input data corresponding to the at least one sub-cell area identifier; wherein the input data is sensed data of an environment of the at least one sub-cell area; as well as A depth map of the at least one sub-cell region is generated by inputting the input data of the at least one sub-cell region identifier into at least one trained ML model.
10. The network node (110) according to claim 9, wherein: Training at least one ML model for at least one sub-cell area includes: sending a location request for the at least one sub-cell area identifier of the at least one sub-cell area to the UE (120); receiving a location response including the at least one sub-cell area identifier of the UE (120); sending the at least one untrained ML model corresponding to the at least one sub-cell area identifier to the UE (120) for training; and At least one trained model corresponding to the at least one sub-cell area identifier is received from the UE (120).
11. The network node (110) according to claim 9, wherein: The at least one untrained ML model is trained by the UE (120) by updating at least one parameter based on the sensing data.
12. The network node (110) according to claim 9, wherein: The network node (110) stores a plurality of trained ML models for a plurality of sub-cell areas, wherein the network node (110) uses the stored plurality of trained ML models to perform virtual beam selection without transmitting actual beams in a physical environment.
13. A user equipment UE (120) for generating an environment depth map in a wireless communication system, the UE comprising: Memory; processor; as well as A UE depth map controller (124) in communication with the memory and the processor, configured to: Requesting at least one trained machine learning (ML) model for at least one sub-cell area by sending at least one sub-cell area identifier to a network node (110); receiving, from the network node (110), the at least one trained ML model for the at least one sub-cell area corresponding to the at least one sub-cell area identifier; determining input data by sensing the environment of the at least one sub-cell area; as well as A depth map of the at least one sub-cell area is generated by inputting the input data into the at least one trained ML model.
14. The network node (110) according to claim 13, wherein: Determining the input data by sensing the environment of the at least one sub-cell area includes: receiving at least one channel of data in a bistatic format; wherein the bistatic format includes data of an RF signal received directly or indirectly from at least one transmitter; and The input data is determined by converting the channel data from a bistatic format to a monostatic format.
15. The network node (110) according to claim 13, wherein: The UE (120) trains at least one untrained ML model in a training phase.