System and method for dynamically adjusting the level of detail of a point cloud
By using neural network models in point cloud capture and distribution systems to dynamically adjust the detail level of point cloud data, the problem of insufficient detail level of point cloud data in the existing technology is solved, and efficient data transmission and improved user experience are achieved.
Patent Information
- Application Number
- CN201980031100.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2018-03-20
- Filing Date
- 2019-03-12
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2039-03-12
AI Technical Summary
The prior art when capturing and distributing immersive spatial 3D content, the resolution and accuracy of point cloud data are limited, resulting in insufficient level of detail, affecting user experience, and large-scale point cloud data transmission will increase bandwidth requirements and limit the actual level of detail.
By receiving point cloud data, tracking viewpoints, selecting selected objects, retrieving neural network models, generating detailed level data, replacing points in point cloud data, and rendering views. This method can increase the details of point cloud data through neural network models, simulate missing details, and dynamically adjust the detail level of point clouds to match viewing positions and perspectives.
It realizes dynamic adjustment of the detail level of point cloud data, improves the sampling density and user experience of point cloud data, reduces the bandwidth requirement for data transmission, and can effectively transmit large-scale point cloud data while maintaining high-quality visualization.
Smart Images

Figure CN112106063B_ABST
Abstract
Description
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application is a non-provisional filing of U.S. Provisional Patent Application Serial No. 62 / 645,618, filed on March 20, 2018, entitled “SYSTEM AND METHOD FOR DYNAMICALLY ADJUSTING LEVEL OF DETAIL OF A POINT CLOUD,” and claims the benefits thereof under 35 U.S.C. §119(e), which application is incorporated herein by reference in its entirety. Background Art
[0003] As virtual reality (VR) and augmented reality (AR) platforms are moving toward increasing consumer adoption, the demand for full three-dimensional (3D) spatial content viewing is growing. The traditional de facto standard for such full 3D content is polygonal 3D graphics, manually produced through modeling and rendered using tools and techniques used to create real-time 3D games. However, emerging mixed reality (MR) displays and content capture technologies, such as RGB-D sensors and light field cameras, have set best practices for new potential developments in the way immersive spatial 3D content can be produced and distributed.
[0004] Spatial capture systems that collect point cloud data from real-world environments can produce large amounts of point cloud data that can be used as content for immersive experiences. However, capture devices that produce point clouds of real-world environments typically have limited accuracy and may work best for viewing at relatively close distances. This potential limitation may often result in limited resolution and accuracy of the captured point clouds. Representing high levels of detail with point clouds may require large amounts of data that may burden or overwhelm the data distribution bandwidth, thereby potentially reducing the actual or usable level of detail of the point cloud. Summary of the invention
[0005] An example method according to some embodiments may include: receiving point cloud data representing one or more three-dimensional objects; tracking a viewpoint of the point cloud data; selecting a selected object from the one or more three-dimensional objects using the viewpoint; retrieving a neural network model for the selected object; generating level of detail data for the selected object using the neural network model; replacing points corresponding to the selected object within the point cloud data with the level of detail data; and rendering a view of the point cloud data.
[0006] In some embodiments of the example method, generating level of detail data for the selected object may include hallucinating additional detail for the selected object.
[0007] In some embodiments of the example method, hallucinating additional detail for the selected object may increase the sampling density of the point cloud data.
[0008] In some embodiments of the example method, generating level of detail data for the selected object may include using a neural network to infer details lost due to limited sampling density for the selected object.
[0009] In some embodiments of the example method, using a neural network to infer detail loss for selected objects may increase the sampling density of the point cloud data.
[0010] An example method according to some embodiments may further include: selecting a training data set for the selected object; and generating a neural network model using the training data set.
[0011] In some embodiments of the example method, the level of detail data for the selected object may have a lower resolution than the points replaced within the point cloud data.
[0012] In some embodiments of the example method, the level of detail data for the selected object may have a higher resolution than the points replaced within the point cloud data.
[0013] In some embodiments of the example method, selecting a selected object from the one or more three-dimensional objects using the viewpoint may include: in response to determining that a point distance between an object picked from the one or more three-dimensional objects and the viewpoint is less than a threshold, selecting the object as the selected object.
[0014] Example methods according to some embodiments may further include detecting, at a viewing client, one or more objects within the point cloud data, wherein selecting the selected object may include selecting the selected object from among the one or more objects detected within the point cloud data.
[0015] Example methods according to some embodiments may further include: capturing data indicative of user movement at a viewing client; and setting the viewpoint using, at least in part, the data indicative of user movement.
[0016] An example method according to some embodiments may further include: capturing motion data of a head mounted display (HMD); and setting the viewpoint using, at least in part, the motion data.
[0017] In some embodiments of the example method, retrieving the neural network model for the selected object may include retrieving the neural network model for the selected object from a neural network server in response to determining that the viewing client lacks the neural network model.
[0018] In some implementations of the example method, retrieving the neural network model may include retrieving an updated neural network model for the selected object.
[0019] An example method according to some embodiments may further include: identifying a second server at the first server; and transmitting an identification of the second server to the client device, wherein retrieving the neural network model may include retrieving the neural network model from the second server.
[0020] In some embodiments of the example method, retrieving the neural network model may include: requesting the neural network model at a point cloud server; receiving the neural network model at the point cloud server; and transmitting the neural network model to a client device.
[0021] In some embodiments of the example method, receiving the point cloud data may include receiving the point cloud data from a sensor.
[0022] An example system according to some embodiments may include: one or more processors; and one or more non-transitory computer-readable media storing instructions that, when executed by the processors, are operable to perform any of the methods described above.
[0023] Example systems according to some embodiments may also include one or more graphics processors.
[0024] Example systems according to some embodiments may also include one or more sensors.
[0025] Another example method according to some embodiments may include: receiving point cloud data representing one or more three-dimensional objects; tracking a viewpoint of the point cloud data; selecting a selected object from the one or more three-dimensional objects using the viewpoint; retrieving a neural network model for the selected object; generating increased level of detail data for the selected object using the neural network model; replacing points corresponding to the selected object within the point cloud data with the increased level of detail data; and rendering a view of the point cloud data.
[0026] Another example system according to some embodiments may include: one or more processors; and one or more non-transitory computer-readable media storing instructions that, when executed by the processors, are operable to perform any of the methods described above.
[0027] Another example method according to some embodiments may include: receiving point cloud data representing one or more three-dimensional objects; receiving a viewpoint of the point cloud data; selecting a selected object from the one or more three-dimensional objects using the viewpoint; retrieving a neural network model for the selected object; generating level of detail data for the selected object using the neural network model; and replacing points corresponding to the selected object within the point cloud data with the level of detail data.
[0028] Another example system according to some embodiments may include: one or more processors; and one or more non-transitory computer-readable media storing instructions that, when executed by the processors, are operable to perform any of the methods described above. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] A more detailed understanding can be obtained from the following description presented by way of example in conjunction with the accompanying drawings.In addition, the same reference numerals in the figures represent the same elements.
[0030] Figure 1A is a system diagram of an example system of an example communication system in accordance with some embodiments.
[0031] Figure 1B is a system diagram of an example system according to some embodiments, which illustrates that Figure 1A An example wireless transmit / receive unit (WTRU) for use within a communication system is shown.
[0032] Figure 2 is a system interface diagram illustrating an exemplary set of interfaces for two servers and one client according to some embodiments.
[0033] Figure 3A is a schematic perspective diagram illustrating an example point cloud of a scene that may be captured and processed by some embodiments.
[0034] Figure 3B is a schematic perspective diagram illustrating an example point cloud of an object that may be captured and processed by some embodiments.
[0035] Figure 4 is a system interface diagram illustrating an exemplary set of interfaces for a set of servers and clients according to some embodiments.
[0036] Figure 5 is a message sequence chart illustrating an example process according to some embodiments.
[0037] Figure 6 is a flow diagram illustrating an example process performed by a server according to some embodiments.
[0038] Figure 7 is a flow diagram illustrating an example process performed by a server according to some embodiments.
[0039] Figure 8 is a flow diagram illustrating an example process performed by a client according to some embodiments.
[0040] Fig. 9 is a flow chart illustrating an example process according to some embodiments.
[0041] The entities, connections, arrangements, etc. depicted in and described in connection with the various figures are presented as examples and not as limitations. Therefore, any and all statements or other indications as to what a particular figure "depicts," "a particular element or entity in a particular figure" is "or "has," and any and all similar statements (which may be read as absolute in isolation and out of context), and thus limitations may only be properly read as constructively preceding them with statements such as "In at least one embodiment, . . . ," For the sake of brevity and clarity of presentation, this implicit main clause is not repeated in the detailed description.
[0042] Example network for implementation of embodiments
[0043] A detailed description of illustrative embodiments will now be described with reference to the various drawings. Although this specification provides detailed examples of possible implementations, it should be noted that these details are intended to be exemplary and in no way limit the scope of the present application.
[0044] Figure 1A 1 is a diagram illustrating an example communication system 100 in which one or more disclosed embodiments may be implemented. The communication system 100 may be a multiple access system that provides content such as voice, data, video, messaging, broadcast, etc. to multiple wireless users. The communication system 100 may enable multiple wireless users to access such content by sharing system resources including wireless bandwidth. For example, the communication system 100 may employ one or more channel access methods, such as code division multiple access (CDMA), time division multiple access (TDMA), frequency division multiple access (FDMA), orthogonal FDMA (OFDMA), single carrier FDMA (SC-FDMA), zero tail unique word DFT spread OFDM (ZT-UW DTS-s OFDM), unique word OFDM (UW-OFDM), resource block filtered OFDM, filter bank multi-carrier (FBMC), etc.
[0045] like Figure 1AAs shown, the communication system 100 may include wireless transmit / receive units (WTRUs) 102a, 102b, 102c, 102d, RAN 104 / 113, CN 106 / 115, public switched telephone network (PSTN) 108, Internet 110, and other networks 112, but it should be understood that the disclosed embodiments contemplate any number of WTRUs, base stations, networks, and / or network elements. Each WTRU 102a, 102b, 102c, 102d may be any type of device configured to operate and / or communicate in a wireless environment. By way of example, the WTRUs 102a, 102b, 102c, 102d (any of which may be referred to as a “station” and / or “STA”) may be configured to transmit and / or receive wireless signals and may include user equipment (UE), a mobile station, a fixed or mobile subscriber unit, a subscription-based unit, a pager, a cellular phone, a personal digital assistant (PDA), a smart phone, a laptop, a netbook, a personal computer, a wireless sensor, a hotspot or MiFi device, an Internet of Things (IoT) device, a watch or other wearable device, a head-mounted display (HMD), a vehicle, a drone, medical devices and applications (e.g., remote surgery), industrial devices and applications (e.g., robots and / or other wireless devices operating in an industrial and / or automated process chain environment), a consumer electronic device, a device operating on a commercial and / or industrial wireless network, etc. Any of the WTRUs 102a, 102b, 102c, and 102d may be interchangeably referred to as a UE.
[0046] The communication system 100 may also include a base station 114a and / or a base station 114b. Each of the base stations 114a, 114b may be any type of device that is configured to wirelessly interface with at least one of the WTRUs 102a, 102b, 102c, 102d to facilitate access to one or more communication networks (e.g., CN 106 / 115, Internet 110, and / or other networks 112). By way of example, the base stations 114a, 114b may be a base transceiver station (BTS), a Node B, an eNode B, a Home Node B, a Home eNode B, a gNB, an NR Node B, a site controller, an access point (AP), a wireless router, and the like. Although the base stations 114a, 114b are each depicted as a single element, it will be appreciated that the base stations 114a, 114b may include any number of interconnected base stations and / or network elements.
[0047] The base station 114a may be part of the RAN 104 / 113, which may also include other base stations and / or network elements (not shown), such as a base station controller (BSC), a radio network controller (RNC), a relay node, etc. The base station 114a and / or the base station 114b may be configured to send and / or receive wireless signals on one or more carrier frequencies (which may be referred to as cells (not shown)). These frequencies may be in a licensed spectrum, an unlicensed spectrum, or a combination of licensed and unlicensed spectrums. A cell may provide coverage of wireless services to a specific geographic area, which may be relatively fixed or may change over time. The cell may be further divided into cell sectors. For example, a cell associated with the base station 114a may be divided into three sectors. Therefore, in one embodiment, the base station 114a may include three transceivers, i.e., one transceiver for each sector of the cell. In an embodiment, the base station 114a may employ multiple-input multiple-output (MIMO) technology, and may use multiple transceivers for each sector of the cell. For example, beamforming may be used to transmit and / or receive signals in a desired spatial direction.
[0048] The base stations 114a, 114b may communicate with one or more of the WTRUs 102a, 102b, 102c, 102d via an air interface 116, which may be any suitable wireless communication link (e.g., radio frequency (RF), microwave, centimeter wave, micrometer wave, infrared (IR), ultraviolet (UV), visible light, etc.). The air interface 116 may be established using any suitable radio access technology (RAT).
[0049] More specifically, as described above, the communication system 100 may be a multiple-access system and may employ one or more channel access schemes such as CDMA, TDMA, FDMA, OFDMA, SC-FDMA, etc. For example, the base station 114a in the RAN 104 / 113 and the WTRUs 102a, 102b, 102c may implement a radio technology such as Universal Mobile Telecommunications System (UMTS) Terrestrial Radio Access (UTRA), which may use Wideband CDMA (WCDMA) to establish the air interface 115 / 116 / 117. WCDMA may include communication protocols such as High Speed Packet Access (HSPA) and / or Evolved HSPA (HSPA+). HSPA may include High Speed Downlink (DL) Packet Access (HSDPA) and / or High Speed UL Packet Access (HSUPA).
[0050] In one embodiment, the base station 114a and the WTRUs 102a, 102b, 102c may implement a radio technology such as Evolved UMTS Terrestrial Radio Access (E-UTRA), which may establish the air interface 116 using Long Term Evolution (LTE) and / or LTE-Advanced (LTE-A) and / or LTE-Advanced Pro (LTE-A Pro).
[0051] In one embodiment, the base station 114a and the WTRUs 102a, 102b, 102c may implement a radio technology such as NR radio access, which may establish the air interface 116 using new radio (NR).
[0052] In one embodiment, the base station 114a and the WTRUs 102a, 102b, 102c may implement multiple radio access technologies. For example, the base station 114a and the WTRUs 102a, 102b, 102c may implement LTE radio access and NR radio access together, for example using dual connectivity (DC) principles. Thus, the air interface used by the WTRUs 102a, 102b, 102c may be characterized by multiple types of radio access technologies and / or transmissions to / from multiple types of base stations, such as eNBs and gNBs.
[0053] In other embodiments, the base station 114a and the WTRUs 102a, 102b, 102c may implement wireless technologies such as IEEE 802.11 (i.e., Wireless Fidelity (WiFi), IEEE 802.16 (i.e., Worldwide Interoperability for Microwave Access (WiMAX)), CDMA2000, CDMA2000 1X, CDMA2000 EV-DO, Internet Standard 2000 (IS-2000), Internet Standard 95 (IS-95), Internet Standard 856 (IS-856), Global System for Mobile Communications (GSM), Enhanced Data Rates for GSM Evolution (EDGE), GSM EDGE (GERAN), etc.
[0054] Figure 1AThe base station 114b in the example may be, for example, a wireless router, a Home NodeB, a Home eNodeB, or an access point, and may utilize any suitable RAT to facilitate wireless connectivity in a local area, such as a business location, a home, a vehicle, a campus, an industrial facility, an air corridor (e.g., for use by drones), a road, and the like. In one embodiment, the base station 114b and the WTRUs 102c, 102d may implement radio technologies such as IEEE 802.11 to establish a wireless local area network (WLAN). In one embodiment, the base station 114b and the WTRUs 102c, 102d may implement radio technologies such as IEEE 802.15 to establish a wireless personal area network (WPAN). In another embodiment, the base station 114b and the WTRUs 102c, 102d may utilize a cellular-based RAT (e.g., WCDMA, CDMA2000, GSM, LTE-A, LTE-A, LTE-A Pro, NR, etc.) to establish a picocell or a femtocell. Figure 1A As shown, the base station 114b may have a direct connection to the Internet 110. Therefore, the base station 114b may not need to access the Internet 110 via the CN 106 / 115.
[0055] The RAN 104 / 113 may be in communication with the CN 106 / 115, which may be any type of network configured to provide voice, data, applications and / or Voice over Internet Protocol (VoIP) services to one or more of the WTRUs 102a, 102b, 102c, 102d. The data may have varying quality of service (QoS) requirements, such as different throughput requirements, latency requirements, error tolerance requirements, reliability requirements, data throughput requirements, mobility requirements, etc. The CN 106 / 115 may provide call control, billing services, mobile location-based services, prepaid calls, Internet connectivity, video distribution, etc., and / or perform advanced security functions (e.g., user authentication). Although in Figure 1A Although not shown, it will be appreciated that the RAN 104 / 113 and / or the CN 106 / 115 may be in direct or indirect communication with other RANs that employ the same RAT as the RAN 104 / 113 or a different RAT. For example, in addition to being connected to the RAN 104 / 113 that may utilize NR radio technology, the CN 106 / 115 may also communicate with another RAN (not shown) that employs GSM, UMTS, CDMA2000, WiMAX, E-UTRA, or WiFi radio technology.
[0056] The CN 106 / 115 may also serve as a gateway for the WTRUs 102a, 102b, 102c, 102d to access the PSTN 108, the Internet 110, and / or other networks 112. The PSTN 108 may include a circuit-switched telephone network that provides plain old telephone service (POTS). The Internet 110 may include a global system of interconnected computer networks and devices that use common communication protocols, such as the Transmission Control Protocol (TCP), the User Datagram Protocol (UDP), and / or the Internet Protocol (IP) in the TCP / IP Internet protocol suite. The networks 112 may include wired and / or wireless communication networks owned and / or operated by other service providers. For example, the networks 112 may include another CN connected to one or more RANs, which may use the same RAT as the RAN 104 / 113 or a different RAT.
[0057] Some or all of the WTRUs 102a, 102b, 102c, 102d in the communication system 100 may include multi-mode capabilities (e.g., the WTRUs 102a, 102b, 102c, 102d may include multiple transceivers to communicate with different wireless networks via different wireless links). Figure 1A The WTRU 102c shown may be configured to communicate with the base station 114a, which may employ a cellular-based radio technology, and with the base station 114b, which may employ an IEEE 802 radio technology.
[0058] Figure 1B is a system diagram illustrating an example WTRU 102. Figure 1B As shown, the WTRU 102 may include, among other things, a processor 118, a transceiver 120, a transmit / receive element 122, a speaker / microphone 124, a keyboard 126, a display / touchpad 128, non-removable memory 130, removable memory 132, a power source 134, a global positioning system (GPS) chipset 136, and / or other peripherals 138. It will be appreciated that the WTRU 102 may include any sub-combination of the foregoing elements while remaining consistent with an embodiment.
[0059] The processor 118 may be a general purpose processor, a special purpose processor, a conventional processor, a digital signal processor (DSP), a plurality of microprocessors, one or more microprocessors associated with a DSP core, a controller, a microcontroller, an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) circuit, any other type of integrated circuit (IC), a state machine, etc. The processor 118 may perform signal decoding, data processing, power control, input / output processing, and / or any other functionality that enables the WTRU 102 to operate in a wireless environment. The processor 118 may be coupled to the transceiver 120, which may be coupled to the transmit / receive element 122. Although Figure 1B The processor 118 and the transceiver 120 are depicted as separate components, but it will be appreciated that the processor 118 and the transceiver 120 may be integrated together in an electronic package or chip.
[0060] The transmit / receive element 122 may be configured to transmit or receive signals to or from a base station (e.g., base station 114a) via the air interface 116. For example, in one embodiment, the transmit / receive element 122 may be an antenna configured to transmit and / or receive RF signals. In one embodiment, the transmit / receive element 122 may be a transmitter / detector configured to transmit and / or receive, for example, IR, UV, or visible light signals. In yet another embodiment, the transmit / receive element 122 may be configured to transmit and / or receive both RF and light signals. It should be appreciated that the transmit / receive element 122 may be configured to transmit and / or receive any combination of wireless signals.
[0061] Although the transmit / receive element 122 is Figure 1B Although depicted as a single element in the embodiment, the WTRU 102 may include any number of transmit / receive elements 122. More specifically, the WTRU 102 may employ MIMO technology. Thus, in one embodiment, the WTRU 102 may include two or more transmit / receive elements 122 (e.g., multiple antennas) for transmitting and receiving wireless signals over the air interface 116.
[0062] The transceiver 120 may be configured to modulate signals to be transmitted by the transmit / receive element 122 and to demodulate signals received by the transmit / receive element 122. As described above, the WTRU 102 may have multi-mode capabilities. Thus, the transceiver 120 may include multiple transceivers to enable the WTRU 102 to communicate via multiple RATs (e.g., NR and IEEE 802.11).
[0063] The processor 118 of the WTRU 102 may be connected to and may receive user input data from a speaker / microphone 124, a keyboard 126, and / or a display / touchpad 128 (e.g., a liquid crystal display (LCD) display unit or an organic light emitting diode (OLED) display unit). The processor 118 may also output user data to the speaker / microphone 124, the keyboard 126, and / or the display / touchpad 128. In addition, the processor 118 may access information from and store data in any type of suitable memory, such as a non-removable memory 130 and / or a removable memory 132. The non-removable memory 130 may include a random access memory (RAM), a read-only memory (ROM), a hard disk, or any other type of memory storage device. The removable memory 132 may include a subscriber identity module (SIM) card, a memory stick, a secure digital (SD) memory card, and the like. In other embodiments, the processor 118 may access information from and store data in a memory that is not physically located on the WTRU 102, such as on a server or a home computer (not shown).
[0064] The processor 118 may receive power from the power source 134, and may be configured to distribute and / or control power to other components in the WTRU 102. The power source 134 may be any suitable device for powering the WTRU 102. For example, the power source 134 may include one or more dry cell batteries (e.g., nickel-cadmium, nickel-zinc, nickel-metal hydride (NiMH), lithium-ion (Li-ion), etc.), solar cells, fuel cells, etc.
[0065] The processor 118 may also be coupled to the GPS chipset 136, which may be configured to provide location information (e.g., longitude and latitude) regarding the current location of the WTRU 102. In addition to or in lieu of the information from the GPS chipset 136, the WTRU 102 may receive location information from a base station (e.g., base stations 114a, 114b) over the air interface 116 and / or determine its location based on the timing of signals received from two or more neighboring base stations. It will be appreciated that the WTRU 102 may acquire location information by any suitable location-determination method while remaining consistent with an embodiment.
[0066] The processor 118 may also be coupled to other peripherals 138, which may include one or more software and / or hardware modules that provide additional features, functionality, and / or wired or wireless connectivity. For example, the peripherals 138 may include an accelerometer, an electronic compass, a satellite transceiver, a digital camera (for photos and / or video), a universal serial bus (USB) port, a vibration device, a television transceiver, a hands-free headset, module, a frequency modulation (FM) radio unit, a digital music player, a media player, a video game player module, an Internet browser, a virtual reality and / or augmented reality (VR / AR) device, an activity tracker, etc. The peripheral device 138 may include one or more sensors, which may be one or more of a gyroscope, an accelerometer, a Hall effect sensor, a magnetometer, an orientation sensor, a proximity sensor, a temperature sensor, a time sensor; a geolocation sensor; an altimeter, a light sensor, a touch sensor, a magnetometer, a barometer, a posture sensor, a biometric sensor, and / or a humidity sensor.
[0067] The WTRU 102 may include a full-duplex radio for which transmission and reception of some or all signals (e.g., signals associated with particular subframes for UL (e.g., for transmission) and downlink (e.g., for reception)) may be concurrent and / or simultaneous. A full-duplex radio may include an interference management unit that reduces and / or substantially eliminates self-interference by means of hardware (e.g., chokes) or by signal processing by a processor (e.g., a separate processor (not shown) or by the processor 118). In an embodiment, the WTRU 102 may include a half-duplex radio that transmits and receives some or all signals (e.g., signals associated with particular subframes for UL (e.g., for transmission) or downlink (e.g., for reception)).
[0068] Given that Figure 1A-1B and Figure 1A-1B As described above, one or more or all of the functions described herein with respect to one or more of the following may be performed by one or more simulation devices (not shown): the WTRUs 102a-d, the base stations 114a-b, and / or any other device(s) described herein. These simulation devices may be one or more devices configured to simulate one or more or all of the functions described herein. For example, the simulation devices may be used to test other devices and / or simulate network and / or WTRU functions.
[0069] The simulation device can be designed to implement one or more tests of other devices in a laboratory environment and / or an operator network environment. For example, one or more simulation devices can perform one or more or all functions while being fully or partially implemented and / or deployed as part of a wired and / or wireless communication network in order to test other devices within the communication network. One or more simulation devices can perform one or more or all functions while being temporarily implemented / deployed as part of a wired and / or wireless communication network. For the purpose of testing and / or performing tests using over-the-air wireless communications, the simulation device can be directly coupled to another device.
[0070] One or more simulation devices can perform one or more functions (including all functions) without being implemented / deployed as part of a wired and / or wireless communication network. For example, the simulation device can be used in a test lab and / or in a test scenario in which a wired and / or wireless communication network is not deployed (e.g., testing) to implement testing of one or more components. One or more simulation devices can be test devices. The simulation device can transmit and / or receive data using direct RF coupling and / or via RF circuits (e.g., which can include one or more antennas) and / or wireless communications. DETAILED DESCRIPTION
[0071] For some embodiments, the level of detail of point cloud-based content may be dynamically adjusted based on the viewing position and perspective. In some embodiments, the dynamic adjustment of point cloud data may reduce or increase the level of detail based on the original unprocessed point cloud features, thereby achieving efficient point cloud data transmission and an enhanced user experience. An enhanced solution may be achieved through an exemplary disclosed method of neural network hallucination according to some embodiments. The use of a non-uniform point density scheme according to some embodiments allows the sampling density to be matched to the visualization requirements for the active viewing position. Some embodiments disclosed herein may include a viewing client that combines point cloud data streamed by a point cloud server with a neural network obtained by the viewing client from a neural network server to, for example, achieve dynamic adjustment of the level of detail over a wide range. In some embodiments, the point cloud server may reduce the level of detail of the original point cloud data, and the viewing client may use a neural network model to simulate the restoration of detailed point cloud elements.
[0072] An example method according to some embodiments of providing three-dimensional (3D) rendering of point cloud data may include: setting a viewpoint for point cloud data; requesting the point cloud data; in response to receiving the point cloud data, selecting an object within the point cloud data for which a level of detail is to be increased; requesting a neural network model for the object; in response to receiving the neural network model for the object, increasing the level of detail for the object; replacing points within the received point cloud data corresponding to the object with the increased detail data from the neural network model; and rendering a 3D view of the point cloud data. According to the example method, increasing the level of detail for the object may include causing an illusion of additional detail for the object. The method may also include detecting the object within the received point cloud data at a viewing client. The method may also include capturing user navigation or movement of a head mounted display (HMD) at the viewing client; and setting a viewpoint based at least in part on the captured user navigation. The method may also include determining at the viewing client that the object requires a new model, and wherein requesting the neural network model for the object includes: in response to determining that the object requires a new model, requesting the neural network model for the object.
[0073] An example method according to some embodiments of providing 3D rendering of point cloud data may include receiving point cloud data; segmenting the point cloud data; detecting objects within the point cloud data; assigning level of detail labels to points within the point cloud; in response to receiving a request from a client to provide the point cloud data; identifying a viewpoint associated with the received request; and selecting an area of the point cloud to reduce resolution based on the viewpoint; reducing the resolution for the selected area; and streaming the point cloud data with the reduced resolution and an identification of the detected object to the client. According to the example method, receiving the point cloud data may include receiving the point cloud data from a sensor. The method may also include: identifying a neural network server at the point cloud server; and transmitting the identification of the neural network server to the client. The method may also include, at the point cloud server, retrieving a neural network model; and transmitting the neural network model to the client.
[0074] A system may include a processor (e.g., one or more processors) and a non-transitory computer-readable medium storing instructions that, when executed by the processor, are operable to perform, for example, any of the methods described herein according to some embodiments. In some embodiments, an example system may include a graphics processor.
[0075] The increasing adoption of virtual reality (VR), augmented reality (AR), and mixed reality (MR) systems is increasingly requiring full three-dimensional (3D) spatial content, such as point cloud data that can be collected from real-world environments (in real time or near real time) using RGB-D sensors, light field cameras, and light detection and ranging (LIDAR) systems.
[0076] In addition to the development of MR-driven reality capture technology, human-computer interaction systems may become increasingly spatially aware. Environmental sensing technologies that enable structural understanding of the environment are increasingly embedded in smart home environments, automotive systems, and mobile devices. These spatially aware systems can collect point cloud data from the environment to produce large amounts of point cloud data that can be used as content for immersive experiences. Unfortunately, the resolution and sensor accuracy of the point cloud may limit the level of detail that the content may have, and therefore limit how close the captured geometry can be examined without suffering from the quality of the visualization of the performance of low-resolution artifacts. The bandwidth required for point cloud data transmission can also limit the display resolution. However, in some cases, this can be mitigated by dynamically adjusting the point cloud resolution based on the viewing distance. A blog post from the Umbra website (available at https: / / web.archive.org / web / 20160512052944 / http: / / umbra3d.com / visualization-of-large-point-clouds / ) published on April 7, 2016, previously available at https: / / umbra3d.com / blog / visualization-of-large-point-clouds / , titled "Visualization of Large Point Clouds" (available at https: / / web.archive.org / web / 20160512052944 / http: / / umbra3d.com / visualization-of-large-point-clouds) (last accessed February 28, 2019) describes a method for point cloud streaming for use in the Umbra optimization system. The raw data capture is understood to place an upper limit on the true resolution of fine detail, which can then be lowered with any adjustments based on point reduction.
[0077] Deep learning neural networks, including convolutional neural networks, have been rapidly and extensively adopted for image analysis and computer vision. Methods developed for analyzing 2D image data can generally be adapted to process spatial 3D data. Although early methods for applying machine learning to 3D data required transforming the 3D data into a volumetric format (i.e., volumetric pixels (voxels) or binary spatial segmentation representations), more recent work allows neural networks to process point cloud data in their native format. For example, the journal article QI (Charles R et al., PointNet: Deep Learning of Point Sets for 3D Classification and Segmentation, 2016, Journal of Computer Vision and Pattern Recognition (CVPR), 4(1)(2) ("Qi")) describes such a system. This recent work, such as Qi, is understood to include the use of neural networks for 3D analysis for point cloud data segmentation as well as object detection and recognition.
[0078] Unfortunately, point cloud data generated by spatial capture of real-world physical scenes can result in high memory consumption (even with low resolution) and sensor noise. High memory consumption can limit the resolution at which complex scenes can be transmitted to the client that is consuming the data. However, because some point cloud data may lack topology, in some cases the level of detail of the point cloud data may be reduced by reducing the sampling density based on the viewpoint. Adjusting the sampling density can enable more efficient transmission of large point clouds. However, a potential negative consequence is that the resolution may result in a reduced quality of viewing experience when examining individual elements in the point cloud scene from a close distance.
[0079] When content is captured from a relatively close distance, the increase in content quality can be used to provide a more realistic viewing experience. In some embodiments, the level of detail of the point cloud can be adjusted based on the viewpoint, for example, (1) if the viewing distance allows, the level of detail is reduced based on the original unprocessed point cloud data, and (2) to enable dynamic simulation (illusion) of increased detail levels. For some embodiments, a six-degree-of-freedom (6DoF) viewing experience and efficient bandwidth usage can be used, such as if the viewing distance is close enough to the object.
[0080] Some embodiments implement dynamic adjustment of point cloud data, for example in two directions, which can reduce and increase the level of detail of the point cloud based on the original collected unprocessed point cloud features. In some embodiments, the level of detail can be dynamically adjusted based on the viewpoint, thereby achieving efficient point cloud data transmission and non-uniform point density, so that the point density of the point cloud matches the point density that can be used in some applications to produce high-quality visualizations for the active viewpoint. In some embodiments, the viewing client can combine the point cloud data streamed by the point cloud server with a neural network obtained from a neural network server (which can be a different or the same server). In some embodiments, the point cloud server can reduce the level of detail of the original point cloud data for transmission to the client, and the viewing client can use one or more neural networks that the viewing client has received to simulate the increase in the level of detail. The result is a dynamic adjustment of the level of detail over a wide range. For some embodiments, the term "neural network model" may be used instead of the term "neural network" used herein.
[0081] Figure 22 is a system interface diagram showing an exemplary set of interfaces for two servers and clients according to some embodiments. According to this example, one server 202 may be a point cloud server, and another server 210 may be a neural network server. In some embodiments, the servers may be consistent. Both servers are connected to the Internet 110 and other networks 112. The client 218 is also connected to the Internet 110 and other networks 112, thereby enabling communication between all three nodes 202, 210, 218. Each node 202, 210, 218 includes a processor 204, 212, 220, a non-temporary computer-readable memory storage medium 206, 214, 224, and executable instructions 208, 216, 226 contained in the storage medium 206, 214, 224, which can be executed by the processor 204, 212, 220 to implement the method or part of the method disclosed herein. As shown, for some embodiments, the client may include a graphics processor 222 for rendering 3D video for a display (such as a head-mounted display (HMD) 228). Any or all of the nodes may include a WTRU and communicate over a network, as described above with respect to Figure 1A-1B described.
[0082] For some embodiments, the system 200 may include a point cloud server 202, a neural network server 210, and / or a client 218, which includes one or more processors 204, 212, 220; and one or more non-transitory computer-readable media 206, 214, 224 storing instructions 208, 216, 226, which when executed by the processors 204, 212, 220 operate to perform the methods disclosed herein. For some embodiments, the node 218 may include one or more graphics processors 222. For some embodiments, the nodes 202, 210, 218 may include one or more sensors.
[0083] Figure 3A is a schematic perspective view showing an example point cloud of a scene that may be captured and processed by some embodiments. Scene 300 includes a plurality of buildings at a distance and some closer objects imaged from a viewer viewpoint having a certain apparent height. As the viewer viewpoint changes, such as by moving below or closer to the buildings, the relative angles to the points within the point cloud may change. The point cloud may be detected in a real world scene, generated using virtual objects, or any combination of these or other applicable techniques.
[0084] Figure 3B is a schematic perspective diagram illustrating an example point cloud of an object that may be captured and processed by some embodiments. Figure 3Bis a two-dimensional black and white line drawing of a three-dimensional point cloud object 350. Within a three-dimensional display environment, a point cloud object has points representing three-dimensional coordinates where a portion of an object has been detected to exist. Such detection can be performed using, for example, 3D sensors such as light detection and ranging (LIDAR), stereo video, and RGB-D cameras. The point cloud data may include, for example, 3D position and radiometric image data or voxels.
[0085] Figure 4 406 . 408 . 409 . 410 . 411 . 412 . 413 . 414 . 415 . 416 . 417 . 418 . 419 . 420 . 421 . 422 . 423 . 424 . 425 . 426 . 427 . 428 . 430 . 431 . 432 . 433 . 434 . 435 . 436 . 437 . 438 . 440 . 441 . 442 . 443 . 444 . 445 . 446 . 447 . 448 . 449 . 450 . 451 . 452 . 453 . 454 . 455 . 456 . 457 . 458 . 459 . 460 . 461 . 462
[0086] For some embodiments, the point cloud data rendered for the viewer may undergo a data construction process from at least two different sources to reduce and / or increase the level of detail. One of the sources may be point cloud data 408 streamed from a point cloud server 402 to a viewing client 406. The point cloud server 402 streams the point cloud data 408 at a resolution already provided by spatial capture, or for some embodiments, downsamples to comply with, for example, bandwidth constraints or viewing distance tolerances. The point cloud server 402 may dynamically reduce the level of detail. Another of the sources may be a repository of neural networks 410 that may be used to simulate an increased level of detail in the point cloud data by hallucinating or inferring additional geometric details. In some embodiments, a neural network for increasing the level of detail by hallucination may be trained on a per-object basis using synthetic data available from an existing 3D object model repository.
[0087] For some embodiments, efficiencies may be gained by optimizing system-level operations; in addition to streaming point cloud data and reducing resolution (as needed or allowed), in some embodiments, the point cloud server may also segment the point cloud data and identify objects within the point cloud. If the point cloud server provides segmentation and object recognition information to a viewing client, the client may request a specific corresponding neural network model that may be customized to cause the illusion of added (or subtracted) detail to that particular geometry (which may increase or decrease the level of detail for the object). This process allows even large point cloud scenes to be navigated efficiently while viewing individual elements of the point cloud from a close distance at a sufficiently high level of detail. In some embodiments, points within the point cloud data corresponding to a selected object may be replaced with lower resolution data.
[0088] Figure 5 is a message sequence diagram illustrating an example process according to some embodiments. In the example message sequence 500, there are three nodes, one client and two servers, for example: viewing client 502, point cloud server 504, and neural network server 506. Figure 6 , 7 The internal functions at each of the three nodes 502, 504, 506 are explained in more detail in the description of and 8, respectively, but the communication between the nodes 502, 504, 506 is shown.
[0089] The point cloud server 504 processes 508 the point cloud to allow for processing of customized levels of detail, taking into account, for example, streaming constraints (time and bandwidth), viewing distance, and detection of objects using a neural network model that may allow for rapid reconstruction at the viewing client. The neural network server 506 collects 510 the high-resolution 3D model, simulates a low-resolution captured version from the model, and trains a neural network to reconstruct a higher level of detail for later use in simulating an increase in the level of detail for a similar low-resolution model. The point cloud server and neural network server wait for the viewing client 502 to start a session.
[0090] The viewing client 502 sends 512 a request for a point cloud and an indication of a viewpoint (e.g., "view from location") to a point server. In some embodiments, the viewpoint may be specified as an origin and a pointing unit vector. For some embodiments, other canonical formats may be used. The point cloud server 504 returns 514 point cloud data at a selected resolution. The point cloud data may be accompanied by an indication of detected objects. The client 502 requests 516 a neural network model for selected objects based on content that the client can use to increase the viewing resolution for a certain set of objects. The objects may include objects recognized by the point cloud server and may also include objects detected by the client itself. The neural network server 506 returns 518 an applicable neural network model. The client 502 may use the received neural network model to increase 520 the level of detail and render 522 the point cloud content for viewing by the user.
[0091] Figure 6 is a flow chart illustrating an example process performed by a server according to some embodiments. Figure 4 An example process 600 performed by a point cloud content server 402 of the point cloud content server 402 may include: 602 receiving point cloud sensor data and processing the point cloud sensor data to obtain detail level information; 604 segmenting the point cloud data; 606 detecting objects within the point cloud data; 608 assigning detail level labels to points (e.g., placing the labels in, for example, a point cloud database 620); 610 waiting for a viewing client request; 612 receiving a client request including a viewpoint; 612 adjusting the point cloud data (using data stored in the point cloud database) and streaming the adjusted point cloud data to the client; 616 monitoring conditions for terminating streaming (e.g., whether a request to end streaming has been received); and 618 monitoring conditions for terminating processing (e.g., whether a request to end processing has been received) 622.
[0092] In some embodiments, a point cloud server may store and distribute point cloud data. The point cloud data may include, for example, an unordered point sample of a scene, a geometric structure represented as the coordinate position of each point, and may also include, for example, the characteristics of each point (such as color, brightness, surface normal vector, and any additional data that may be applied). In some cases, these data may be typical data generated by some LIDAR systems. Point clouds may be obtained from multiple sources, including, for example, capture from local or remotely connected sensors, and downloading from other distribution sources. If the point cloud server has received new point cloud data, the point cloud server may process the data based on client-specific level of detail adjustments to enable distribution optimization. The point cloud server may stream point cloud data to a viewing client so that the level of detail (or, for some embodiments, the point density in some volumes or regions) may depend on the virtual distance of the data point from the client's specified viewpoint. In the case of enabling streaming based on the level of detail, the point cloud data may be dynamically reduced by omitting details that are not visible to the viewing client. The result according to this example is a streaming level of detail lower than the level of detail that the point cloud server has stored. For some embodiments, selecting a three-dimensional object within the point cloud using a viewer's viewpoint may include selecting the object if a point distance between the viewpoint and the object is less than a threshold. For some embodiments, one or more objects may be detected within the point cloud, and selecting the three-dimensional object within the point cloud may include selecting one of the detected objects.
[0093] During level of detail processing, for some embodiments, the point cloud server estimates the density of the stored point cloud area and assigns a level of detail label to each point so that, depending on the viewing distance, oversampling can be avoided. That is, according to the example, details that are not visible from a given viewpoint are removed and sampling of multiple points per final rendered image pixel can be avoided. According to some embodiments, in addition to assigning level of detail labels to points, the point cloud can also be identified as segments for more efficient streaming. For example, a blog post from the Umbra website titled “Visualization of Large Point Clouds” (available at https: / / web.archive.org / web / 20160512052944 / http: / / umbra3d.com / visualization-of-large-point-clouds / ) published on April 7, 2016 and previously available at https: / / umbra3d.com / blog / visualization-of-large-point-clouds / (last accessed on February 28, 2019) describes a method for point cloud streaming used in the Umbra optimization system.
[0094] In some embodiments, the point cloud server can also segment the point cloud data for streaming and object detection and recognition. The point cloud segments can, for example, be assigned with detected object labels so that viewing clients can request neural networks corresponding to those specific object labels. Figure 6 4 is a flow chart showing an example method that may be performed by a point cloud server (such as point cloud content server 402) according to some embodiments. Figure 6 The flowchart can be executed by Figure 7 , 8 And / or the device executing the flowchart of 9, for some embodiments, retrieving the neural network model may include: requesting the neural network model at the point cloud server; receiving the neural network model at the point cloud server; and transmitting the neural network model to the client device.
[0095] Figure 7 700 is a flowchart illustrating an example process performed by a server according to some embodiments. The example process 700 performed by a server (e.g., a neural network server, such as the example server 404) may include: 702 collecting synthetic and high-resolution 3D scan models of objects, and 718 storing these objects in a database; 704 generating a high-resolution point cloud from the 3D model; 706 simulating low-resolution scan results based on the high-resolution objects, and 720 storing the simulated low-resolution scan results; 708 compiling a training data set featuring high / low resolution pairs, which are collected in an object category group, and 722 storing these as a training set; 710 using the training set to train a neural network model for hallucinating a higher level of detail based on the low-resolution data, and storing the neural network model; 712 waiting to watch a client request; 724 receiving a client request for a neural network; 714 sending the requested neural network from the stored neural network model; and 716 monitoring 726 the conditions that terminate the process.
[0096] In some embodiments, a neural network server may be responsible for building a neural network that can be used to (e.g., simulate or otherwise adjust or generate, e.g., dynamically simulate, adjust, and / or generate) create an illusion of an increased level of detail for an object represented as a point cloud. The neural network may be trained using high-resolution point cloud models collected from existing 3D models and high-resolution 3D scans available in multiple 3D model repositories. In some embodiments, the neural network server collects and classifies object models and generates training data by creating a low-resolution point cloud version of the object (to simulate a low-resolution 3D capture). Using the training data, the neural network may be trained to create an illusion of an increased level of detail for an object class represented by or similar to the training data set.
[0097] For model training, existing 3D models and high-resolution 3D scans of selected object categories can be collected into a training database from existing 3D model repositories and available model and 3D scan datasets. For neural network training, a low-resolution version of the collected data can be generated by simulating a 3D scan of the object with a spatial data capture method that produces a lower resolution result. This process can be done by sampling the object at a lower sample density with limited visibility and adding simulated typical sensor noise to the sample. For example, a lower resolution point cloud generated by a spatial sensor capture simulation can also be stored in the training database with a link to the high-resolution data.
[0098] For training, captured high-resolution object point clouds and low-resolution object point clouds obtained from spatial capture simulations can be used as training pairs. Training pairs from the same or sufficiently similar objects can be collected into a training set. Object type similarity can be determined by merging object types with sufficient similarity.
[0099] According to some embodiments, a neural network model for increased detail level illusion can be constructed using a convolutional neural network (CNN) architecture. The trained neural network can be used to generate a high-resolution version of a low-resolution point cloud input. For example, in order to build a CNN neural network architecture and train the model, a deep learning framework or a higher-level neural network API running on top of a deep learning framework can be used. Liang, Shu et al., in their conference paper 3D Face Hallucination from a Single Depth Frame published at the 2014 Second International Conference on 3D Vision (3DV) 31-38 (2014), describe methods for, for example, hallucinating high-resolution 3D models from low-resolution depth frames. In some embodiments, if a client requests a neural network model for increasing the level of detail of a "head" or "face", these methods use examples of the type of neural network model that can be transmitted from a neural network server to a viewing client. Some embodiments of methods for constructing and training networks for specific object categories can utilize illusion techniques that are similar in at least some respects to, for example, illusion techniques that produce an increased level of detail illusion of a 3D face.
[0100] For some embodiments, a low-resolution point cloud is used as input. For example, a neural network model can be trained to infer surface details, thereby increasing the density of the input point cloud. The journal article PU-Net: Point Cloud Upsampling Network (Point Cloud Upsampling Network,) published by Yu, Lequan et al. in Computer Vision and Pattern Recognition (CVPR) (2018) on pages 2790-2799 in 2018 ("Yu") describes an example technique that can be used in accordance with some embodiments to increase the resolution of a point cloud by upsampling the data. According to Yu's example implementation, the point cloud server process is understood to learn multi-level features for each point and expand the point set via multi-branch convolutions in the feature space. According to Yu's example implementation, the expanded features are divided into multiple features, which are reconstructed into an upsampled point set.
[0101] In addition to certain example techniques that may be used or utilized in some embodiments, for some embodiments, generating a level of detail for a selected object may include causing an illusion of additional detail for the selected object. In some embodiments, points within the point cloud data corresponding to the selected object may be replaced with higher resolution data. In some embodiments, the viewing client may capture user navigation data (or, for some embodiments, data indicating user movement), and the viewing client may set a viewpoint to render data using, at least in part, the user navigation data. For some embodiments, the viewpoint may be adjusted by a vector corresponding to a user movement vector, such that the user movement vector is a vector determined relative to an origin in the real world environment, and the user movement vector is expressed as user navigation data. For some embodiments, retrieving a neural network model may include retrieving an updated neural network model for the selected object. For some embodiments, causing an illusion of additional detail for a three-dimensional object may include, for example, inferring (or generating) a level of detail data for the object using a neural network model. For example, a set of points within a set of point cloud data may be selected such that the selected set of points represents a three-dimensional object. The selected set of points may correspond to a first level of resolution. The selected points may be upsampled (or downsampled) using the neural network model to generate a set of points corresponding to a second level of resolution. Upsampling (or downsampling) increases (or decreases) the density of sampling points of the point cloud data. For some embodiments, the second resolution level may represent a higher resolution than the first resolution level. For some embodiments, the second resolution may represent a lower resolution than the first resolution.
[0102] Figure 8 is a flow chart illustrating an example process performed by a client according to some embodiments. Figure 4An exemplary method 800 performed by a viewing client (such as a viewing client 406 of the viewing client 406) may include: 802 requesting point cloud content from a server; 804 capturing user navigation or device tracking to determine a viewpoint; 806 sending a viewpoint relative to the point cloud to a point cloud server; 808 receiving point cloud data, which for some embodiments can be adjusted relative to the viewpoint; 810 selecting an object whose level of detail can be increased based on the user-controlled viewpoint; 812 determining whether the object corresponds to a new neural network model (or for some embodiments, a new object type or object category) or whether the client already has a model; 814 in response to determining that the client lacks a new model, requesting a new model from a neural network server and receiving the requested neural network model; 816 using a neural network model (most recently received or previously existing) to increase resolution and create the illusion of a higher level of detail for the selected object; replacing object fragments in the streamed point cloud with the high-resolution version generated using the neural network model; 818 rendering point cloud data to a viewer; and 820 monitoring 822 for an indication that processing is terminated.
[0103] In some embodiments, an example process performed by a client may begin with a user launching a point cloud viewing application. If the user starts the application, he / she also indicates the point cloud content to be viewed. The point cloud content may be a link to content residing on a point cloud server. In some embodiments, the link to the content may be a URL identifying the point cloud server and the specific content. The viewing client application may be launched, for example, by an explicit command from the user, or automatically by the operating system based on an identification content type request and an application associated with a specific content type. The viewing client application may be a stand-alone program, may be integrated with a web browser or social media client, or may be part of the client device operating system.
[0104] In some embodiments, in addition to making the initial point cloud request, the viewing client may continuously send updated viewpoint information (relative to the point cloud) to the point cloud server. As previously described, the point cloud server, for example, dynamically adjusts the level of detail of the point cloud streamed to the client. If the viewer client starts the process, the viewer client may start device tracking, which may facilitate viewpoint adjustment for VR, MR, and AR use cases. For example, if the client includes a wearable display device such as a head-mounted display (HMD), a sensor that detects the position and / or motion of the HMD may update the viewpoint (e.g., in response to the user's head movement (one or more)) as the user's head moves. For some embodiments, receiving point cloud data may include receiving point cloud data from a sensor. In some embodiments, in addition to device tracking, the user may utilize some additional navigation input to control the viewpoint relative to the point cloud. For some embodiments, if the viewing client initially requests point cloud data from the point cloud server, the viewing client may signal a default or specific viewpoint. If the viewing client starts receiving point cloud data, the viewing client may update the viewpoint based on user input and / or device tracking used to render the received data. In some embodiments, rendering may be performed on a dedicated graphics processor, running in parallel with processing performed by the viewing client on the received point cloud data.
[0105] If the viewing client has received point cloud data, the viewing client may identify object tags assigned to the point set by the point cloud server. In some embodiments, the client may additionally identify the objects. Using the object tags, the viewing client may determine which objects are represented in the point cloud and are close enough to merit higher resolution. For objects viewed at a closer range, the viewer client may request a neural network model from the neural network server. The neural network model may be adjusted for the user's viewpoint.
[0106] For some embodiments, a client sends a neural network request to a neural network server indicating the object class to which the object most closely matches. Using the requested object class, the neural network server may send a neural network to a viewing client to increase the level of detail. The level of detail of the selected object may be increased by feeding the receiving neural network a fragment of a point cloud that is labeled as part of the object. The neural network may create new geometry of the object at a higher resolution by hallucinating typical details seen, for example, in similar objects at a higher point cloud resolution (based on a training data set). The details hallucinated for the point cloud may be typical geometric patterns that add to the overall shape of the object. The details may be created based on features extracted by lower layers of the neural network from objects that have been fed into the network during training. The ability of certain neural networks to hallucinate details is understood.
[0107] If the illusion of geometry is created, the output may be new points to replace the input point cloud segments. The new points may feature altered geometry with an increased level of detail. If a neural network is used to generate high-resolution point cloud data for a selected object, the viewing client may replace the original data points with the corresponding high-resolution points generated by the neural network. The point cloud data may be rendered using the high-resolution points.
[0108] For some embodiments, an example process may include capturing data indicating user movement (or, for some embodiments, user navigation data) at a viewing client; and using, at least in part, the data indicating user movement (or, for some embodiments, user navigation data) to set a viewpoint. For some embodiments, the process may include capturing motion data of a head mounted display (HMD); and using, at least in part, the motion data to set a viewpoint. For certain embodiments of the process, retrieving a neural network model for a selected object may include: in response to determining that the viewing client lacks a neural network model, retrieving a neural network model for the selected object from a neural network server.
[0109] In some embodiments, the level of detail of the point cloud data can be dynamically adjusted in both directions to less than the full resolution of the point cloud data while also being able to increase the level of detail above the detail and sampling density present in the original, unprocessed point cloud. In some embodiments, a point cloud server (examples of which are described above) enables adjustment of the point cloud level of detail to a resolution less than the original, unprocessed point cloud data, while a neural network server enables adjustment of the level of detail above the original, unprocessed point cloud resolution. This approach allows for a wide range of variations in the range of data that can be inspected, allowing the viewer to more freely navigate the content while providing a consistent viewing experience.
[0110] For some embodiments, if the point distance from the viewpoint allows such a reduction, the reduced data can be transmitted from the point cloud server to the viewing client. Thus, the bandwidth requirements can be controlled for large environments captured as point clouds, and the level of detail of the data can be adjusted for rendering and viewer viewing.
[0111] In some embodiments, rather than the viewing client searching for a suitable neural network server, the point cloud server may identify a (e.g., suitable) neural network server as part of data preprocessing, for example based on object detection performed by the point cloud server on the received point cloud. For some embodiments, the point cloud server streams the point cloud to the viewing client and also sends a neural network (or an identification of the server) that the client can use to increase the level of detail. The neural network and point cloud data may be streamed based on a viewpoint specified by the client relative to the point cloud. In some embodiments, the point cloud server may use a neural network to increase the level of detail of the point cloud data and stream the enhanced point cloud data.
[0112] In some embodiments, a point cloud server may use raw point cloud data to train a neural network to create the illusion of additional detail by generating low-resolution versions of selected object fragments from a full-resolution point cloud. For example, such an example method may be performed where the point cloud data has a dramatically varying sampling density. Such processing may alleviate general resolution requirements while leaving high-resolution portions to be reproduced by a viewing client using a neural network. For some embodiments, a process may include: identifying a neural network server at a point cloud server; and sending an identification of the neural network server to a client device, wherein retrieving the neural network model may include retrieving the neural network model from the neural network server.
[0113] Fig. 9 is a flow chart illustrating an example process according to some embodiments. For some embodiments, process 900 may be performed, which includes receiving point cloud data representing one or more three-dimensional objects at 902. Process 900 may further include tracking a viewpoint of the point cloud data at 904. Process 900 may also include selecting a selected object from the one or more three-dimensional objects using the viewpoint at 906. Process 900 may also include retrieving a neural network model for the selected object at 908. Process 900 may also include generating level of detail data for the selected object using the neural network model at 910, and replacing points corresponding to the selected object with the level of detail data within the point cloud data at 912. Process 900 may also include rendering a view of the point cloud data at 914. For some embodiments, generating the level of detail data for the selected object may generate increased level of detail data for the selected object.
[0114] Although methods and systems according to some embodiments are discussed in the context of virtual reality (VR), some embodiments may also be applicable to mixed reality (MR) / augmented reality (AR) contexts. In addition, although the term "head mounted display (HMD)" is used herein, for some embodiments, some embodiments may be applicable to wearable devices (which may or may not be attached to the head) capable of, for example, VR, AR, and / or MR.
[0115] An example method according to some embodiments may include: receiving point cloud data representing one or more three-dimensional objects; tracking a viewpoint of the point cloud data; selecting a selected object from the one or more three-dimensional objects using the viewpoint; retrieving a neural network model for the selected object; generating level of detail data for the selected object using the neural network model; replacing points corresponding to the selected object within the point cloud data with the level of detail data; and rendering a view of the point cloud data.
[0116] In some embodiments of the example method, generating level of detail data for the selected object may include hallucinating additional detail for the selected object or inferring additional detail for the selected object.
[0117] In some embodiments of the example method, hallucinating or inferring additional details for a selected object may increase the sampling density of the point cloud data.
[0118] In some embodiments of the example method, generating level of detail data for the selected object may include using a neural network to infer details lost due to limited sampling density for the selected object.
[0119] In some embodiments of the example method, using a neural network to infer missing details for a selected object may increase the sampling density of the point cloud data.
[0120] An example method according to some embodiments may further include: selecting a training data set for the selected object; and generating a neural network model using the training data set.
[0121] In some embodiments of the example method, the level of detail data for the selected object may have a lower resolution than the replaced points within the point cloud data.
[0122] In some embodiments of the example method, the level of detail data for the selected object may have a higher resolution than the replaced points within the point cloud data.
[0123] In some embodiments of the example method, selecting the selected object from the one or more three-dimensional objects using the viewpoint may include: in response to determining that a point distance between the object picked from the one or more three-dimensional objects and the viewpoint is less than a threshold, selecting the object as the selected object.
[0124] Example methods according to some embodiments may further include detecting, at a viewing client, one or more objects within the point cloud data, wherein selecting the selected object may include selecting the selected object from the one or more objects detected within the point cloud data.
[0125] Example methods according to some embodiments may further include: capturing data indicative of user movement at a viewing client; and setting a viewpoint using, at least in part, the data indicative of user movement.
[0126] An example method according to some embodiments may further include: capturing motion data of a head mounted display (HMD); and setting the viewpoint using, at least in part, the motion data.
[0127] In some embodiments of the example method, retrieving the neural network model for the selected object may include: in response to determining that the viewing client lacks the neural network model, retrieving the neural network model for the selected object from the neural network server.
[0128] In some implementations of the example method, retrieving the neural network model may include retrieving an updated neural network model for the selected object.
[0129] An example method according to some embodiments may further include: identifying a second server at the first server; and transmitting an identification of the second server to the client device, wherein retrieving the neural network model may include retrieving the neural network model from the second server. In some embodiments, the first server may be a point cloud server and the second server may be a neural network server.
[0130] In some embodiments of the example method, retrieving the neural network model may include: requesting the neural network model at a point cloud server; receiving the neural network model at the point cloud server; and transmitting the neural network model to a client device.
[0131] In some embodiments of the example method, receiving the point cloud data may include receiving the point cloud data from a sensor.
[0132] An example system according to some embodiments may include: one or more processors; and one or more non-transitory computer-readable media storing instructions that, when executed by the processors, are operable to perform any of the methods described above.
[0133] Example systems according to some embodiments may also include one or more graphics processors.
[0134] Example systems according to some embodiments may also include one or more sensors.
[0135] Another example method according to some embodiments may include: receiving point cloud data representing one or more three-dimensional objects; tracking a viewpoint of the point cloud data; selecting a selected object from the one or more three-dimensional objects using the viewpoint; retrieving a neural network model for the selected object; generating increased level of detail data for the selected object using the neural network model; replacing points corresponding to the selected object within the point cloud data with the increased level of detail data; and rendering a view of the point cloud data.
[0136] Another example system according to some embodiments may include: one or more processors; and one or more non-transitory computer-readable media storing instructions that, when executed by the processors, are operable to perform any of the methods described above.
[0137] Another example method according to some embodiments may include: receiving point cloud data representing one or more three-dimensional objects; receiving a viewpoint of the point cloud data; selecting a selected object from the one or more three-dimensional objects using the viewpoint; retrieving a neural network model for the selected object; generating level of detail data for the selected object using the neural network model; and replacing points corresponding to the selected object within the point cloud data with the level of detail data.
[0138] Another example system according to some embodiments may include: one or more processors; and one or more non-transitory computer-readable media storing instructions that, when executed by the processors, are operable to perform any of the methods described above.
[0139] An example method for providing three-dimensional (3D) rendering of point cloud data according to some embodiments may include: setting a viewpoint for the point cloud data; requesting the point cloud data; in response to receiving the point cloud data, selecting an object whose level of detail within the point cloud data is to be increased; requesting a neural network model for the object; in response to receiving the neural network model of the object, increasing the level of detail for the object; replacing points within the received point cloud data corresponding to the object with the increased detail data from the neural network model; and rendering a 3D view of the point cloud data.
[0140] In some embodiments of the example method, increasing the level of detail for the object may include creating an illusion of additional detail for the object.
[0141] Some embodiments of the example method may further include, at a viewing client, detecting an object within the received point cloud data.
[0142] Some embodiments of the example method may further include, at the viewing client, capturing user navigation; and setting the viewpoint based at least in part on the captured user navigation.
[0143] Some embodiments of the example method may further include: capturing motion of a head mounted display (HMD); and setting the viewpoint based at least in part on the captured motion.
[0144] Some embodiments of the example method may also include determining, at the viewing client, that the object requires a new model, wherein requesting the neural network model for the object includes: requesting the neural network model for the object in response to determining that the object requires a new model.
[0145] An example system according to some embodiments may include: a processor; and a non-transitory computer-readable medium storing instructions that, when executed by the processor, are operable to perform one of the example methods described above.
[0146] Some embodiments of the example system may also include a graphics processor.
[0147] An example method for providing three-dimensional (3D) rendering of point cloud data according to some embodiments may include: receiving point cloud data; segmenting the point cloud data; detecting objects within the point cloud data; assigning fine-level labels to point nodes within the point cloud; in response to receiving a request from a client to provide the point cloud data; identifying a viewpoint associated with the received request; selecting an area of the point cloud to reduce resolution based on the viewpoint; reducing the resolution for the selected area; and transmitting a stream of the point cloud data with the reduced resolution and identification of the detected objects to the client.
[0148] In some embodiments of the example method, receiving the point cloud data includes receiving the point cloud data from a sensor.
[0149] Some embodiments of the example method may also include: identifying, at the point cloud server, a neural network server; and transmitting the identification of the neural network server to the client.
[0150] Some embodiments of the example method may also include: retrieving the neural network model at the point cloud server; and transmitting the neural network model to the client.
[0151] An example system according to some embodiments may include: a processor; and a non-transitory computer-readable medium storing instructions that, when executed by the processor, are operable to perform one of the methods described above.
[0152] Example methods according to some embodiments may include presenting a low-resolution view of an object to a client.
[0153] Example methods according to some embodiments may further include requesting, by the client from the server, point cloud data identifying a selected viewpoint relative to the point cloud.
[0154] Example methods according to some embodiments may also include selecting objects whose detail or resolution may be increased based on the viewpoint to the point cloud data.
[0155] Example methods according to some embodiments may also include enabling the illusion of an added level of detail for the selected object.
[0156] Example methods according to some embodiments may also include feeding the received point cloud fragments labeled as part of the selected object to the neural network to produce a higher resolution version of the selected object.
[0157] Example methods according to some embodiments may also include replacing object fragments in the point cloud streamed from the point cloud server with a higher resolution version of the selected object.
[0158] Example methods according to some embodiments may also include presenting point cloud data with a dynamically adjusted level of detail to a user.
[0159] Note that the various hardware elements of one or more described embodiments are referred to as "modules," which implement (i.e., execute, implement, etc.) the various functions described herein in conjunction with the corresponding modules. As used herein, a module includes hardware that a person skilled in the relevant art deems suitable for a given implementation (e.g., one or more processors, one or more microprocessors, one or more microcontrollers, one or more microchips, one or more application specific integrated circuits (ASICs), one or more field programmable gate arrays (FPGAs), one or more memory devices). Each described module may also include executable instructions for performing one or more functions described as being performed by the corresponding module, and note that these instructions may take the form of hardware (i.e., hardwired) instructions, firmware instructions, software instructions, etc., or include various instructions above, and may be stored in any suitable non-transitory computer-readable medium or media (such as what is commonly referred to as RAM, ROM, etc.).
[0160] Although the features and elements are described above in specific combinations, it will be understood by those of ordinary skill in the art that each feature or element may be used alone or in any combination with other features and elements. In addition, the methods described herein may be implemented in a computer program, software, or firmware that is incorporated into a computer-readable medium for execution by a computer or processor. Examples of computer-readable storage media include, but are not limited to, read-only memory (ROM), random access memory (RAM), registers, cache memory, semiconductor memory devices, magnetic media such as internal hard disks and removable disks, magneto-optical media, and optical media such as CD-ROM disks and digital versatile disks (DVDs). A processor associated with the software may be used to implement a radio frequency transceiver used in a WTRU, UE, terminal, base station, RNC, or any host computer.
Claims
1. A method performed by a client device (406, 502), comprising: Receiving point cloud data representing a plurality of three-dimensional objects from a first server (402, 504), wherein the point cloud data is segmented into segments respectively representing the three-dimensional objects and are assigned tags respectively associated with the three-dimensional objects; determining a viewpoint indicating a location where a user is viewing the point cloud data; selecting an object from the three-dimensional objects based on a distance between the object and the determined viewpoint; Retrieving a neural network model from a second server (404, 506) based on the tag associated with the selected object, wherein the neural network model has been previously trained to increase the level of detail of the selected object; generating an increased level of detail for the selected object using the acquired neural network model; as well as A view of the point cloud data is rendered, the point cloud data including the increased level of detail. 2 . The method of claim 1 , wherein generating the increased level of detail for the selected object comprises creating an illusion of additional geometric detail for the selected object. 3 . The method of claim 2 , wherein creating the illusion of additional geometric detail for the selected object increases a sampling density of the point cloud data.
4. The method according to claim 1, further comprising: sending the viewpoint to the first server; as well as Updated point cloud data having a level of detail based on the viewpoint is received from the first server. The method of claim 1 , wherein the point cloud data is received at a resolution provided by spatial capture. The method of claim 5 , wherein the resolution of the received point cloud is downsampled to comply with bandwidth constraints or viewing distance tolerances. 7 . The method of claim 1 , further comprising replacing points corresponding to the selected object within the point cloud data with an increased level of detail.
8. The method of claim 1 , wherein selecting the object from the three-dimensional objects using the viewpoint comprises: In response to determining that a point distance between the viewpoint and an object picked up from the three-dimensional object is less than a threshold, the object is selected as the selected object.
9. The method according to claim 1, further comprising: detecting one or more objects within the point cloud data, Wherein selecting the object further comprises selecting the object from the one or more objects detected in the point cloud data.
10. The method according to claim 1, further comprising: capturing data indicative of user movement; as well as The viewpoint is set using, at least in part, the data indicative of the user's movement.
11. The method according to claim 1, further comprising: Capturing motion data for head-mounted displays (HMDs); as well as The viewpoint is set using, at least in part, the motion data.
12. The method according to claim 1, further comprising: An identification of the second server is received from the first server.
13. The method of claim 1, wherein the first server and the second server are the same server.
14. The method of claim 1, wherein receiving the point cloud data comprises receiving the point cloud data from a local or remote sensor.
15. The method of claim 1, wherein acquiring the neural network model is based on mapping corresponding objects to information of available neural network models.
16. The method of claim 15, wherein the information mapping the corresponding object to the available neural network model includes a corresponding object label indicating a corresponding object category to which the corresponding object most closely matches.
17. The method according to claim 1, Wherein selecting the object comprises: selecting an object from the three-dimensional objects for which a corresponding level of detail can be increased, and Wherein obtaining the corresponding neural network model corresponding to the selected object includes determining whether the selected object corresponds to a new neural network model that does not exist on the client device.
18. An apparatus of a client device, comprising: one or more processors; as well as One or more non-transitory computer-readable media storing instructions that, when executed by the one or more processors, are operable to perform the method of any one of claims 1-17.