Dynamic overlay of mobile objects with real and virtual scenes for video see-through extended reality

By using an imaging sensor and a machine learning model to separate human skin pixels, the VST XR device achieves efficient overlay of moving objects with static scene content, solves the visual artifact problem, and enhances user interaction.

CN122122637APending Publication Date: 2026-05-29SAMSUNG ELECTRONICS CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SAMSUNG ELECTRONICS CO LTD
Filing Date
2024-11-06
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

When VST XR devices process moving objects, it is difficult to effectively overlay static scene content without producing obvious visual artifacts, especially when the user's hand is moving. Furthermore, existing technologies make it difficult to achieve convenient interaction between the user and the device.

Method used

Image frames and depth data are captured using an imaging sensor. A machine learning model is used to generate masks to separate human skin pixels, reconstruct images of moving objects and static scene content, and then combine virtual features for overlay and rendering.

Benefits of technology

It effectively reconstructs the overlay of moving objects and static scene content, reduces visual artifacts, and promotes convenient interaction between users and devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122122637A_ABST
    Figure CN122122637A_ABST
Patent Text Reader

Abstract

According to embodiments of the present disclosure, a method performed by a video see-through extended reality device can include obtaining image frames captured using at least one imaging sensor and depth data associated with a scene; generating a mask associated with a moving object using an artificial intelligence model trained to separate pixels of the scene corresponding to human skin from other portions; reconstructing an image of the moving object based on the image frames, the depth data, and the mask; reconstructing an image of the static scene content based on the image frames and the depth data; combining the image of the moving object, the image of the static scene content, and at least one virtual feature; and rendering the combined image on at least one display.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure generally relates to extended reality (XR) systems and processing. More specifically, this disclosure relates to the dynamic overlay of moving objects with real-world and virtual scenes for video perspective (VST) XR. Background Technology

[0002] Extended reality (XR) systems have become increasingly popular over time, and numerous applications have been developed and are being developed for XR systems. Some XR systems (such as augmented reality or "AR" systems and mixed reality or "MR" systems) enhance the user's view of the current environment by overlaying digital content (such as information or virtual objects) onto the user's current view of their surroundings. For example, some XR systems can often seamlessly blend computer-generated virtual objects with real-world scenes. Summary of the Invention

[0003] Technical solution According to embodiments of this disclosure, a method performed by a Video See-Through (VST) Extended Reality (XR) device may include: obtaining image frames of a scene captured using at least one imaging sensor and depth data associated with the scene.

[0004] According to embodiments of this disclosure, the method may include: using an artificial intelligence model to generate a mask associated with a moving object, the artificial intelligence model being trained to separate pixels of the scene corresponding to human skin from other parts.

[0005] According to embodiments of this disclosure, the method may include: reconstructing an image of the moving object based on the image frame, the depth data, and the mask.

[0006] According to embodiments of this disclosure, the method may include reconstructing an image of the static scene content based on the image frame and the depth data.

[0007] According to embodiments of this disclosure, the method may include combining an image of the moving object, an image of the static scene content, and at least one virtual feature.

[0008] According to embodiments of this disclosure, the method may include rendering the combined image on at least one display.

[0009] According to embodiments of this disclosure, a video see-through (VST) extended reality (XR) device may include: at least one imaging sensor; a memory; at least one display; and at least one processor communicatively coupled to the memory.

[0010] According to embodiments of this disclosure, the at least one processor executes a program or at least one instruction stored in the memory to cause the VST XR device to obtain image frames of a scene captured using at least one imaging sensor and depth data associated with the scene.

[0011] According to embodiments of this disclosure, the at least one processor executes the program or at least one instruction stored in the memory to cause the VST XR device to generate a mask associated with a moving object using an artificial intelligence model trained to separate pixels of the scene corresponding to human skin from other parts.

[0012] According to embodiments of this disclosure, the at least one processor executes the program or at least one instruction stored in the memory to cause the VST XR device to reconstruct an image of the moving object based on the image frame, the depth data, and the mask.

[0013] According to embodiments of this disclosure, the at least one processor executes the program or at least one instruction stored in the memory to cause the VST XR device to reconstruct an image of the static scene content based on the image frame and the depth data.

[0014] According to embodiments of this disclosure, the at least one processor executes the program or at least one instruction stored in the memory to cause the VST XR device to combine an image of the moving object, an image of the static scene content, and at least one virtual feature.

[0015] According to embodiments of this disclosure, the at least one processor executes the program or at least one instruction stored in the memory to cause the VST XR device to render a combined image on at least one display. Attached Figure Description

[0016] To gain a more complete understanding of this disclosure and its advantages, reference is now made to the following description taken in conjunction with the accompanying drawings, wherein: Figure 1 An example network configuration including electronic devices is shown according to this disclosure; Figure 2 An example architecture is shown that supports the dynamic overlay of moving objects with real-world and virtual scenes for Video Perspective (VST) Extended Reality (XR) according to this disclosure; Figure 3a and 3b An example architecture supporting the dynamic overlay of moving objects with real-world and virtual scenes for VST XR is shown according to this disclosure; Figure 4 and Figure 5 An example mask generation function for moving objects for VST XR according to this disclosure is shown; Figure 6 An example processing for scene reconstruction using VST XR according to this disclosure is shown; Figure 7 An example overlay process between a moving object and static scene content for VST XR is shown according to this disclosure; Figures 8a to 8c Example use cases of dynamic overlay of moving objects with real-world and virtual scenes for VST XR are shown according to this disclosure; and Figure 9 An example method for dynamically overlaying moving objects with real-world and virtual scenes for VST XR, according to this disclosure, is shown. Detailed Implementation

[0017] This disclosure relates to the dynamic overlay of moving objects with real-world and virtual scenes for Video Perspective (VST) Extended Reality (XR).

[0018] According to embodiments of this disclosure, a method includes: obtaining image frames of a scene captured using one or more imaging sensors of a VST XR device and depth data associated with the scene. The image frames capture moving objects and static scene content, and the moving objects include a portion of a user's body. The method further includes: using a machine learning model to generate a mask associated with the moving objects, the machine learning model being trained to separate pixels of the scene corresponding to human skin from other parts. The method further includes: reconstructing an image of the moving objects based on the image frames, the depth data, and the mask, and reconstructing an image of the static scene content based on the image frames and the depth data. Furthermore, the method includes combining the image of the moving objects, the image of the static scene content, and one or more virtual features to generate a combined image, and rendering the combined image to be displayed on at least one display of the VST XR device.

[0019] According to embodiments of this disclosure, a VST XR device includes: one or more imaging sensors, at least one display, and at least one processing device. The at least one processing device is configured to: acquire image frames of a scene captured using the one or more imaging sensors and depth data associated with the scene. The image frames capture moving objects and static scene content, and the moving objects include a portion of a user's body. The at least one processing device is further configured to use a machine learning model to generate a mask associated with the moving objects, the machine learning model being trained to separate pixels of the scene corresponding to human skin from other parts. The at least one processing device is further configured to: reconstruct an image of the moving objects based on the image frames, the depth data, and the mask, and reconstruct an image of the static scene content based on the image frames and the depth data. Additionally, the at least one processing device is configured to: combine the image of the moving objects, the image of the static scene content, and one or more virtual features to generate a combined image and render the combined image to be displayed on the at least one display.

[0020] According to embodiments of this disclosure, a non-transitory machine-readable medium includes instructions, when executed, to cause at least one processor of a VST XR device to perform the following operations: obtain image frames of a scene captured using one or more imaging sensors of the VST XR device and depth data associated with the scene. The image frames capture moving objects and static scene content, and the moving objects include a portion of a user's body. The non-transitory machine-readable medium also includes instructions, when executed, to cause the at least one processor to perform the following operations: generate a mask associated with the moving object using a machine learning model trained to separate pixels of the scene corresponding to human skin from other parts. The non-transitory machine-readable medium also includes instructions, when executed, to cause the at least one processor to perform the following operations: reconstruct an image of the moving object based on the image frames, the depth data, and the mask, and reconstruct an image of the static scene content based on the image frames and the depth data. Additionally, the non-transitory machine-readable medium includes instructions, when executed, to cause the at least one processor to perform the following operations: combine the image of the moving object, the image of the static scene content, and one or more virtual features to generate a combined image and render the combined image to be displayed on at least one display of the VST XR device.

[0021] Other technical features will be apparent to those skilled in the art from the following figures, description and claims.

[0022] It may be advantageous to define certain words and phrases used throughout this patent document. The terms “send,” “receive,” and “communicate,” and their derivatives, include both direct and indirect communication. The terms “comprise,” “include,” and their derivatives, mean including but not limited to. The term “or” is inclusive, meaning and / or. The phrase “associated with,” and its derivatives, mean including, being included in, interconnected with, containing, being contained within, connected to or connected with, coupled to or coupled with, communicable with, cooperating with, intertwined, juxtaposed, proximate, bound to or bound with, having, possessing the nature of, having a relationship to or with, etc.

[0023] Furthermore, the various functions described below can be implemented or supported by one or more computer programs, each of which is formed and implemented in a computer-readable medium by computer-readable program code. The terms "application" and "program" refer to one or more computer programs, software components, instruction sets, procedures, functions, objects, classes, instances, associated data, or portions thereof adapted to be implemented in suitable computer-readable program code. The phrase "computer-readable program code" includes any type of computer code, including source code, object code, and executable code. The phrase "computer-readable medium" includes any type of medium accessible by a computer, such as read-only memory (ROM), random access memory (RAM), hard disk drive, compact disc (CD), digital video optical disc (DVD), or any other type of storage. "Non-transitory" computer-readable media does not include wired, wireless, optical, or other communication links that transmit transient electrical or other signals. Non-transitory computer-readable media includes media that can permanently store data and media that can store data and later rewrite it, such as rewritable optical discs or erasable memory devices.

[0024] It should be understood that the boxes in each flowchart and the combination of flowcharts can be executed by one or more computer programs that include computer-executable instructions. The entirety of one or more computer programs can be stored in a single memory, or one or more computer programs can be divided into different parts stored in multiple different memories.

[0025] Any function or operation described herein can be processed by a processor or a combination of processors. A processor or combination of processors is circuitry that performs processing and includes circuitry such as an application processor (AP), a communication processor (CP), a graphics processing unit (GPU), a neural processing unit (NPU), a microprocessor unit (MPU), a system-on-a-chip (SoC), an IC, etc.

[0026] As used herein, terms and phrases such as “having,” “may have,” “comprising,” or “may include” indicate the presence of a feature (such as a number, function, operation, or component, like a part) and do not exclude the presence of other features. Furthermore, as used herein, the phrases “A or B,” “at least one of A and / or B,” or “one or more of A and / or B” can include all possible combinations of A and B. For example, “A or B,” “at least one of A and B,” and “at least one of A or B” can indicate all of the following: (1) including at least one A, (2) including at least one B, or (3) including at least one A and at least one B. Furthermore, as used herein, the terms “first” and “second” can modify various components regardless of their importance and do not limit the components. These terms are used only to distinguish one component from another. For example, a first user device and a second user device can refer to user devices that are different from each other, regardless of the order or importance of these devices. Without departing from the scope of this disclosure, a first component can be referred to as a second component, and vice versa.

[0027] It will be understood that when an element (such as the first element) is referred to as being "coupled" / "coupled to" another element (such as the second element) or "connected" / "connected to" another element (such as the second element), it can be directly coupled or connected / coupled to or connected to the other element, either directly or via a third element. Conversely, it will be understood that when an element (such as the first element) is referred to as being "directly coupled" / "directly coupled to" another element (such as the second element) or "directly connected" / "directly connected to" another element (such as the second element), no other element (such as a third element) intervenes between that element and the other element.

[0028] As used herein, the phrase “configured (or set) to” may be used interchangeably with the phrases “suitable for,” “capable of,” “designed to,” “adapted to,” “manufactured as,” or “capable of.” The phrase “configured (or set) to” does not inherently mean “specifically designed in hardware.” Rather, the phrase “configured to” can mean that a device can perform operations in conjunction with another device or component. For example, the phrase “processor configured (or set) to perform A, B, and C” can refer to a general-purpose processor (such as a CPU or application processor) that can perform operations by executing one or more software programs stored in a memory device, or a special-purpose processor (such as an embedded processor) for performing operations.

[0029] The terms and phrases used herein are provided to describe only some embodiments of this disclosure and are not intended to limit the scope of other embodiments of this disclosure. It will be understood that, unless the context clearly indicates otherwise, the singular forms “a,” “an,” and “the” include plural references. All terms and phrases used herein, including technical and scientific terms and phrases, have the same meaning as commonly understood by one of ordinary skill in the art to which the embodiments of this disclosure pertain. It will be further understood that terms and phrases such as those defined in common dictionaries should be interpreted as having the meaning consistent with their meaning in the context of the relevant art and will not be interpreted in an idealized or overly formal sense unless expressly so defined herein. In some cases, the terms and phrases defined herein may be interpreted as excluding embodiments of this disclosure.

[0030] Examples of "electronic devices" according to embodiments of this disclosure may include at least one of a smartphone, tablet PC, mobile phone, video phone, e-book reader, desktop PC, laptop computer, netbook computer, workstation, personal digital assistant (PDA), portable multimedia player (PMP), MP3 player, mobile medical device, camera, or wearable device (such as smart glasses, head-mounted display (HMD), electronic clothing, electronic bracelet, electronic necklace, electronic accessory, electronic tattoo, smart mirror, or smartwatch). Other examples of electronic devices include smart home appliances. Examples of smart home appliances may include at least one of a television, digital video disc (DVD) player, audio player, refrigerator, air conditioner, vacuum cleaner, oven, microwave oven, washing machine, dryer, air purifier, set-top box, home automation control panel, security control panel, TV box (such as Samsung HOMESYNC, Apple TV, or Google TV), smart speaker or speaker with integrated digital assistant (such as Samsung GALAXY HOME, Apple HOMEPOD, or Amazon ECHO), game console (such as XBOX, PLAYSTATION, or NINTENDO), electronic dictionary, electronic key, camera, or electronic photo frame. Other examples of electronic devices include at least one of various medical devices (such as various portable medical measuring devices (e.g., blood glucose measuring devices, heart rate measuring devices, or body temperature measuring devices), magnetic resonance angiography (MRA) devices, magnetic resonance imaging (MRI) devices, computed tomography (CT) devices, imaging devices, or ultrasound devices), navigation devices, global positioning system (GPS) receivers, event data recorders (EDR), flight data recorders (FDR), automotive infotainment devices, marine electronic devices (such as marine navigation devices or gyrocompasses), avionics equipment, safety devices, vehicle head units, industrial or household robots, automated teller machines (ATMs), point-of-sale (POS) devices, or Internet of Things (IoT) devices (such as light bulbs, various sensors, electricity or gas meters, sprinklers, fire alarms, thermostats, streetlights, toasters, fitness equipment, hot water tanks, heaters, or boilers). Other examples of electronic devices include at least a portion of furniture or building / structure, electronic boards, electronic signature receivers, projectors, or various measuring devices (such as devices for measuring water, electricity, gas, or electromagnetic waves). Note that, according to embodiments of this disclosure, the electronic device may be one or a combination of the devices listed above. According to embodiments of this disclosure, the electronic device may be a flexible electronic device. The electronic devices disclosed herein are not limited to those listed above and may include any other electronic devices now known or developed hereafter.

[0031] In the following description, an electronic device according to embodiments of the present disclosure is described with reference to the accompanying drawings. As used herein, the term "user" may refer to a person using the electronic device or another device (such as an artificial intelligence electronic device).

[0032] Definitions for certain other words and phrases may be provided throughout this patent document. Those skilled in the art will understand that, in many cases (if not most), such definitions apply to the prior and future use of the words and phrases defined in this way.

[0033] Nothing described in this application should be construed as implying that any particular element, step, or function is an essential element that must be included within the scope of the claims. The scope of the patent subject matter is defined solely by the claims. Any other terms used in the claims (including, but not limited to, “mechanism,” “module,” “device,” “unit,” “component,” “element,” “building,” “equipment,” “machine,” “system,” “processor,” or “controller”) shall be understood to refer to structures known to a person skilled in the art.

[0034] The following discussion is described with reference to the accompanying drawings. Figures 1 to 9 Various embodiments of this disclosure are also described. However, it should be understood that this disclosure is not limited to these embodiments, and all changes and / or equivalents or substitutions thereof are also within the scope of this disclosure. Throughout the specification and drawings, the same or similar reference numerals may be used to refer to the same or similar elements.

[0035] As mentioned above, extended reality (XR) systems have become increasingly popular over time, and numerous applications have been and are being developed for XR systems. Some XR systems (such as augmented reality or "AR" systems and mixed reality or "MR" systems) enhance the user's view of the current environment by overlaying digital content (such as information or virtual objects) onto the user's current view of their surroundings. For example, some XR systems can often seamlessly blend computer-generated virtual objects with real-world scenes.

[0036] Optical see-through (OST) XR systems refer to XR systems where users directly view real-world scenes through a head-mounted display (HMD). Unfortunately, OST XR systems face several challenges that may limit their adoption. Some of these challenges include a limited field of view, limited space for use (such as indoor use only), the inability to display completely opaque black objects, and the potential use of complex optical pipelines that may require projectors, waveguides, and other optical components. In contrast to OST XR systems, video see-through (VST) XR systems (also known as "pass-through" XR systems) present users with a generated video sequence of real-world scenes. VST XR systems can be built using virtual reality (VR) technology and can offer various advantages over OST XR systems. For example, VST XR systems can provide a wider field of view and can offer improved contextual augmented reality.

[0037] Unfortunately, VST XR devices can encounter various problems when moving objects (including the user's hand) are present in the imaged scene. For example, a moving object often occludes part of static content in one image frame and another part in another, and VST XR devices struggle to compensate for this without producing noticeable visual artifacts. Furthermore, VST XR devices often incorrectly render virtual features (such as virtual objects) for moving objects. In some cases, users of VST XR devices can provide input by interacting with a physical or virtual keyboard, and the inability to effectively generate images showing the user's hand movements and the hand superimposed on the keyboard can interfere with convenient interaction with the VST XR device.

[0038] This disclosure provides various techniques for supporting the dynamic overlay of moving objects with real-world and virtual scenes for VST XR. As described in more detail below, image frames of a scene can be acquired using one or more imaging sensors of a VST XR device, and depth data associated with the scene can be acquired. The image frames can capture moving objects and static scene content, and the moving object may include a part of the user's body (such as one or more of the user's hands). A mask associated with the moving object can be generated using a machine learning model trained to separate pixels of the scene corresponding to human skin from other parts. An image of the moving object can be reconstructed based on the image frames, depth data, and mask, and an image of the static scene content can be reconstructed based on the image frames and depth data. The image of the moving object, the image of the static scene content, and one or more virtual features can be combined to generate a combined image, and the combined image can be rendered for presentation on at least one display of the VST XR device.

[0039] In this way, the disclosed technology provides an efficient mechanism for resolving the overlay of moving objects (such as a user's hand) onto captured static scene content. For example, moving objects and static scene content can be reconstructed efficiently, and the moving object can be overlaid onto the static scene content while achieving parallax correction and hole filling. Unoccluded areas (areas in the scene that are overlaid by moving objects in some image frames but not in others) can be effectively identified so that hole filling can be performed, significantly reducing visual artifacts. Therefore, a final view of the scene can be created using correct parallax for moving objects, static scene content, and one or more virtual objects. As a specific example, the disclosed technology can be used to efficiently reconstruct a keyboard in a captured scene and overlay the user's reconstructed hand onto the keyboard, which can facilitate easier and more efficient interaction between the user and the VST XR device.

[0040] Figure 1 An example network configuration 100 including electronic devices according to this disclosure is shown. Figure 1 The embodiment of network configuration 100 shown is for illustrative purposes only. Other embodiments of network configuration 100 may be used without departing from the scope of this disclosure.

[0041] According to embodiments of this disclosure, electronic device 101 is included in network configuration 100. Electronic device 101 may include at least one of bus 110, processor 120, memory 130, input / output (I / O) interface 150, display 160, communication interface 170, and sensor 180. According to embodiments of this disclosure, electronic device 101 may not include at least one of these components, or at least one other component may be added. Bus 110 includes circuitry for connecting components 120 to 180 to each other and for transmitting communication (such as control messages and / or data) between components.

[0042] Processor 120 includes one or more processing devices, such as one or more microprocessors, microcontrollers, digital signal processors (DSPs), application-specific integrated circuits (ASICs), or field-programmable gate arrays (FPGAs). According to embodiments of this disclosure, processor 120 includes one or more of a central processing unit (CPU), application processor (AP), communication processor (CP), graphics processing unit (GPU), or neural processing unit (NPU). Processor 120 is capable of performing control and / or performing operations or data processing related to communication or other functions on at least one of the other components of electronic device 101. As described below, processor 120 can perform one or more functions related to the dynamic overlay of moving objects with real-world and virtual scenes for VST XR.

[0043] Memory 130 may include volatile and / or non-volatile memory. For example, memory 130 may store commands or data associated with at least one other component of electronic device 101. According to embodiments of this disclosure, memory 130 may store software and / or programs 140. Program 140 includes, for example, a kernel 141, middleware 143, an application programming interface (API) 145, and / or application programs (or “applications”) 147. At least a portion of kernel 141, middleware 143, or API 145 may be represented as an operating system (OS).

[0044] Kernel 141 can control or manage system resources (such as bus 110, processor 120, or memory 130) used to perform operations or functions implemented in other programs (such as middleware 143, API 145, or application 147). Kernel 141 provides interfaces that allow middleware 143, API 145, or application 147 to access various components of electronic device 101 to control or manage system resources. Application 147 may include one or more applications that, among other functions, perform dynamic overlays of moving objects with real and virtual scenes for VST XR. These functions may be performed by a single application or multiple applications, each of which performs one or more of these functions. For example, middleware 143 may act as a relay to allow API 145 or application 147 to communicate data with kernel 141. Multiple applications 147 may be provided. Middleware 143 can control job requests received from application 147, for example, by allocating priority of system resources (such as bus 110, processor 120, or memory 130) using electronic device 101 to at least one of multiple applications 147. API 145 is an interface that allows application 147 to control functionality provided from kernel 141 or middleware 143. For example, API 145 includes at least one interface or function (such as commands) for archive control, window control, image processing, or text control.

[0045] I / O interface 150 serves as an interface through which commands or data input from a user or other external device can be transmitted to other components of electronic device 101, for example. I / O interface 150 can also output commands or data received from other components of electronic device 101 to the user or other external device.

[0046] Display 160 includes, for example, a liquid crystal display (LCD), a light-emitting diode (LED) display, an organic light-emitting diode (OLED) display, a quantum dot light-emitting diode (QLED) display, a microelectromechanical system (MEMS) display, or an electronic paper display. Display 160 can also be a depth-sensing display, such as a multi-focal display. Display 160 is capable of displaying various content (such as text, images, videos, icons, or symbols) to a user. Display 160 may include a touchscreen and can receive input such as touch, gestures, proximity, or hover using an electronic pen or a user's body part.

[0047] For example, communication interface 170 can establish communication between electronic device 101 and external electronic devices (such as first electronic device 102, second electronic device 104, or server 106). For example, communication interface 170 can be connected to network 162 or 164 via wireless or wired communication to communicate with external electronic devices. Communication interface 170 can be a wired transceiver or a wireless transceiver or any other component for sending and receiving signals.

[0048] Wireless communication can use at least one of the following as a communication protocol: WiFi, LTE, LTE-A, 5G, millimeter wave or 60 GHz wireless communication, wireless USB, CDMA, WCDMA, UMTS, Wi-Fi, or GSM. Wired connections may include at least one of the following: USB, HDMI, RS-232, or POTS. Network 162 or 164 includes at least one communication network, such as a computer network (e.g., a local area network (LAN) or wide area network (WAN)), the Internet, or a telephone network.

[0049] Electronic device 101 also includes one or more sensors 180 capable of measuring physical quantities or detecting the activation state of electronic device 101 and converting the measured or detected information into electrical signals. For example, sensor 180 may include a camera or other imaging sensor that can be used to capture images of a scene. Sensor 180 may also include one or more buttons for touch input, one or more microphones, depth sensors, gesture sensors, gyroscopes or gyroscope sensors, barometric pressure sensors, magnetic sensors or magnetometers, accelerometers or accelerometers, grip sensors, proximity sensors, color sensors (such as red-green-blue (RGB) sensors), biometric physical sensors, temperature sensors, humidity sensors, illuminance sensors, ultraviolet (UV) sensors, electromyography (EMG) sensors, electroencephalography (EEG) sensors, electrocardiography (ECG) sensors, infrared (IR) sensors, ultrasound sensors, iris sensors, or fingerprint sensors. Furthermore, sensor 180 may include one or more position sensors, such as inertial measurement units that may include one or more accelerometers, gyroscopes, and other components. Additionally, sensor 180 may include control circuitry for controlling at least one of the sensors included herein. Any of these sensors 180 can be located within the electronic device 101.

[0050] According to embodiments of this disclosure, electronic device 101 may be a wearable device or a wearable device (such as an HMD) on which electronic devices can be installed. For example, electronic device 101 may represent an XR wearable device, such as headphones or smart glasses. According to embodiments of this disclosure, a first external electronic device 102 or a second external electronic device 104 may be a wearable device or a wearable device (such as an HMD) on which electronic devices can be installed. According to embodiments of this disclosure, when electronic device 101 is installed in electronic device 102 (such as an HMD), electronic device 101 can communicate with electronic device 102 via communication interface 170. Electronic device 101 can be directly connected to electronic device 102 to communicate with electronic device 102 without involving a separate network.

[0051] Each of the first external electronic device 102, the second external electronic device 104, and the server 106 may be a device of the same or different type as electronic device 101. According to embodiments of this disclosure, server 106 comprises a group of one or more servers. Furthermore, according to embodiments of this disclosure, all or some of the operations performed on electronic device 101 may be performed on another or more other electronic devices (such as electronic device 102 and electronic device 104 or server 106). Furthermore, according to embodiments of this disclosure, when electronic device 101 is required to automatically or upon request perform certain functions or services, electronic device 101 may request another device (such as electronic device 102 and electronic device 104 or server 106) to perform at least some of its associated functions, rather than performing the functions or services itself, or additionally performing said functions or services. The other electronic device (such as electronic device 102 and electronic device 104 or server 106) is capable of performing the requested functions or additional functions and transmitting the execution results to electronic device 101. Electronic device 101 can provide the requested function or service by processing the received results as is or additionally. For this purpose, cloud computing, distributed computing, or client-server computing technologies can be used, for example. Although Figure 1 The electronic device 101 is shown to include a communication interface 170 for communicating with an external electronic device 104 or a server 106 via a network 162 or 164; however, according to embodiments of this disclosure, the electronic device 101 may operate independently without separate communication functionality.

[0052] Server 106 may include components that are the same as or similar to those of electronic device 101 (or a suitable subset thereof). Server 106 may support driving electronic device 101 by performing at least one of the operations (or functions) implemented on electronic device 101. For example, server 106 may include a processing module or processor that can support processor 120 implemented in electronic device 101. As described below, server 106 may perform one or more functions related to the dynamic overlay of moving objects with real-world and virtual scenes for VST XR.

[0053] although Figure 1 An example of a network configuration 100 including electronic device 101 is shown, but more can be found elsewhere. Figure 1 Various changes can be made. For example, network configuration 100 can include any number of each component in any suitable arrangement. Typically, computing and communication systems have a wide variety of configurations, and Figure 1 This disclosure is not intended to limit the scope to any particular configuration. Furthermore, although… Figure 1 An operating environment is shown that can use the various features disclosed in this patent document, but these features can be used in any other suitable system.

[0054] Figure 2 A first example architecture 200 supporting the dynamic overlay of moving objects with real-world and virtual scenes for VST XR is shown according to this disclosure. For ease of explanation, Figure 2 The architecture 200 is described as using Figure 1 The network configuration 100 is implemented using electronic device 101. However, architecture 200 can be implemented using any other suitable device and in any other suitable system.

[0055] like Figure 2 As shown, architecture 200 includes: image and depth data capture operation 202, which generally operates to obtain captured image frames of a scene and depth data associated with the scene. For example, image and depth data capture operation 202 can be used to obtain perspective image frames captured using one or more perspective cameras or other imaging sensors 180 of a VST XR device. In some cases, image and depth data capture operation 202 can be used to obtain perspective image frames at a desired frame rate (such as 30, 60, 90, or 120 frames per second). Each perspective image frame can have any suitable size, shape, and resolution, and includes image data in any suitable domain. As a specific example, each perspective image frame may include RGB image data, YUV image data, or Bayer or other raw image data. Image and depth data capture operation 202 can also be used to obtain depth data associated with the captured image frames. For example, at least one depth sensor 180 used in or with a VST XR device can capture depth data within a scene imaged using a perspective camera. Any suitable type of depth sensor 180 can be used, such as a light detection and ranging (LIDAR) or time-of-flight (ToF) depth sensor. In some cases, the acquired depth data may have a resolution smaller than (and possibly significantly smaller than) the resolution of the captured image frames. For example, the depth data may have a resolution equal to or less than half the resolution of each of the captured image frames. As a specific example, the captured image frames may have a 3K or 4K resolution, and the depth data may have a resolution of 320 depth values ​​multiplied by 320.

[0056] Architecture 200 also includes a head pose tracking and prediction operation 204, which generally operates to obtain information related to the head pose of a user using the VSTXR device and to predict how the user's head pose might change over time. For example, when capturing image frames, the head pose tracking and prediction operation 204 can obtain input from an IMU sensor, a head pose tracking camera, or other sensors 180 of the electronics 101. The head pose tracking and prediction operation 204 can also use at least one model of the user's head pose to predict how the user's head pose might change in the future based on the user's previous and current head poses or changes in head pose. This can be useful because there is typically a delay between the capture of image frames and the display of a rendered image based on those captured image frames, and the user may move his or her head during this intermediate time period.

[0057] The acquired image frames and depth data are provided to moving object reconstruction operation 206 and real-world scene reconstruction operation 208. Moving object reconstruction operation 206 generally operates to generate images of one or more moving objects captured in the image frames. For example, moving object reconstruction operation 206 may generate a mask associated with each moving object in the image frame, where the mask identifies the boundaries of each moving object in the image frame. In some cases, moving objects in the image frames may include one or more hands or other body parts of a user, and moving object reconstruction operation 206 may use a skin color model or other machine learning model trained to separate pixels corresponding to human skin from other parts of the scene. The mask effectively allows moving object reconstruction operation 206 to track the boundaries of each moving object across image frames, which allows the generation of images that include only the moving objects. Real-world scene reconstruction operation 208 generally operates to generate images of static scene content (such as one or more static or background objects) captured in the image frames. For example, real-world scene reconstruction operation 208 may use the same mask as moving object reconstruction operation 206 to generate images omitting moving objects, which allows the generation of images that include only static scene content.

[0058] The virtual scene construction operation 210 generally operates to create one or more virtual features that will be included in an image displayed to a user of a VST XR device. For example, the virtual scene construction operation 210 may generate one or more virtual objects that will be placed at one or more desired locations in an image displayed to a user of a VST XR device. One or more virtual objects may represent or include any suitable virtual content to be presented to the user, and the virtual content may vary depending on the application. Generally, this disclosure is not limited to use with any particular type of virtual object or other virtual content. The virtual content generated by the virtual scene construction operation 210 can be said to form a virtual scene.

[0059] The moving object overlay and parallax correction operation 212 generally operates to determine how the images of moving objects should be overlaid with the images of static scene content. For example, the moving object overlay and parallax correction operation 212 can determine where each image of a moving object should be placed relative to the static scene content in its corresponding image. The moving object overlay and parallax correction operation 212 also operates to adjust the images of the moving objects to provide the correct parallax. For example, the moving object overlay and parallax correction operation 212 can estimate the depth of each moving object (or the depth of individual points of each moving object) to ensure that the images of the moving objects are placed to provide the correct parallax from the user's perspective.

[0060] The 2D-to-3D warp operation 214 generally operates to superimpose and warp an image of a moving object and an image of static scene content. For example, the 2D-to-3D warp operation 214 can use the user's predicted head pose received from the head pose tracking and prediction operation 204 to determine how to warp the superimposed image of the moving object and static scene content to compensate for the user's predicted motion. The 2D-to-3D warp operation 214 can also perform hole filling, which may involve identifying locations where image data is missing in the 3D version of the scene and filling that image data (such as by using image data from one or more other image frames). The 2D-to-3D warp operation 214 can also perform parallax correction based on the user's predicted head pose so that the resulting warped image has appropriate parallax. Similarly, the 2D-to-3D warp operation 216 generally operates to warp the virtual content generated by the virtual scene construction operation 210. For example, the 2D-to-3D warp operation 216 can use the user's predicted head pose received from the head pose tracking and prediction operation 204 to determine how to warp the virtual scene generated by the virtual scene construction operation 210 to compensate for the user's predicted motion. The 2D-to-3D warp operation 216 can also perform parallax correction based on the user's predicted head pose so that the resulting warped virtual scene has appropriate parallax.

[0061] This results in the generation of a reconstructed image of the actual / real-world scene via warping operation 214, and the generation of a reconstructed image of the virtual scene via warping operation 216. Scene compositing operation 218 is performed in general to combine the reconstructed images of the actual and virtual scenes. For example, scene compositing operation 218 may overlay the reconstructed image of the virtual scene onto the reconstructed image of the actual scene. This results in the generation of a combined image, which may also be referred to as the composite image. Image rendering and display operation 220 is performed in general to render the combined image and begin rendering the rendered image on one or more displays 160 of the VST XR device. According to the implementation, the rendered image may be displayed on separate displays 160 (such as the left and right display panels) or separate portions of the same display 160 (such as the left and right portions of a common display panel).

[0062] although Figure 2 This illustrates a first example of an architecture 200 that supports the dynamic overlay of moving objects with real-world and virtual scenes for VST XR, but it is possible to... Figure 2 Make various changes. For example, Figure 2 The various components or functions can be combined, further subdivided, copied, omitted or rearranged, and additional components or functions can be added as needed.

[0063] Figure 3a and Figure 3b A second example architecture 300 supporting the dynamic overlay of moving objects with real-world and virtual scenes for VST XR is shown according to this disclosure. For ease of explanation, Figure 3a and Figure 3b The architecture 300 is described as using Figure 1 The network configuration 100 is implemented using electronic device 101. However, architecture 300 can be implemented using any other suitable device and in any other suitable system.

[0064] like Figure 3aAs shown, data capture operation 302 generally operates to obtain image frames and associated depth maps. For example, data capture operation 302 may include: an image frame capture function 304 for acquiring perspective image frames captured using one or more perspective cameras or other imaging sensors 180 of the VST XR device. The image frames can be captured at any suitable frame rate, have any suitable size, shape, and resolution, and include image data in any suitable domain. Data capture operation 302 may also include: a depth map capture function 306 for acquiring depth maps or other depth data from one or more depth sensors 180 of the VST XR device. The depth data can have any suitable size, shape, and resolution. Any suitable type of depth sensor 180 can be used, such as a LiDAR or ToF depth sensor.

[0065] The head pose tracking operation 308 generally operates to obtain information related to the head pose of the user using the VST XR device and how the user's head pose changes over time. For example, while an image frame is being captured, the head pose tracking operation 308 can obtain input from one or more of the head pose tracking camera 310, head pose tracking sensor 312, and / or IMU sensor 314 of the electronic device 101. The head pose tracking operation 308 can also track how these inputs change over time.

[0066] Depth processing operation 316 generally operates to process the acquired image frames and the acquired depth maps or other depth data. For example, contour / boundary extraction function 318 can process image frames and depth data to estimate the contours and boundaries of objects or other content captured in the image frames. In some cases, contour / boundary extraction function 318 can process image frames and depth data to identify estimated outer contours or outer boundaries of objects within the scene. One or more of the objects detected here may represent one or more moving objects. Depth map verification and super-resolution function 320 can process stereo image frame pairs and depth data to verify depth measurements in the depth data and improve the resolution of the depth data. For example, depth map verification and super-resolution function 320 can use the disparity in the stereo image frame pairs to estimate the depth within the scene, combine those estimated depths with depths generated using one or more depth sensors 180, perform hole filling to fill missing depths, and filter the resulting set of depths. This is often referred to as depth densification and can result in the generation of denser depth maps or other depth data with higher resolution.

[0067] Mask creation operation 322 generally operates to detect moving objects captured in an image frame and generate a mask associated with the moving object. For example, moving object detection function 324 may process image frames, dense depth maps, or other information to identify one or more moving objects captured in the image frame. Mask generation function 326 may generate at least one mask for the moving object detected in each of the image frames. In some cases, each mask may represent a binary mask, in which each pixel has a value indicating whether the corresponding pixel in the image frame is associated with or not associated with a moving object. According to embodiments of this disclosure, mask generation function 326 may perform segmentation to distinguish moving objects in an image frame from other parts of the image frame. According to embodiments of this disclosure, mask creation operation 322 may use a trained skin color model or other trained machine learning model to separate pixels in a scene corresponding to human skin from other parts of the scene. The machine learning model may represent any suitable machine learning architecture that can be trained to generate masks, such as a deep neural network. Note that regardless of how the machine learning model operates (such as identifying areas of visible skin on a user or performing body part segmentation), the machine learning model effectively operates to separate the image pixels of the scene corresponding to human skin from the rest of the scene.

[0068] The dense depth map or other depth data generated by depth processing operation 316, and the mask generated by mask creation operation 322, are provided to moving object reconstruction operation 328, which generally operates to generate reconstructed images of any moving objects in the image frame. For example, boundary / edge estimation and tracking function 330 can be used to estimate the edges or boundaries of each detected moving object within the captured image frame. In some cases, boundary / edge estimation and tracking function 330 can use results from contour / boundary extraction function 318 and / or from mask generation function 326 to identify the edges or boundaries of each detected moving object, or boundary / edge estimation and tracking function 330 can identify the edges or boundaries of each detected moving object individually. Boundary / edge estimation and tracking function 330 can also track the edges or boundaries of each detected moving object over time, such as by identifying edges or boundaries in each image frame of the captured image frame sequence. The moving object reconstruction and de-occlusion filling function 332 can be used to reconstruct an image of each moving object associated with each captured image frame. For example, the moving object reconstruction and de-occlusion filling function 332 can isolate each moving object in each captured image frame and perform hole filling. As described above, hole filling can involve recreating image content in at least one area that was previously occluded (but is no longer occluded) due to the movement of the moving object. Any suitable technique can be used to perform hole filling, such as recreating image content in one image frame based on image content in one or more other image frames. The transformation function 334 can be used to transform each image of the moving object to provide parallax correction. For example, the transformation function 334 can be used to distort the image of the moving object to provide the desired parallax.

[0069] The captured image frames, depth data, and reconstructed images of moving objects are provided to a real-scene reconstruction operation 336, which generally operates to generate an image of the reconstructed static scene content captured in the image frames and overlays the reconstructed images of the moving objects onto the reconstructed static scene image. For example, scene reconstruction function 338 can generate an image of the static scene content captured in the image frames, such as by excluding portions of the image frames that include the moving objects. A moving object / static scene overlay function 340 can be used to place an image of the moving object on top of an image of the static scene content. In some cases, for example, the moving object / static scene overlay function 340 can overlay images of the moving objects to place them in a suitable position within a larger image of the static scene content.

[0070] The virtual scene generation operation 342 generally operates to generate a virtual scene that will be combined with an image of the real scene. For example, the virtual object rendering function 344 can be used to generate one or more virtual objects or other virtual content to be inserted into the image. Furthermore, any suitable virtual content can be generated here. The parallax correction function 346 can be used to transform or otherwise modify one or more virtual objects or other virtual content to provide the desired parallax for the virtual content. In some cases, a dense depth map or other depth data provided by the depth processing operation 316 can be used to obtain the correct parallax for the virtual content.

[0071] like Figure 3bAs shown, head pose prediction operation 348 generally operates to estimate what the user's head pose will be when a rendered image based on captured image frames is displayed to the user. For example, latency estimation function 350 can be used to identify the estimated latency between the capture of an image frame and the display of the rendered image. As a specific example, latency estimation function 350 can estimate the amount of time required for various operations to be performed in architecture 300, such as by identifying the time between capturing an image frame and completing other functions of architecture 300. Head motion model estimation function 352 can be used to process head pose information associated with the user in order to construct a model of the user's head pose. For example, the model can identify how the user's head pose changes over time based on different starting positions and movements. Head pose prediction function 354 can apply this model to information from head pose tracking operation 308 to estimate what the user's head pose might be after a period of time equal to the estimated latency. In other words, head pose prediction function 354 can use the model and information associated with the user's head pose to estimate the user's head pose when the rendered image is presented to the user.

[0072] The post-processing operation 356 generally operates to modify the overlay image of moving objects and static scene content. For example, the head pose change compensation function 358 can modify the overlay image of moving objects and static scene content based on the user's predicted head pose. As a specific example, the head pose change compensation function 358 can rotate and / or translate the overlay image so that the overlay image is adapted to be presented with the user's predicted head pose.

[0073] The Geometric Distortion Correction (GDC) / Cchromatic Aberration Correction (CAC) function 360 can modify the overlaid image to correct distortions that occur in the displayed image. For example, in many VST XR devices, a rendered image is presented on one or more displays 160, and the rendered image is typically viewed by a user through left and right display lenses located between the user's eyes and the displays 160. However, when viewing the displayed image, the display lenses may produce geometric distortion, and chromatic aberration may occur as light passes through them. The GDC / CAC function 360 can adjust the overlaid image so that the resulting image pre-compensates for the expected geometric distortion and chromatic aberration. Therefore, the GDC / CAC function 360 can determine how the image should be pre-distorted to compensate for subsequent geometric distortion and chromatic aberration that occur when the image is displayed and viewed through the display lenses. In some cases, the GDC / CAC function 360 can operate based on a display lens GDC and CAC model, which mathematically represents the geometric distortion and chromatic aberration caused by the display lenses.

[0074] Similarly, the virtual scene post-processing operation 362 generally operates to modify the virtual scene. For example, the head pose change compensation function 364 can modify the virtual scene based on the user's predicted head pose. As a specific example, the head pose change compensation function 364 can rotate and / or translate the virtual scene so that it is adapted to be presented with the user's predicted head pose. The GDC / CAC function 366 can modify the virtual scene to correct distortions that occur in the displayed image. For example, the GDC / CAC function 366 can adjust the virtual scene so that the resulting image pre-compensates for expected geometric distortions and chromatic aberrations caused by the display lens of the VST XR device. In some cases, the GDC / CAC function 366 can operate based on the display lens GDC and CAC models.

[0075] Scene compositing operation 368 generally operates to produce a final view of the scene for presentation to the user. For example, reality / virtual integration function 370 can be used to combine a reconstructed image of the actual scene with a reconstructed image of the virtual scene to generate a composite image. As a specific example, reality / virtual integration function 370 can overlay a reconstructed image of the virtual scene onto a reconstructed image of the actual scene. Final view generation function 372 can process the composite image and perform any required or desired additional refinement or modification, and the resulting image can represent the final view of the scene. 3D to 2D distortion function 374 can be used to distort the final view of the scene into a 2D image. Display operation 376 generally operates to present a 2D image to the user. For example, final view rendering function 378 can render the 2D image in a form suitable for transmission to at least one display 160. Final view display function 380 can initiate the display of the rendered image, such as by providing the rendered image to one or more displays 160.

[0076] although Figure 3a and Figure 3b A second example of an architecture 300 supporting the dynamic overlay of moving objects with real-world and virtual scenes for VST XR is shown, but it is possible to... Figure 3a and Figure 3b Make various changes. For example, Figure 3a and 3b The various components or functions within can be combined, further subdivided, copied, omitted, or rearranged, and additional components or functions can be added as needed. As a specific example, architecture 300 is shown applying head pose change compensation to both the reconstructed image of the real scene and the reconstructed image of the virtual scene. In other cases, head pose change compensation can be applied to the composite image generated by scene compositing operation 368.

[0077] Figure 4 and Figure 5An example mask generation function 326 for moving objects for VST XR according to this disclosure is shown. Figure 4 As shown, the mask generation function 326 uses a trained skin color model 402 representing a machine learning model trained to recognize user skin within an image frame. In this example, the mask generation function 326 includes or is used in conjunction with a skin color model training operation 404, which generally operates to train the skin color model 402. Note that the skin color model training operation 404 can be performed on the VST XR device, or it can be performed remotely (such as on server 106), and the resulting trained skin color model 402 is deployed to the VST XR device for use.

[0078] The skin color model 402 can be trained using one or more training datasets 406. For example, training dataset 406 may include training images showing where parts of a person's body are visible, and ground truth values ​​identifying which regions of the training images contain visible skin of the person in the training images. During skin color model training operation 404, skin color model 402 can process the training images and generate estimates of where a person's skin is visible, and these estimates can be compared to ground truth values. The difference between the estimates and ground truth values ​​is used to calculate a loss, and if the loss exceeds a threshold, the weights or other parameters 408 of skin color model 402 can be adjusted. The adjusted skin color model 402 can process the training images again, the resulting loss can be compared to the threshold again, and additional adjustments can be made to the weights or other parameters 408 of skin color model 402. This can be repeated any number of times, typically until the loss no longer exceeds the threshold, indicating that skin color model 402 has been trained to generate estimates with at least the desired level of accuracy. If skin color model training operation 404 is performed on a device other than a VST XR device, deployment operation 410 can be used, for example, to send the trained skin color model 402 to a VST XR device via the Internet.

[0079] The mask generation function 326 receives and processes the image frame 412 and the depth map 414. In some cases, the image frame 412 may be provided by the data capture operation 302, and the depth map 414 may represent a dense depth map generated by the depth processing operation 316. The image processing function 416 may perform one or more preprocessing functions involving the image frame 412, such as denoising and image enhancement functions.

[0080] Model application function 418 applies the trained skin color model 402 to the processed image frames to determine whether any image frame in the processed image frames includes one or more parts of the user's body, such as the user's hand. This results in image segmentation 420, which can segment the user's hand or other body parts from other parts in each image frame 412. Moving object reconstruction function 422 can process the image segmentation 420 for each image frame 412 and generate an initial mask for each object detected as part of the user's body (defined by the visible user skin). For example, model application function 418 can generate the probability that pixels or other parts of image frame 412 include different skin colors, and moving object reconstruction function 422 can generate the initial mask based on those probabilities. Image boundary refinement function 424 can use the initial mask to extract and refine the image boundaries of each detected object. For example, image boundary refinement function 424 can smooth the image boundaries of each detected object and treat independent objects that are close enough together as a single object.

[0081] Depth map 414 is provided to moving object depth map extraction function 426, which can extract depth map or other depth data associated with each moving object from depth map 414. For example, moving object depth map extraction function 426 can isolate the depth-related portions of depth map 414 for each moving object. Depth map reconstruction function 428 can process the extracted depth map to reconstruct any missing depth data from the extracted depth map, such as by filling in any holes or other missing data. Depth map boundary refinement function 430 can refine the object boundaries of each detected object in the depth map. For example, depth map boundary refinement function 430 can smooth the depth map boundaries of each detected object and treat independent objects that are sufficiently close together as a single object.

[0082] Image boundaries identified by image boundary refinement function 424 and depth map boundaries identified by depth map boundary refinement function 430 are provided to mask refinement function 432. Mask refinement function 432 processes the image boundaries and depth map boundaries to generate a final object mask 434 for moving objects. For example, mask refinement function 432 can combine image boundaries and depth map boundaries to accurately identify the boundaries of a user's hand or other moving objects within the scene. As described above, the generated object mask 434 can be used for moving object reconstruction and final view generation.

[0083] Figure 5An example illustrating how the skin color model 402 can be trained and used is shown. In this example, the skin color model training operation 404 includes a color normalization function 502, which normalizes the pixel colors (such as red, green, and blue) in one or more training datasets 406. In some cases, this normalization can be represented as follows:

[0084] in:

[0085] Here, R represents the value of each red pixel, G represents the value of each green pixel, and B represents the value of each blue pixel. Furthermore, R... n G represents the value of each normalized red pixel. n Let B represent the value of each normalized green pixel, and B... n This represents the value of each normalized blue pixel. Color vector generation function 504 generates color vectors based on the normalized pixel colors and fits the color vectors to a Gaussian distribution. Each color vector can be defined as X = (R... n B n ), and the Gaussian distribution can be defined as In some cases, the fit of the color vector generation function 504 can be represented as follows:

[0086] here, Represents the mean vector Maximum likelihood estimation, Represents the covariance matrix Maximum likelihood estimation, This represents the estimated Gaussian model of skin color.

[0087] Mask generation function 326 applies a skin color Gaussian model (trained skin color model 402) to image frame 412. According to embodiments of this disclosure, this may involve constructing a skin color filter based on estimated skin color probabilities. In some cases, the estimated skin color probabilities may be expressed as follows:

[0088] here, This represents the maximum likelihood estimate of the trained skin color model 402. By applying the trained skin color model 402 to the color image frame 412 using a defined threshold, the skin color pixels of the image frame 412 can be separated from the rest of the image frame 412. Using the segmented skin color regions, hole filling and edge thinning guided by the image frame 412 can be performed, and an object mask 434 can be generated as described above.

[0089] although Figure 4 and Figure 5 An example of mask generation function 326 for moving objects in VST XR is shown, but it is possible to modify it further. Figure 4 and Figure 5 Various modifications can be made. For example, mask generation function 326 can use any other suitable technique to generate a mask for a moving object. As a specific example, mask generation function 326 can use any suitable machine learning model to separate pixels corresponding to human skin from other parts of the scene.

[0090] Figure 6 An example process 600 for scene reconstruction for VST XR according to this disclosure is shown. More specifically, Figure 6 This demonstrates how various operations of the architecture 300 can be used to reconstruct an image of a scene. For ease of explanation, Figure 6 The processing 600 is described as using Figure 1 The network configuration 100 is implemented using electronic device 101. However, the processing 600 can be implemented using any other suitable device and in any other suitable system.

[0091] like Figure 6 As shown, image frame capture function 602 generally operates to capture or otherwise acquire image frame 604. Depth capture function 606 generally operates to capture or otherwise acquire depth data from one or more depth sensors 180 (such as from one or more LiDAR or ToF depth sensors). Depth estimation function 608 generally operates to estimate the depth within a scene, such as based on a pair of stereo image frames 604. According to embodiments of this disclosure, functions 602, 606, and 608 may represent or be included in data capture operation 302. Depth fusion and integration function 610 generally operates to combine the depth data from depth capture function 606 and depth estimation function 608. Depth hole filling function 612 generally operates to fill the depth of an image frame, such as by interpolation to fill missing depth values ​​or based on the depth of one or more other image frames. According to embodiments of this disclosure, functions 610 and 612 may represent or be included in depth processing operation 316.

[0092] Depth map separation function 614 can divide the resulting dense depth map or other dense depth data into a depth map 616 associated with any moving object and a depth map 618 associated with static scene content. Depth map 616 is provided to object mask generation function 620, which can generate an object mask for each moving object in image frame 604. In some cases, this can be done using a trained skin color model 402. Object mask thinning function 622 can thin the object mask, for example, by smoothing it. Depth map 618 may include holes caused by de-occlusion, meaning that at least one moving object moves and exposes one or more previously occluded portions of the scene. To help compensate for these holes, reprojection function 624 generally operates to reproject at least one previous depth map (such as one or more previous depth maps associated with one or more previous image frames 604) onto the current depth map of the current image frame 604. This reprojection can be used by the depth map hole-filling function 626 to fill the holes in the depth map 618 of the current image frame 604 based on the reprojection of the previous image frame. According to embodiments of this disclosure, functions 614, 620, 622, 624, and 626 may represent or be included in the mask creation operation 322.

[0093] Image frame 604 and the refined object mask generated by object mask refinement function 622 are provided to moving object image reconstruction function 628, which generates an initial image for each moving object. Moving object boundary refinement function 630 processes the initial image and the refined object mask to more accurately define the boundaries of the moving object. Moving object boundary refinement function 630 can thus generate image 632 and depth map 634 of the moving object. According to embodiments of this disclosure, functions 628 and 630 may represent or be included in moving object reconstruction operation 328.

[0094] Image frame 604 and the depth map generated by the depth map hole-filling function 626 are provided to the static scene image reconstruction function 636, which generates an initial image of the static scene content in each image frame 604. The image hole-filling function 638 can fill holes included in the initial image of the static scene content. Similar to the depth map 618, the initial image of the static scene content may include holes caused by de-occlusion, and the image hole-filling function 638 can be used to fill these holes. In some cases, for example, the image hole-filling function 638 may use image data from one or more other image frames 604 (possibly after reprojection) to fill one or more holes in the initial image of the static scene content of the current image frame 604. The static scene boundary refinement function 640 processes the initial image of the static scene content to more accurately define the boundaries of the static scene content. The static scene boundary refinement function 640 can thus generate an image 642 and a depth map 644 of the static scene content. According to embodiments of this disclosure, functions 636, 638, and 640 may represent or be included in the real-world scene reconstruction operation 336.

[0095] although Figure 6 An example of a 600-level scene reconstruction process for VST XR is shown, but it is possible to perform a different process. Figure 6 Make various changes. For example, Figure 6 The various components or functions can be combined, further subdivided, copied, omitted or rearranged, and additional components or functions can be added as needed.

[0096] Figure 7 An example overlay process 700 for VST XR between a moving object and static scene content, according to this disclosure, is shown. More specifically, Figure 7 This illustrates how the real-world scene reconstruction operation 336 of architecture 300 can be used to overlay images of moving objects and static scene content. For ease of explanation, Figure 7 The processing 700 is described as using Figure 1 The network configuration 100 is implemented using electronic device 101. However, the processing 700 can be implemented using any other suitable device and in any other suitable system.

[0097] like Figure 7As shown, as described above, the skin color model training operation 404 can be used to train the skin color model 402. At least one other training operation 704 can also be used to train at least one other machine learning (ML) model 702. According to embodiments of this disclosure, an ML model can refer to an artificial intelligence model. The machine learning model 702 can be trained to detect at least one type of object within an image frame. In this example, the machine learning model 702 can be trained to detect a keyboard in an image frame. The training operation 704 can use one or more training datasets 706 and adjust the weights or other parameters 708 of the machine learning model 702. For example, one or more training datasets 706 can include training images containing and not containing keyboards, and ground truth values ​​identifying which training images contain keyboards and the locations where keyboards are present in the training images. The weights or other parameters 708 of the machine learning model 702 can be adjusted until the machine learning model 702 accurately identifies keyboards in the training images, at least within the desired level of accuracy. The machine learning model 702 can represent any suitable machine learning architecture, such as a deep neural network, that can be trained to recognize keyboards or other objects. Note that the training operation 704 can be performed on the VST XR device, or the training operation 704 can be performed remotely (such as on server 106), and the resulting trained machine learning model 702 is deployed to the VST XR device for use.

[0098] The real-scene reconstruction operation 336 can obtain image frame 710 and depth map 712. In some cases, image frame 710 can be provided by data capture operation 302, and depth map 712 can represent a dense depth map generated by depth processing operation 316. Hand region detection and extraction function 714 can be used to identify the region of the captured user's hand in image frame 710, which can be done using a trained skin color model 402. Hand mask reconstruction and thinning function 716 can be used to generate a thinned mask associated with the user's hand, such as by smoothing any identified hand mask. Hand image edge detection and thinning function 718 can be used to clearly identify the edges or boundaries of the hand mask associated with the user's hand, such as by making the identified hand mask conform to the actual content of image frame 710.

[0099] Keyboard detection and extraction function 720 can be used to identify regions of the captured keyboard in image frame 710, which can be accomplished using a trained machine learning model 702. Keyboard mask reconstruction and thinning function 722 can be used to generate a thinned mask associated with the keyboard, such as by smoothing any identified keyboard mask. Keyboard image edge detection and thinning function 724 can be used to clearly define the edges or boundaries of the keyboard mask associated with the keyboard, such as by making the identified keyboard mask conform to the actual content of image frame 710.

[0100] Image compositing function 726 can process image frame 710, hand mask, and keyboard mask to generate overlay image 728. Overlay image 728 here includes an image of the user's hand detected within image frame 710 and overlaid with the keyboard. In some cases, the keyboard itself can be extracted and reconstructed from image frame 710, for example, using the same techniques described above. According to embodiments of this disclosure, the reconstructed keyboard image can be placed over an image of reconstructed static scene content, and the reconstructed image of the user's hand can be placed over the reconstructed keyboard image.

[0101] although Figure 7 An example of overlay processing 700 between moving objects and static scene content for VST XR is shown, but it is possible to... Figure 7 Make various changes. For example, Figure 7 The various components or functions can be combined, further subdivided, copied, omitted or rearranged, and additional components or functions can be added as needed.

[0102] Figures 8a to 8c An example use case is shown for dynamically overlaying moving objects with real-world and virtual scenes for VST XR according to this disclosure. More specifically, Figures 8a to 8c Examples of various operations that can occur within the aforementioned architectures 200 and 300 are shown, wherein architectures 200 and 300 can be used to generate overlay images of moving objects and static scene content.

[0103] like Figure 8a As shown, a scene 800 imaged by a VST XR device includes a user's hand 802 moving in front of various objects 804 that form part of the static scene content. As part of the VST XR pipeline, architecture 200 or 300 can operate to identify the boundaries 806 of the user's hand 802 and track the boundaries 806 of the user's hand 802 across multiple image frames. Among other functions, this allows architecture 200 or 300 to efficiently overlay an image of the user's hand 802 onto an image of the static scene content. Furthermore, boundary tracking tends to have lower per-pixel computational intensity than tracking the user's hand 802 across image frames.

[0104] like Figure 8bAs shown, scene 808 imaged by the VST XR device includes a user's moving hand 810. At least one perspective camera 812 can be used to generate image frames capturing the user's hand 810, and these image frames can be used to generate an image presented to the user. However, the user's head is also moving, so the user may have a head pose 814a when the image frames are captured, and a head pose 814b when the rendered image based on those image frames is displayed. Note that the change between head pose 814a and head pose 814b is exaggerated for illustration and can be much smaller. The change in head pose can alter the parallax of the user's hand 810. Architecture 200 or 300 can perform parallax correction as described above, such that the rendered image of the user's hand 810 can be presented with the correct parallax at head pose 814b.

[0105] like Figure 8c As shown, scene 816 imaged by a VST XR device includes a user's hand 818 moving in front of various objects 820 that form part of the static scene content. As part of the VST XR pipeline, architecture 200 or 300 can identify when holes 822 are present in the reconstructed image of the static scene content. As described above, holes 822 can be formed when the user's hand 818 moves so that an occluded portion of the static scene content is no longer occluded. Architecture 200 or 300 can perform hole filling to replace holes 822 with actual image data. In some cases, this can be achieved by using image data of the same portion of the static scene content from other image frames (possibly after reprojection or other processing).

[0106] although Figures 8a to 8c An example of a use case for dynamically overlaying moving objects with real-world and virtual scenes in VST XR is shown, but it is possible to... Figures 8a to 8c Various modifications can be made. For example, the specific use case shown here is merely an example, and architectures 200 and 300 can be used to perform any other suitable operations as described above. Furthermore, the imaged scene can vary widely, and Figures 8a to 8c This disclosure is not intended to limit the scope to any particular type of scenario.

[0107] Figure 9 An example method 900 for dynamically overlaying moving objects with real-world and virtual scenes for VST XR according to this disclosure is shown. For ease of explanation, Figure 9 Method 900 is described as using Figure 1 The network configuration 100 can achieve Figure 2The method 900 can be implemented using the electronic device 101 of architecture 200 or architecture 300 of Figure 3. However, the method 900 can be implemented using any other suitable device and in any other suitable system, and the method 900 can be executed using any other suitable architecture.

[0108] like Figure 9 As shown, in step 902, image frames and depth data are acquired using the VST XR device. This may include, for example, the processor 120 of the electronic device 101 acquiring multiple image frames captured using one or more imaging sensors 180 of the VST XR device. The image frames may capture one or more moving objects and static scene content within a scene, and the moving objects may include at least a portion of the user's body. In some cases, the image frames may be preprocessed, such as by denoising and image enhancement. This may also include the processor 120 of the electronic device 101 receiving, generating, or otherwise acquiring depth data associated with the image frames. In step 904, a depth map may be generated using the depth data. This may include, for example, the processor 120 of the electronic device 101 performing depth densification to generate a dense depth map associated with the captured image frames.

[0109] In step 906, masks are generated for one or more moving objects in the scene. This may include, for example, the processor 120 of electronic device 101 generating one or more binary masks for the moving objects in each image frame. At least some of the masks may be generated using a machine learning model trained to separate pixels corresponding to human skin from other parts of the scene (such as a trained skin color model 402). In step 908, images of the moving objects and images of the static scene content are reconstructed. This may include, for example, the processor 120 of electronic device 101 generating the reconstructed images of the moving objects based on image frames, depth maps or other depth data, and the masks. This may also include the processor 120 of electronic device 101 generating images of the reconstructed static scene content based on image frames and depth maps or other depth data. A portion of this step may include modifying the images to correct for parallax. A portion of this step may also include performing head pose change compensation based on estimated latency from the VST XR pipeline.

[0110] In step 910, the reconstructed image of the moving object and the reconstructed image of the static scene content are combined. This may include, for example, the processor 120 of electronic device 101 overlaying the image of the reconstructed moving object onto the reconstructed static scene content. It may also include the processor 120 of electronic device 101 overlaying one or more virtual features (such as one or more virtual objects or other virtual content) onto the images of the reconstructed moving object and the reconstructed static scene content. This can result in the generation of a combined image. A portion of this step may include performing head pose change compensation based on estimated latency of the VST XR pipeline. In step 912, the combined image is rendered, and in step 914, the display of the resulting rendered image begins. This may include, for example, the processor 120 of electronic device 101 rendering the combined image and displaying the rendered image on at least one display 160 of the VST XR device.

[0111] although Figure 9 An example of method 900 for dynamically overlaying moving objects with real and virtual scenes in VST XR is shown, but it is possible to modify... Figure 9 Various changes were made. For example, although it was shown as a series of steps, Figure 9 The steps in the process can overlap, occur in parallel, occur in different orders, or occur any number of times (including zero times).

[0112] It should be noted that, Figures 2 to 9 Shown or about Figures 2 to 9 The described functionality can be implemented in any suitable manner in electronic device 101, electronic device 102, electronic device 104, server 106, or other devices. For example, according to embodiments of this disclosure, one or more software applications or other software instructions executed by the processor 120 of electronic device 101, electronic device 102, electronic device 104, server 106, or other devices can be used to implement or support the functionality. Figures 2 to 9 Shown or about Figures 2 to 9 At least some of the functions described. According to embodiments of this disclosure, dedicated hardware components can be used to implement or support [the functions]. Figures 2 to 9 Shown or about Figures 2 to 9 At least some of the functions described. Typically, this can be performed using any suitable hardware or any suitable combination of hardware and software / firmware instructions. Figures 2 to 9 Shown or about Figures 2 to 9 The described function. Furthermore, in Figures 2 to 9 Shown or about Figures 2 to 9 The described functions can be performed by a single device or multiple devices.

[0113] Although this disclosure has been described with reference to exemplary embodiments, various changes and modifications may be suggested to those skilled in the art. It is intended that this disclosure cover such changes and modifications that fall within the scope of the appended claims.

[0114] According to embodiments of this disclosure, a method performed by a Video See-Through (VST) Extended Reality (XR) device may include obtaining image frames of a scene captured using at least one imaging sensor and depth data associated with the scene.

[0115] According to embodiments of this disclosure, the method may include using an artificial intelligence model trained to separate pixels of the scene corresponding to human skin from other parts to generate a mask associated with a moving object.

[0116] According to embodiments of this disclosure, the method may include reconstructing an image of the moving object based on the image frame, the depth data, and the mask.

[0117] According to embodiments of this disclosure, the method may include reconstructing an image of the static scene content based on the image frame and the depth data.

[0118] According to embodiments of this disclosure, the method may include combining an image of the moving object, an image of the static scene content, and at least one virtual feature.

[0119] According to embodiments of this disclosure, the method may include rendering the combined image on at least one display.

[0120] According to embodiments of this disclosure, moving objects and static scene content are captured in the image frame.

[0121] According to embodiments of this disclosure, the moving object includes a part of the user's body.

[0122] According to embodiments of this disclosure, the method may include generating a first depth map associated with the moving object based on the depth data.

[0123] According to embodiments of this disclosure, the method may include generating a second depth map associated with the static scene content based on the depth data.

[0124] According to embodiments of this disclosure, the mask is generated based on a first depth map.

[0125] According to embodiments of this disclosure, the image of the moving object is reconstructed based on the image frame, the first depth map, and the mask.

[0126] According to embodiments of this disclosure, the image of the static scene content is reconstructed based on the image frame and the second depth map.

[0127] According to embodiments of this disclosure, the method may include determining the depth of each portion of the scene behind the moving object based on a previous depth map corresponding to a previous image frame.

[0128] According to embodiments of this disclosure, the method may include adjusting the parallax associated with the moving object in an image of the moving object and the parallax associated with the static scene content in an image of the static scene content.

[0129] According to embodiments of this disclosure, the method may include: overlaying an image of the moving object onto an image of the static scene content based on an estimated boundary of the moving object within the scene.

[0130] According to embodiments of this disclosure, the method may include performing hole filling to generate image content for each portion of the scene behind the moving object.

[0131] According to embodiments of this disclosure, the method may include modifying at least one of the image of the moving object, the image of the static scene content, or the combined image based on the user's estimated head pose.

[0132] According to embodiments of this disclosure, a video see-through (VST) extended reality (XR) device may include at least one imaging sensor, a memory, at least one display, and at least one processor communicatively coupled to the memory.

[0133] According to embodiments of this disclosure, the at least one processor executes a program or at least one instruction stored in the memory to cause the VST XR device to obtain image frames of a scene captured using the at least one imaging sensor and depth data associated with the scene.

[0134] According to embodiments of this disclosure, the at least one processor executes the program or at least one instruction stored in the memory to cause the VST XR device to generate a mask associated with a moving object using an artificial intelligence model trained to separate pixels of the scene corresponding to human skin from other parts.

[0135] According to embodiments of this disclosure, the at least one processor executes the program or at least one instruction stored in the memory to cause the VST XR device to reconstruct an image of the moving object based on the image frame, the depth data, and the mask.

[0136] According to embodiments of this disclosure, the at least one processor executes the program or at least one instruction stored in the memory to cause the VST XR device to reconstruct an image of the static scene content based on the image frame and the depth data.

[0137] According to embodiments of this disclosure, the at least one processor executes the program or at least one instruction stored in the memory to cause the VST XR device to combine an image of the moving object, an image of the static scene content, and at least one virtual feature.

[0138] According to embodiments of this disclosure, the at least one processor executes the program or at least one instruction stored in the memory to cause the VST XR device to render a combined image on at least one display.

[0139] According to embodiments of this disclosure, moving objects and static scene content are captured in the image frame. According to embodiments of this disclosure, the moving object includes a portion of the user's body. According to embodiments of this disclosure, the at least one processor executes the program or at least one instruction stored in the memory to cause the VST XR device to generate a first depth map associated with the moving object based on the depth data.

[0140] According to embodiments of this disclosure, the at least one processor executes the program or at least one instruction stored in the memory to cause the VST XR device to generate a second depth map associated with the static scene content based on the depth data.

[0141] According to embodiments of this disclosure, the mask is generated based on a first depth map. According to embodiments of this disclosure, an image of the moving object is reconstructed based on the image frame, the first depth map, and the mask. According to embodiments of this disclosure, an image of the static scene content is reconstructed based on the image frame and a second depth map. According to embodiments of this disclosure, the at least one processor executes the program or at least one instruction stored in the memory to cause the VST XR device to determine the depth of each portion of the scene behind the moving object based on a previous depth map corresponding to a previous image frame.

[0142] According to embodiments of this disclosure, the at least one processor executes the program or at least one instruction stored in the memory to cause the VST XR device to adjust the parallax associated with the moving object in the image of the moving object and the parallax associated with the static scene content in the image of the static scene content.

[0143] According to embodiments of this disclosure, the at least one processor executes the program or at least one instruction stored in the memory to cause the VST XR device to overlay an image of the moving object with an image of the static scene content based on an estimated boundary of the moving object within the scene.

[0144] According to embodiments of this disclosure, the at least one processor executes the program or at least one instruction stored in the memory to cause the VST XR device to perform hole filling to generate image content for each portion of the scene behind the moving object.

[0145] According to embodiments of this disclosure, the at least one processor executes the program or at least one instruction stored in the memory to cause the VST XR device to modify at least one of the image of the moving object, the image of the static scene content, or the combined image based on the user's estimated head pose.

Claims

1. A method performed by a Video Perspective (VST) Extended Reality (XR) device, comprising: Obtain image frames of the scene captured using at least one imaging sensor and depth data associated with the scene; An artificial intelligence model is used to generate a mask associated with a moving object, the model being trained to separate pixels in the scene that correspond to human skin from other parts. The image of the moving object is reconstructed based on the image frame, the depth data, and the mask; Reconstruct an image of the static scene content based on the image frame and the depth data; Combine the image of the moving object, the image of the static scene content, and at least one virtual feature; as well as Render the combined image on at least one monitor; Moving objects and static scene content are captured in the image frame, and The moving object includes a part of the user's body.

2. The method according to claim 1, further comprising: A first depth map associated with the moving object is generated based on the depth data; as well as A second depth map associated with the static scene content is generated based on the depth data; The mask is generated based on a first depth map. Wherein, the image of the moving object is reconstructed based on the image frame, the first depth map, and the mask, and The image of the static scene content is reconstructed based on the image frame and the second depth map.

3. The method according to claim 2, wherein, Generating the second depth map includes: The depth of each portion of the scene behind the moving object is determined based on a previous depth map corresponding to a previous image frame.

4. The method according to any one of the preceding claims further includes: Adjust the parallax associated with the moving object in the image of the moving object and the parallax associated with the static scene content in the image of the static scene content, respectively.

5. The method according to any one of the preceding claims, wherein, The combination of the image of the moving object, the image of the static scene content, and at least one virtual feature includes: Based on the estimated boundary of the moving object within the scene, the image of the moving object is overlaid with the image of the static scene content.

6. The method according to any one of the preceding claims, wherein, The combination of the image of the moving object, the image of the static scene content, and at least one virtual feature further includes: Perform hole filling to generate image content for each part of the scene behind the moving object.

7. The method according to any one of the preceding claims further comprises: The image of the moving object, the image of the static scene content, or the combined image is modified based on the user's estimated head pose.

8. A video perspective (VST) extended reality (XR) device, comprising: At least one imaging sensor; Memory; At least one display; as well as At least one processor is communicatively coupled to the memory, wherein the at least one processor executes a program or at least one instruction stored in the memory to cause the VST XR device to: Obtain image frames of the scene captured using the at least one imaging sensor and depth data associated with the scene; An artificial intelligence model is used to generate a mask associated with a moving object, the model being trained to separate pixels in the scene that correspond to human skin from other parts. The image of the moving object is reconstructed based on the image frame, the depth data, and the mask; Reconstruct an image of the static scene content based on the image frame and the depth data; Combining the image of the moving object, the image of the static scene content, and at least one virtual feature; and Render the combined image on the at least one display; Moving objects and static scene content are captured in the image frame, and The moving object includes a part of the user's body.

9. The VST XR device according to claim 8, wherein, The at least one processor executes the program or at least one instruction stored in the memory to cause the VST XR device to: A first depth map associated with the moving object is generated based on the depth data; as well as A second depth map associated with the static scene content is generated based on the depth data; The mask is generated based on a first depth map; Wherein, the image of the moving object is reconstructed based on the image frame, the first depth map, and the mask; and The image of the static scene content is reconstructed based on the image frame and the second depth map.

10. The VST XR device according to claim 9, wherein, The at least one processor executes the program or at least one instruction stored in the memory to cause the VST XR device to: The depth of each portion of the scene behind the moving object is determined based on a previous depth map corresponding to a previous image frame.

11. The VST XR device according to any one of the preceding claims, wherein, The at least one processor executes the program or at least one instruction stored in the memory to cause the VST XR device to: Adjust the parallax associated with the moving object in the image of the moving object and the parallax associated with the static scene content in the image of the static scene content, respectively.

12. The VST XR device according to any one of the preceding claims, wherein, The at least one processor executes the program or at least one instruction stored in the memory to cause the VST XR device to: Based on the estimated boundary of the moving object within the scene, the image of the moving object is overlaid with the image of the static scene content.

13. The VST XR device according to any one of the preceding claims, wherein, The at least one processor executes the program or at least one instruction stored in the memory to cause the VST XR device to: Perform hole filling to generate image content for each part of the scene behind the moving object.

14. The VST XR device according to any one of the preceding claims, wherein, The at least one processor executes the program or at least one instruction stored in the memory to cause the VST XR device to: The image of the moving object, the image of the static scene content, or the combined image is modified based on the user's estimated head pose.

15. A computer-readable medium comprising at least one instruction, which, when executed, causes at least one processor of the apparatus to perform an operation corresponding to the method of any one of claims 1 to 7.