Mask generation with object and scene segmentation for transmissive augmented reality (XR)
By training the object segmentation model to generate object segmentation prediction and depth or disparity map, the problem of inaccurate segmentation mask in VST XR system is solved, more accurate object and scene segmentation is achieved, and the accuracy of digital content superposition is improved.
Patent Information
- Application Number
- CN202380088169.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-07-27
- Filing Date
- 2023-10-31
- Publication Date
- 2025-07-29
AI Technical Summary
Existing segmentation algorithms produce poor quality segmentation masks in penetrating extended reality (VST XR) systems, resulting in inaccurate superposition of digital content, which may cover parts that should not be obscured or are difficult to properly superimpose within the scene.
By training the object segmentation model, the object segmentation prediction and depth or disparity map are generated using the first and second image frames, the object boundaries are determined based on these predictions, and a virtual view is generated for presentation on the XR device.
More accurate object and scene segmentation is achieved, a complete segmentation mask is generated, and the overlay accuracy of digital content in the VST XR system is improved.
Smart Images

Figure CN120390941A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure generally relates to extended reality (XR) systems and processes. More specifically, the present disclosure relates to mask generation for penetrative XR utilizing object and scene segmentation. Background Art
[0002] Over time, extended reality (XR) systems have become increasingly popular, and many applications have been developed for XR systems. Some XR systems (e.g., augmented reality or "AR" systems and mixed reality or "MR" systems) can enhance a user's view of his or her current environment by overlaying digital content (e.g., information or virtual objects) on the user's view of the current environment. For example, some XR systems can typically seamlessly blend virtual objects generated by computer graphics with real-world scenes. Summary of the Invention
[0003] Technical Solution
[0004] The present disclosure relates to mask generation for penetrative extended reality (XR) utilizing object and scene segmentation.
[0005] In a first embodiment, a method includes: obtaining a first image frame and a second image frame of a scene. The method further includes: providing the first image frame as an input to an object segmentation model, wherein the object segmentation model is trained to generate a first object segmentation prediction for an object in the scene and a depth or disparity map based on the first image frame. The method further includes: generating a second object segmentation prediction for an object in the scene based on the second image frame. The method further includes: determining a boundary of an object in the scene based on the first object segmentation prediction and the second object segmentation prediction. Additionally, the method includes: generating a virtual view for presentation on a display of an XR device based on the boundary of the object in the scene. In a related embodiment, a non-transitory machine-readable medium stores instructions that, when executed, cause at least one processor to perform the method of the first embodiment.
[0006] In a second embodiment, an XR device includes a plurality of imaging sensors configured to capture a first image frame and a second image frame of a scene. The XR device further includes at least one processing device configured to provide the first image frame as an input to an object segmentation model, wherein the object segmentation model is trained to generate a first object segmentation prediction for an object in the scene and a depth or disparity map based on the first image frame. The at least one processing device is further configured to generate a second object segmentation prediction for an object in the scene based on the second image frame. The at least one processing device is further configured to determine a boundary of an object in the scene based on the first object segmentation prediction and the second object segmentation prediction. Additionally, the at least one processing device is configured to generate a virtual view based on the boundary of the object in the scene. The XR device further includes at least one display configured to present the virtual view.
[0007] In a third embodiment, a method includes: obtaining a first training image frame and a second training image frame of a scene and extracting features of the first training image frame. The method further includes: providing the extracted features of the first training image frame as an input to an object segmentation model being trained, wherein the object segmentation model is configured to generate an object segmentation prediction for an object in the scene and a depth or disparity map. The method further includes: reconstructing the first training image frame based on the depth or disparity map and the second training image frame. Further, the method includes: updating the object segmentation model based on the first training image frame and the reconstructed first training image frame. In a related embodiment, an electronic device includes at least one processing device configured to perform the method of the third embodiment. In another related embodiment, a non-transitory machine-readable medium stores instructions that, when executed, cause at least one processor to perform the method of the third embodiment.
[0008] Based on the following drawings, description, and claims, other technical features may be apparent to those skilled in the art.
[0009] Before presenting the following detailed description, it may be advantageous to set forth definitions of certain words and phrases used throughout this patent document. The terms “send,” “receive,” and “communicate,” and derivatives thereof, include both direct and indirect communication. The terms “include” and “comprise,” and derivatives thereof, mean including but not limited to. The term “or” is inclusive, meaning and / or. The phrase “associated with,” and derivatives thereof, means including, included within, interconnected with, containing, contained within, connected to or coupled with, capable of communicating with, cooperating with, interlacing, juxtaposing, proximate to, bound to or bound with, having, having the attribute of, having a relationship or relationship with, etc.
[0010] In addition, the various functions described below can be implemented or supported by one or more computer programs, each of which is formed of computer-readable program code and implemented in a computer-readable medium. The terms "application" and "program" refer to one or more computer programs, software components, instruction sets, procedures, functions, objects, classes, instances, related data, or portions thereof that are suitable for implementation in appropriate computer-readable program code. The phrase "computer-readable program code" includes any type of computer code, including source code, object code, and executable code. The phrase "computer-readable medium" includes any type of medium that can be accessed by a computer (e.g., read-only memory (ROM), random access memory (RAM), hard disk drive, optical disk (CD), digital video disk (DVD), or any other type of memory). A "non-transitory" computer-readable medium does not include wired, wireless, optical, or other communication links that transmit transitory electrical signals or other signals. Non-transitory computer-readable media include media that can permanently store data, as well as media that can store data and then overwrite the data (e.g., rewritable optical disks or erasable storage devices).
[0011] As used herein, terms and phrases such as "has", "may have", "includes", or "may include" a feature (such as a number, a function, an operation, or a component such as a part) indicate the presence of the feature and do not exclude the presence of other features. In addition, as used herein, the phrases "A or B", "at least one of A and / or B", or "one or more of A and / or B" may include all possible combinations of A and B. For example, "A or B", "at least one of A and B", and "at least one of A or B" may each indicate any of the following: (1) including at least one A, (2) including at least one B, or (3) including at least one A and at least one B. In addition, as used herein, the terms "first" and "second" may modify various components regardless of importance and do not limit the components. These terms are only used to distinguish one component from another. For example, a first user device and a second user device may indicate different user devices from each other, regardless of the importance or order of the devices. Without departing from the scope of the present disclosure, the first component may be represented as the second component, and vice versa.
[0012] It will be understood that when an element (e.g., a first element) is referred to as being “coupled” / “coupled to” / “connected” / “connected to” another element (e.g., a second element) (operatively or communicatively), it can be coupled or connected / coupled or connected to the other element directly or via a third element. In contrast, it will be understood that when an element (e.g., a first element) is referred to as being “directly coupled” / “directly coupled to” another element (e.g., a second element) or “directly connected” / “directly connected to” another element, there is no other element (e.g., a third element) intervening between the element and the other element.
[0013] As used herein, the phrase “configured (or set) to” may be used interchangeably with the phrases “suitable for,” “capable of,” “designed to,” “adapted to,” “made to,” or “able to,” as the context may require. The phrase “configured (or set) to” essentially means “specially designed in hardware to.” In contrast, the phrase “configured to” may mean that a device can perform an operation in conjunction with another device or component. For example, the phrase “a processor configured (or set) to perform A, B, and C” may mean a general-purpose processor (e.g., a CPU or an application processor) that can perform the operations by executing one or more software programs stored in a storage device, or a dedicated processor (e.g., an embedded processor) for performing the operations.
[0014] The terms and phrases used herein are provided only to describe some embodiments of the present disclosure and are not intended to limit the scope of other embodiments of the present disclosure. It should be understood that, unless the context clearly indicates otherwise, the singular forms “a,” “an,” and “the” include plural references. All terms and phrases used herein (including technical and scientific terms and phrases) have the same meaning as that commonly understood by one of ordinary skill in the art to which the embodiments of the present disclosure pertain. It will be further understood that terms and phrases (e.g., those defined in a commonly used dictionary) should be interpreted as having a meaning consistent with their meaning in the context of the relevant art and will not be interpreted in an idealized or overly formal sense unless expressly so defined herein. In some cases, the terms and phrases defined herein may be interpreted as excluding embodiments of the present disclosure.
[0015] According to embodiments of the present disclosure, examples of an "electronic device" may include at least one of the following: a smart phone, a tablet personal computer (PC), a mobile phone, a video phone, an e-book reader, a desktop PC, a laptop computer, a netbook computer, a workstation, a personal digital assistant (PDA), a portable multimedia player (PMP), an MP3 player, a mobile medical device, a camera, or a wearable device (e.g., smart glasses, a head-mounted device (HMD), electronic clothing, an electronic bracelet, an electronic necklace, electronic accessories, an electronic tattoo, a smart mirror, or a smart watch). Other examples of the electronic device include smart home appliances. Examples of the smart home appliances may include at least one of the following: a television, a digital video disc (DVD) player, an audio player, a refrigerator, an air conditioner, a cleaner, an oven, a microwave oven, a washing machine, a dryer, an air purifier, a set-top box, a home automation control panel, a security control panel, a TV box (e.g., SAMSUNG HOMESYNC, APPLE TV, or GOOGLE TV), a smart speaker, or a speaker with an integrated digital assistant (e.g., SAMSUNG GALAXY HOME, APPLE HOMEPOD, or AMAZON ECHO), a game console (e.g., XBOX, PLAYSTATION, or NINTENDO), an electronic dictionary, an electronic key, a video camera, or an electronic photo frame. Other examples of the electronic device include at least one of the following: various medical devices (e.g., various portable medical measurement devices such as a blood glucose measurement device, a heartbeat measurement device, or a body temperature measurement device, a magnetic source angiography (MRA) device, a magnetic source imaging (MRI) device, a computed tomography (CT) device, an imaging device, or an ultrasonic device), a navigation device, a global positioning system (GPS) receiver, an event data recorder (EDR), a flight data recorder (FDR), an in-vehicle infotainment device, marine electronic devices (e.g., marine navigation devices or gyrocompasses), avionics, security devices, vehicle audio units, industrial or household robots, an automated teller machine (ATM), a point of sale (POS) device, or an Internet of Things (IoT) device (e.g., a light bulb, various sensors, a water meter, a gas meter, a sprinkler, a fire alarm, a thermostat, a street lamp, a toaster, a fitness device, a hot water tank, a heater, or a kettle). Other examples of the electronic device include at least a portion of the following: furniture or a building / structure, an electronic board, an electronic signature receiving device, a projector, or various measurement devices (e.g., devices for measuring water, electricity, gas, or electromagnetic waves). Note that, according to various embodiments of the present disclosure, the electronic device may be one or a combination of the devices listed above. According to some embodiments of the present disclosure, the electronic device may be a flexible electronic device. The electronic device disclosed herein is not limited to the devices listed above and may include any other electronic device known now or developed later.
[0016] In the following description, electronic devices are described with reference to the accompanying drawings in accordance with various embodiments of the present disclosure. As used herein, the term "user" may refer to a person using the electronic device or another device (e.g., an artificial intelligence electronic device).
[0017] Definitions may be provided for certain words and phrases throughout this patent document. Those of ordinary skill in the art should understand that in many, if not most, instances, such definitions apply to the prior as well as future use of such defined words and phrases.
[0018] The description in this application should not be construed as implying that any particular element, step, or function is an essential element that must be included in the scope of the claims. The scope of the patent subject matter is defined only by the claims. Additionally, the claims are not intended to invoke 35 U.S.C. § 112(f) unless the exact phrase "means for" is followed by a participle. The applicant understands that the use of any other term (including but not limited to "mechanism", "module", "device", "unit", "component", "element", "member", "apparatus", "machine", "system", "processor", or "controller") within the claims refers to a structure known to those of ordinary skill in the relevant art and is not intended to invoke 35 U.S.C. § 112(f). BRIEF DESCRIPTION OF THE DRAWINGS
[0019] To more fully understand the present disclosure and its advantages, reference is now made to the following description taken in conjunction with the accompanying drawings, in which:
[0020] Figure 1 An example network configuration including an electronic device in accordance with the present disclosure is shown;
[0021] Figure 2 An example architecture for training an object segmentation model to support mask generation for object and scene segmentation for see-through extended reality (XR) in accordance with the present disclosure is shown;
[0022] Figure 3 An example process for determining a reconstruction loss within the architecture of Figure 2 in accordance with the present disclosure is shown;
[0023] Figure 4 An example architecture for using an object segmentation model that supports mask generation for object and scene segmentation for see-through XR in accordance with the present disclosure is shown;
[0024] Figure 5 and Figure 6 An example process for boundary refinement and virtual view generation within the architecture of Figure 4 in accordance with the present disclosure is shown;
[0025] Figure 7 Shows an example relationship between segmentation results associated with a stereoscopic image pair according to the present disclosure;
[0026] Figure 8 Shows an example image-guided segmentation reconstruction using stereo consistency according to the present disclosure;
[0027] Figures 9 to 11 Shows an example result obtainable using an object segmentation model according to the present disclosure, the object segmentation model supporting mask generation for object and scene segmentation for see-through XR;
[0028] Figures 12A to 13D Shows other example results obtainable using an object segmentation model according to the present disclosure, the object segmentation model supporting mask generation for object and scene segmentation for see-through XR;
[0029] Figure 14 Shows an example method for training an object segmentation model to support mask generation for object and scene segmentation for see-through XR according to the present disclosure; and
[0030] Figure 15 Shows an example method for using an object segmentation model according to the present disclosure, the object segmentation model supporting mask generation for object and scene segmentation for see-through XR. DETAILED DESCRIPTION
[0031] The following discussion is described with reference to the accompanying drawings Figures 1 to 15 and various embodiments of the present disclosure. However, it should be understood that the present disclosure is not limited to these embodiments, and all changes and / or equivalents or substitutions thereto also fall within the scope of the present disclosure. Throughout the specification and the drawings, the same or similar reference numerals may be used to refer to the same or similar elements.
[0032] As mentioned above, over time, extended reality (XR) systems have become increasingly popular, and many applications have been developed for XR systems. Some XR systems (e.g., augmented reality or "AR" systems and mixed reality or "MR" systems) can enhance a user's view of his or her current environment by overlaying digital content (e.g., information or virtual objects) on the user's view of the current environment. For example, some XR systems can typically seamlessly blend virtual objects generated by computer graphics with real-world scenes.
[0033] Optical See-Through (OST) XR systems typically allow a user to directly view his or her environment, where light from the user's environment is passed to the user's eyes. One or more panels or other structures through which the light from the user's environment passes can be used to overlay digital content onto the user's view of the environment. In contrast, Video See-Through (VST) XR systems (also referred to as "passthrough" XR systems) typically use perspective cameras to capture images of the user's environment. Digital content can be blended with the captured images and the blended images can be displayed to the user for viewing. Both methods can provide the user with a full-scale contextual extended reality experience.
[0034] Mask generation is typically a useful or desired function performed by VST XR systems (e.g., for supporting scene reconstruction and recognition). For example, when blending digital objects or other digital content with an image of a captured scene, it can be useful or desired to avoid overlaying the digital content on one or more types of objects within the scene, or to overlay the digital content on one or more types of objects within the scene. Unfortunately, current segmentation algorithms typically produce segmentation masks of poor quality. A segmentation mask typically represents a mask in which pixel values indicate which pixels in one or more image frames are associated with different objects within the scene. For example, in an image of an indoor scene, the segmentation mask can identify which pixels in the image are associated with windows, furniture, plants, walls, floors, or other types of objects. Current segmentation algorithms may produce segmentation masks with incomplete or inaccurate boundaries of objects or other artifacts. In addition to this, these artifacts can cause the VST XR system to overlay digital content on parts of an object that should not be occluded, or make it more difficult for the VST XR system to appropriately overlay digital content within an image of the scene.
[0035] The present disclosure provides various techniques for mask generation for passthrough XR that utilize object and scene segmentation. As described in more detail below, during the training of an object segmentation model (a machine learning model), a first training image frame and a second training image frame of a scene can be obtained, and features of the first training image frame can be extracted. The extracted features of the first training image frame can be provided as input to the object segmentation model being trained, and the object segmentation model can be configured to generate object segmentation predictions for objects in the scene as well as a depth or disparity map. The first training image frame can be reconstructed based on the depth or disparity map and the second training image frame. The object segmentation model can be updated based on the first training image frame and the reconstructed first training image frame.
[0036] During the use of a trained object segmentation model, a first image frame and a second image frame of a scene can be obtained. The first image frame can be provided as input to the object segmentation model, and the object segmentation model has been trained to generate a first object segmentation prediction for objects in the scene and a depth or disparity map based on the first image frame. A second object segmentation prediction for objects in the scene can be generated based on the second image frame, and the boundaries of the objects in the scene can be determined based on the first object segmentation prediction and the second object segmentation prediction. A virtual view for presentation on a display of the XR device can be generated based on the boundaries of the objects in the scene.
[0037] As described below, these techniques can support the generation of segmentation masks based on panoptic segmentation, which refers to a combined task of semantic segmentation and instance segmentation. In other words, these techniques can allow the generation of a segmentation mask that identifies both (i) the objects within a scene and (ii) the types of the objects identified within the scene. These techniques can use efficient algorithms to perform object and scene segmentation. In some cases, depth information can be applied to the segmentation process without the need for depth ground truth data for machine learning model training, which can provide convenience and result in more accurate segmentation. These techniques can also use efficient methods for boundary refinement in order to generate complete object and scene regions. These complete object and scene regions can be used to generate more accurate segmentation masks. Thus, these techniques provide an efficient method for segmenting objects and scenes captured using a perspective camera of an XR device, and depth information can be applied in order to achieve better segmentation results, perform boundary refinement, and make the segmented regions complete and enhanced.
[0038] Figure 1 An example network configuration 100 including an electronic device is shown in accordance with the present disclosure. Figure 1 The illustrated embodiments of the network configuration 100 are for illustrative purposes only. Other embodiments of the network configuration 100 may be used without departing from the scope of the present disclosure.
[0039] According to an embodiment of the present disclosure, the electronic device 101 is included in the network configuration 100. The electronic device 101 may include at least one of a bus 110, a processor 120, a memory 130, an input / output (I / O) interface 150, a display 160, a communication interface 170, and a sensor 180. In some embodiments, the electronic device 101 may not include at least one of these components, or may add at least one other component. The bus 110 includes circuitry for connecting the components 120 to 180 to each other and for transmitting communications (e.g., control messages and / or data) between the components.
[0040] The processor 120 includes one or more processing devices (e.g., one or more microprocessors, microcontrollers, digital signal processors (DSPs), application specific integrated circuits (ASICs), or field programmable gate arrays (FPGAs)). In some embodiments, the processor 120 includes one or more of a central processing unit (CPU), an application processor (AP), a communication processor (CP), a graphics processing unit (GPU), or a neural processing unit (NPU). The processor 120 is capable of controlling at least one of the other components of the electronic device 101, and / or performing operations or data processing related to communication or other functions. As described below, the processor 120 may perform one or more functions related to mask generation for object and scene segmentation for see-through XR.
[0041] The memory 130 may include volatile and / or non-volatile memory. For example, the memory 130 may store commands or data related to at least one of the other components of the electronic device 101. According to an embodiment of the present disclosure, the memory 130 may store software and / or a program 140. The program 140 includes, for example, a kernel 141, middleware 143, an application programming interface (API) 145, and / or an application (or “app”) 147. At least a part of the kernel 141, middleware 143, or API 145 may be represented as an operating system (OS).
[0042] The kernel 141 may control or manage system resources (e.g., the bus 110, the processor 120, or the memory 130) for performing operations or functions implemented in other programs (e.g., the middleware 143, the API 145, or the app 147). The kernel 141 provides an interface that allows the middleware 143, the API 145, or the app 147 to access the respective components of the electronic device 101 to control or manage system resources. In addition to this, the app 147 may include one or more apps that perform mask generation for object and scene segmentation for see-through XR. These functions may be performed by a single app or multiple apps, with each app performing one or more of these functions. For example, the middleware 143 may act as a relay to allow the API 145 or the app 147 to pass data to and from the kernel 141. Multiple apps 147 may be provided. The middleware 143 is capable of controlling work requests received from the app 147, for example, by assigning priorities for using system resources of the electronic device 101, such as the bus 110, the processor 120, or the memory 130, to at least one of the multiple apps 147. The API 145 is an interface that allows the app 147 to control functions provided from the kernel 141 or the middleware 143. For example, the API 145 includes at least one interface or function (e.g., a command) for file control, window control, image processing, or text control.
[0043] The I / O interface 150 serves as an interface that can transfer, for example, commands or data input from a user or other external devices to other components of the electronic device 101. The I / O interface 150 can also output commands or data received from other components of the electronic device 101 to the user or other external devices.
[0044] The display 160 includes, for example, a liquid crystal display (LCD), a light-emitting diode (LED) display, an organic light-emitting diode (OLED) display, a quantum dot light-emitting diode (QLED) display, a microelectromechanical systems (MEMS) display, or an electronic paper display. The display 160 can also be a depth perception display (e.g., a multi-focus display). The display 160 is capable of displaying various contents (e.g., text, images, videos, icons, or symbols) to the user. The display 160 can include a touch screen and can receive, for example, touches, gestures, proximity, or hovering input using an electronic pen or a user's body part.
[0045] The communication interface 170 can, for example, establish communication between the electronic device 101 and an external electronic device (e.g., the first electronic device 102, the second electronic device 104, or the server 106). For example, the communication interface 170 can be connected to the network 162 or 164 through wireless or wired communication to communicate with the external electronic devices 102 or 104. The communication interface 170 can be a wired or wireless transceiver, or any other component for transmitting and receiving signals.
[0046] Wireless communication can use at least one of the following as a communication protocol: for example, WiFi, Long-Term Evolution (LTE), Advanced Long-Term Evolution (LTE-A), Fifth Generation Wireless System (5G), millimeter wave or 60 GHz wireless communication, Wireless USB, Code Division Multiple Access (CDMA), Wideband Code Division Multiple Access (WCDMA), Universal Mobile Telecommunications System (UMTS), Wireless Broadband (WiBro), or Global System for Mobile Communications (GSM). Wired connections can include at least one of the following: for example, Universal Serial Bus (USB), High-Definition Multimedia Interface (HDMI), Recommended Standard 232 (RS-232), or Plain Old Telephone Service (POTS). The network 162 or 164 includes at least one communication network (e.g., a computer network (such as a local area network (LAN) or a wide area network (WAN)), the Internet, or a telephone network).
[0047] The electronic device 101 also includes one or more sensors 180 that can measure a physical quantity or detect an activation state of the electronic device 101 and convert the measured or detected information into an electrical signal. For example, the sensor 180 includes one or more cameras, or other imaging sensors that can be used to capture an image of a scene. The sensor 180 may also include one or more buttons for touch input, one or more microphones, a depth sensor, a gesture sensor, a gyroscope or gyroscopic sensor, a barometric pressure sensor, a magnetic sensor or magnetometer, an acceleration sensor or accelerometer, a grip sensor, a proximity sensor, a color sensor (e.g., a red, green, blue (RGB) sensor), a bio-physical sensor, a temperature sensor, a humidity sensor, an illuminance sensor, an ultraviolet (UV) sensor, an electromyogram (EMG) sensor, an electroencephalogram (EEG) sensor, an electrocardiogram (ECG) sensor, an infrared (IR) sensor, an ultrasonic sensor, an iris sensor, or a fingerprint sensor. In addition, the sensor 180 may include one or more position sensors (e.g., an inertial measurement unit that may include one or more accelerometers, gyroscopes, and other components). Additionally, the sensor 180 may include a control circuit for controlling at least one of the sensors included herein. Any one of these sensors 180 may be located within the electronic device 101.
[0048] In some embodiments, the electronic device 101 may be a wearable device or a wearable device on which an electronic device can be mounted (e.g., an HMD). For example, the electronic device 101 may represent an XR wearable device (e.g., a headset or smart glasses). In other embodiments, the first external electronic device 102 or the second external electronic device 104 may be a wearable device or a wearable device on which an electronic device can be mounted (such as an HMD). In those other embodiments, when the electronic device 101 is mounted in the electronic device 102 (e.g., an HMD), the electronic device 101 may communicate with the electronic device 102 through the communication interface 170. The electronic device 101 may be directly connected to the electronic device 102 to communicate with the electronic device 102 without involving a separate network.
[0049] The first external electronic device 102, the second external electronic device 104, and the server 106 can each be a device of the same or different type as the electronic device 101. According to certain embodiments of the present disclosure, the server 106 includes a group of one or more servers. Additionally, according to certain embodiments of the present disclosure, all or some of the operations performed on the electronic device 101 can be performed on another or multiple other electronic devices (e.g., electronic devices 102 and 104 or the server 106). Further, according to certain embodiments of the present disclosure, when the electronic device 101 should automatically or upon request perform some functions or services, the electronic device 101 can request another device (e.g., electronic devices 102 and 104 or the server 106) to perform at least some functions associated therewith, rather than performing the functions or services by itself or additionally. Another electronic device (e.g., electronic devices 102 and 104 or the server 106) is capable of performing the requested functions or additional functions and transmitting the execution results to the electronic device 101. The electronic device 101 can provide the requested functions or services by processing the received results as is or additionally. For this purpose, for example, cloud computing, distributed computing, or client-server computing technologies can be used. Although Figure 1 FIG. shows that the electronic device 101 includes a communication interface 170 to communicate with the external electronic devices 102 and 104 or the server 106 via the networks 162 or 164, but according to some embodiments of the present disclosure, the electronic device 101 can operate independently without having a separate communication function.
[0050] The server 106 can include components that are the same as or similar to those of the electronic device 101 (or a suitable subset thereof). The server 106 can support driving the electronic device 101 by performing at least one of the operations (or functions) implemented on the electronic device 101. For example, the server 106 can include a processing module or a processor that can support the processor 120 implemented in the electronic device 101. As described below, the server 106 can perform one or more functions related to mask generation for object and scene segmentation for the utilization of see-through XR.
[0051] Although Figure 1 FIG. shows an example of the network configuration 100 including the electronic device 101, various changes can be made to Figure 1 it. For example, the network configuration 100 can include any number of each component in any suitable arrangement. Generally, computing and communication systems have a wide variety of configurations, and Figure 1 the scope of the present disclosure is not limited to any specific configuration. Additionally, although Figure 1 FIG. shows an operating environment in which various features disclosed in this patent document can be used, these features can be used in any other suitable system.
[0052] Figure 2 Shown is an example architecture 200 for training an object segmentation model to support mask generation for object and scene segmentation for see-through XR. For ease of explanation, Figure 2 the architecture 200 is described as being implemented using Figure 1 the server 106 in the network configuration 100 of
[0053] As Figure 2 shown, the architecture 200 includes or can access a training data store 202. The training data store 202 stores training data that can be used to train one or more object segmentation models. The training data store 202 includes any suitable structure configured to store and facilitate data retrieval. In some cases, the training data store 202 can be local to other components of the architecture 200, such as when the training data store 202 is maintained by the same entity that is training the object segmentation model. In other cases, the training data store 202 can be remote from other components of the architecture 200, such as when the training data store 202 is maintained by one entity and the training data is obtained by another entity that is training the object segmentation model. The training data can represent labeled or annotated training images with known segmentations of objects captured within the training images.
[0054] In this example, the training data store 202 provides a left perspective image frame 204 and a right perspective image frame 206. The image frames 204 and 206 form a plurality of stereo image pairs, which indicates that the left perspective image frame 204 is associated with the right perspective image frame 206, and both represent images of a common scene captured from slightly different positions. Note that, for convenience only, these image frames 204 and 206 are referred to as the left image frame and the right image frame. Each image frame 204 and 206 represents a scene image with a known segmentation. Each image frame 204 and 206 can have any suitable resolution and size. In some cases, for example, each image frame 204 and 206 can have a 2K, 3K, or 4K resolution. Each image frame 204 and 206 can also include image data in any suitable format. In some embodiments, for example, each image frame 204 and 206 includes RGB image data, which typically includes image data in three color channels (i.e., red, green, and blue channels). However, each image frame 204 and 206 can include image data with any other suitable resolution, form, or arrangement.
[0055] The left perspective image frame 204 is provided to a feature extraction function 208, which generally operates to extract specific features from the image frame 204. For example, the feature extraction function 208 may include one or more convolutional layers, other neural network layers, or other machine learning layers that process the image data to identify specific features that the machine learning layer has been trained to recognize. In this example, the feature extraction function 208 is configured to generate lower resolution features 210 (which may also be referred to as "single-stage" features in some embodiments) and higher resolution features 212. As the name implies, the features 212 have a higher resolution than the features 210. For example, the higher resolution features 212 may have a resolution that matches the resolution of the image frame 204, while the lower resolution features 210 may have a resolution of 600×600 or other lower resolution. In some cases, the feature extraction function 208 may generate the higher resolution features 212 and perform downsampling to generate the lower resolution features 210. However, the lower resolution features 210 and the higher resolution features 212 may be generated in any other suitable manner.
[0056] The lower resolution features 210 are provided to a classification function 214, a mask kernel generation function 216, and a depth or disparity kernel generation function 218. The classification function 214 generally operates to process the lower resolution features 210 in order to identify and classify the objects captured in the left perspective image frame 204. For example, the classification function 214 may analyze the lower resolution features 210 in order to: (i) detect one or more objects within each left perspective image frame 204, and (ii) classify each detected object as one of a plurality of object categories or types. The classification function 214 may use any suitable technique to detect and classify the objects in the image. Various classification algorithms are known in the art, and other classification algorithms will surely be developed in the future. The present disclosure is not limited to any particular technique for detecting and classifying the objects in the image.
[0057] The mask kernel generation function 216 generally operates to process the lower resolution features 210 and the classification results from the classification function 214 in order to generate mask kernels for the different objects detected within the left perspective image frame 204. Each mask kernel may represent an initial lower resolution mask that identifies the boundaries of the associated object within the associated left perspective image frame 204. The mask kernel generation function 216 may use any suitable technique to generate the mask kernels associated with the objects in the image frame. In some cases, for example, the mask kernel generation function 216 may include one or more convolutional layers that are trained to convolve the lower resolution features 210 in order to generate the mask kernels.
[0058] The depth or disparity kernel generation function 218 generally operates to process the lower resolution features 210 in order to generate depth or disparity kernels for different objects detected within the left perspective image frame 204. Each depth or disparity kernel may represent an initial lower resolution estimate of one or more depths or disparities of the associated object within the associated left perspective image frame 204. As described below, depth and disparity are related, and thus knowledge of disparity values can be used to generate depth values (and vice versa). The depth or disparity kernel generation function 218 may use any suitable technique to generate depth or disparity kernels associated with objects in the image frame. In some cases, for example, the depth or disparity kernel generation function 218 may include one or more convolutional layers trained to convolve the lower resolution features 210 in order to generate depth or disparity kernels.
[0059] The higher resolution features 212 are provided to the mask embedding generation function 220 and the depth or disparity embedding generation function 222. The mask embedding generation function 220 generally operates to process the higher resolution features 212 in order to create a mask embedding that may represent the embedding of the higher resolution features 212 within a mask embedding space associated with the left perspective image frame 204. The mask embedding space represents an embedding space within which masks defining object boundaries can be defined at a higher resolution. Similarly, the depth or disparity embedding generation function 222 generally operates to process the higher resolution features 212 in order to create a depth or disparity embedding that may represent the embedding of the higher resolution features 212 within a depth or disparity embedding space associated with the left perspective image frame 204. The depth or disparity embedding space represents an embedding space within which depths or disparities can be defined at a higher resolution. Each of the mask embedding generation function 220 and the depth or disparity embedding generation function 222 may use any suitable technique to generate an embedding within the associated embedding space.
[0060] The mask kernel generated by the mask kernel generation function 216 and the mask embedding generated by the mask embedding generation function 220 are provided to the instance mask generation function 224, which generally operates to produce instance masks associated with different objects within the left perspective image frame 204. Each instance mask represents a higher resolution mask identifying the boundaries of the associated object within the associated left perspective image frame 204. For example, the mask embedding generation function 220 may use the mask kernel generated by the mask kernel generation function 216 for a particular object in order to identify a subset of the mask embedding associated with that particular object generated by the mask embedding generation function 220.
[0061] Similarly, the depth or disparity kernel generated by the depth or disparity kernel generation function 218 and the depth or disparity embedding generated by the depth or disparity embedding generation function 222 are provided to the instance depth or disparity map generation function 226, which typically operates to produce an instance depth or disparity map associated with different objects in the left perspective image frame 204. Each instance depth or disparity map represents a higher-resolution estimate of one or more depths or disparities of the associated object within the associated left perspective image frame 204. For example, the instance depth or disparity map generation function 226 can use the depth or disparity kernel generated by the depth or disparity kernel generation function 218 for a specific object to identify a subset of the depth or disparity embeddings associated with that specific object generated by the depth or disparity embedding generation function 222. The instance depth or disparity maps associated with each left perspective image frame 204 can collectively represent the depth or disparity map associated with all detected objects within that left perspective image frame 204.
[0062] The various functions 214 to 226 described above can represent functions implemented or performed by an object segmentation model representing a machine learning model. Thus, various weights or other hyperparameters of the object segmentation model can be adjusted during training performed using the architecture 200. In machine learning model training, a loss function is typically used to calculate the loss for the machine learning model, where the loss is identified based on the difference or error between: (i) the actual output or other values generated by the machine learning model, and (ii) the expected or desired output or other values that should be generated by the machine learning model. A common goal in machine learning model training is to minimize the loss for the machine learning model by adjusting the weights or other hyperparameters of the machine learning model. For example, assume the loss function is defined as follows.
[0063] (1)
[0064] Loss function Example types can include cross-entropy loss, log-likelihood loss, mean square loss, and mean absolute loss. Here, represents the value generated by the machine learning model, represents the actual value (ground truth data) provided by the training sample, and represents the weights or other hyperparameters of the machine learning model. An optimization algorithm can be used to minimize the loss to obtain the optimal weights or other hyperparameters for the machine learning model, which can be represented as follows.
[0065] (2)
[0066] The minimization algorithm can minimize the loss by adjusting hyperparameters, and when the minimum loss is reached, the optimal hyperparameters can be obtained. At this time, it can be assumed that the machine learning model generates accurate outputs.
[0067] In Figure 2 the example, the loss function used to train the object segmentation model represents a combination of two losses. One of these losses is referred to as the segmentation loss, which relates to the extent to which the object segmentation model segments an object or identifies the boundaries of the object captured in the image frames 204 and 206. Here, the segmentation loss calculation function 228 is used to determine the segmentation loss, and the segmentation loss calculation function 228 compares the instance mask generated by the instance mask generation function 224 with the segmentation ground truth 230. As described above, the instance mask generated by the instance mask generation function 224 represents the estimated boundaries of the objects contained in the left perspective image frame 204. The segmentation ground truth 230 represents the known segmentation of the left perspective image frame 204, which represents that the segmentation ground truth 230 identifies the actual or correct object segmentation that should be generated by the object segmentation model. Therefore, the segmentation loss calculation function 228 can compare the object segmentation defined by the instance mask generated by the instance mask generation function 224 with the correct object segmentation defined by the segmentation ground truth 230. The error between the two can be used to calculate the segmentation loss for the object segmentation model.
[0068] For this reason, in
[0069] the example, the training data memory 202 may lack ground truth data related to the depth or disparity of the image frames 204 and 206. That is, the training data memory 202 may lack the correct depth or disparity map that should be generated by the object segmentation model. This situation may commonly exist because accurately identifying depth in training image frames can be very time-consuming and expensive. However, the extent to which the object segmentation model estimates the disparity or depth can be determined by using the instance depth or disparity map generated by the instance depth or disparity map generation function 226 to reconstruct the left perspective image frame 204 using the right perspective image frame 206. As described below, when the depth or disparity is known, one image frame associated with a first viewpoint can be projected or otherwise transformed into another image frame associated with a different viewpoint. Therefore, for example, when the depth or disparity in the scene is known, the image frame captured at the image plane of the right imaging sensor can be converted into the image frame at the image plane of the left imaging sensor (and vice versa). Figure 2Among them, the other loss used to train the object segmentation model is called the reconstruction loss, which involves the degree to which the object segmentation model generates depth or disparity values for reconstructing the left perspective image frame 204 using the right perspective image frame 206. In this example, the reconstruction loss calculation function 232 is used to determine the reconstruction loss. The reconstruction loss calculation function 232 projects or otherwise transforms the right perspective image frame 206 based on the depth or disparity values included in the instance depth or disparity map generated by the instance depth or disparity map generation function 226. This realizes the creation of the reconstructed left perspective image frame. If the instance depth or disparity map generated by the instance depth or disparity map generation function 226 is accurate, there should be less error between the reconstructed left perspective image frame and the actual left perspective image frame 204. If the instance depth or disparity map generated by the instance depth or disparity map generation function 226 is inaccurate, there should be more error between the reconstructed left perspective image frame and the actual left perspective image frame 204. Therefore, the reconstruction loss calculation function 232 can compare the reconstructed left perspective image frame with the actual left perspective image frame 204. The error between the two can be used to calculate the reconstruction loss for the object segmentation model.
[0070] One possible advantage of this method is that it does not require or demand ground truth depth or disparity information. That is to say, the object segmentation model can be trained without using the ground truth depth or disparity information associated with the image frames 204 and 206. Therefore, this can simplify the training of the object segmentation model and reduce the cost associated with training. However, this is not necessary. For example, the training data memory 202 can include ground truth depth or disparity values. In this case, the reconstruction loss calculation function 232 can identify the error between the instance depth or disparity map generated by the instance depth or disparity map generation function 226 and the ground truth depth or disparity values.
[0071] Although Figure 2 shows an example of the architecture 200 for training an object segmentation model to support mask generation for object and scene segmentation for see-through XR, various changes can be made to Figure 2 For example, various components and functions in Figure 2 can be combined, further subdivided, replicated, omitted, or rearranged, and additional components and functions can be added according to specific needs. As a specific example, although different operations are shown to be performed using the left perspective image frame 204 and the right perspective image frame 206, these operations can be reversed. In other words, the right perspective image frame 206 can be provided to the feature extraction function 208, while the left perspective image frame 204 can be used for reconstruction purposes.
[0072] Figure 3 shows according to the present disclosure Figure 2An example process 300 for determining a reconstruction loss within the architecture 200. More specifically, Figure 3 The illustrated process 300 may involve the use of the above-described reconstruction loss calculation function 232. Here, a possible goal may be to calculate a reconstruction loss during object segmentation model training in order to apply depth or disparity information during training, which can help improve the segmentation process and generate more accurate segmentation results.
[0073] As described above, one way to incorporate a reconstruction loss into object segmentation model training is to construct a depth or disparity map, apply the depth or disparity map to guide the segmentation process, and determine the reconstruction loss based on ground truth depth or disparity data. Feedback generated using the reconstruction loss can be used during the model training process to adjust the object segmentation model being trained, which ideally results in lower losses over time. In these embodiments, the model training process will require both: (i) segmentation ground truth 230, and (ii) ground truth depth or disparity data. As described above, ground truth depth or disparity data may not be available because (among other reasons) it may be difficult to obtain ground truth depth or disparity data in various situations.
[0074] Figure 3 The illustrated process 300 does not require ground truth depth or disparity data to identify the reconstruction loss. Instead, Figure 3 The illustrated process 300 constructs depth or disparity information and applies the depth or disparity information to create a reconstructed image frame. The reconstructed image frame is compared with the actual image frame to determine the reconstruction loss. Additionally, this allows the use of stereo image pairs that are spatially consistent with each other to determine the reconstruction loss. Thus, process 300 uses depth or disparity information but does not require ground truth depth or disparity data during model training.
[0075] As Figure 3 illustrated, each left perspective image frame 204 can be used to generate one or more instance depth or disparity maps 302, which can be produced using the instance depth or disparity map generation function 226 as described above. The one or more instance depth or disparity maps 302 are used to generate a reconstructed left perspective image frame 304, which can be achieved by projecting, warping, or otherwise transforming the right perspective image frame 206 associated with the left perspective image frame 204 based on the one or more instance depth or disparity maps 302. Similarly, when the depth or disparity associated with a scene is known, an image frame of the scene at one image plane can be transformed into a different image frame of the scene at a different image plane based on the depth or disparity.
[0076] Here, the reconstruction loss calculation function 232 can take the right perspective image frame 206 and generate a reconstructed version of the left perspective image frame 204 based on the depth or disparity generated by the instance depth or disparity map generation function 226. This makes the reconstructed left perspective image frame 304 and the left perspective image frame 204 form a stereo image pair with known consistency in terms of their depth or disparity. Therefore, any difference between the image frames 204 and 304 may be due to inaccurate disparity or depth estimation by the object segmentation model. As a result, the error determination function 306 can identify the difference between the image frames 204 and 304 to calculate the reconstruction loss 308, and can use the reconstruction loss 308 to adjust the object segmentation model during training. For example, the reconstruction loss 308 can be combined with the segmentation loss determined by the segmentation loss calculation function 228, and the combined loss can be compared with a threshold. The object segmentation model can be adjusted during training until the combined loss drops below the threshold or until some other one or more criteria are met (e.g., a specified number of training iterations have occurred or a specified amount of training time has elapsed).
[0077] Although Figure 3 illustrates Figure 2 an example of the process 300 for determining the reconstruction loss within the architecture 200 of Figure 3 various changes can be made to Figure 3 For example, various components and functions in
[0078] Figure 4 illustrates an example architecture 400 for using an object segmentation model that supports mask generation for object and scene segmentation for see-through XR. For ease of explanation, Figure 4 the architecture 400 of Figure 1 is described as being implemented using the electronic device 101 in the network configuration 100 of
[0079] As Figure 4As shown, the architecture 400 receives a left perspective image frame 402 and a right perspective image frame 404. These image frames 402 and 404 form a stereo image pair, which means that the image frames 402 and 404 represent images of a common scene captured from slightly different positions (again, for convenience, referred to as the left position and the right position). In some cases, the imaging sensor 180 of the electronic device 101 can be used to capture the image frames 402 and 404, such as when using the imaging sensor 180 of an XR headset or other XR device to capture the image frames 402 and 404. Each of the image frames 402 and 404 can have any suitable resolution and size. In some cases, for example, each of the image frames 402 and 404 can have a resolution of 2K, 3K, or 4K depending on the capabilities of the imaging sensor 180. Each of the image frames 402 and 404 can also include image data in any suitable format. In some embodiments, for example, each of the image frames 402 and 404 includes RGB image data that typically includes image data in the red, green, and blue channels. However, each of the image frames 402 and 404 can include image data with any other suitable resolution, form, or arrangement.
[0080] The left perspective image frame 402 is provided to the trained machine learning model 406. The trained machine learning model 406 represents an object detection model that has been trained using the above-described architecture 200. The trained machine learning model 406 is used to generate a segmentation 408 of the left image frame 402 and a depth or disparity map 410 of the left image frame 402. The segmentation 408 of the left image frame 402 identifies the different objects contained in the left perspective image frame 402. For example, the segmentation 408 of the left image frame 402 can include various instance masks generated by the instance mask generation function 224 of the trained machine learning model 406, or is formed by various instance masks generated by the instance mask generation function 224 of the trained machine learning model 406. The segmentation 408 of the left image frame represents the object segmentation prediction associated with the left image frame 402. The depth or disparity map 410 of the left image frame 402 identifies the depth or disparity associated with the scene imaged by the left perspective image frame 402. The depth or disparity map 410 of the left image frame 402 can include various instance depth or disparity maps generated by the instance depth or disparity map generation function 226 of the trained machine learning model 406, or is formed by various instance depth or disparity maps generated by the instance depth or disparity map generation function 226 of the trained machine learning model 406.
[0081] Object segmentation predictions associated with the right image frame 404 can be generated in different ways. For example, in some cases, the image-guided segmentation reconstruction function 412 can be used to generate a segmentation 414 of the right image frame 404. In these embodiments, the image-guided segmentation reconstruction function 412 can project or otherwise transform the segmentation 408 of the left image frame 402 to produce a segmentation 414 of the right image frame 404. As a specific example, the image-guided segmentation reconstruction function 412 can project the segmentation 408 of the left image frame 402 onto the right image frame 404. This transformation can be based on the depth or disparity map 410 of the left image frame 402. This can be similar to the process described above for reconstruction loss calculation, but here, the transformation is used to apply the segmentation 408 of the left image frame 402 to the right image frame 404. This method supports spatial consistency between the segmentation 408 of the left image frame 402 and the segmentation 414 of the right image frame 404. Since only one rather than two image frames 402 and 404 are processed using the trained machine learning model 406, this method also simplifies the segmentation process and saves computational power. However, this is not necessary. In other embodiments, for example, the trained machine learning model 406 can also be used to process the right image frame 404 and generate a segmentation 414 of the right image frame 404.
[0082] The boundary refinement function 416 generally operates to process the left segmentation 408 of the left image frame 402 and the right segmentation 414 of the right image frame 404. The boundary refinement function 416 can be used to refine and correct the left segmentation 408 and the right segmentation 414 at the positions where needed, which can enable the generation of the final segmentation 418 for the image frames 402 and 404. For example, the boundary refinement function 416 can detect regions where objects overlap and determine appropriate boundaries for the overlapping objects (possibly based on known information or prior experience of a specific type of object). The boundary refinement function 416 can also clarify and validate the segmentation results to provide noise reduction and improved results. The final segmentation 418 can represent or include at least one segmentation mask that identifies or isolates one or more objects within the image frames 402 and 404. The final segmentation 418 can be used in any suitable way (e.g., by processing the final segmentation 418, as Figure 5 shown and described below).
[0083] As can be seen in Figure 4 it is possible to obtain both the segmentation 408 of the image frame and the depth or disparity map 410 simultaneously. This can be very useful in certain applications (e.g., scene reconstruction and recognition in perspective XR applications). As a specific example, this is very useful in applications where a physical keyboard is captured in the image frames 402 and 404. As described below, the XR system can identify user input by identifying the physical keyboard and identifying which keys of the physical keyboard the user presses.
[0084] although Figure 4 One example of an architecture 400 for using an object segmentation model that supports mask generation with object and scene segmentation for see-through XR is shown, but can be used for Figure 4 For example, you can make various changes to Figure 4 The various components and functions in the image processing unit 404 may be combined, further subdivided, duplicated, omitted, or rearranged, and additional components and functions may be added as needed. As a specific example, although different operations are shown as being performed using left perspective image frame 402 and right perspective image frame 404, these operations can be reversed. In other words, right perspective image frame 404 can be provided to trained machine learning model 406, while left perspective image frame 402 can be provided to image-guided segmentation and reconstruction function 412.
[0085] Figure 5 and Figure 6 Shown according to the present disclosure Figure 4 An example process 500 for boundary refinement and virtual view generation within the architecture 400 of FIG. More specifically, Figure 5 The illustrated process 500 may involve the use of the boundary refinement functionality 416 described above, as well as additional functionality for generating, rendering, and displaying a virtual view of a scene. Here, one possible goal of boundary refinement may be to finalize the boundaries of detected objects and regions within a captured scene, allowing the segmented regions to accurately represent the objects and regions in the scene.
[0086] like Figure 5 As shown, the object boundary extraction and boundary region expansion function 502 processes the segmentation 408 of the left image frame 402 and the segmentation 414 of the right image frame 404 to identify the boundaries of the objects captured in the image frames 402 and 404 and expand the identified boundaries. An example of this is shown in FIG. Figure 6 , where the boundary region includes a portion of a boundary 602 of a detected object identified within image frames 402 and 404. The boundary region can be expanded to produce an extended boundary region defined by two boundaries 604 and 606. The extended boundary region defines the space where the boundary on one object may intersect the boundary of another object. Here, the boundary 602 is expanded by amounts of +δ and -δ to produce boundaries 604 and 606 defining the extended boundary region. The object boundary extraction and boundary region expansion function 502 can use any suitable technique to identify and expand boundaries and boundary regions. For example, each boundary can be identified based on the boundaries of the objects included in the segmentation 408 and 414, and a fixed value or a variable value for δ can be used to generate the extended boundary region.
[0087] The classified image-guided boundary refinement function 504 can be used to process the identified boundaries and boundary regions to correct certain problems using object segmentation. For example, the classified image-guided boundary refinement function 504 can perform classification of pixels within the extended boundary region defined by boundaries 604 and 606 to complete one or more incomplete regions associated with at least one object in the scene. That is, the classified image-guided boundary refinement function 504 can determine which pixels in the extended boundary region belong to which objects within the scene. This can be useful when there may be multiple objects within the extended boundary region. Here, the classified image-guided boundary refinement function 504 can use the original image frames 402 and 404 to support boundary refinement. If there are multiple objects, the classified image-guided boundary refinement function 504 can separate the objects (based on the classification of pixels in the extended boundary region) to enhance the boundaries of the objects and complete the boundaries of the objects. An example of this is described below with reference to Figures 9 to 11 to describe an example of it.
[0088] Note that when objects overlap, the classified image-guided boundary refinement function 504 can identify the boundaries of the upper object and may or may not be able to estimate the boundaries of the lower object. For example, in some cases, the boundaries of the lower object can be estimated by: (i) identifying one or more boundaries of one or more visible parts of the lower object, and (ii) estimating one or more boundaries of one or more other parts of the lower object that are occluded by the upper object. The estimation of the boundaries of the occluded parts of the lower object can be based on known information or prior experience of a specific type of object (e.g., when a specific type of object typically has a known shape). An example of this is described below with reference to Figures 12A to 13D to describe two examples of it.
[0089] At this time, at least one post-processing function 506 can be performed. For example, at least one post-processing function 506 can be used to finally determine the boundaries of the segmented objects by creating edge connections and modifying the edge thickness to remove noise. As a specific example of this, boundary 602 can include one or more gaps 608 as shown in Figure 6 which may be caused by any number of factors. Here, the post-processing function 506 can be performed to connect the segments of boundary 602 and create a continuous boundary for the associated object. This can be performed for all objects to ensure that each object has a continuous boundary. At least one post-processing function 506 can be used to create a refined panoramic segmentation 508 that represents the segmentation of a three-dimensional (3D) scene. Here, the refined panoramic segmentation 508 can represent the final segmentation 418 as shown in Figure 4
[0090] The refined panoramic segmentation 508 can be used in any suitable manner. In this example, the refined panoramic segmentation 508 is provided to the 3D object and scene reconstruction function 510. The 3D object and scene reconstruction function 510 generally operates to process the refined panoramic segmentation 508 in order to generate a 3D model of the scene and one or more objects within the scene captured in the image frames 402 and 404. For example, in some cases, the refined panoramic segmentation 508 can be used to define one or more masks based on the boundaries of one or more objects in the scene. Additionally, the masks can help separate the individual objects in the scene from the background of the scene. The 3D object and scene reconstruction function 510 can also use the 3D models of the objects and the scene to perform object and scene reconstruction (e.g., by reconstructing each object using the 3D model of the object and separately reconstructing the background of the scene).
[0091] The reconstructed objects and scene can be provided to the left and right view generation function 512, which generally operates to produce a left virtual view and a right virtual view of the scene. For example, the left and right view generation function 512 can perform viewpoint matching and disparity correction in order to create virtual views for the left and right eyes to be presented to the user. The distortion and aberration correction function 514 generally operates to process the left virtual view and the right virtual view in order to correct for various distortions, aberrations, or other problems. As a specific example, the distortion and aberration correction function 514 can be used to pre-compensate the left virtual view and the right virtual view for geometric distortion caused by the display lenses of the XR device worn by the user. In this example, the user can generally view the left virtual view and the right virtual view through the display lenses of the XR device, and these display lenses may introduce geometric distortion due to the shape of the display lenses. Therefore, the distortion and aberration correction function 514 can pre-compensate the left virtual view and the right virtual view in order to reduce or substantially eliminate the geometric distortion in the left virtual view and the right virtual view viewed by the user. As another specific example, the distortion and aberration correction function 514 can be used to correct chromatic aberration.
[0092] The left - view and right - view rendering function 516 can be used to render the corrected virtual view, and this left - view and right - view rendering function can generate the actual image data to be presented to the user. The rendered view is presented on one or more displays of the XR device via the left - view and right - view display function 518 (e.g., via one or more displays 160 of the electronic device 101). Note that multiple separate displays 160 (e.g., a left display and a right display that can be separately viewed by the user's eyes) or a single display 160 (e.g., a display where the left and right portions of the display can be separately viewed by the user's eyes) can be used to present the rendered view. In addition, this can allow the user to view the transformed and integrated image streams from multiple perspective cameras, in which a graphics pipeline performing the various functions described above is used to generate the images.
[0093] Although Figure 5 and Figure 6 shows Figure 4 an example of the process 500 for boundary refinement and virtual - view generation within the architecture 400 of Figure 5 and Figure 6 various changes can be made to Figure 5 . For example, various components and functions in Figure 6 can be combined, further subdivided, replicated, omitted, or rearranged, and additional components and functions can be added according to specific requirements. In addition,
[0094] It should be noted that Figures 2 to 6 the functions shown in Figures 2 to 6 or described with respect to Figures 2 to 6 can be implemented in any suitable manner in the electronic devices 101, 102, 104, the server 106, or other devices. For example, in some embodiments, one or more software applications or other software instructions executed by the processors 120 of the electronic devices 101, 102, 104, the server 106, or other devices can be used to implement or support Figures 2 to 6 at least some of the functions shown in Figures 2 to 6 or described with respect to Figures 2 to 6 . In other embodiments, dedicated hardware components can be used to implement or support at least some of the functions shown in Figures 2 to 6 or described with respect to Figures 2 to 6 . Generally, the functions shown in Figures 2 to 6 or described with respect to Figures 2 to 6The described functionality may be performed by a single device or by multiple devices. For example, server 106 may implement architecture 200 to train an object segmentation model, and electronic device 101 may implement architecture 400 to use the trained object segmentation model deployed to electronic device 101 after training.
[0095] Among the various functions described above, an image frame or a projection or other transformation of a segmentation is described as being performed from left to right (or vice versa) based on depth or disparity information. For example, the reconstruction loss calculation function 232 may project the right perspective image frame 206 based on the depth or disparity values included in the instance depth or disparity map generated by the instance depth or disparity map generation function 226. As another example, the image-guided segmentation reconstruction function 412 may project the segmentation 408 of the left image frame 402 onto the right image frame 404 based on the depth or disparity map 410 of the left image frame 402. Examples of how these projections or other transformations may occur are now described below.
[0096] Figure 7 An example relationship 700 between segmentation results associated with a stereo image pair according to the present disclosure is shown. As Figure 7 shown, an object 702 (representing a tree in this example) is captured in two image frames 704 and 706 associated with different image planes. The left image frame 704 may be viewed by the user's left eye 708, and the right image frame 706 may be viewed by the user's right eye 710. Here, B represents the inter-pupillary distance (IPD) between the user's eyes 708 and 710, f represents the focal length of the imaging sensor that captures the image frames 704 and 706, and d represents the depth associated with the object 702.
[0097] Disparity refers to the distance between two points in the left and right image frames of a stereo image pair, where the two points correspond to the same point in the scene. For example, a point 712 of the object 702 appears at position (x l , f) in the left image frame 704, and the same point 712 of the object 702 appears at position (x r , f) in the right image frame 706. Based on this, the disparity (p) can be calculated as follows:
[0098] (3)
[0099] Therefore, it can be seen that:
[0100] (4)
[0101] (5)
[0102] Thus, if the depth or disparity associated with a pixel is known, pixel data can be obtained at one image plane and converted to pixel data at another image plane.
[0103] Figure 8 An example image-guided segmentation reconstruction 800 using stereo consistency according to the present disclosure is shown. Figure 8 The reconstruction 800 shown therein can be performed by an image-guided segmentation reconstruction function 412, for example, when generating a segmentation 414 of a right image frame 404 based on a segmentation 408 of a left image frame 402.
[0104] As Figure 8 shown, a mask 802 of an object 702 can be generated, for example, as described above by using a left image frame 704. The mask 802 is associated with or defined by respective points 804, and only one of them is shown here for simplicity. When the depth or disparity associated with the scene is known, the mask 802 of the object 702 for the left image frame 704 can be converted to a corresponding mask 806 for the right image frame 706. For example, each point 804 can be converted to a corresponding point 808 using equation (4) above. Alternatively, equation (5) above can be used to start from the mask 806 for the right image frame 706 and reconstruct the mask 802 for the left image frame 704. As a result, the method can be used to Figure 4 convert the segmentation 408 of the left image frame 402 to the segmentation 414 of the right image frame 404 (and vice versa). This can enable the generation of segmentations 408 and 414 that are known to be spatially consistent with each other. Since the trained machine learning model 406 may only need to generate one of the segmentations 408 and 414, this can also help reduce the computational load required to generate the segmentations 408 and 414 for different image frames 402 and 404.
[0105] Although Figure 7 an example of the relationship 700 between segmentation results associated with a stereo image pair is shown, and Figure 8 an example of an image-guided segmentation reconstruction 800 using stereo consistency is shown, various changes can be made to Figure 7 and Figure 8 them. For example, the scene imaged here is only for illustration and can vary widely depending on the circumstances.
[0106] The ability to generate masks for objects captured in image frames can be used in many applications. The following are examples of applications where this functionality can be used. However, note that these are non-limiting examples, and the ability to generate masks for objects captured in image frames can be used in any other suitable application.
[0107] As a first example application, the functionality can be used for nearby object mask generation, which involves generating one or more masks for one or more objects that are nearby or otherwise close to a VST XR headset or other VST XR device. The proximity of the nearby objects to the VST XR device can vary due to the movement of the objects, the movement of the VST XR device (e.g., when the user moves around), or both. If and when it is determined that any detected object is too close to the VST XR device (e.g., based on a threshold distance) or approaching the VST XR device too quickly (e.g., based on a threshold speed), a warning can be generated for the user of the VST XR device. The warning can indicate to the user to not hit or bump into the nearby objects.
[0108] As a second example application, the functionality can be used for mask generation used in 3D object reconstruction, which can involve reconstructing 3D objects detected in a scene captured using a perspective camera. Here, separate masks can be generated to separate the objects in the foreground of the scene from the background of the scene. After reconstructing these objects, the 3D objects and the background can be re-projected separately in order to effectively generate a high-quality final view.
[0109] As a third example application, the functionality can be used for keyboard mask generation, which can involve generating a mask for a physical keyboard detected within a scene. For example, as described above, a physical keyboard can be captured in an image frame, and the XR device can identify the physical keyboard and which keys of the keyboard the user has pressed to identify the input from the user. Here, the keyboard can be detected in the image frame captured by the XR device. After the keyboard is captured in the image frame, the keyboard object can be segmented, and a mask for the keyboard object can be generated. Then, the XR device can avoid rendering any digital content that occludes any part of the keyboard object. Using the depth or disparity map associated with the keyboard object, it can also be estimated which keys of the keyboard the user has pressed and use that input without a physical or wireless connection to the keyboard. Two examples of this use case are described below with reference to Figures 12A to 13D Describe two examples of this use case.
[0110] Figures 9 to 11 Shows example results obtainable using an object segmentation model according to the present disclosure, which supports mask generation for use in see-through XR with object and scene segmentation. For ease of explanation, Figures 9 to 11 The results shown in Figure 4 Are described as being obtained using an object segmentation model (such as the trained machine learning model 406 in the architecture 400 of Figure 2 ), where the object segmentation model can be trained using the architecture 200 of
[0111] As Figure 9As shown, an image 900 of a captured scene. As can be seen in Figure 9 , the scene includes multiple objects (e.g., a round table, a window, a plant, a wall, and a floor). As Figure 10 shown, a segmentation mask 1000 represents a segmentation mask that can be generated using known methods. As can be seen in Figure 10 , the segmentation mask 1000 is not particularly accurate in identifying different objects in the image 900. For example, the boundary between the window and the plant is not particularly accurate, the table is incomplete and has large gaps, and the wall is misclassified as two different objects. Here, part of the difficulty may come from the fact that the shadows produced by the sun on various objects in the image 900 potentially cause artifacts in the segmentation mask 1000.
[0112] As Figure 11 shown, a segmentation mask 1100 represents a segmentation mask that can be generated using the above techniques. As can be seen in Figure 11 , the segmentation mask 1100 is more accurate in identifying different objects in the image 900. For example, the boundary between the window and the plant is more clearly defined, the table is complete, and the wall is correctly classified as a single object. In addition, the use of the classified image-guided boundary refinement function 504 of the boundary refinement function 416 can help fill the gaps in the table mask, which allows the table to be identified without large gaps.
[0113] The segmentation mask 1100 also more accurately identifies the object categories. For example, in Figure 10 , the plant in the scene is identified with 49% certainty, which may be due to the boundary of the plant identified in the segmentation mask 1000. In Figure 11 , the plant in the scene is identified with a higher certainty, which can be due to the more accurate boundary of the plant in the segmentation mask 1100.
[0114] Figures 12A to 13D Other example results that can be obtained using an object segmentation model according to the present disclosure are shown, where the object segmentation model supports mask generation for use in object and scene segmentation for see-through XR. In the example of Figures 12A to 12D , image frames 1200 and 1202 have been captured, where each of the image frames 1200 and 1202 captures a keyboard. The above methods can be used to identify regions 1204 and 1206 within the image frames 1200 and 1202, where each of the regions 1204 and 1206 includes a keyboard. The regions 1204 and 1206 can be refined to generate boundaries 1208 and 1210 of the keyboard, as Figure 12C and Figure 12DAs shown. Each of these boundaries 1208 and 1210 can be used to define a mask that isolates the keyboard in each image frame 1200 and 1202. For example, one mask can identify the pixels within boundary 1208 of image frame 1200 and mask all other pixels, and another mask can identify the pixels within boundary 1210 of image frame 1202 and mask all other pixels.
[0115] In Figures 12A to 12D , even if the keyboard physically overlaps with another object (in this example, a book), the keyboard can still be recognized. In the absence of additional known information about the book, it may not be possible to identify the exact boundaries of the book in each image frame 1200 and 1202. For example, unless the height of the book is known (e.g., learned from prior experience), it may not be possible to define the boundaries of the occluded part of the book. Thus, in this example, only the boundaries of the keyboard are identified, but if needed or desired, the boundaries of the exposed part of the book can be identified.
[0116] In Figures 13A to 13D 's example, image frames 1300 and 1302 have been captured, where each of image frames 1300 and 1302 captures the keyboard. Similarly, the above method can be used to identify regions 1304a to 1304b and 1306a to 1306b within image frames 1300 and 1302, where each of regions 1304a to 1304b includes the keyboard, and each of regions 1306a to 1306b includes the book. The regions 1304a to 1304b and 1306a to 1306b can be refined to generate boundaries 1308a to 1308b and 1310a to 1310b, as Figure 13C and Figure 13D shown. Each of these boundaries 1308a to 1308b and 1310a to 1310b can be used to define a mask that isolates the keyboard or the book in each image frame 1300 and 1302.
[0117] In Figures 13A to 13D , even if the keyboard physically overlaps with another object, the keyboard can be recognized again. However, here, the possible boundaries of the book can be identified in each image frame 1300 and 1302. This can be based on known information about the underlying object (e.g., the known information that most books are rectangular). Based on this known information, even if the book is partially occluded by the keyboard, the possible boundaries of the book can be estimated. Note that this represents an example where sufficient information about the overlapping object is available. The system can learn and estimate the shapes of various objects, for example, based on prior experience or training.
[0118] Although Figures 9 to 13DShows an example of the results that can be obtained using an object segmentation model that supports mask generation for object and scene segmentation for see-through XR, but various changes can be made to Figures 9 to 13D it. For example, the specific examples of the image frames, objects, and associated segmentation results are for illustration and explanation only and can vary easily based on the environment (e.g., the specific scene being imaged).
[0119] Figure 14 Shows an example method 1400 for training an object segmentation model to support mask generation for object and scene segmentation for see-through XR according to the present disclosure. For ease of explanation, method 1400 is described as being executed by Figure 1 the server 106 in the network configuration 100, where the server 106 can implement Figure 2 the architecture 200. However, method 1400 can be executed using any other suitable device (e.g., the electronic device 101) with any other suitable architecture and in any other suitable system.
[0120] As Figure 14 shown, at step 1402, a first training image frame and a second training image frame for the scene are obtained. This can include, for example, the processor 120 of the server 106 obtaining the left perspective image frame 204 and the right perspective image frame 206 from the training data memory 202, for example. Note that while the training of a machine learning model typically includes many image pairs, for simplicity, the use of one left perspective image frame 204 and one right perspective image frame 206 is described here.
[0121] At step 1404, higher-resolution features and lower-resolution features are extracted from the first image frame. This can include, for example, the processor 120 of the server 106 executing the feature extraction function 208 to extract the lower-resolution features 210 and the higher-resolution features 212 from the image frame 204. At step 1406, the extracted features are provided to the object segmentation model being trained. This can include, for example, the processor 120 of the server 106 providing the lower-resolution features 210 and the higher-resolution features 212 as inputs to the object segmentation model being trained.
[0122] At step 1408, object classification is performed by the object segmentation model using lower-resolution features. This can include, for example, the processor 120 of server 106 executing the classification function 214 to identify the objects in image frames 204 and 206 and classifying the detected objects, for example, by classifying the detected objects into different object categories or types. At step 1410, the object segmentation model generates a mask kernel and a depth or disparity kernel using lower-resolution features. This can include, for example, the processor 120 of server 106 executing the mask kernel generation function 216 to generate a kernel mask based on the lower-resolution features 210. This can also include the processor 120 of server 106 executing the depth or disparity kernel generation function 218 to generate a depth or disparity kernel based on the lower-resolution features 210.
[0123] At step 1412, the object segmentation model generates a mask embedding and a depth or disparity embedding using higher-resolution features. This can include, for example, the processor 120 of server 106 executing the mask embedding generation function 220 to generate a mask embedding based on the higher-resolution features 212. This can also include the processor 120 of server 106 executing the depth or disparity embedding generation function 222 to generate a depth or disparity embedding based on the higher-resolution features 212. At step 1414, the object segmentation model uses the mask kernel and the mask embedding to generate instance masks. This can include, for example, the processor 120 of server 106 executing the instance mask generation function 224 to generate instance masks for the objects in image frames 204 and 206 based on the mask kernel and the mask embedding. At step 1416, the object segmentation model uses the depth or disparity kernel and the depth or disparity embedding to generate an instance depth or disparity map. This can include, for example, the processor 120 of server 106 executing the instance depth or disparity map generation function 226 to generate an instance depth or disparity map for the objects in image frames 204 and 206 based on the depth or disparity kernel and the depth or disparity embedding.
[0124] At step 1418, the first training image frame is reconstructed using the second training image frame. This can include, for example, the processor 120 of the server 106 executing the reconstruction loss calculation function 232 to generate the reconstructed image frame 304. As a specific example, the reconstructed image frame 304 can be generated by projecting or otherwise transforming the image frame 206 based on one or more instance depth or disparity maps associated with the image frame 204. At step 1420, a loss associated with the object segmentation model is determined. This can include, for example, the processor 120 of the server 106 executing the reconstruction loss calculation function 232 to calculate a reconstruction loss based on the error between the image frame 204 and the reconstructed image frame 304. This can also include the processor 120 of the server 106 executing the segmentation loss calculation function 228 to determine a segmentation loss. The server 106 can combine the reconstruction loss and the segmentation loss (or one or more other or additional losses) to identify the total loss associated with the object segmentation model. Again, note that the total loss here can be based on the error associated with multiple pairs of image frames 204 and 206. At step 1422, the loss associated with the object segmentation model is minimized and the best hyperparameters of the object segmentation model are identified. This can include, for example, the processor 120 of the server 106 using a minimization algorithm that attempts to minimize the total loss by adjusting the weights or other hyperparameters of the object segmentation model. Once the training is complete, the object segmentation model can be used in any suitable manner, such as when the object segmentation model is put into use by the server 106 or deployed to one or more other devices (e.g., the electronic device 101) for use.
[0125] Although Figure 14 illustrates an example of a method 1400 for training an object segmentation model to support mask generation for object and scene segmentation for see-through XR, various changes can be made to Figure 14 it. For example, although Figure 14 is shown as a series of steps, the steps in Figure 14 can overlap, occur in parallel, occur in a different order, or occur any number of times (including zero times). Additionally, although it is assumed above that the first training image frame is the left perspective image frame 204 and the second training image frame is the right perspective image frame 206, the left and right image frames can be reversed.
[0126] Figure 15 illustrates an example method 1500 for using an object segmentation model that supports mask generation for object and scene segmentation for see-through XR. For ease of explanation, method 1500 is described as being executed by the electronic device 101 in the network configuration 100 of Figure 1 where the electronic device 101 can implement Figure 4architecture 400. However, any other suitable device with any other suitable architecture (e.g., server 106) can be used and method 1500 can be executed in any other suitable system.
[0127] As Figure 15 shown, at step 1502, a first image frame and a second image frame of a scene captured by an XR device are obtained. This can include, for example, the processor 120 of the electronic device 101 obtaining a left perspective image frame 402 and a right perspective image frame 404 using a plurality of imaging sensors 180 of the electronic device 101. At step 1504, the first image frame is provided as an input to a trained object segmentation model. This can include, for example, the processor 120 of the electronic device 101 providing the left perspective image frame 402 to the trained machine learning model 406. The trained machine learning model 406 can be trained according to Figure 14 method 1400.
[0128] At step 1506, a first object segmentation prediction is generated using the trained machine learning model, and at step 1508, a depth or disparity map is generated using the trained machine learning model. This can include, for example, the processor 120 of the electronic device 101 using the trained machine learning model 406 to simultaneously generate a segmentation 408 of the left image frame 402 and a depth or disparity map 410 of the left image frame 402. At step 1510, a second object segmentation prediction is generated using the second image frame. In some cases, this can include the processor 120 of the electronic device 101 using the trained machine learning model 406 to generate a segmentation 414 of the right image frame 404. In other cases, this can include the processor 120 of the electronic device 101 performing an image-guided segmentation reconstruction function 412 to project or otherwise transform the segmentation 408 of the left image frame 402 based on the depth or disparity map 410 of the left image frame 402 in order to produce a segmentation 414 of the right image frame 404.
[0129] At step 1512, the first object segmentation prediction and the second object segmentation prediction are used to determine the boundaries of the objects in the scene. This can include, for example, the processor 120 of the electronic device 101 performing a boundary refinement function 416 in order to generate a final segmentation 418 for the image frames 402 and 404. As a specific example, this can include the processor 120 of the electronic device 101 performing an object boundary extraction and boundary region expansion function 502, a classified image-guided boundary refinement function 504, and a post-processing function 506 in order to generate a refined panoramic segmentation 508. A part of this process can include identifying objects that overlap or are close to each other and classifying the pixels in the extended boundary regions associated with the objects, and this process can be completed to complete one or more incomplete regions associated with at least one object in the scene.
[0130] At step 1514, a virtual view can be generated for presentation on at least one display of the XR device. This can include, for example, the processor 120 of the electronic device 101 executing the 3D object and scene reconstruction function 510, the left and right view generation function 512, and the distortion and aberration correction function 514. This can enable the generation of left and right virtual views suitable for presentation. At step 1516, the virtual view is presented on the display of the XR device. This can include, for example, the processor 120 of the electronic device 101 executing the left and right view rendering function 516 and the left and right view display function 518 to render and display the left and right virtual views.
[0131] Optionally, at step 1518, it can be determined whether an input is received from a user of the XR device via at least one object in the scene. This can include, for example, the processor 120 of the electronic device 101 determining whether any segmented object represents a physical keyboard and whether the user appears to be typing on the physical keyboard. If so, at step 1520, the input based on the user's interaction with the keyboard can be recognized and, at step 1522, used to perform one or more actions. This can include, for example, the processor 120 of the electronic device 101 using a mask associated with the keyboard and a depth or parallax map to estimate which buttons on the keyboard the user has selected. This can also include: the processor 120 of the electronic device 101 using the recognized keyboard buttons to identify the user input and performing one or more actions requested by the user based on the user input. This can be done in the absence of any physical or wireless connection for sending data from the keyboard to the XR device. However, note that the recognition of the keyboard or other user input devices can be used in any other suitable way, such as to ensure that no digital content overlays the keyboard or other user input devices when generating the virtual view.
[0132] Although Figure 15 An example of a method 1500 for using an object segmentation model that supports mask generation using object and scene segmentation for see-through XR is shown, various changes can be made to Figure 15 For example, although Figure 15 is shown as a series of steps, the steps in Figure 14 can overlap, occur in parallel, occur in a different order, or occur any number of times (including zero times). Additionally, although it is assumed above that the first image frame is the left perspective image frame 402 and the second image frame is the right perspective image frame 404, the left and right image frames can be reversed. Although the present disclosure has been described using example embodiments, various changes and modifications can be suggested to those skilled in the art. The present disclosure is intended to cover such changes and modifications that fall within the scope of the appended claims.
Claims
1. A method, comprising: Obtaining a first image frame and a second image frame of a scene; Providing the first image frame as an input to an object segmentation model, the object segmentation model being trained to generate a first object segmentation prediction for an object in the scene and a depth or disparity map based on the first image frame; Generating a second object segmentation prediction for the object in the scene based on the second image frame; Determining a boundary of the object in the scene based on the first object segmentation prediction and the second object segmentation prediction; And Generating a virtual view for presentation on a display of an extended reality (XR) device based on the boundary of the object in the scene.
2. The method according to claim 1, wherein, Generating the second object segmentation prediction includes: Performing image-guided segmentation reconstruction to generate the second object segmentation prediction based on the second image frame, the first object segmentation prediction, and the depth or disparity map, the second object segmentation prediction being spatially consistent with the first object segmentation prediction.
3. The method according to claim 1, wherein Generating the second object segmentation prediction includes: Providing the second image frame as an input to the object segmentation model, the object segmentation model being configured to generate the second object segmentation prediction for the object in the scene based on the second image frame.
4. The method according to claim 1, wherein, Determining the boundary of the object in the scene includes: Performing boundary refinement based on the first object segmentation prediction and the second object segmentation prediction.
5. The method according to claim 4, wherein Performing the boundary refinement includes, for each boundary in at least one of the boundaries: Identifying a boundary region associated with the boundary; Expanding the identified boundary region; and Classifying pixels within the expanded boundary region to complete one or more incomplete regions associated with at least one object among the objects in the scene.
6. The method according to claim 1, wherein Generating the virtual view includes: Performing object and scene reconstruction using one or more masks based on the boundaries of one or more objects in the scene to generate a three-dimensional (3D) model of the scene; and Using the 3D model to generate the virtual view.
7. The method according to claim 1, wherein: One of the objects in the scene includes a keyboard; and The method further includes identifying a user input to the XR device based on a physical or virtual interaction of a user with the keyboard.
8. An extended reality (XR) device, comprising: A plurality of imaging sensors configured to capture a first image frame and a second image frame of a scene; At least one processing device configured to: Provide the first image frame as an input to an object segmentation model, the object segmentation model being trained to generate a first object segmentation prediction for an object in the scene and a depth or disparity map based on the first image frame; Generate a second object segmentation prediction for the object in the scene based on the second image frame; Determine a boundary of the object in the scene based on the first object segmentation prediction and the second object segmentation prediction; And Generate a virtual view based on the boundary of the object in the scene; And At least one display configured to present the virtual view.
9. The XR device according to claim 8, wherein, To generate the second object segmentation prediction, the at least one processing device is configured to: perform image-guided segmentation reconstruction to generate the second object segmentation prediction based on the second image frame, the first object segmentation prediction, and the depth or disparity map, the second object segmentation prediction being spatially consistent with the first object segmentation prediction.
10. The XR device according to claim 8, wherein, To generate the second object segmentation prediction, the at least one processing device is configured to: provide the second image frame as an input to the object segmentation model, the object segmentation model being configured to generate the second object segmentation prediction for the objects in the scene based on the second image frame.
11. The XR device according to claim 8, wherein, To determine the boundaries of the objects in the scene, the at least one processing device is configured to: perform boundary refinement based on the first object segmentation prediction and the second object segmentation prediction.
12. The XR device according to claim 11, wherein, To perform the boundary refinement, the at least one processing device is configured to: identify a boundary region associated with the boundary; expand the identified boundary region; and classify pixels within the expanded boundary region to complete one or more incomplete regions associated with at least one of the objects in the scene.
13. The XR device according to claim 8, wherein, To generate the virtual view, the at least one processing device is configured to: perform object and scene reconstruction using one or more masks based on the boundaries of one or more of the objects in the scene to generate a three-dimensional (3D) model of the scene; and use the 3D model to generate the virtual view.
14. The XR device according to claim 8, wherein: one of the objects in the scene includes a keyboard; and the at least one processing device is further configured to: avoid placing one or more virtual objects on the keyboard in the virtual view.
15. A computer-readable storage medium including a program that, when executed by a processor, causes the processor to perform the method according to any one of claims 1 to 7.