Visual search refinement in computer-generated rendering environments
By performing coarse identification on the device and transmitting feature information to the server for fine identification, the problem of large latency in detecting physical objects in image or video data is solved, enabling faster CGR environment updates and object recognition, and improving the user experience.
Patent Information
- Application Number
- CN202411081178.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-07-09
- Filing Date
- 2020-07-31
- Publication Date
- 2026-01-06
- Estimated Expiration
- 2040-07-31
AI Technical Summary
In existing technologies, the information processing delay for detecting physical objects in image or video data is relatively long, resulting in low efficiency of real-time object recognition.
By performing coarse identification on the device and transmitting feature information to a server or cloud computing device for fine identification, combined with local neural networks and privacy standards, the computer-generated rendering environment can be updated rapidly.
It enables faster and more efficient real-time physical object identification and CGR environment updates, reduces computing resource consumption, and improves user experience.
Smart Images

Figure CN119045659B_ABST
Abstract
Description
[0001] This application is a divisional application of invention patent application 202010757211.5 entitled "Visual Search Refinement in Computer-Generated Rendering Environment" filed on July 31, 2020. Technical Field
[0002] The present invention relates to providing a computer-generated rendering (CGR) environment, and more particularly, to a system, method, and apparatus for identifying physical objects and updating the CGR environment based on information associated with the identified physical objects. Background Technology
[0003] Providing information about physical objects detected by a device (e.g., a portable electronic device) in image or video data is computationally expensive and time-consuming, typically resulting in significant processing delays associated with performing image or video data segmentation and object recognition. For example, a user of the device might walk into a room and experience a considerable delay before the device can identify all the objects in the room. Summary of the Invention
[0004] Various embodiments of this invention disclose devices, systems, and methods capable of faster and more efficient real-time physical object identification, information retrieval, and CGR environment updates. In some embodiments, a device (e.g., a mobile computing device or head-mounted display (HMD)) with a processor, display, and camera implements a method for acquiring images or videos of a physical setting via the camera. For example, the device may capture a real-time video stream. The device classifies physical objects into categories based on features associated with detected physical objects in the image or video and provides a CGR environment on the display according to those categories. In some embodiments, the device classifies physical objects based on a coarse identification, for example, based on relatively little information or using processing techniques with lower processing requirements than other fine-grained identification techniques. Furthermore, a specific portion of the image or video may be considered to contain the object. For example, the device may detect that an object is a laptop computer based on its shape, color, size, volume, markings, or any number of simple or complex features associated with the object using a local neural network, and may display a generic laptop computer from an online store based on classifying the physical object as a laptop computer. In some specific implementations, the device may, based on this category, utilize generic virtual content from the resource store residing on the device to replace, enhance, or supplement the physical object.
[0005] In some implementations, the device transmits a portion (e.g., some or all) of an image or video to a second device (e.g., a server, desktop computer, cloud computing device, etc.). For example, the device may send a portion of an image or video that includes the physical object to the second device. In some implementations, the device determines whether the image or video meets privacy standards. Furthermore, the device may determine whether to transmit the portion of the image or video that includes the physical object to the second device based on whether the privacy standards are met. In some implementations, the device transmits one or more detected features or identified categories to the second device.
[0006] In some implementations, the device receives a response associated with an object and updates the CGR environment on the display based on that response. This response is determined based on the identification of the physical object performed by the second device. The second device identifies the physical object based on a fine-grained identification, for example, based on relatively more information, or using processing techniques that require more processing than coarse identification techniques used for object classification. For example, the second device may access a stable object information database not available on the first device to identify a specific object or obtain specific information about that object. For example, the response associated with the object may include: identification of a subcategory (e.g., a laptop computer of brand X), identification of a specific type of item (e.g., model number), supplementary information (e.g., a user manual), an animation associated with the object, a 3D pose, a computer-aided drawing (CAD) model, etc. Having received a response from the second device, the device may modify the depiction of the object based on that identification (e.g., displaying a depiction of a laptop to match a specific model, displaying a user manual adjacent to the laptop, etc.), or may generate an experience associated with the object (e.g., triggering an animation, providing a 3D pose, etc.). In some implementations, the received response associated with the object includes identification data (e.g., vectors of attributes, descriptors, etc.) that the device will use to identify the physical object in future images or videos. Furthermore, in some implementations, the received response associated with the object includes an assessment of the physical object's condition (e.g., whether the device is damaged, undamaged, wet, dry, vertical, or horizontal, etc.).
[0007] In some implementations, the device can replace generic virtual content displayed based on the category with object-specific content based on the received responses associated with that object.
[0008] In some embodiments, a non-transitory computer-readable storage medium stores instructions that are computer-executable to perform or cause to perform any of the methods described herein. In some embodiments, an apparatus includes one or more processors, non-transitory memory, and one or more programs; the one or more programs are stored in the non-transitory memory and configured to be executed by the one or more processors, and the one or more programs include instructions for performing or causing to perform any of the methods described herein. Attached Figure Description
[0009] Therefore, this disclosure will be understood by those skilled in the art, and a more detailed description can be made with reference to some exemplary embodiments, some of which are shown in the accompanying drawings.
[0010] Figure 1 The CGR environment is shown as the physical environment set on the device's display according to some specific implementations.
[0011] Figure 2 It shows the implementation of some specific methods by Figure 1 The provided CGR environment.
[0012] Figure 3 The illustration shows a CGR environment on the device, according to some specific implementations, including object-specific content based on responses received from a second device associated with physical objects.
[0013] Figure 4 The diagram illustrates a CGR environment based on some specific implementations.
[0014] Figure 5 It is a block diagram based on some specific implementations of exemplary devices.
[0015] Figure 6 It is a block diagram based on some specific implementations of exemplary devices.
[0016] Figure 7 This is a flowchart illustrating an exemplary method for providing a CGR environment according to some specific implementations.
[0017] As is customary, the various features shown in the accompanying drawings may not be drawn to scale. Therefore, for clarity, the dimensions of various features may be arbitrarily expanded or reduced. Additionally, some drawings may not depict all components of a given system, method, or apparatus. Finally, similar reference numerals may be used throughout the specification and drawings to denote similar features. Detailed Implementation
[0018] Cross-reference to related applications
[0019] This application claims the benefit of U.S. Provisional Application Serial No. 62 / 881,476, filed August 1, 2019, the entire contents of which are incorporated herein by reference.
[0020] Numerous details have been described to provide a thorough understanding of the exemplary embodiments illustrated in the accompanying drawings. However, the drawings illustrate only some exemplary aspects of this disclosure and should not be considered limiting. Those skilled in the art will recognize that other effective aspects or variations do not include all the specific details set forth herein. Furthermore, well-known systems, methods, components, devices, and circuits have not been described exhaustively so as not to obscure further relevant aspects of the exemplary embodiments described herein.
[0021] refer to Figure 1 and Figure 2 This disclosure illustrates an exemplary operating environment 100 according to some specific embodiments. Although relevant features are shown, those skilled in the art will recognize from this disclosure that various other features are not shown for brevity and to avoid obscuring further relevant aspects of the exemplary embodiments disclosed herein. Therefore, as a non-limiting example, operating environment 100 includes device 120 (e.g., a mobile electronic device or head-mounted display (HMD)) and device 145 (e.g., a server, cloud computing device, personal computer, or mobile electronic device), one or both of which may be situated within a physical setting 105. Physical setting 105 refers to a world that an individual can perceive or interact with without the assistance of electronic systems. A physical setting (e.g., a physical forest) includes physical objects (e.g., physical trees, physical structures, and physical animals). An individual may directly interact with or perceive the physical setting, such as through touch, vision, smell, hearing, and taste.
[0022] In some implementations, device 120 is configured to manage and coordinate a user's computer-generated rendering (CGR) environment 125, and in some implementations, device 120 is configured to present the CGR environment 125 to the user. In some implementations, device 120 includes a suitable combination of software, firmware, or hardware. References are provided below. Figure 5 The device 120 is described in more detail.
[0023] In some embodiments, device 120 is a computing device located locally or remotely relative to physical set 105. In some embodiments, the functionality of device 120 is provided by or combined with a controller (e.g., a local server located within physical set 105 or a remote server located outside physical set 105). In some embodiments, device 120 is communicatively coupled to other devices or peripherals via one or more wired or wireless communication channels (e.g., Bluetooth, IEEE 802.11x, IEEE 802.16x, IEEE 802.3x, etc.).
[0024] In some specific implementations, when a user is present within the physical set 105, the device 120 presents the CGR environment 125 to the user. Compared to a physical set, a CGR environment refers to a set that is entirely or partially created by a computer, which an individual can perceive or interact with via electronic systems.
[0025] In some specific implementations, the CGR environment 125 provided includes virtual reality (VR). A VR scene is a simulated scene designed to include only computer-created sensory input for at least one sense. A VR scene includes multiple virtual objects that an individual can interact with or perceive. An individual can interact with or perceive virtual objects in a VR scene by simulating a subset of their own actions within the computer-created scene or by simulating their own presence or their presence within the computer-created scene.
[0026] In some specific implementations, the CGR environment 125 provided includes Mixed Reality (MR). An MR scene is a simulated scene designed to integrate computer-created sensory input (e.g., virtual objects) with sensory input or representations from a physical scene. In the reality spectrum, mixed reality scenes lie between (but do not include) VR scenes at one end and fully physical scenes at the other.
[0027] In some MR (Multi-Real-Time) scenes, computer-generated sensory input can adapt to changes in sensory input from the physical scene. Additionally, some electronic systems used to present MR scenes can monitor orientation or position relative to the physical scene, enabling virtual objects to interact with real objects (i.e., physical objects from the physical scene or their representations). For example, the system can monitor motion so that virtual plants appear stationary relative to physical buildings.
[0028] An example of mixed reality is augmented reality (AR). An AR scene refers to a simulated scene in which at least one virtual object is superimposed on a physical scene or its representation. For example, an electronic system may have an opaque display and at least one imaging sensor for capturing images or videos of the physical scene, which are representations of the physical scene. The system combines the images or videos with virtual objects and displays the combination on the opaque display. An individual uses the system to indirectly view the physical scene via images or videos of the physical scene and observes the virtual objects superimposed on the physical scene. When the system uses one or more image sensors to capture images of the physical scene and uses those images to present the AR scene on an opaque display, the displayed images are referred to as video pass-through. Alternatively, the electronic system for displaying the AR scene may have a transparent or semi-transparent display through which an individual can directly view the physical scene. The system may display virtual objects on the transparent or semi-transparent display, allowing an individual to observe the virtual objects superimposed on the physical scene using the system. As another example, the system may include a projection system that projects virtual objects onto the physical scene. Virtual objects can be projected, for example, onto a physical surface or as holograms, allowing individuals to use the system to observe virtual objects superimposed on a physical setting.
[0029] Augmented reality scenes can also refer to simulated scenes in which the representation of a physical scene is altered by sensory information created by a computer. For example, a portion of the representation of a physical scene can be graphically altered (e.g., magnified) such that the altered portion still represents one or more initially captured images, but is not a faithful reproduction. As another example, when providing video pass-through, the system can alter at least one of the sensor images to impose a specific viewpoint different from the viewpoint captured by one or more image sensors. Furthermore, the representation of a physical scene can be altered by graphically blurring or removing portions of it.
[0030] Another example of mixed reality is augmented virtual (AV). An AV scene refers to a computer-created or virtual scene incorporating at least one sensory input from a physical scene. The one or more sensory inputs from the physical scene can be a representation of at least one feature of the physical scene. For example, virtual objects can exhibit the colors of physical objects captured by one or more imaging sensors. Similarly, virtual objects can exhibit features consistent with actual weather conditions in the physical scene, such as those identified via weather-related imaging sensors or online weather data. In another example, an augmented reality forest can have virtual trees and structures, but the animals can have features accurately reproduced from images taken of physical animals.
[0031] Many electronic systems enable individuals to interact with or perceive various forms of mixed environments. One example includes a head-mounted system. A head-mounted system may have an opaque display and one or more speakers. Alternatively, the head-mounted system may be designed to receive external displays (e.g., smartphones). The head-mounted system may have one or more imaging sensors or microphones, respectively for capturing images / videos of a physical scene or capturing audio of a physical scene. The head-mounted system may also have a transparent or translucent display. The transparent or translucent display may be combined with a substrate through which light representing the image is guided to the individual's eyes. The display may incorporate LEDs, OLEDs, digital light projectors, laser scanning light sources, liquid crystal on silicon, or any combination of these technologies. The substrate transmitting light may be an optical waveguide, an optical combiner, a light reflector, a holographic substrate, or any combination of these substrates. In one specific implementation, the transparent or translucent display may selectively switch between an opaque state and a transparent or translucent state. As another example, the electronic system may be a projection-based system. A projection-based system may use retinal projection to project images onto the individual's retina. Alternatively, the projection system can also project virtual objects onto a physical scene (e.g., onto a physical surface or as a hologram). Other examples of SR systems include head-up displays, car windshields capable of displaying graphics, windows capable of displaying graphics, lenses capable of displaying graphics, headphones or earphones, speaker arrangements, input mechanisms (e.g., controllers with or without haptic feedback), tablet computers, smartphones, and desktop or laptop computers.
[0032] In some implementations, device 120 displays a CGR environment 125, allowing user 115 to simultaneously view a physical environment 105 on the display of device 120. For example, the device displays the CGR environment 125 in a real-world coordinate system with real-world content. In some implementations, such viewing modes include visual content that combines CGR content with real-world content of the physical environment 105. Furthermore, the CGR environment 125 may include video perspective (e.g., where real-world content is captured by a camera and displayed on the display along with a 3D model) or optical perspective content (e.g., where real-world content is viewed directly or through glass and supplemented by the display of a 3D model).
[0033] For example, the CGR environment 125 can provide user 115 with video perspective CGR on the display of a consumer cellular phone by integrating rendered three-dimensional (3D) graphics into a real-time video stream captured by an onboard camera. Alternatively, the CGR environment 125 can provide user 115 with optical perspective CGR by overlaying rendered 3D graphics onto a wearable perspective HMD, thereby electronically enhancing the user's optical view of the real world using the overlaid 3D model.
[0034] In some embodiments, the physical environment 10 includes at least one physical object 130. For example, the physical object 130 may be a consumer electronic device (e.g., a laptop computer), a piece of furniture or artwork, a photograph, etc. In some embodiments, the device 120 receives image or video data, detects the physical object 130, and identifies a portion of the image or video (e.g., the portion 135 of the image or video depicting the physical environment 105) as including the physical object 130.
[0035] In some implementations, device 120 classifies physical object 130 using a local neural network based on one or more simple or complex features (e.g., shape, color, size, volume, markings, etc.) associated with it, and provides a CGR environment 125 that combines physical environment 105 (e.g., locally captured images or videos via physical environment 105) with content corresponding to physical object 130. For example, device 120 may coarsely identify physical object 130 (e.g., a laptop computer) using a local neural network based on one or more simple or complex features such as shape, color, size, volume, markings, etc., associated with it. In some implementations, device 120 transmits one or more features (e.g., shape, color, size, volume, markings, etc.) associated with the physical object to device 145. For example, device 120 can identify a tag associated with physical object 30, transmit the tag to device 145, and device 145 can identify physical object 130 based on the tag (e.g., the tag can be associated with a specific brand and model of a laptop).
[0036] In some implementations, device 120 uses generic virtual content 140 to replace, enhance, or supplement physical objects 130 in CGR environment 125. For example, generic virtual content 140 (e.g., laptop identifier) can be obtained from a resource store on device 120 based on a determined category of physical object 130.
[0037] Various implementations enable device 120 to receive responses or additional information about physical object 130 associated with it. In some implementations, device 120 transmits a portion of an image or video (e.g., the portion 135 of an image or video depicting the physical environment 105 including physical object 130) to device 145 via link 150. For example, device 120 may send a portion of an image or video that includes a laptop computer to device 145 (e.g., a remote server, personal computer, or cloud computing device).
[0038] However, for privacy reasons, user 115 may not want data associated with physical environment 105 sent to device 145 (e.g., a remote server). Therefore, in some implementations, device 120 only sends image or video data associated with physical environment 105 to device 145 if the data meets predetermined privacy criteria. Furthermore, in some implementations, device 120 determines whether to send a portion of an image or video to device 145 by determining whether that portion meets the privacy criteria. Such privacy criteria may include whether the image or video contains a person, name, identification number, financial information, etc. In some implementations, user 115 can specify that they do not want image or video data sent to device 145. For example, privacy criteria may include whether configuration settings on device 120 are set to allow data to be sent to device 145. In some implementations, privacy criteria may include whether the data is associated with a specific category of physical objects or collected at a specific physical location (e.g., GPS location of a home or workplace). Many other privacy criteria, either user-specified or configured at the device level, can also be used to determine whether to send image or video data to device 145.
[0039] In some implementations, device 145 identifies physical object 130. For example, device 145 can identify physical object 130 at a more refined level than the coarse identification performed by device 120. Furthermore, device 120 can transfer classification categories to device 145. For example, device 120 can classify physical object 130 as a laptop, and device 145 can further identify physical object 130 (e.g., a specific brand and model of the laptop).
[0040] like Figure 1 As shown, according to some embodiments, device 120 is a mobile electronic device, and according to some embodiments, device 120 is an HMD configured to be worn on the head of user 115. Such an HMD may surround the user 115's field of view. Device 120 includes one or more screens or other displays configured to display a CGR environment 125. In some embodiments, device 120 includes one or more screens or other displays to display virtual elements having real-world content within the user 115's field of view. In some embodiments, device 120 is worn such that one or more screens are positioned to display a CGR environment 125 having real-world content of a physical environment 105 within the user 115's field of view. In some embodiments, device 120 providing the CGR environment 125 is configured to present a chamber, housing, or room containing the CGR environment 125, in which user 115 is not wearing or holding device 120.
[0041] like Figure 3 and Figure 4 As shown, according to some specific implementations, device 120 receives a response 160 associated with physical object 130, which is determined by device 145 identifying physical object 130. For example, the received response 160 may include identification of a subcategory, identification of a specific type item, or supplementary information.
[0042] In some implementations, information 160 is presented to user 115 within CGR environment 125. For example, information 160 regarding physical object 130 may include: identification of a subcategory (e.g., laptop computer), identification of a specific type of item (e.g., model number), supplementary information (e.g., user manual, maintenance information, auxiliary information, etc.), or an assessment of the condition of physical object 130. In some implementations, generic virtual content 140 is replaced with object-specific content based on information 160. For example, a generic laptop description may be changed to a description of a specific model.
[0043] In some implementations, device 120 may receive object recognition data from device 145, which will be used by device 120 to identify physical objects 130 in future images or videos of the physical environment 105. For example, device 120 may receive a 3D model, vector, or descriptor associated with the physical object 130 from device 145. In some implementations, device 120 may generate an experience based on a received response 160. For example, the received response 160 may enable or otherwise trigger device 120 to display an animation, 3D pose, or CAD model associated with the physical object 130.
[0044] like Figure 3 As shown, according to some embodiments, device 120 is a mobile electronic device, and according to some embodiments, device 120 is an HMD configured to be worn on the head of user 115.
[0045] Figure 5This is a block diagram of an example of a device 120 according to some specific implementations. Although some specific features are shown, those skilled in the art will recognize from this disclosure that various other features are not shown for the sake of brevity and in order not to obscure further relevant aspects of the specific implementations disclosed herein. Therefore, as a non-limiting example, in some specific implementations, device 120 includes one or more processing units 502 (e.g., microprocessors, ASICs, FPGAs, GPUs, CPUs, processing cores, etc.), one or more input / output (I / O) devices and sensors 506, one or more communication interfaces 508 (e.g., USB, Firewire, Thunderbolt, IEEE 802.3x, IEEE 802.11x, IEEE 802.16x, GSM, CDMA, TDMA, GPS, IR, BlueTooth, ZigBee, SPI, I2C, or similar interfaces), one or more programming (e.g., I / O) interfaces 510, one or more displays 512, one or more internal or external image sensors 514, memory 520, and one or more communication buses 504 for interconnecting these components and various other components.
[0046] In some embodiments, one or more communication buses 504 include circuitry for interconnecting and communicating between control system components. In some embodiments, the one or more I / O devices and sensors 506 include at least one of the following: an inertial measurement unit (IMU), an accelerometer, a magnetometer, a gyroscope, a thermometer, one or more physiological sensors (e.g., a blood pressure monitor, a heart rate monitor, a blood oxygen sensor, a blood glucose sensor, etc.), one or more microphones, one or more speakers, a haptic engine, or one or more depth sensors (e.g., structured light, time-of-flight, etc.).
[0047] In some embodiments, one or more displays 512 are configured to present a CGR environment to a user. In some embodiments, one or more displays 512 correspond to holographic, digital light processing (DLP), liquid crystal display (LCD), liquid crystal on silicon (LCoS), organic light-emitting field-effect transistor (OLET), organic light-emitting diode (OLED), surface-conducting electron emitter display (SED), field emission display (FED), quantum dot light-emitting diode (QD-LED), microelectromechanical system (MEMS), or similar display types. In some embodiments, one or more displays 512 correspond to waveguide displays such as diffraction, reflection, polarization, and holography. For example, device 120 includes a single display. As another example, device 120 includes displays for each of the user's eyes.
[0048] In some embodiments, the one or more image sensor systems 514 are configured to acquire image or video data corresponding to at least a portion of a user's face, including the user's eyes. For example, the one or more image sensor systems 514 may include one or more RGB cameras (e.g., having a complementary metal-oxide-semiconductor (CMOS) image sensor or a charge-coupled device (CCD) image sensor), monochrome cameras, IR cameras, event-based cameras, etc. In various embodiments, the one or more image sensor systems 514 may also include an illumination source, such as a flash or flash source, that emits light to that portion of the user's face.
[0049] Memory 520 includes high-speed random access memory, such as DRAM, SRAM, DDR RAM, or other random access solid-state memory devices. In some embodiments, memory 520 includes non-volatile memory, such as one or more disk storage devices, optical disk storage devices, flash memory devices, or other non-volatile solid-state storage devices. Memory 520 optionally includes one or more storage devices remotely located to one or more processing units 502. Memory 520 includes a non-transitory computer-readable storage medium. In some embodiments, memory 520 or the non-transitory computer-readable storage medium of memory 520 stores programs, modules, and data structures, or subsets thereof, including optional operating system 530 and CGR environment module 540.
[0050] Operating system 530 includes processes for handling various basic system services and for performing hardware-related tasks. In some implementations, CGR environment module 540 is configured to create, edit, or experience CGR environments. Content creation unit 542 is configured to create and edit CGR content that will be used as part of a CGR environment for one or more users (e.g., a single SR environment for one or more users, or multiple SR environments for a corresponding group of one or more users). Although these modules and units are shown residing on a single device (e.g., device 120), it should be understood that in other implementations, any combination of these modules and units may reside in separate computing devices.
[0051] also, Figure 5 More often, it serves as a functional description of various features present in a specific embodiment, which differ from the structural schematic diagram of the specific embodiment described herein. As those skilled in the art will recognize, items shown individually can be combined, and some items can be separated. For example, Figure 5Some functional modules shown individually can be implemented in a single module, and the various functions of a single functional block can be implemented in various specific implementations through one or more functional blocks. The actual number of modules and the division of specific functions, as well as how features are allocated therein, will vary depending on the specific implementation, and in some specific implementations, it depends in part on the specific combination of hardware, software, or firmware chosen for that particular implementation.
[0052] Figure 6 This is a block diagram of an example of a device 145 according to some specific implementations. Although some specific features are shown, those skilled in the art will recognize from this disclosure that various other features are not shown for the sake of brevity and in order not to obscure further relevant aspects of the specific implementations disclosed herein. Therefore, as a non-limiting example, in some specific implementations, device 145 includes one or more processing units 602 (e.g., microprocessors, ASICs, FPGAs, GPUs, CPUs, processing cores, etc.), one or more input / output (I / O) devices and sensors 606, one or more communication interfaces 608 (e.g., USB, Firewire, Thunderbolt, IEEE 802.3x, IEEE 802.11x, IEEE 802.16x, GSM, CDMA, TDMA, GPS, IR, BlueTooth, ZigBee, SPI, I2C, or similar interfaces), one or more programming (e.g., I / O) interfaces 610, one or more displays 612, one or more internal or external image sensors 614, memory 620, and one or more communication buses 604 for interconnecting these components and various other components.
[0053] In some embodiments, one or more communication buses 604 include circuitry for interconnecting and communicating between system components. In some embodiments, one or more I / O devices and sensors 606 include at least one of the following: an IMU, an accelerometer, a magnetometer, a gyroscope, a thermometer, one or more physiological sensors (e.g., a blood pressure monitor, a heart rate monitor, a blood oxygen sensor, a blood glucose sensor, etc.), one or more microphones, one or more speakers, a haptic engine, and / or one or more depth sensors (e.g., structured light, time-of-flight, etc.).
[0054] In some embodiments, one or more displays 612 are configured to present a CGR environment to a user. In some embodiments, one or more displays 612 correspond to holographic, DLP, LCD, LCoS, OLET, OLED, SED, FED, QD-LED, MEMS, or similar display types. In some embodiments, one or more displays 612 correspond to waveguide displays such as diffraction, reflection, polarization, and holography.
[0055] In some embodiments, the one or more image sensor systems 614 are configured to acquire image or video data corresponding to at least a portion of a user's face, including the user's eyes. For example, the one or more image sensor systems 614 may include one or more RGB cameras (e.g., having a complementary metal-oxide-semiconductor (CMOS) image sensor or a charge-coupled device (CCD) image sensor), monochrome cameras, IR cameras, event-based cameras, etc. In various embodiments, the one or more image sensor systems 614 may also include an illumination source, such as a flash or flash source, that emits light to that portion of the user's face.
[0056] Memory 620 includes high-speed random access memory, such as DRAM, SRAM, DDR RAM, or other random access solid-state memory devices. In some embodiments, memory 620 includes non-volatile memory, such as one or more disk storage devices, optical disk storage devices, flash memory devices, or other non-volatile solid-state storage devices. Memory 620 optionally includes one or more storage devices remotely located to one or more processing units 602. Memory 620 includes a non-transitory computer-readable storage medium. In some embodiments, memory 620 or the non-transitory computer-readable storage medium of memory 620 stores programs, modules, and data structures, or subsets thereof, including optional operating system 630, object database 640, and identification module 650.
[0057] In some embodiments, object database 640 includes characteristics (e.g., shape, color, size, volume, or markings) associated with a particular physical object or object category. For example, object database 640 may include any information associated with an object that facilitates the identification of the physical object. Furthermore, in some embodiments, object database 640 includes a library of models or virtual content associated with physical objects or categories of physical objects.
[0058] Operating system 630 includes processes for handling various basic system services and for performing hardware-related tasks. In some embodiments, identification module 650 is configured to identify physical objects. Object identification unit 652 is configured to access, create, and edit object identification data that will be used to identify or classify physical objects. Although these modules and units are shown residing on a single device (e.g., device 145), it should be understood that in other embodiments, any combination of these modules and units may reside in a separate computing device.
[0059] also, Figure 6More often, it serves as a functional description of various features present in a specific embodiment, which differ from the structural schematic diagram of the specific embodiment described herein. As those skilled in the art will recognize, items shown individually can be combined, and some items can be separated. For example, Figure 6 Some functional modules shown individually can be implemented in a single module, and the various functions of a single functional block can be implemented in various specific implementations through one or more functional blocks. The actual number of modules and the division of specific functions, as well as how features are allocated therein, will vary depending on the specific implementation, and in some specific implementations, it depends in part on the specific combination of hardware, software, or firmware chosen for that particular implementation.
[0060] Figure 7 This is a flowchart illustrating an exemplary method for providing a CGR environment according to some specific implementations. In some specific implementations, method 700 is provided by a device (e.g., Figures 1 to 6 The method 700 may be executed by device 120 or device 145. It may be executed on a mobile device, HMD, desktop computer, laptop computer, server device, or by multiple devices communicating with each other. In some embodiments, the method 700 is executed by processing logic components (including hardware, firmware, software, or a combination thereof). In some embodiments, the method 700 is executed by a processor that executes code stored in a non-transitory computer-readable medium (e.g., memory).
[0061] At box 702, in method 700, an image or video of the physical scene is acquired via a camera. In one example, the image or video is acquired based on user input or selection, and in another example, the image or video is acquired automatically by the device.
[0062] At box 704, method 700 classifies a physical object into a category based on features associated with a physical object detected in an image or video. In one example, the physical object is an electronic device (e.g., a laptop computer), and the classification is based on a coarse identification of any or a combination of features including shape, color, size, volume, markings, etc. Furthermore, in another example, a specific portion of an image or video is identified as containing a physical object.
[0063] At box 706, method 700 provides a computer-generated rendering (CGR) environment on the display based on this category. In one example, the CGR environment includes a general representation or identification of physical objects, as well as other depictions of the physical scene.
[0064] At box 708, method 700 transmits a portion of an image or one or more features associated with a physical object. In one example, an image or video, or a portion of an image or video, is transmitted to a remote server or cloud computing device. Alternatively, features associated with a physical object are transmitted to another local device. Furthermore, the transmission of any image or video data may be based on determining that the image or video meets privacy standards (e.g., user-defined privacy settings, identification of personal data, etc.). In one example, method 700 may also transmit the category determined at box 704.
[0065] At box 710, method 700 receives a response associated with an object, wherein the response is determined based on the identification of the physical object. In one example, the response is received from a remote or local device, and the response includes identification of a subcategory, identification of a specific type item (e.g., model number), supplementary information (e.g., user manual), or physical conditions. Furthermore, the response may generate an experience associated with the object (e.g., triggering an animation, providing a 3D pose, etc.).
[0066] At box 712, method 700 updates the computer-generated rendering (CGR) environment on the display based on the received response. In one example, method 700 changes a general representation or identification of a physical object to a more specific representation or identification of the physical object. Alternatively, method 700 displays additional information or resources that are close to the representation of the physical object. For example, the classification of the depicted physical object may be changed to the identification of the physical object. In yet another example, a user manual may be displayed adjacent to the physical object. Furthermore, method 700 may display animations associated with the object, such as 3D poses or CAD models.
[0067] The various specific implementations disclosed herein provide techniques for classifying and recognizing physical objects that save computational resources and enable faster processing. For example, a user wearing an HMD might want to obtain information about their surroundings. However, providing such information may require the HMD to perform segmentation and object recognition on the input video stream, which is computationally expensive and time-consuming. Therefore, if a user enters a room and experiences a considerable delay before anything is recognized, it results in a poor user experience.
[0068] However, depending on the specific implementation, coarse object identification is performed on the HMD to quickly provide basic information about the objects. Furthermore, the coarse object identification is sent to a server or cloud computing device, where more detailed analysis is performed to obtain and return a detailed identification. The server or cloud computing device may also send information to the HMD that will facilitate future object identification by the HMD. For example, information sent by the HMD to the server or cloud computing device may include a simplified version of a visual search. Information received by the HMD from the server or cloud computing device may include a learned version of an abstract representation of the objects or a Simultaneous Localization and Mapping (SLAM) point cloud mixed with meta-information. Additionally, the HMD may maintain memory persistence for objects owned by a specific user, or in other words, associated with a specific user.
[0069] In one example, the device can store information about previously identified physical objects for later use or reference. In an exemplary use case, a retail store user can identify physical objects such as tables. The user can then check whether the table is suitable for their home or office based on the pre-identification of physical objects in their home or office.
[0070] Furthermore, in MR applications, the view of a physical object can be replaced by a virtual object. Initially, the physical object can be replaced by a generic version of the object, but subsequently, the view can be refined using a more specific 3D representation of the object based on information received from a server or cloud computing device.
[0071] This document sets forth numerous specific details to provide a comprehensive understanding of the claimed subject matter. However, those skilled in the art will understand that the claimed subject matter can be practiced without these specific details. In other instances, methods, apparatus, or systems known to a person of ordinary skill in the art have not been described in detail so as not to obscure the claimed subject matter.
[0072] Unless otherwise specifically stated, it should be understood that throughout this specification, discussions using terms such as “processing,” “calculating,” “computing,” “determining,” and “identifying” refer to the actions or processes of computing devices, such as one or more computers or similar electronic computing devices, which manipulate or convert data representing physical electronic or magnetic quantities within the memory, registers, or other information storage, transmission, or display devices of a computing platform.
[0073] The one or more systems discussed herein are not limited to any particular hardware architecture or configuration. A computing device may include any suitable arrangement of components that provide results conditioned on one or more inputs. Suitable computing devices include computer systems based on multi-purpose microprocessors that access stored software that programs or configures the computing system from a general-purpose computing device to a special-purpose computing device that implements one or more specific embodiments of the subject matter of this invention. The teachings contained herein can be implemented in the software used for programming or configuring the computing device using any suitable programming, scripting, or other type of language or combination of languages.
[0074] Specific implementations of the methods disclosed herein can be performed in the operation of such computing devices. The order of the boxes presented in the above examples can be changed; for example, the boxes can be reordered, grouped, or divided into sub-boxes. Some boxes or procedures can be executed in parallel.
[0075] The use of “applies to” or “configured to” in this document implies open and inclusive language, which does not exclude applicability to or configuration to devices performing additional tasks or steps. Similarly, the use of “based on” implies openness and inclusivity, as processes, steps, calculations, or other actions “based on” one or more of the stated conditions or values may in practice be based on additional conditions or values beyond those stated. The headings, lists, and numbering included in this document are for illustrative purposes only and are not intended to be restrictive.
[0076] It will also be understood that while terms such as "first," "second," etc., may be used in this document to describe various objects, these objects should not be limited by these terms. These terms are merely used to distinguish one object from another. For example, a first node can be called a second node, and similarly, a second node can be called a first node, changing the meaning of the description, provided that all occurrences of "first node" are consistently renamed and all occurrences of "second node" are consistently renamed. First nodes and second nodes are both nodes, but they are not the same node.
[0077] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the claims. As used in the description of these embodiments and the appended claims, the singular forms “a” and “the” are intended to also cover the plural forms unless the context clearly indicates otherwise. It will also be understood that the term “or” as used herein refers to and covers any and all possible combinations of one or more of the associated listed items. It will also be understood that the terms “comprising” or “including” as used in this specification specify the presence of the stated features, integers, steps, operations, objects, or components, but do not exclude the presence or addition of one or more other features, integers, steps, operations, objects, components, or groups thereof.
[0078] As used herein, the term "if" can be interpreted as meaning "when the prerequisite is true" or "when the prerequisite is true" or "in response to determination" or "according to determination" or "in response to detection" that the prerequisite is true, depending on the context. Similarly, the phrases "if it is determined [the prerequisite is true]" or "if [the prerequisite is true]" or "when [the prerequisite is true]" are interpreted as meaning "when it is determined that the prerequisite is true" or "in response to determination" or "according to determination" that the prerequisite is true or "when the prerequisite is detected" or "in response to detection" that the prerequisite is true, depending on the context.
[0079] The foregoing description and summary of the present invention should be understood as illustrative and exemplary in every respect, and not restrictive, and the scope of the invention disclosed herein is determined not only by the detailed description of the illustrative specific embodiments, but also by the full extent permitted by patent law. It should be understood that the specific embodiments shown and described herein are merely illustrative of the principles of the invention, and various modifications can be made by those skilled in the art without departing from the scope and spirit of the invention.
Claims
1. A method comprising: at a first device having a processor: detecting a physical object based on obtained image content; determining generic virtual content associated with the physical object based on detected features associated with the physical object; providing the generic virtual content within a computer-generated rendered (CGR) environment; transmitting a portion of the obtained image content or the detected features to a second device; receiving a response from the second device, the response based on an identification of the physical object performed by the second device; and updating the CGR environment by replacing the generic virtual content with object-specific content.
2. The method of claim 1, wherein, providing the generic virtual content within the CGR environment includes replacing, augmenting, or supplementing the physical object with the generic virtual content.
3. The method of claim 1, wherein, the generic virtual content is determined from a resource store residing on the first device.
4. The method of claim 1, wherein, the object-specific content is based on information about the physical object.
5. The method of claim 1, wherein, the generic virtual content associated with the physical object is determined based on a coarse identification of the physical object.
6. The method of claim 1, wherein, the portion of the obtained image content does not include a second portion of the obtained image content.
7. The method of claim 1, wherein, the detected features associated with the physical object are a shape, a color, a size, a volume, or a marking.
8. The method of claim 1, further comprising: determining whether the obtained image content satisfies a privacy criterion; and determining the portion of the obtained image content based on satisfaction of the privacy criterion.
9. The method of claim 1, further comprising: determining whether to transmit the portion of the obtained image content based on a privacy criterion.
10. The method of claim 1, further comprising: transmitting the detected features to the second device.
11. The method of claim 1, wherein, the response associated with the physical object includes identification data to be used by the first device to identify the physical object in future image content.
12. The method of claim 1, wherein, the response associated with the physical object includes an assessment of a condition of the physical object.
13. The method of claim 1, wherein, the first device or the second device is a head-mounted device (HMD).
14. The method of claim 1, wherein, the first device or the second device is a mobile electronic device.
15. A system comprising: a first device having a processor; and a computer-readable storage medium comprising instructions that, when executed by the processor, cause the system to perform operations comprising: detecting a physical object based on obtained image content; determining generic virtual content associated with the physical object based on detected features associated with the physical object; providing the generic virtual content within a computer-generated rendered (CGR) environment; transmitting a portion of the obtained image content or the detected features to a second device; receiving a response from the second device, the response based on an identification of the physical object performed by the second device; and updating the CGR environment by replacing the generic virtual content with object-specific content.
16. The system of claim 15, wherein, providing the generic virtual content within the CGR environment includes replacing, augmenting, or supplementing the physical object with the generic virtual content.
17. The system of claim 16, wherein, the generic virtual content is determined from a resource store that resides on the first device.
18. The system of claim 16, wherein, the object-specific content is based on information about the physical object.
19. The system of claim 15, wherein, the generic virtual content associated with the physical object is determined based on a coarse identification of the physical object.
20. A non-transitory computer-readable storage medium storing program instructions executable by a processor for performing operations comprising: detecting, by a first device, a physical object based on obtained image content; determining, based on detected features associated with the physical object, generic virtual content associated with the physical object; providing the generic virtual content within a computer-generated rendered CGR environment; transmitting, to a second device, a portion of the detected features or the obtained image content; receiving a response from the second device, the response based on an identification of the physical object performed by the second device; and updating the CGR environment by replacing the generic virtual content with object-specific content.
Citation Information
Patent Citations
Method and system for providing virtual display of a physical environment
CN107850936A
KR20190045679A