Methods, computer systems, and computer-readable media for providing virtual coaching

By integrating cameras and processors into electronic devices, real-time rendering of user gestures, and combining this with a server-client interaction model, the problem of low efficiency in user interaction with home appliances in a virtual environment is solved, enabling intuitive virtual guidance and product recommendations.

CN115016638BActive Publication Date: 2025-11-18MIDEA GROUP CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210556151.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2018-10-15
Publication Date
2025-11-18
Estimated Expiration
2038-10-15

AI Technical Summary

Technical Problem

Existing online sales platforms cannot meet users' needs to try out home appliances, understand their functions, and interact with them virtually. Traditional VR/AR technologies are inefficient and user input detection is not intuitive.

Method used

By integrating cameras, processors, and memory into electronic devices, the system renders user gestures in real-time within augmented and virtual reality environments, providing virtual assistance templates and virtual guidance, and combining server-client interaction models for virtual image processing and rendering.

Benefits of technology

It enables users to have an intuitive interactive experience with home appliances in a virtual environment, improves the efficiency of VR/AR technology, and provides virtual guidance and product recommendations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115016638B_ABST
    Figure CN115016638B_ABST
Patent Text Reader

Abstract

A method of providing virtual guidance, comprising: capturing one or more images of a physical environment using one or more cameras; rendering a 3-D virtual environment in real-time based on the one or more images as the camera captures the images; capturing a first gesture in the physical environment through the one or more cameras; in response to the first gesture: converting the first gesture into a first operation that displays a virtual aid template associated with a physical object in the virtual environment; rendering the virtual aid template in real-time on a display in association with the physical object adjacent to a represented location of the physical object in the 3-D virtual environment; capturing a second gesture in the physical environment through the one or more cameras; in response to the second gesture; converting the second gesture into a first interaction with the representation of the physical object in the 3-D virtual environment; determining a second operation on the virtual aid template based on the first interaction with the representation of the physical object; and rendering the second operation on the virtual aid template in real-time on the display.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Case Analysis

[0002] This application is a divisional application filed within the scope of Chinese patent application No. 201880097904.7, filed on October 15, 2018, entitled "System and Method for Providing Real-Time Product Interaction Assistance". Technical Field

[0003] This application relates to household appliances, and more particularly to methods for providing virtual guidance, systems for providing virtual guidance, and computer-readable storage media. Background Technology

[0004] In today's era of widespread e-commerce, many product suppliers invest heavily in developing and using online sales platforms that display images of listed products and promote sales by providing product descriptions, online reviews, and informatized videos on individual product pages. While online sales platforms also provide a channel for selling home appliances, conventional platforms fail to meet users' desires to try out appliances, understand their functions, and virtually interact with them or watch them operate in a virtual environment that mimics their expected home environment. Virtual / augmented reality (VR / AR) environments include at least some virtual elements that represent or enhance corresponding objects in the physical world. However, traditional VR / AR technologies are inefficient and fail to integrate virtual products well into the virtual environment. Furthermore, user input used to detect user interactions with virtual objects (e.g., detected by various sensors) is not intuitive and is inefficient.

[0005] Therefore, there is a need for an effective and more intuitive method or system to provide real-time virtual experiences associated with user and object interactions. Summary of the Invention

[0006] Therefore, there is a need for computer systems with improved methods and interfaces to render user interactions in real-time within augmented and virtual reality (VR / AR) environments using user gestures. These methods and interfaces can optionally complement or replace conventional methods for interacting with VR / AR environments. The computer systems disclosed herein can reduce or eliminate the aforementioned deficiencies and other problems associated with user interfaces used for VR / AR. For example, the methods and interfaces provide users with a vivid virtual experience of interacting with one or more objects in an AR / VR environment using gestures. The methods and interfaces also provide users with virtual auxiliary templates displayed simultaneously with a virtual view of the product to help place the product or assemble the product using the user's gestures.

[0007] In some embodiments, this disclosure provides a method for providing virtual guidance (e.g., on-site troubleshooting and repair). The method includes: in an electronic device (e.g., a user device) having a display, one or more cameras, one or more processors, and memory, using one or more images captured by the one or more cameras, the physical environment including physical objects arranged at a first location; while the one or more cameras capture the one or more images, rendering a three-dimensional (3-D) virtual environment in real time based on the one or more images in the physical environment, wherein the 3-D virtual environment includes representations of the physical objects in the virtual environment at locations corresponding to the first location in the physical environment; capturing a first gesture in the physical environment via the one or more cameras; and responding to the capture of the first gesture by the one or more cameras: The first gesture is converted into a first operation that displays a virtual auxiliary template associated with the physical object in the virtual environment; the virtual auxiliary template associated with the physical object and its position adjacent to the physical object's representation in the 3D virtual environment is rendered on the display in real time; a second gesture is captured in the physical environment by the one or more cameras; in response to the second gesture being captured by the one or more cameras: the second gesture is converted into a first interaction with the physical object's representation in the 3D virtual environment; a second operation is determined on the virtual auxiliary template associated with the physical object based on the first interaction with the physical object's representation; the second operation on the virtual auxiliary template associated with the physical object is rendered on the display in real time.

[0008] According to some embodiments, an electronic device includes a display; one or more cameras; one or more processors; and a memory storing one or more programs; said one or more programs are configured to be executed by said one or more processors, said one or more programs including instructions for performing or causing operation of any method of this disclosure. According to some embodiments, a computer-readable storage medium storing instructions thereon, when executed by an electronic device, causes the device to perform or cause operation of any method of this disclosure. According to some embodiments, an electronic device includes means for performing or causing operation of any method of this disclosure. According to some embodiments, an information processing device for an electronic device includes means for performing or causing operation of any method of this disclosure.

[0009] The various additional advantages of this application will be apparent from the following description. Attached Figure Description

[0010] The foregoing features and advantages, as well as their additional features and advantages, of the disclosed technology will become clearer below through a detailed description of preferred embodiments taken in conjunction with the accompanying drawings.

[0011] To more clearly describe the embodiments of the present disclosure or the technical solutions in the prior art, the accompanying drawings required for describing the embodiments or the prior art are briefly introduced below. Obviously, the drawings in the following description only illustrate some embodiments of the present disclosure, and those skilled in the art can still derive other drawings from these drawings without creative effort.

[0012] Figure 1 This is a block diagram illustrating an operating environment that provides a real-time virtual experience and virtual guidance for user interaction with objects, based on some embodiments.

[0013] Figure 2A This is a block diagram of a server system according to some embodiments.

[0014] Figure 2B This is a block diagram of a client device according to some embodiments.

[0015] Figures 3A to 3D This is a flowchart of a method for providing a real-time virtual experience for user interaction with objects, according to some embodiments.

[0016] Figures 4A to 4L Examples of systems and user interfaces that provide a real-time virtual experience for user interaction with objects, according to some embodiments, are shown.

[0017] Figure 5 This is a flowchart of a method for providing virtual guidance for user-object interaction according to some embodiments.

[0018] Figures 6A to 6E Examples of systems and user interfaces that provide virtual guidance for user interactions with objects, according to some embodiments, are shown.

[0019] In all the accompanying drawings, similar reference numerals refer to the corresponding parts. Detailed Implementation

[0020] Reference will now be made in detail to embodiments, examples of which are illustrated in the accompanying drawings. Numerous specific details are set forth in the following detailed description to provide a thorough understanding of the subject matter presented herein. However, it will be apparent to those skilled in the art that the subject matter can be practiced without these specific details. In other instances, well-known methods, processes, components, and circuits have not been described in detail to avoid unnecessarily obscuring various aspects of the embodiments.

[0021] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. The described embodiments are only a part of, and not all of, the embodiments of the present application. All other embodiments obtained by those skilled in the art based on the embodiments of the present application without creative effort are within the scope of protection of the present invention.

[0022] like Figure 1 As shown, according to some embodiments, a virtual image processing and rendering system 100 (e.g., including a server system 104 and one or more user devices 102) is implemented based on a server-client interaction model. According to some embodiments, the server-client interaction model includes client modules (not shown) executing on corresponding user devices 102-1, 102-2…102-n, which are deployed in various locations (e.g., physical stores, roadshow booths, product demonstration sites, product testing sites, product design and manufacturing sites, showrooms, users' homes, users' kitchens, users' offices, and scenarios for on-site machine troubleshooting). In some embodiments, the server-client interaction model also includes various server-side modules 106 (also referred to as "back-end modules 106") executing on the server system 104. The client modules (not shown) communicate with the server modules 106 via one or more networks 110. The client modules provide client-side functionality to the virtual image processing and rendering system 100 and communicate with the server-side modules 106. Server-side module 106 provides server-side functionality for a virtual image processing and rendering system 100 that exists on any number of client modules on user equipment 102 (e.g., user's mobile phone 102-1, head-mounted display (HMD) 102-2... user's tablet computer 102-n, etc.).

[0023] In some embodiments, server system 104 includes one or more processing modules 106 (e.g., including but not limited to image processing modules, 3D rendering modules, gesture analysis modules, recommendation modules, measurement modules, and troubleshooting modules), one or more processors 112, one or more databases 130 storing data and models (e.g., gesture data and gesture recognition models, facial expression data and facial expression recognition models, troubleshooting data and machine error recognition models, customer transaction data, user profile data, and product recommendation models), input / output (I / O) interfaces 118 for one or more client devices 102, and I / O interfaces 120 for one or more external services (not shown) (e.g., machine manufacturers, component suppliers, e-commerce or social networking platforms) or other types of online interaction (e.g., users interacting with online sales / marketing channels (e.g., e-commerce applications or social networking applications 105) on their respective user devices 103 (e.g., smartphones, tablets, and PCs in retail stores). In some embodiments, the I / O interface 118 for the client module facilitates client input and output processing for the client module on the respective client device 102 and the module on the in-store device 103. In some embodiments, one or more server-side modules 106 utilize various real-time data obtained through various internal and external services, real-time data received from the client device (e.g., acquired image data), and existing data stored in various databases to render 3D virtual images while interacting with virtual object gestures, and / or use gestures to guide the user to interact with virtual auxiliary templates and generate product recommendations for the user at various deployment locations of the user device 102 (e.g., in the user's home or in a store).

[0024] Examples of user equipment 102 include, but are not limited to, cellular phones, smartphones, handheld computers, wearable computing devices (e.g., HMDs, head-mounted displays), personal digital assistants (PDAs), tablets, laptops, desktop computers, Enhanced General Packet Radio Service (EGPRS) mobile phones, media players, navigation devices, game consoles, televisions, remote controls, point-of-sale (POS) terminals, in-vehicle computers, e-book readers, field computer kiosks, mobile sales robots, humanoid robots, or any combination of two or more of these data processing devices or other data processing devices. (See reference...) Figure 2BThe user equipment 102 discussed may include one or more client modules for performing functions similar to those discussed in server module 106. The user equipment 102 may also include one or more databases storing various types of data, similar to database 130 in server system 104.

[0025] Examples of one or more networks 110 include Local Area Networks (LANs) and Wide Area Networks (WANs) such as the Internet. Optionally, one or more networks 110 may be implemented using any known network protocol, including various wired or wireless protocols such as Ethernet, Universal Serial Bus (USB), FireWire, Long Term Evolution (LTE), Global System for Mobile Communications (GSM), Enhanced Data Rate for GSM Evolution (EDGE), Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Bluetooth, Wi-Fi, Voice over Internet Protocol (VoIP), Wi-MAX, or any other suitable communication protocol.

[0026] In some embodiments, server system 104 is implemented on a distributed network of one or more independent data processing devices or computers. In some embodiments, server system 104 also employs various virtual devices and / or services from third-party service providers (e.g., third-party cloud service providers) to provide the underlying computing resources and / or infrastructure resources of the backend information exchange system 108. In some embodiments, server system 104 includes, but is not limited to, handheld computers, tablets, laptops, desktop computers, server computers, or any combination of two or more of these data processing devices or other data processing devices.

[0027] In some embodiments, the server system 104 also implements various modules for supporting user interaction and providing product recommendations to users. In some embodiments, the server system 104 includes services based on various statistical techniques, rule-based techniques, and artificial intelligence techniques, such as audio / video processing services, natural language processing services, model building services, statistical analysis services, data mining services, data collection services, and product recommendation services.

[0028] Figure 1 The illustrated virtual image processing and rendering system 100 includes a client portion (e.g., a client module on client device 102) and a server portion (e.g., server module 106). In some embodiments, data processing is implemented as a standalone application installed on client device 102, which is deployed at a location physically displaying multiple real-world products (e.g., home appliances, furniture, heavy equipment, and vehicles, etc.), where a user is physically located and interacts directly with the client device and the products. In some other embodiments, the user is located away from the physical location displaying multiple real-world products (e.g., a user is conducting online virtual shopping from home). Furthermore, the functional division between the client and server portions of the virtual image processing and rendering system 100 can vary in different embodiments. For example, in some embodiments, the client module is a simplified client that only provides user interface input (e.g., capturing user gestures using a camera) and output (e.g., image rendering) processing functions, delegating all other data processing functions to a backend server (e.g., server system 104). Although many aspects of this technology have been described from the perspective of the backend system, it will be apparent to those skilled in the art that the corresponding actions are performed by the frontend system without any inventive effort. Similarly, although many aspects of this technology have been described from the perspective of the client system, it will be apparent to those skilled in the art that the corresponding actions are performed by the backend server system without any inventive effort. Furthermore, some aspects of this technology can be performed by the server, client devices, or a combination of both. In some embodiments, databases storing various types of data are distributed across multiple locations local to certain frontend systems, enabling faster data access and reduced local data processing time.

[0029] Figure 2A A block diagram of a representative server system 104 according to some embodiments is shown. Server 104 typically includes one or more central processing units (CPUs) 202 (e.g., Figure 1The server 104 includes a processor 112, one or more network interfaces 204, memory 206, and one or more communication buses 208 (sometimes referred to as chipsets) for interconnecting these components. Optionally, the server 104 also includes a user interface 201. The user interface 201 includes one or more output devices 203 capable of presenting media content, including one or more speakers and / or one or more visual displays. The user interface 201 also includes one or more input devices 205, which include user interface components that facilitate user input, such as a keyboard, mouse, voice command input unit or microphone, touch screen display, touch-sensitive input pad, gesture capture camera, or other input buttons or controls. The memory 206 includes high-speed random access memory, such as dynamic random access memory (DRAM), static random access memory (SRAM), double data rate synchronous dynamic random access memory (DDR RAM), or other random access solid-state storage devices; optionally, it includes non-volatile memory, such as one or more disk storage devices, one or more optical disk storage devices, one or more flash memory devices, or one or more other non-volatile solid-state storage devices. Optionally, memory 206 includes one or more storage devices remote from one or more processing units 202. Memory 206 or the non-volatile memory within memory 206 includes a non-transitory computer-readable storage medium. In some embodiments, memory 206 or the non-transitory computer-readable storage medium of memory 206 stores programs, modules, and data structures, or subsets or supersets thereof:

[0030] • Operating system 210, which includes programs for handling various basic system services and performing hardware-related tasks;

[0031] • Network communication module 212, which is used to connect server 104 to other computing devices (e.g., client device 102 or third-party service network) connected to one or more networks 110 via one or more network interfaces 204 (wired or wireless);

[0032] • Presentation module 213, which is used to present information (e.g., application user interface, widgets, web pages, audio and / or video content, text, etc.) at server 104 through one or more output devices 203 (e.g., display, speakers, etc.) associated with user interface 210;

[0033] • Input processing module 214, which is used to detect one or more user inputs or interactions from one or more input devices 205 and interpret the detected inputs or interactions;

[0034] • One or more applications 216 that are executed by server 104;

[0035] • Server-side module 106 that provides server-side data processing and functions, including but not limited to:

[0036] Image processing module 152 is used to process user gesture data, facial expression data, object data and / or camera data during calibration and real-time virtual image rendering. The image processing module can perform real-time image segmentation, real-time depth analysis, object position / movement analysis, etc.

[0037] Augmented Reality (AR) and Virtual Reality (VR) processing and rendering module 222, the augmented reality (AR) and virtual reality (VR) processing and rendering module is used to generate AR and VR experiences for users based on products or virtual representations of products that users interact with, products recommended to users, products requested by users, and user characteristics, preferences, interaction styles, etc.

[0038] The gesture analysis module 224 is used to analyze gesture data based on gesture data, position / depth data and contour data to identify various gestures. The gesture analysis module 224 can also build gesture models based on gesture data obtained through a calibration process, and these gesture analysis modules can update these gesture models during real-time virtual image processing and rendering.

[0039] The recommendation module 226 is used to recommend products based on factors such as product, space, environmental size, appearance, color, theme, and user facial expressions, and to build and maintain a corresponding recommendation model using appropriate data.

[0040] Measurement module 228, the measurement module is used to measure the size of one or more objects, spaces and environments (e.g., a user's kitchen) using camera data (e.g., depth information) and / or image data (comparing the number of pixels of an object of known size with that of an object of unknown size);

[0041] The troubleshooting module 230 is used to identify product errors / defects using various models, establish and maintain troubleshooting models based on common problems with machine error characteristics, and select maintenance guidance to be provided to facilitate user maintenance.

[0042] other modules used to perform the other functions described above in this disclosure;

[0043] • A server-side database 130 for storing data and related models, including but not limited to:

[0044] о Gesture data (e.g., including but not limited to hand contour data, hand position data, hand size data, and hand depth data associated with various gestures) acquired by a camera and processed by an image processing module, and gesture recognition model 232 (e.g., established during calibration and updated when users interact with a virtual environment in real time using gestures);

[0045] о Facial expression data and facial expression recognition model constructed based on user facial expression data for one or more products 234;

[0046] Troubleshooting data, including image data related to mechanical errors, malfunctions, electronic component defects, circuit errors, compressor failures, etc., as well as problem identification models (e.g., machine error identification models);

[0047] User transaction and profile data 238 (e.g., customer name, age, income level, color preference, previously purchased products, product category, product bundle / bundle, previously queried products, past delivery locations, interaction channels, interaction locations, purchase time, delivery time, special requests, identity data, demographic data, social relationships, social network account names, social network posts or comments, interaction records with sales representatives, customer service representatives or delivery personnel, likes, dislikes, emotions, beliefs, superstitions, personality, temperament, interaction style, etc.); and

[0048] Recommendation model 240 includes various types of recommendation models, such as size-based product recommendation models, user facial expression-based product recommendation models, and user data and purchase history-based recommendation models.

[0049] Each element of the aforementioned identifier can be stored in one or more of the aforementioned storage devices and corresponds to a set of instructions for performing the aforementioned functions. The modules or programs (i.e., instruction sets) of the aforementioned identifiers do not need to be implemented as separate software programs, processes, modules, or data structure programs; therefore, various subsets of these modules can be combined or otherwise rearranged in various implementations. In some embodiments, memory 406 may optionally store subsets of the modules and data structures of the aforementioned identifiers. Furthermore, memory 406 may optionally store additional modules and data structures not described above.

[0050] In some embodiments, at least some functions of the server system 108 are performed by the client device 104, and corresponding submodules of these functions may reside within the client device 104 rather than within the server system 108. The client device 104 and server system 108 shown in the figures are merely illustrative, and different configurations of modules for implementing the functions described in this disclosure are possible in various embodiments.

[0051] While specific embodiments have been described above, it should be understood that they are not intended to limit this application to these specific embodiments. Rather, this application includes alternatives, modifications, and equivalents within the spirit and scope of the appended claims. Numerous specific details have been set forth to provide a thorough understanding of the subject matter presented in this application. However, it will be apparent to those skilled in the art that the subject matter can be practiced without these specific details. In other instances, well-known methods, processes, components, and circuits have not been described in detail to avoid unnecessarily obscuring various aspects of the embodiments.

[0052] Figure 2BA block diagram of a representative client device 102 according to some embodiments is shown. Client device 102 typically includes one or more processing units (CPUs) 252 (e.g., processors), one or more network interfaces 254, memory 256, and one or more communication buses 258 for interconnecting these components (sometimes referred to as chipsets). Client device 102 also includes a user interface 260. User interface 260 includes one or more output devices 262 for presenting media content, including one or more speakers and / or one or more visual displays. User interface 260 also includes one or more input devices 264, including user interface components that facilitate user input, such as a keyboard, mouse, voice command input unit or microphone, touchscreen display, touch-sensitive input pad, gesture capture camera, one or more cameras, depth camera, or other input buttons or controls. Furthermore, some client devices 102 use microphones and voice recognition or cameras and gesture recognition as a supplement to or alternative to a keyboard. In some embodiments, client device 102 also includes sensors that provide background information about the current state of client device 102 or environmental conditions associated with client device 102. Sensors include, but are not limited to, one or more microphones, one or more cameras, ambient light sensors, one or more accelerometers, one or more gyroscopes, GPS positioning systems, Bluetooth or BLE systems, temperature sensors, one or more motion sensors, one or more biosensors (e.g., skin conductance sensors, pulse oximeters, etc.), and other sensors. Memory 256 includes high-speed random access memory, such as DRAM, SRAM, DDR RAM, or other random access solid-state memory devices; optionally, it includes non-volatile memory, such as one or more disk storage devices, one or more optical disk storage devices, one or more flash memory devices, or one or more other non-volatile solid-state storage devices. Optionally, memory 256 includes one or more storage devices remote from one or more processing units 252. Memory 256 or the non-volatile memory within memory 256 includes non-transitory computer-readable storage media. In some embodiments, memory 256 or the non-transitory computer-readable storage media of memory 256 stores programs, modules, and data structures, or subsets or supersets thereof:

[0053] • Operating system 266, the operating system including programs for handling a variety of basic system services and performing hardware-related tasks;

[0054] • Network communication module 268, which is used to connect client device 102 to other computing devices (e.g., server system 104) connected to one or more networks 110 via one or more network interfaces 254 (wired or wireless);

[0055] • Presentation module 270, which is used to present information at client device 102 (e.g., a user interface for presenting text, images, videos, web pages, audio, etc.) via one or more output devices 262 (e.g., displays, speakers, etc.) associated with user interface 260;

[0056] • Input processing module 272, which is used to detect one or more user inputs or interactions from one or more input devices 264 and interpret the detected inputs or interactions;

[0057] • One or more applications 274 executed by client device 102 (e.g., payment platforms, media players, and / or other World Wide Web (Web) based or non-Web based applications);

[0058] • Client module 106 provides client data processing and functionality, including but not limited to:

[0059] Image processing module 275, which is used to process user gesture data, facial expression data, object data and / or camera data, etc.;

[0060] Augmented Reality (AR) and Virtual Reality (VR) processing and rendering module 276, which generates AR and VR experiences for users based on virtual representations of products or products with which users interact;

[0061] o Gesture analysis module 277, which is used to analyze gesture data based on gesture data, position / depth data and contour data to recognize various gestures;

[0062] The recommendation module 278 is used to recommend products based on factors such as product, space, environmental dimensions, appearance, color, theme, and user facial expressions.

[0063] Measurement module 279 is used to measure the dimensions of one or more objects, spaces, and environments (e.g., a user's kitchen) using camera data (e.g., depth information) and / or image data (comparing the number of pixels of an object of known size with that of an object of unknown size);

[0064] The troubleshooting module 280 is used to identify product errors / defects using various models and select the repair guidance to provide to facilitate user repair.

[0065] other modules used to perform the other functions described above in this disclosure; and

[0066] • Client database 271 used to store data and models, including but not limited to:

[0067] о Gesture data (e.g., including but not limited to hand contour data, hand position data, hand size data, and hand depth data associated with various gestures) acquired by a camera and processed by an image processing module, and gesture recognition model 281 (e.g., established during calibration and updated when users interact with the virtual environment in real time using gestures);

[0068] о Facial expression data and facial expression recognition model constructed based on user facial expression data for one or more products 282;

[0069] Troubleshooting data, including image data related to mechanical errors, faults, electronic component defects, circuit errors, compression faults, etc., as well as problem identification models 283 (e.g., machine error identification models);

[0070] User transaction and profile data 284 (e.g., customer name, age, income level, color preference, previously purchased products, product category, product bundle / bundle, previously queried products, past delivery locations, interaction channels, interaction locations, purchase time, delivery time, special requests, identity data, demographic data, social relationships, social network account names, social network posts or comments, interaction records with sales representatives, customer service representatives or delivery personnel, likes, dislikes, emotions, beliefs, superstitions, personality, temperament, interaction style, etc.); and

[0071] Recommendation model 285 includes various types of recommendation models, such as size-based product recommendation models, user facial expression-based product recommendation models, and user data and purchase history-based recommendation models.

[0072] Each element of the aforementioned identifier can be stored in one or more of the aforementioned storage devices and corresponds to a set of instructions for performing the aforementioned functions. The modules or programs (i.e., instruction sets) of the aforementioned identifiers do not need to be implemented as separate software programs, processes, modules, or data structure programs; therefore, various subsets of these modules can be combined or otherwise rearranged in various implementations. In some embodiments, memory 256 may optionally store subsets of the modules and data structures of the aforementioned identifiers. Furthermore, memory 256 may optionally store additional modules and data structures not described above.

[0073] In some embodiments, at least some functions of the server system 104 are performed by the client device 102, and corresponding submodules of these functions may reside within the client device 102 rather than within the server system 104. The client device 102 and server system 104 shown in the figures are merely illustrative, and different configurations of modules for implementing the functions described in this disclosure are possible in various embodiments.

[0074] In some embodiments, image processing module 220 or 275 includes multiple machine learning models for analyzing images (e.g., images of gestures) from one or more cameras and providing parameters derived from image analysis of the images (e.g., the outline of a user's hand, hand size, hand shape, hand movement, depth information). Optionally, in some embodiments, the image processing module includes some components local to the client device 102 and some remote components on the server 104. In some embodiments, the image processing module resides entirely on the server 104.

[0075] In some embodiments, the virtual image processing and rendering system 100 continuously collects image data (e.g., data related to user gestures during calibration processes and real-time virtual image rendering and user interaction processes), processes the image data, and mines the data to improve the accuracy of models, statistics, and intelligent decision-making. During specific interactions with customers, the virtual image processing and rendering system 100 utilizes feedback and information received from individual customers to modify the selection and prioritization of models and decision logic used to generate predictions, interactions, and recommendations, thereby improving the speed and efficiency of data processing and enhancing the accuracy and effectiveness of predictions, interactions, and recommendations. For example, individual users' facial expressions, reactions, emotions, and intentions (e.g., obtained via captured images including facial expressions, gestures, postures, etc.) are fed back to the virtual image processing and rendering system 100 in real time to add additional parameters to the analysis, prediction, and recommendation models, or to reselect the model set used to perform analysis, prediction, and recommendation and / or redirect intelligent decision / logic, etc. (e.g., model deletion, replacement, and / or addition), etc.

[0076] In some embodiments, the recommendation module or troubleshooting module uses various artificial intelligence techniques to build models. For example, the corresponding module integrates knowledge and conclusions from different data sources and analytical methods, such as various machine learning algorithms and specially designed decision logic and algorithms, and / or combinations thereof (e.g., various types of neural networks, deep neural networks, search and optimization analysis, rule-based decision logic, probabilistic methods and fuzzy logic, Bayesian networks, hidden Markov models, classifiers and statistical learning models, cybernetics, etc.) to determine product recommendations or identify machine errors, and uses the above to identify subsets of models and analytical tools to further generate appropriate responses for users, and to provide the most relevant recommendations as quickly as possible using as few computational resources as possible.

[0077] In some embodiments, the virtual image processing and rendering system 100 is configured to provide augmented reality and / or virtual reality experiences (e.g., using various AR / VR technologies). In some embodiments, user reactions to AR and VR experiences (e.g., verbal and facial expressions) are processed, and the results are then used to modify product recommendations and / or AR and VR experiences. For example, if a user initially requests to try a first model of a washing machine (which has a virtual reality setting) and is unable to figure out how to use the washing machine correctly (e.g., manipulating multiple buttons and parts of the virtual washing machine without a clear purpose for more than a certain amount of time) and exhibits frustration (e.g., facial expressions captured by a camera), the virtual image processing and rendering system 100 uses this information as new input and generates a new recommendation for another model with simplified functionality but similar characteristics (e.g., similar colors and sizes). Alternatively, if the user has a profile including facial expression data indicating that the user is generally satisfied with products having many features, the virtual image processing and rendering system 100 recommends products that the user was previously satisfied with. In some embodiments, the client device 102 generates a virtual representation of a physical environment (e.g., a kitchen) within the AR / VR environment, and simultaneously generates a virtual representation of the user's gestures within the AR / VR environment. In some embodiments, the virtual image processing and rendering system 100 also generates virtual auxiliary templates to demonstrate how to repair, use, or install products in the AR / VR environment. In some embodiments, the virtual image processing and rendering system 100 enables a user to visually experience one or more home appliances recommended to the user in a simulated home built within the AR / VR environment.

[0078] Figures 3A to 3DThis is a flowchart of method 300 according to some embodiments, which performs real-time image processing on user gestures captured by a camera and simultaneously renders the representation of the user gestures and the movement caused by the user's hand interacting with virtual objects in a virtual environment. In some embodiments, one or more steps of method 300 are performed by a server system (e.g., Figure 1 The server system 104) performs the method. In some embodiments, one or more steps of method 300 are performed by the client device 102 (e.g., Figure 1 The method is executed in a smartphone 102-1, HMD 102-2, or tablet computer 102-n. In some embodiments, method 300 is controlled by instructions stored in a non-transitory computer-readable storage medium, and these instructions are executed by one or more processors of the client and / or server system. See below for reference. Figures 4A to 4L The user interface (UI) of method 300 is discussed.

[0079] In such Figure 3A In some embodiments shown, method 300 includes: in an electronic device (e.g., a client device, such as a mobile phone 102-1, a head-mounted display (HMD) 102-2, or a tablet computer 102-n) having a display, one or more cameras, one or more processors, and memory: based on one or more images of the physical environment (e.g., ... Figure 4A Images 402-1…402-n in the image are used to render a 3D virtual environment on the display (e.g., Figure 4A The virtual image 404 in the image is shown. In some embodiments, the 3-D virtual environment includes physical objects arranged in the physical environment (e.g., virtual images 404). Figure 4A One or more of the characteristics of 406-1, 406-2 and 406-3 in the model.

[0080] In some embodiments, such as Figure 4AAs shown, a 3D virtual environment 404 can be rendered in one or more virtual reality (VR) images, VR videos (including multiple image frames), one or more augmented reality (AR) images, and AR videos. In some embodiments, taking virtual shopping on an electronic device for appliances in a physical environment (e.g., a user's kitchen) as an example, one or more images 402-1…402-n are one or more photographs of the user's kitchen, or a video containing multiple image frames showing various items in the kitchen. In some embodiments, the user is located in a physical environment, for example, in the user's kitchen at home, and holds the electronic device 102 to capture one or more images of the kitchen using one or more cameras of the electronic device. In some embodiments, the electronic device 102 has the ability to process images and generate AR / VR images of the 3D virtual environment. In some embodiments, the electronic device 102 works in conjunction with a server system 104 to process images and generate AR / VR images of the 3D virtual environment. For example, the electronic device 102 captures images and uploads them to the server 104 for processing and generating AR / VR images. The generated AR / VR images are then downloaded to the electronic device 102 for display to the user. In some alternative embodiments, the user is located away from the kitchen. For example, a user is in a physical store selling various kitchen utensils. In one example, the user can take images (e.g., photos or videos) of their kitchen at home before leaving. In another example, the user can ask someone else to take images of their kitchen at home, which are then uploaded via a link to a server (e.g., server system 104 with image processing and rendering modules) for processing and to generate a VR / AR image 404. The VR / AR image 404 can then be sent to the store's electronic device 102 for display to the user. Figure 4A As shown, the 3D virtual environment 404 includes one or more representations of physical objects in the kitchen, cabinets 406-1 and 406-2 with corresponding work surfaces, cups 406-4 arranged on the work surfaces, and a wine cabinet 406-3. Although not shown, one or more representations of the physical objects may also include one or more kitchen utensils, such as a stove, microwave oven, etc.

[0081] Referring back to Figure 3, method 300 further includes receiving (304) user input to arrange a first preset virtual object (e.g., preset virtual object 412-1) in a space of the 3D virtual environment 404 corresponding to a space in the physical environment (e.g., the user's kitchen). Figure 4A (space 408 in the 3-D virtual environment). Method 300 further includes: in response to user input, rendering (306) a first preset virtual object arranged in the space of the 3-D virtual environment.

[0082] In some embodiments, the input device of electronic device 102 may be used to receive user input. For example, input may be received directly on a touchscreen (e.g., from a selection in a displayed product catalog) or on a physical button of the electronic device. In some alternative embodiments, such as... Figure 4B As shown, user input is a gesture 410 captured by one or more cameras 409 (e.g., stereo camera, depth camera, time-of-flight (ToF) camera) or any other type of imaging sensor capable of measuring depth information from electronic device 102. In some embodiments, gesture 410 (e.g., from position 410-1 to position 410-2) is a predetermined gesture. Alternatively, gesture 410 indicates a pick-up and put-down action in interaction with a virtual product catalog 410 displayed on the user interface. For example, selecting a first preset virtual object from a product catalog 410 comprising multiple preset virtual objects (e.g., a 3-D virtual image of a first model of refrigerator 412-1, a 3-D virtual image of a second model of refrigerator 412-2, and a 3-D virtual image of island countertop 412-3). In some embodiments, representations of user gestures 414-1 to 414-2 are displayed in real time in a 3-D virtual image 404. For example, as... Figure 4B As shown, the camera 409 of the electronic device 102 captures real-time user gestures 410-1 to 410-2. The 3-D virtual environment 404 displays a representation of the user's gestures. A 3-D virtual image of a first model of refrigerator 412-1 is selected from the virtual catalog 412, and the virtual refrigerator 412-1 is arranged (414-2) in the space 408 between cabinets 406-1 and 406-2 in the virtual environment 404, the space corresponding to the physical space between two corresponding cabinets in the kitchen. In some embodiments, the orientation of the virtual refrigerator 412-1 can also be adjusted manually or automatically to align with the representation of the space and physical object 406.

[0083] Next, refer to again Figure 3A Method 300 includes using one or more cameras ( Figures 4C to 4F The camera 409 in the middle detects (308) the 3-D virtual environment (e.g., Figure 4C The first preset virtual object (e.g., in the 3-D virtual image 404) in the image 404 Figure 4C User gestures for interaction with the virtual refrigerator (412-1) (e.g., Figure 4C Gestures in (416). Figure 4C In some embodiments shown, the user's gesture 416 corresponds to a hand movement for opening the upper door of the refrigerator. In response to detecting (310) a user gesture, method 300 includes transferring the user's gesture (e.g., ...) to a register. Figure 4C The gesture 416) is converted (312) into a first preset virtual object in the 3D virtual environment (e.g., Figure 4C Interactions with the virtual refrigerator (412-1) in the system (e.g., Figure 4C The virtual hand interaction 418 for opening the door of the virtual refrigerator. In the 3-D virtual environment, the recognition (314) of the virtual environment (e.g. Figure 4C The first preset virtual object in the 3-D virtual environment (404) (e.g., Figure 4C The first part of the virtual refrigerator 412-1 in the system (e.g., Figure 4C In the on-site 420), the first part of the first preset virtual object is affected by the first preset virtual object in the virtual environment (e.g., Figure 4C Interactions with the virtual refrigerator (412-1) in the system (e.g., Figure 4C The method 300 also includes simultaneously displaying (316): (318) the representation of the user's gesture in the 3-D virtual environment (e.g., virtual hand interaction 418) on the display in real time. Figure 4C (318) virtual hand interaction in the 3D virtual environment; and (320) movement of the first part of the first preset virtual object caused by the interaction 418 with the first preset virtual object 412-1 in the 3D virtual environment (e.g., Figure 4C The virtual door on the virtual refrigerator rotates 420 degrees to open.

[0084] In some embodiments, the user gesture includes moving the user's hand from a first position to a second position. Camera 409 captures the gesture 416, and server 104 or electronic device 102 processes the image to calculate changes in the user's hand position, as well as changes in its contour and size. In some embodiments, depth information associated with the hand movement may be determined based on the captured image. Based on the determined depth information, electronic device 102 renders a 3D virtual image in a 3D virtual environment to show the user's hand appearing in front of an object (e.g., ...). Figure 4E In the middle, the representation of hand 428 is arranged in front of the representation of cup 406-4) or located behind the object (e.g., Figure 4F In the middle, the representation of hand 432 is arranged behind the representation of cup 406-4.

[0085] In some embodiments, user gestures include interactions between the user's hand and a preset virtual object or a portion thereof. Figure 3C (338), for example, opening the door of a virtual refrigerator and allowing the user to reach into the compartments to see how easy and stressful it is to place or retrieve items within the virtual refrigerator compartments. For example, as... Figures 4C to 4D As shown, a user's gestures include using his or her hand to interact with virtual objects displayed in the 3-D virtual environment (e.g., a virtual refrigerator 412-1). For example, as... Figure 4CAs shown, camera 409 captures the user's gesture 416 in mid-air, and this gesture is determined to be opening the door 420 of the virtual refrigerator 412-1 (418). In... Figure 4D In another example shown, one or more cameras 409 capture a user's gesture 422, which is determined to be extending (424) the user's hand further away from the one or more cameras 409 and into the upper compartment of the virtual refrigerator 412-1. This system can provide a vivid virtual user experience of using a refrigerator in one's own kitchen without physically placing the refrigerator in the kitchen and interacting with it.

[0086] In some embodiments, user gestures include using the user's hand to pin and move virtual objects from a first position in the virtual kitchen (e.g., ...). Figure 4A The space 408 in the middle is moved to the second position (e.g., Figure 4A The space on the left side of the middle cabinet 406-1, or Figure 4A The space to the right of cabinet 406-2 in the middle section is used to view the virtual results in a 3D virtual environment. This helps to provide users with direct visual results to evaluate the placement of the refrigerator in different locations and orientations in the kitchen, without having to place the refrigerator in the kitchen and then actually move it around to test different locations.

[0087] In some embodiments, a user can use gestures to place different types of virtual products into space 408, such as within the same space 408 ( Figure 4A In the process, the preset virtual refrigerator 412-1 is swapped with different virtual objects (e.g., virtual refrigerators 412-2 with different colors and / or sizes). Figure 3C (336 in the text). Users can also use gestures to swap the virtual refrigerator 412-1 with different types of virtual products (such as virtual stoves) to see the appropriate result.

[0088] In some embodiments, user gestures include initiation (e.g., Figure 3C(340) The function of a product corresponding to a preset virtual object. For example, in response to a user's gesture, the electronic device causes the user's hand to rotate a knob or press a button on the virtual product, which triggers the virtual product to perform the corresponding function in the virtual environment, such as preheating an oven. In some embodiments, the user can use gestures to exchange a portion of the representation of a physical object with a virtual part that can replace a portion of the physical object (e.g., a refrigerator compartment, a stove display panel). This virtual function testing and / or virtual part exchange can provide users with a direct visual effect of the product design, such as panel button design, handle design, etc., before building a real product or product demonstration. Therefore, users do not have to build multiple product demos with different knob shapes, colors, and sizes, thus saving time, material, and labor costs and improving the user experience.

[0089] In some embodiments, based on real-time image processing (e.g., Figures 2A to 2B The image processing module 220 and / or 275 performs the conversion (312) of the user's gestures into interactions with a preset virtual object. In some embodiments, prior to real-time user interaction with the virtual object using gesture processing, the electronic device 102 (or the electronic device 102 cooperating with the server system 104) performs a calibration process that is also customized for the individual user. For example, the user's hand is marked, or the user wears gloves marked with multiple feature points for marking the user's hand to be captured by one or more cameras 409. In some embodiments, these feature points are used to define the contour of the user's hand. Various parameters of the hand contour, such as area, perimeter, centroid, bounding box, and / or other suitable parameters, can be analyzed to understand the user's gestures. In some embodiments, changes in the shape of the user's hand contour can be analyzed to determine the action of the virtual object, such as opening the door or lid of a virtual object (e.g., a rice cooker or refrigerator). In some embodiments, detected changes in the position of the user's hand are used to determine the movement path of the user's hand (e.g., including the distance and displacement between a first position and a second position). In some embodiments, the size variation of the user's hand can be used to determine depth data related to the movement of the user's hand, in conjunction with other types of data (e.g., camera data (e.g., depth-related information), changes in hand shape and / or hand position). In some embodiments, depth data can be obtained based on camera data and / or changes in hand size and position. For example, it can be determined whether the user's hand is in front of an object in the 3-D virtual environment 404 (e.g., a representation of a virtual or physical object in a 3-D virtual image) by comparing the depth data of the user's hand with the depth data of a physical object in the 3-D virtual environment or the depth data of a virtual object. Figure 4E In the middle, hand 426 and the corresponding hand representation 428 are arranged in front of and behind the representation of cup 406-4 (e.g., Figure 4F In the middle, the hand 430 and the corresponding hand representation 432 are arranged behind the representation of the cup 406-4 or inside the representation of the virtual object or physical object.

[0090] In some embodiments, during calibration, the user may be instructed to perform a set of predetermined gestures within a predetermined distance from the camera, such as clenching a fist, holding a ball, pulling a handle, and opening a door. Relationships are established and stored between the predetermined gestures and corresponding datasets associated with the user's gestures captured by one or more cameras. For example, the dataset associated with the predetermined gestures may include data on the user's hand position, contour region data, contour shape factor, and depth data when performing the predetermined gestures. In some embodiments, such data, alone or in combination with other gesture data, may be further used to build machine learning models to analyze and determine various user gestures (e.g., ...). Figures 2A to 2B Gesture model 232 or 281 in the text.

[0091] In some embodiments, following the calibration process, when a user virtually experiences interaction with a virtual product using an electronic device (e.g., virtual shopping or virtual product design and testing experience), one or more cameras on the electronic device capture the user's gestures in real time. The contours defined by the markings on the user's hand are analyzed in real time to perform image segmentation. Then, based on the user gesture data obtained from the calibration process, the user's hand interactions with the virtual object can be determined in real time.

[0092] In some embodiments, the representation of the user's gestures and the movement of a preset virtual object or a portion thereof caused by the interaction are simultaneously displayed in real time on the display of the electronic device. In some embodiments, the identified hand and the preset virtual product are integrated to render an updated 3D virtual image of the user's hand interacting with the virtual product, such as when the user opens a refrigerator door (e.g., ...). Figure 4C The user reaches out and places the fruit into the refrigerator compartment (e.g., Figure 4D In some embodiments, user gestures are analyzed solely based on real-time image processing. For example, depth information, size changes, positional changes, and magnitude changes of the hand are analyzed based on image data and camera data, without the need for additional sensors.

[0093] Reference Figure 3B Method 300 also includes displaying (322) one or more representations of the physical object in a 3-D virtual environment 404 (e.g., Figure 4G The dimensional data associated with the characterization of cabinets 406-1 and 406-2 (e.g., Figure 4GIn some embodiments, the size data associated with a corresponding representation of a physical object corresponds to the size of the corresponding physical object in the physical environment. In some embodiments, the size data is obtained based on image information and / or camera data, wherein the camera data is associated with one or more cameras that capture one or more images. In some embodiments, the size data is obtained from one or more images captured by one or more cameras (e.g., metadata of the image, including length, width, and depth). In some embodiments, the camera may be the camera of electronic device 102 (e.g., when a user is in their kitchen at home), or the camera of another device located away from the user, for example, when another person takes a picture of the kitchen while the user is in a store. In some embodiments, the person taking the image of the kitchen may use an application that provides measuring tools to measure the dimensions (e.g., length, width, and depth) of the physical environment. For example, the application may display a scale in the image that translates to actual dimensions in the physical environment. In another example, the application may identify a physical object in the kitchen with a known size (e.g., an existing product) as a reference (e.g., by retrieving product specifications from a database) and compare the known size (e.g., by pixel count) with one or more representations of the physical object in the image to determine the size of other physical objects. Figure 4G In some of the embodiments shown, size data can be displayed while rendering (302) a 3-D virtual environment based on the physical environment.

[0094] In some embodiments, method 300 further includes simultaneously updating (324) a first preset virtual object (e.g., in real time within a 3-D virtual environment). Figure 4H The size data of the virtual refrigerator 412-1) and its proximity to the space (e.g., Figure 4H The dimensional data associated with one or more representations of physical objects in space 408 (e.g., cabinet representations 406-1 and 406-2) are used to arrange the first preset virtual object according to the interaction with the first preset virtual object caused by user gestures (representations 414-1 to 410-2 of user gestures 410-1 to 410-2). In some embodiments, such as Figure 4H As shown, only the relevant dimensions of the physical objects are displayed, not all dimensions of all objects. For example, if a user places a virtual refrigerator between two workbenches, the distance between adjacent edges of the workbenches is displayed, and the height of adjacent objects can also be shown. However, the length of the workbenches does not necessarily need to be displayed.

[0095] In some embodiments, the electronic device may also simultaneously display a size description of the virtual object, or a representation of a physical object that changes simultaneously with the user's interaction with the virtual object. For example, when a user gestures to change the orientation of a virtual object to place it in space, different surfaces or edges of the virtual object may be displayed, along with corresponding size data for the displayed surfaces or edges. In some embodiments, size data is displayed when a user gestures to pick up a virtual object from a virtual catalog, or when a user gestures to drag a virtual object toward space 408 and bring it sufficiently close to the space. In some embodiments, the electronic device may scan an area in the 3-D virtual environment to generate or highlight one or more surfaces or spaces marked with relevant dimensions for arranging virtual products. In some embodiments, measurement modules 228 or 279 ( Figures 2A to 2B Measurements can be calculated based on distances (or the number of pixels) in an image.

[0096] In some embodiments, method 300 further includes: displaying (326) a virtual placement result of a first preset virtual object in a 3-D virtual environment based on dimensional data associated with one or more representations of a physical object. Figure 4I As shown, when the virtual refrigerator 412-1 is narrower than the width of the space 408 between the representations of cabinets 406-1 and 406-2, the virtual gaps 434 and 436 between the virtual refrigerator and the respective cabinets are highlighted (e.g., in color or bold) to inform the user of the mismatch. In some embodiments, a virtual placement result is displayed when a preset virtual object is rendered (306) in response to user input. In some embodiments, a virtual installation result is displayed when a preset virtual object is rendered (316) in response to one or more user gestures (e.g., when virtual objects are placed in multiple spaces in the kitchen, or multiple different virtual objects are placed in a specific space in the kitchen). In some embodiments, when a user selects one or more devices from a virtual product catalog to place in a specific space, the user can drag the virtual product closer to or further away from the specific space, and a visual virtual placement result can be displayed when the virtual product is near the specific space. In some embodiments, when a specific space cannot accommodate one or more virtual products from the catalog (e.g., the virtual product is too wide for the space), such unplaceable virtual products in the virtual product catalog will be displayed as unsuitable for placement in the space (e.g., a gray shadow on the screen). Figure 4I (412-3 in the middle).

[0097] exist Figure 3BIn some embodiments shown, method 300 further includes selecting (328) one or more products to be placed in one or more spaces in the physical environment from a preset product database, based on the dimensions of one or more physical objects and the physical environment, without receiving any user input. In some embodiments, electronic device 102, operating independently or in conjunction with server system 104, automatically recommends products to the user based on the dimensions of the user's kitchen and the dimensions of one or more existing household appliances and furniture in the kitchen. In some embodiments, recommendation module 226 or 278 ( Figures 2A to 2B The system selects recommended products based on the size, color, and / or style of existing physical objects (e.g., adjacent cabinets) in the physical environment (e.g., a kitchen) and the size, color matching, style matching, theme matching, and user interaction (e.g., detecting which space in the kitchen the user wants to place the products in and the direction the user wants to arrange them in). In some embodiments, the recommendation module ( Figures 2A to 2B It can also refer to the user's historical purchase data, customized preference data, budget and / or other suitable user data.

[0098] Method 300 further includes: updating (330) the 3-D virtual environment to display one or more preset virtual objects of the selected one or more products (for recommended products) in one or more spaces of the 3-D virtual environment corresponding to one or more spaces of the physical environment. Figure 4J As shown, in some embodiments, the electronic device 102 can render a 3-D virtual environment to show a visual result 442 of placing a recommended product (e.g., a virtual refrigerator 438) in a user's kitchen (e.g., a virtual refrigerator arranged between cabinets), and match the recommended product with other items in the kitchen (e.g., representations of cabinets 406-1 and 406-2) in the 3-D virtual view. The electronic device can further display a note 440 "This refrigerator is perfect for this space and is for sale" to promote the recommended product. In some embodiments, relevant dimensions, such as the dimensions of the recommended virtual product 438 and the dimensions of the space 408 where the virtual product is placed, are also displayed in the 3-D virtual view 404. In some embodiments, a measurement module (e.g., Figures 2A to 2B The system scans the kitchen area to obtain dimensions, then automatically compares these dimensions with a product database to display suitable products (e.g., appliances and / or furniture) that can be placed in or match the kitchen space. In some embodiments, the measurement module and the recommendation module ( Figures 2A to 2B The kitchen area is further scanned to generate one or more surfaces for product placement (e.g., creating an island worktop with a sink and / or stove in the center of the kitchen, adding a cabinet with a work surface to house a microwave oven). In some embodiments, the measurement module (e.g., Figures 2A to 2B Measurements can be calculated based on depth-related camera data (e.g., focal length, depth data) from a camera that captured images of the kitchen and / or depth-related image data (e.g., number of pixels) in the image.

[0099] Reference Figure 3D In some embodiments, method 300 further includes: while simultaneously displaying (342) representations of the user's gestures and movement of a first portion of a first preset virtual object caused by interaction with the first preset virtual object in the 3D virtual environment, in response to the user viewing the movement of the first portion of the first preset virtual object caused by interaction with the first preset virtual object, acquiring (344) one or more facial expressions of the user. In some embodiments, facial expressions (e.g., Figure 4K Facial expressions (444) can also be used with user input or user gestures (e.g., Figure 4B Users can view the placement of virtual products in a space, or use user gestures (e.g., Figures 4C to 4D The method 300 further includes: recognizing (346) a negative facial expression of the user (e.g., an unhappy, frustrated, regretful, or disgusted face) in response to a first movement of a first portion of the first preset virtual object caused by the user viewing the interaction with the first preset virtual object in the 3-D virtual environment. The method 300 further includes: automatically selecting (348) a second preset virtual object from a preset product database without receiving any user input; updating (350) the 3-D virtual environment to display the second preset virtual object in the space of the 3-D virtual environment, replacing the first preset virtual object. In some embodiments, the second preset virtual product is displayed in the 3-D virtual environment 404 in response to user confirmation, or automatically without any user input.

[0100] In some embodiments, a user's facial expressions are captured by pointing the camera 409 of the electronic device 102 at the user's face. In some other embodiments, facial expressions are captured by one or more cameras of another device. Facial expression data is stored in a preset database 234 or 282 (e.g., Figures 2A to 2B In this database, the data can be customized for a single user or used to store facial expression data from multiple users. Machine learning algorithms can be used to build facial expression models relating to the relationship between user facial expressions and user reactions / preferences to various products (e.g., like, dislike, neutral, no reaction, happy, excited, regretful, disgusted, etc.). Figure 4KIn some embodiments shown, when a user is viewing a virtual product 412-1 arranged (446) in a 3-D virtual environment 404, a dissatisfied face 444 of the user is captured. Electronic device 102, or working in conjunction with server system 104, can identify that the user dislikes the product. In response to detecting negative feedback from the user, such as... Figure 4L As shown, the recommendation module ( Figures 2A to 2B Based on users' previous feedback on other products (e.g., Figure 4L Feedback (reflected by a positive facial expression 448 associated with virtual product 412-2) may be used to recommend another virtual product (e.g., a different model of virtual refrigerator 412-2). A virtual product may also be recommended because its size (and / or color, style) would better fit into a specific space in the kitchen. In some embodiments, the virtual placement result 452 is rendered to provide the user with a direct visual experience. In some embodiments, the electronic device 102 also displays comments 450 (e.g., “You liked this refrigerator last time; I think it would be more suitable for your kitchen”) to provide suggestions to the user.

[0101] Figure 5 This is a flowchart of a method 500 for rendering a virtual auxiliary template associated with a physical object based on user gestures that interact with the representation of the physical object in a virtual environment, according to some embodiments. In some embodiments, one or more steps of method 500 are performed by a server system (e.g., Figure 1 The server system 104 in the system performs the method. In some embodiments, one or more steps of method 500 are performed by client device 102 (e.g., Figure 1 This is performed on a smartphone 102-1, HMD 102-2, or tablet 102-n. In some embodiments, method 500 is controlled by instructions stored in a non-transitory computer-readable storage medium, and these instructions are executed by one or more processors of the client and / or server system. See below for reference. Figures 6A to 6E The discussion of user interface (UI) methods in 500.

[0102] In some embodiments, method 500 can be used for on-site troubleshooting of a faulty machine. In some embodiments, method 500 can be used to demonstrate assembling multiple components into a piece of furniture. In some embodiments, method 500 can be used to demonstrate how to use, for example, a device with multiple complex functions. Figure 5 In some embodiments shown, method 500 includes: in an electronic device having a display, one or more cameras, one or more processors, and memory (e.g., such as mobile phone 102-1, head-mounted display).

[0103] In client device 102 (HMD) 102-2 or tablet 102-n): using one or more cameras (e.g., Figure 6A The camera 609 captures (502) one or more images (e.g., one or more photographs or video containing multiple image frames) of a physical environment (e.g., kitchen 600), said physical environment including physical objects arranged at the first location (e.g., a damaged refrigerator 602). Figure 6A As shown, the field of view of camera 609 includes at least a portion of the kitchen of the damaged refrigerator 602.

[0104] Method 500 includes: when one or more cameras acquire one or more images, rendering (504) a 3-D virtual environment (e.g., a damaged refrigerator in a kitchen) in real time based on one or more images of a physical environment. Figure 6A 3D virtual images in (604). In such Figure 6A In some embodiments shown, the 3-D virtual environment 604 includes a representation of the physical object in the virtual environment at a location corresponding to a first location in the physical environment 600 (e.g., Figure 6A Characterization of refrigerator 612 in the middle).

[0105] Method 500 also includes: via one or more cameras (e.g., Figure 6A One or more cameras 609) capture (506) a first gesture (e.g., in the physical environment (e.g., kitchen 600) in the physical environment (e.g., kitchen 600). Figure 6A Gesture 606). In some embodiments, the first gesture is a trigger event for triggering a virtual auxiliary display. In some embodiments, method 500 further includes: in response to the first gesture being captured by one or more cameras (508): displaying the first gesture (e.g., Figure 6A The gesture 606) is converted (510) into a display virtual environment (e.g., Figure 6A In the 3-D virtual environment (604), the physical object (e.g., Figure 6A The refrigerator 602 in the middle) is associated with a virtual auxiliary template (e.g., Figure 6A The first operation of the virtual auxiliary template 616 in the text (e.g., Figure 6A (Unscrew the screws on the rear panel to remove the cover); render in real time on the monitor (512) with physical objects in a 3-D virtual environment (e.g., Figure 6A The representation in the 3-D virtual environment (604) (e.g., Figure 6A The virtual auxiliary template associated with the physical objects adjacent to the location of the refrigerator 612 (representation in the image) is (e.g., Figure 6A Virtual auxiliary template 616 in the middle.

[0106] In some embodiments, the electronic device 102, which may work in conjunction with the server system 104, can process images captured by the camera 609 to understand the user's gestures. For example, such as Figure 6A As shown, the camera 609 of the electronic device 102 captures a user gesture 606. After analyzing the captured image, the user's gesture 606 is identified as loosening a screw to remove the back cover of the lower compartment of the refrigerator. In some embodiments, the first gesture (e.g., loosening a screw to remove the back cover of the refrigerator) is a system-predetermined gesture or a user-customized gesture associated with displaying a virtual assistive template. In some embodiments discussed in this application, as the camera 609 captures the gesture 606 in the kitchen, a representation of the gesture 614 is rendered in real time in a 3-D virtual environment 604. In some embodiments, the electronic device simultaneously renders the representation 614 of the first gesture and the representation of movement of a physical object caused by the first gesture (e.g., loosening a screw and removing the back cover to expose the interior of the lower compartment). Figure 6A As shown, in response to a detected user gesture of loosening a screw to remove the back cover of the lower compartment of refrigerator 602, the electronic device renders a virtual auxiliary template 616 side-by-side and adjacent to the representation of the refrigerator. In some other embodiments, the virtual auxiliary template 616 is rendered to overlay the representation of refrigerator 612. In some embodiments, the virtual auxiliary template 616 includes one or more items, each corresponding to a specific diagnostic aid, such as a machine's user manual, design blueprints, exploded views showing the internal structure, and / or the machine's circuit design. In some embodiments, the camera 609 of the electronic device 102 captures machine-readable code (e.g., Quick Response Code, QR code) 608 attached to a physical object (e.g., a damaged refrigerator 602 in kitchen 600). The electronic device can retrieve the identification and model information of the physical object (e.g., refrigerator 601) stored in the machine-readable code. The electronic device can then select a virtual auxiliary template (e.g., virtual auxiliary template 616) based on the identification and model information of the physical object (e.g., the damaged refrigerator 602).

[0107] Method 500 further includes: capturing (514) a second gesture via one or more cameras. For example... Figure 6BAs shown, in some embodiments, the second gesture is a user gesture 618 used to inspect (e.g., directly interact with a physical part of the kitchen) electronic components in the lower compartment of a damaged refrigerator in kitchen 600. As discussed in this disclosure, the representation of the gesture (620) can be rendered in real-time in a 3-D virtual environment 604. In some embodiments, the second gesture is the user's gesture 618, which is performed by the user while viewing the 3-D virtual environment 604 and intending to interact with the representation of the physical object 612. In some embodiments, the representation of the gesture 620 is displayed in real-time as the camera 609 captures the gesture 618 in the kitchen.

[0108] In response to one or more cameras capturing (516) a second gesture, method 500 further includes: transferring the second gesture (e.g., Figure 6B The gesture (618) is converted (518) into a first interaction with the representation of the physical object in the 3-D virtual environment (e.g., inspecting / testing the representation of electronic components in the lower compartment of the refrigerator 620). Method 500 also includes: determining (520) a second operation on a virtual auxiliary template associated with the physical object based on the first interaction with the representation of the physical object (e.g., displaying a virtual circuit diagram 622 associated with the electronic component), and rendering the second operation on the virtual auxiliary template associated with the physical object in real time on a display (e.g., displaying a virtual circuit diagram 622 of the electronic component to provide a visual reference to the user when troubleshooting the lower compartment of the refrigerator).

[0109] In some embodiments, the second gesture is a trigger event for adjusting the 3D virtual view while updating the virtual auxiliary template. In some embodiments, to troubleshoot a faulty machine on-site, the user needs to physically interact with the physical object to view the problem. For example, the second gesture is a physical interaction with a first part of the physical object, such as opening a refrigerator door to check why the freezer light is not on, turning to the side of the machine, removing the lid to view the internal circuitry of the machine, or inspecting and testing electronic components (e.g., ...). Figure 6B (User gesture 618 in the image). In some embodiments, hand position information (including depth information) and hand contour data can be analyzed to determine if a second gesture is interacting with a specific electronic component. Therefore, the 3-D virtual view 604 displays in real-time the representation of opening the refrigerator door or testing the electronic component. Simultaneously, the virtual auxiliary template 622 is updated to show the circuit diagram of the corresponding electronic component, such as... Figure 6B As shown.

[0110] In some embodiments, for troubleshooting or other applications such as product assembly, the second gesture interacts with the representation of a physical object in a 3-D virtual view. For example, the second gesture 620 interacts with the representation of a refrigerator 612 in a 3D virtual view 604 to rotate the viewing perspective of the refrigerator representation from the front to the side without actually rotating the electronic device, such as without rotating a mobile phone or a head-mounted display (HMD) on a user's head, and without actually rotating the physical refrigerator 602 in kitchen 600. In response to the second gesture, the electronic device renders the representation of the refrigerator 612 simultaneously with the rotation of the virtual auxiliary template 616.

[0111] In some embodiments, the second gesture is translated to interact with a specific target part of the machine, and a second operation on the virtual auxiliary template is determined based on that specific target part. For example, after translating the second gesture (e.g., the corresponding representations of gestures 618 and 620) into a first interaction (e.g., inspecting an electronic component) with a portion of the representation of a physical object 620, a second operation on the virtual auxiliary template is performed based on pre-stored relationships between multiple parts of the machine and virtual auxiliary templates of corresponding parts of the machine. For example, circuit diagram 622 is selected based on a pre-stored relationship between an electronic component inspected by the second gesture and a circuit diagram of that electronic component.

[0112] exist Figure 6B In some embodiments shown, in response to a second gesture (e.g., gesture 618) captured by one or more cameras (e.g., camera 609), the electronic device simultaneously renders on the display in real time: (1) a representation of the second gesture (e.g., a representation of gesture 620); (2) a movement of the representation of the physical object caused by a first interaction with the representation of the physical object in a 3-D virtual environment (e.g., any movement of the representation of the refrigerator 612 and / or the representation of the components of the refrigerator (e.g., electronic components) caused by the gesture); and (3) a second operation on a virtual auxiliary template associated with the physical object based on the first interaction with the representation of the physical object (e.g., updating the virtual auxiliary template to display a circuit diagram 622 of the electronic component that interacted with the second gesture).

[0113] In some embodiments, in response to one or more cameras (e.g., camera 609) capturing a second gesture (e.g., ... Figure 6B (Gesture 618) The electronic device renders a second interaction with a representation of a physical object in a virtual environment on a display based on a second operation on a virtual auxiliary template. In some embodiments, the second interaction with the representation of the physical object is a virtual repair process based on the representation of the physical object according to the second operation, such as user gestures when repairing electronic components of a refrigerator while referring to an updated circuit diagram 622 or a virtual video repair demonstration.

[0114] In some embodiments, such as Figure 6C As shown, the camera captures a third gesture (e.g., gesture 624) in the physical environment 600. In some embodiments, gesture 624 is performed when a user views a virtual assistive template and intends to interact with it. For example, the user wants to use... Figure 6C Gesture 624, such as swiping left, to turn to another page or entry of a virtual auxiliary template related to another component of a physical object.

[0115] In some embodiments, in response to a third gesture being captured by one or more cameras, or by an electronic device 102 that can work in conjunction with server system 104, the third gesture is translated into a third action on a virtual auxiliary template associated with a physical object. For example, the third action includes turning a page or switching between entries in the virtual auxiliary template, rotating a design view in the virtual auxiliary template, or zooming in or out of the rendered virtual auxiliary template.

[0116] In some embodiments, electronic device 102 and / or server system 104 determine a second representation of the physical object in a 3-D virtual environment based on a third operation of a virtual auxiliary template associated with the physical object (e.g., Figure 6C The representation of the refrigerator 629). In some embodiments, a second representation of the physical object is determined to include specific parts of the physical object based on the current view of the virtual auxiliary template. In some embodiments, this is done in conjunction with rendering the virtual auxiliary template on the display (e.g., navigating to a virtual auxiliary template that demonstrates how to repair such a refrigerator). Figure 6C Simultaneously with the third operation of the compressor shown, the electronic device 102 renders a second representation of the physical object in the virtual environment in real time based on the third operation on the virtual auxiliary template (e.g., Figure 6C (This includes a magnified view of the physical object 629, including a representation of the compressor 630). In one example, a second representation of the physical object can be rendered to override the previous first representation of the physical object in the 3-D virtual environment.

[0117] For example, such as Figure 6C As shown, while camera 609 captures gesture 624 in physical environment 600, a representation of gesture 626 is rendered in 3D virtual view 604. After gesture 624 is translated into a swipe left to turn to another page of the virtual auxiliary template to view a video demonstration 628 on how to repair a compressor, the representation of physical object 629 is updated based on the current view of the virtual auxiliary template (e.g., video demonstration 628). For example, the representation of physical object 629 is updated to show a magnified view of a specific part (e.g., compressor 630) associated with video demonstration 628.

[0118] In some embodiments, electronic device 102, which may work in conjunction with server system 104, identifies one or more recommended options associated with a physical object in the physical environment (e.g., such as correcting potential defects in parts, resolving potential problems with the machine, performing possible assembly steps from the current stage, or performing possible functions of the panel in the current view) based on camera data and / or image data from one or more images, without receiving any user input. In response to a first gesture captured by one or more cameras, the electronic device renders a virtual auxiliary template of circuit diagram 642 associated with the physical object (e.g., electronic components of a refrigerator) based on the identified one or more recommended options (e.g., correcting errors in circuit board 636).

[0119] For example, such as Figure 6D As shown, after opening the back cover, based on images acquired of the circuit board 636, fan, and / or compressor behind the refrigerator's back cover, the system can provide a display in a virtual auxiliary template (e.g., Figure 6D The system provides possible troubleshooting recommendations for the circuit design diagram 642. Users can repair the electronic components on circuit board 636 while referring to the circuit diagram 642. In some embodiments, the characterization of refrigerator 612 is updated in relevant sections (e.g., the characterization of lower compartment 638) to simultaneously display the characterization of the faulty electronic component 640. In some embodiments, a continuously updated database stores image data of common defects / errors associated with corresponding parts of a physical object. The system can perform image recognition (e.g., compressor burnout, dust filter blockage, etc.) on images of faulty parts, based on an error recommendation model (e.g., ...). Figures 2A to 2B The troubleshooting model (236 or 283) identifies one or more defects / errors and provides recommendations by rendering a helpful virtual auxiliary template corresponding to the identified error.

[0120] In such Figure 6E In some embodiments shown, gestures (e.g., Figure 6E The interaction with the representation of a physical object in a 3-D virtual environment, derived from the gesture 644 (swiping up to zoom in on the selected portion), includes providing a magnified view (e.g., a magnified view 646 of a circuit board portion) of a first portion of the representation of the physical object in the 3-D virtual environment. In some embodiments, rendering a second operation on a virtual auxiliary template associated with the physical object includes simultaneously presenting on the display in real time: (1) a second operation on one or more virtual auxiliary items of a virtual auxiliary template (e.g., rendering circuit board portion 648) associated with the first portion of the representation of the physical object (e.g., the representation of circuit board portion 646); and (2) a magnified view (e.g., a magnified representation of circuit board portion 646) of the first portion of the representation of the physical object in the 3-D virtual environment. Figure 6E As shown, a 3D virtual view can be used as a "magnifying glass" by rendering a magnified view of a specific part selected by the user's gesture. Furthermore, virtual auxiliary templates for specific parts can be rendered side-by-side, allowing users to easily inspect and repair components using the "magnifying glass" and virtual references rendered alongside the object's representation.

[0121] While specific embodiments have been described above, it should be understood that they are not intended to limit this application to these specific embodiments. Rather, this application includes alternatives, modifications, and equivalents within the spirit and scope of the appended claims. Numerous specific details have been set forth to provide a thorough understanding of the subject matter presented herein. However, it will be apparent to those skilled in the art that the subject matter can be practiced without these specific details. In other instances, well-known methods, processes, components, and circuits have not been described in detail to avoid unnecessarily obscuring various aspects of the embodiments.

[0122] Each element of the aforementioned identifier can be stored in one or more of the aforementioned storage devices and corresponds to a set of instructions for performing the aforementioned functions. The modules or programs (i.e., instruction sets) of the aforementioned identifiers do not need to be implemented as separate software programs, processes, modules, or data structures; therefore, various subsets of these modules can be combined or otherwise rearranged in various implementations. In some embodiments, memory 806 optionally stores a subset of the modules and data structures of the aforementioned identifiers. Furthermore, memory 806 optionally stores additional modules and data structures not described above.

Claims

1. A method for providing virtual guidance, comprising: In a computer system having a display, one or more cameras, one or more processors, and memory: One or more images of a physical environment are captured using the one or more cameras, the physical environment including physical objects arranged at the first location; When the one or more cameras capture the one or more images, a 3-D virtual environment is rendered in real time based on the one or more images in the physical environment, wherein the 3-D virtual environment includes a representation of the physical object in the virtual environment at a position corresponding to the first position in the physical environment; A first gesture is captured in the physical environment via one or more cameras; the first gesture is a system-defined gesture associated with displaying a virtual assistance template; the virtual assistance template includes one or more items, each corresponding to a specific diagnostic assistance item; In response to the first gesture being captured by the one or more cameras: Transform the first gesture into a first operation that displays a virtual auxiliary template associated with the physical object in the virtual environment; A virtual auxiliary template is rendered in real time on the display and associated with physical objects that are adjacent to the physical objects represented in the 3-D virtual environment. A second gesture is captured in the physical environment through one or more cameras; the second gesture is a user gesture that interacts with the representation of a physical object in the 3-D virtual environment. In response to the second gesture being captured by the one or more cameras: The second gesture is translated into a first interaction with the representation of the physical object in the 3-D virtual environment; Based on the first interaction with the representation of the physical object, determine a second operation on the virtual auxiliary template associated with the physical object; and The second operation is rendered in real time on the virtual auxiliary template associated with the physical object on the display. The method further includes: User input is captured by one or more cameras; the user input instructs the placement of a first preset virtual object in the space of the 3-D virtual environment corresponding to the space in the physical environment. In response to the user input, a first preset virtual object is rendered in the space of the 3D virtual environment.

2. The method according to claim 1, further comprising: In response to capturing the second gesture via the one or more cameras, simultaneously render it on the display in real time: The representation of the second gesture and the movement of the representation of the physical object caused by the first interaction with the representation of the physical object in the 3-D virtual environment; and The second operation is performed on a virtual auxiliary template associated with the physical object, based on the first interaction with the representation of the physical object.

3. The method according to claim 1 or 2, further comprising: In response to the second gesture being captured by the one or more cameras, Based on the second operation on the virtual auxiliary template, a second interaction between the representation of the physical object in the virtual environment is rendered on the display.

4. The method of claim 3, further comprising: A third gesture in the physical environment is captured by one or more cameras; In response to the third gesture being captured by the one or more cameras: The third gesture is converted into a third operation on the virtual auxiliary template associated with the physical object; Based on a third operation on the virtual auxiliary template associated with the physical object, a second representation of the physical object in the 3D virtual environment is determined; and While rendering the third operation on the virtual auxiliary template on the display, the second representation of the physical object in the virtual environment is rendered in real time according to the third operation of the virtual auxiliary template.

5. The method according to claim 1 or 2, further comprising: Without receiving any user input, based on camera data and / or image data of the one or more images, identify one or more recommended options associated with the physical object in the physical environment; and In response to the first gesture being captured by the one or more cameras: The virtual auxiliary template associated with the physical object is rendered based on the identified one or more recommended options.

6. The method according to claim 1 or 2, wherein: In the 3D virtual environment, the first interaction with the representation of the physical object includes: Provides a magnified view of the first portion of the representation of a physical object in the 3D virtual environment; and The second operation of rendering a virtual auxiliary template associated with the physical object includes: Simultaneous real-time rendering on the display: The second operation on one or more virtual auxiliary items of the virtual auxiliary template associated with the first portion of the representation of the physical object; and The magnified view of the first part of the representation of the physical object in the 3-D virtual environment.

7. The method according to claim 1 or 2, further comprising: Use one or more cameras to capture machine-readable code attached to the physical object; Retrieve model information of the physical object stored in the machine-readable code; and The virtual auxiliary template is selected based on the model information of the physical object.

8. A computer system, comprising: monitor; One or more cameras; One or more processors; and A memory for storing instructions, which, when executed by one or more processors, enable the processor to perform a method for providing virtual guidance according to any one of claims 1-7.

9. A computer-readable storage medium storing instructions that, when executed by one or more processors, cause the one or more processors to perform a method for providing virtual guidance according to any one of claims 1-7.

Citation Information

Patent Citations

  • Method and system for receiving gesture input via virtual control objects

    CN107995964A

  • Augmented reality e-commerce for home improvement

    US20170132841A1