Augmented reality ergonomic estimation system
Through an ergonomic estimation system, computer vision and sensors are used to identify joint angles and generate feedback to optimize user interaction of augmented reality devices, solving health problems caused by long-term use and improving the safety and efficiency of the device.
Patent Information
- Application Number
- CN202380087507.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-12-21
- Filing Date
- 2023-12-18
- Publication Date
- 2025-07-29
AI Technical Summary
Long-term use of augmented reality/virtual reality devices can lead to postural, repetitive effects, including health problems such as musculoskeletal diseases and physical environmental collisions, and existing technologies have failed to effectively address these safety and health effects.
Through an ergonomic estimation system, computer vision algorithms and sensor data are used to identify joint angles, generate ergonomic feedback, provide posture guidance and device interaction optimization, and improve the user experience and health and safety of the device.
Optimize device user interaction, reduce the risk of musculoskeletal diseases, improve the safety and health of device use, and reduce the demand for computing resources.
Smart Images

Figure CN120390915A_ABST
Abstract
Description
[0001] Priority Claim
[0002] This application claims the priority benefit of U.S. Patent Application Serial No. 18 / 069,779, filed on December 21, 2022, which is hereby incorporated by reference in its entirety. Technical Field
[0003] The subject matter disclosed herein generally relates to an ergonomic estimation system for augmented reality devices. Specifically, the present disclosure relates to systems and methods for simulating the use of an augmented reality device and generating ergonomic feedback. Background Art
[0004] Augmented reality (AR) devices enable a user to observe a scene while seeing relevant virtual content that can be aligned with items, images, objects, or the environment within the device's field of view. Virtual reality (VR) devices provide a more immersive experience than AR devices. VR devices utilize virtual content based on the positioning and orientation of the VR device to occlude the user's field of view.
[0005] Due to the weight and adjustment of the device, the physical effects of long-term use of AR / VR devices can be postural and repetitive. Other effects of VR immersion include disorientation and collisions with the physical environment. Other examples of long-term use of AR devices include more physical activity during user interaction (e.g., gaming, fitness applications). Therefore, it is desirable to address the safety and health impacts of long-term use of AR / VR devices. Brief Description of the Drawings
[0006] To easily identify the discussion of any particular element or action, one or more of the most significant digits in the reference numerals refer to the figure number in which the element was first introduced.
[0007] Figure 1 is a block diagram showing a network environment for operating an AR device according to one example embodiment.
[0008] Figure 2 is a block diagram showing a server according to one example embodiment.
[0009] Figure 3 is a block diagram showing a server according to one example embodiment.
[0010] Figure 4 is a block diagram showing an AR device according to one example embodiment.
[0011] Figure 5 is a block diagram showing the operation of an AR device ergonomic system according to one example embodiment.
[0012] Figure 6 is a flowchart showing a method for generating ergonomic feedback according to an example embodiment.
[0013] Figure 7 is a flowchart showing a method for providing ergonomic recommendations according to an example embodiment.
[0014] Figure 8 shows a network environment in which a head-mounted device can be implemented according to an example embodiment.
[0015] Figure 9 is a block diagram showing a software architecture in which the present disclosure can be implemented according to an example embodiment.
[0016] Figure 10 is a graphical representation of a machine in the form of a computer system within which a set of instructions can be executed to cause the machine to perform any one or more of the methods discussed herein. Detailed Description
[0017] The following description describes systems, methods, techniques, instruction sequences, and computer program products that illustrate example embodiments of the subject matter. In the following description, for purposes of explanation, numerous specific details are set forth in order to provide an understanding of the various embodiments of the subject matter. However, it will be apparent to those skilled in the art that embodiments of the subject matter may be practiced without some or other of these specific details. The examples merely represent possible variations. Unless otherwise explicitly stated, structures (e.g., structural components such as modules) are optional and may be combined or subdivided, and operations (e.g., in a process, algorithm, or other functionality) may vary in order or be combined or subdivided.
[0018] The term “augmented reality” (AR) is used herein to refer to an interactive experience of the real-world environment in which physical objects residing in the real world are “augmented” or enhanced by computer-generated digital content (also referred to as virtual content or synthetic content). AR may also refer to a system that enables a combination of the real world and the virtual world, real-time interaction, and 3D registration of virtual objects and real objects. A user of an AR system perceives virtual content that appears to be connected to or interact with physical objects in the real world.
[0019] The term “virtual reality” (VR) is used herein to refer to a simulated experience of a virtual-world environment that is completely different from the real-world environment. Computer-generated digital content is displayed in the virtual-world environment. VR may also refer to a system that enables a user of a VR system to be fully immersed in the virtual-world environment and interact with virtual objects presented in the virtual-world environment.
[0020] The term "AR application" is used herein to refer to an application of computer operations that implements an AR experience. The term "VR application" is used herein to refer to an application of computer operations that implements a VR experience. The term "AR / VR application" refers to an application of computer operations that implements a combination of an AR experience or a VR experience.
[0021] The terms "visual tracking system" and "visual tracking device" are used herein to refer to an application or system of computer operations that enables the system to track visual features identified in an image captured by one or more camera devices of the visual tracking system. The visual tracking system constructs a model of the real-world environment based on the tracked visual features. Non-limiting examples of visual tracking systems include: visual simultaneous localization and mapping systems (VSLAM) and visual inertial odometry (VIO) systems. VSLAM can be used to construct an object from an environment or scene based on one or more camera devices of the visual tracking system. The VIO system (also referred to as a visual inertial tracking system) determines the latest pose (e.g., localization and orientation) of a device based on data obtained from multiple sensors of the device (e.g., optical sensors, inertial sensors).
[0022] The term "inertial measurement unit" (IMU) is used herein to refer to a device that can report the inertial state of a moving object, where the inertial state includes the acceleration, velocity, orientation, and localization of the moving object. The IMU tracks the movement of the object by integrating the acceleration and angular velocity measured by the IMU. The IMU can also refer to a combination of an accelerometer and a gyroscope, which can respectively determine and quantify linear acceleration and angular velocity. The values obtained from the IMU gyroscope can be processed to obtain the pitch, roll, and heading of the IMU, and thus the pitch, roll, and heading of the object associated with the IMU. The signals from the accelerometer of the IMU can also be processed to obtain the velocity and displacement of the IMU.
[0023] The term "three-degree-of-freedom tracking system" (3DOF tracking system) is used herein to refer to a device that tracks rotational movement. For example, a 3DOF tracking system can track whether a user of a head-mounted device looks left or right, rotates their head up or down, and turns left or right. However, a head-mounted device cannot use a 3DOF tracking system to determine whether a user moves around a scene by moving in the physical world. Therefore, a 3DOF tracking system may not be accurate enough to be used for positioning signals. A 3DOF tracking system can be part of an AR / VR display device that includes an IMU sensor. For example, a 3DOF tracking system uses sensor data from sensors such as accelerometers, gyroscopes, and magnetometers.
[0024] The term "six degrees of freedom tracking system" (6DOF tracking system) is used in this document to refer to a device that tracks rotational and translational movements. For example, a 6DOF tracking system can track whether a user has rotated their head and moved forward or backward, laterally or vertically, and up or down. A 6DOF tracking system can include a SLAM system or a VIO system that relies on data obtained from multiple sensors (e.g., depth cameras, inertial sensors). The 6DOF tracking system analyzes data from the sensors to accurately determine the pose of a display device.
[0025] The term "ergonomics" refers to the design and arrangement of user interface elements of an AR application such that an AR user can interact most effectively and safely. The term can also refer to the design characteristics of visual elements resulting from the application of ergonomic science.
[0026] Proper body posture during physical tasks is important for alleviating health problems such as musculoskeletal disorders (MSDs). This application describes an ergonomics estimation system that provides feedback / guidelines during the AR application user interaction phase and the AR application development phase. In one example, the ergonomics estimation system identifies joint angles (e.g., neck angle / trunk angle / wrist angle) inferred from a CV algorithm and from sensors to estimate the ergonomic risk level and provides guidance on which joint angle(s) is / are at risk and the recommended continuous usage duration.
[0027] This application describes an ergonomics estimation system for AR applications of AR devices. The ergonomics estimation system uses joint angles (e.g., neck angle / trunk angle / wrist angle) inferred from a CV algorithm and from sensors (from a real device or from a simulation) to generate ergonomic feedback (e.g., estimate the ergonomic risk level). The ergonomic feedback can be used to identify the optimal reach envelope for different user groups. The ergonomic feedback can also be provided to AR application creators / developers to guide them in the development of their AR applications. In another example, the ergonomic feedback indicates ergonomic compatibility as an app rating criterion to help AR users find suitable applications. For example, the ergonomic feedback can include user classification.
[0028] In an example embodiment, the method includes: accessing user interface elements of an augmented reality application; accessing a plurality of user models and corresponding simulated augmented reality device sensor data for the plurality of user models; identifying, based on the plurality of user models and the corresponding simulated augmented reality device sensor data for the plurality of user models, simulated user interactions with a simulated augmented reality device that operates the augmented reality application; applying a computer vision algorithm to the simulated user interactions; identifying a user pose and user movement based on the simulated user interactions; and generating first ergonomic feedback based on the user pose and user movement.
[0029] Accordingly, one or more of the methods described herein help address the technical problem of power consumption by optimizing user interaction with a device by improving ergonomics. The currently described methods provide improvements to the functional operation of a computer by recommending configurations and configuring an AR application based on ergonomic feedback. Accordingly, one or more of the methods described herein can eliminate the need for certain efforts or computing resources. Examples of such computing resources include processor cycles, network traffic, memory usage, data storage capacity, power consumption, network bandwidth, and cooling capacity.
[0030] Figure 1 is a network diagram showing a network environment 100 suitable for operating an AR device 110, a server 112, and a content creator client device 118 according to some example embodiments. The network environment 100 includes an AR device 110, a server 112, and a content creator client device 118 communicatively coupled to each other via a network 104. The AR device 110, the content creator client device 118, and the server 112 can each be implemented in whole or in part in a computer system as described below with respect to Figure 10 The server 112 can be part of a network-based system. For example, the network-based system can be or include a cloud-based server system that provides additional information such as ergonomic estimates to the AR device 110 and the content creator client device 118.
[0031] A user 106 operates the AR device 110. The user 106 can be a human user (e.g., a human), a machine user (e.g., a computer configured by a software program to interact with the AR device 110), or any suitable combination thereof (e.g., a human assisted by a machine or a machine supervised by a human). The user 106 is not part of the network environment 100 but is associated with the AR device 110.
[0032] The AR device 110 can be a computing device with a display, such as a smartphone, a tablet computer, or a wearable computing device (e.g., a watch or glasses). The computing device can be handheld or can be movably mounted to the head of the user 106. In one example, the display can be a screen that displays content captured by the imaging device of the AR device 110. In another example, the display of the device can be transparent, such as in the lenses of wearable computing glasses. In other examples, the display can be a transparent display, such as the windshield of a car, an airplane, or a truck. In another example, the display can be opaque and wearable by the user to cover the user's field of view.
[0033] The user 106 operates an application of the AR device 110. The application can include an AR application that is configured to provide the user 106 with an experience triggered by a physical object 108, such as a two-dimensional physical object (e.g., a picture), a three-dimensional physical object (e.g., a statue), a location (e.g., at a factory), or any reference in the real-world physical environment (e.g., a perceived corner of a wall or furniture, a QR code). For example, the user 106 can direct the imaging device of the AR device 110 to capture an image of the physical object 108.
[0034] The AR device 110 includes a tracking system (not shown). The tracking system uses optical sensors (e.g., a depth-enabled 3D imaging device, an image imaging device), inertial sensors (e.g., a gyroscope, an accelerometer), wireless sensors (Bluetooth, Wi-Fi), GPS sensors, and audio sensors to track the pose (e.g., the positioning, orientation, and location) of the AR device 110 relative to the real-world environment 102 to determine the position of the AR device 110 within the real-world environment 102.
[0035] The user 120 operates the content creator client device 118 to create or develop an AR application using the server 112. The content creator client device 118 can be a computing device that communicates with the server 112 to assist the content creator client device 118 in developing the AR application.
[0036] In one example, server 112 includes a lens application 116, a lens creation platform 114, and an ergonomics estimation system 122. The lens application 116 can be used to detect and identify physical objects 108 based on sensor data (e.g., image and depth data) from the AR device 110, determine the poses of the AR device 110 and the physical object 108 based on the sensor data. The lens application 116 can also generate virtual objects based on the poses of the AR device 110 and the physical object 108. The lens application 116 transmits the virtual objects to the AR device 110. Object recognition, tracking, and AR rendering can be performed on the AR device 110, the server 112, or a combination between the AR device 110 and the server 112.
[0037] The lens creation platform 114 enables the user 120 to generate / develop AR applications that can be accessed by the lens application 116 at development time. For example, the lens creation platform 114 enables the user 120 to identify or generate user interface elements, select the features and characteristics of the user interface elements, and identify the placement of the user interface elements.
[0038] The ergonomics estimation system 122 identifies lens applications mapped to user groups. For example, the ergonomics estimation system 122 identifies lens applications using the ergonomic characteristics of the user group corresponding to the user 106. When the user searches for an ARP application, ergonomic estimation results belonging to the user's group are provided for reference.
[0039] In another example, when the user 106 interacts with an AR application, sensor data (images, IMU signals) is fed into computer vision (CV) algorithms (e.g., SLAM, HT, etc.), and the ergonomics estimation system 122 uses these inputs to provide ergonomic recommendations to the user 106 (e.g., reminding them to reduce forward neck bending). Thus, the ergonomics estimation system 122 provides ergonomic feedback (e.g., adjust posture, take a break) to the user 106 based on usage data and sensor data from the AR device 110. In addition, ergonomic estimations and the time the user spends on the AR application are recorded to classify the user 106 and calibrate the ergonomics estimation system 122.
[0040] In another example, the ergonomics estimation system 122 accesses the characteristics of an AR application being developed with the lens creation platform 114. The characteristics indicate the placement of user interface elements, the user interface scale, the hand usage habit configuration, the virtual object size, etc. For example, the ergonomics estimation system 122 identifies the usability of the AR application for users with various capabilities and requirements based on the simulation of the AR application interaction through different human models. The ergonomics estimation can include suggestions for UI adjustment according to the target user group and an estimation of the change in the user's time spent on the AR application when adjusting the characteristics (e.g., hand usage habit adaptation +5%).
[0041] In another example, the ergonomics estimation system 122 identifies the physical tasks (e.g., mid-air gestures) when interacting with virtual objects, the physical load of the user, and provides ergonomics feedback to the user 120 based on the posture / movement inferred from the CV algorithm.
[0042] Figure 1 Any one of the machines, databases, or devices shown in can be implemented in a general-purpose computer that is modified by software (e.g., configured or programmed) to be a special-purpose computer to perform one or more of the functions described herein for that machine, database, or device. For example, a computer system capable of implementing any one or more of the methods described herein is discussed below. As used herein, a "database" is a data storage resource and can store data structured as text files, tables, spreadsheets, relational databases (e.g., object-relational databases), triple stores, hierarchical data stores, or any suitable combination thereof. Additionally, Figure 6 Any two or more of the machines, databases, or devices shown in can be combined into a single machine, and the functions described herein for any single machine, database, or device can be subdivided among multiple machines, databases, or devices. Figure 1 The network 104 can be any network that enables communication between or among machines (e.g., the server 112), databases, and devices (e.g., the AR device 110, the content creator client device 118). Thus, the network 104 can be a wired network, a wireless network (e.g., a mobile or cellular network), or any suitable combination thereof. The network 104 can include one or more portions that constitute a private network, a public network (e.g., the Internet), or any suitable combination thereof.
[0043]
[0044] Figure 2 is a block diagram showing a server 112 according to an example embodiment. The server 112 includes a lens creation platform 114, a computer vision algorithm 204, a lens application 116, and a data simulation platform 208. A content creator client device 118 communicates with the lens creation platform 114 to develop AR applications (e.g., lens application 116). The lens creation platform 114 includes content creation tools having an interactive user interface and AR application / lens application templates to allow content creators (e.g., user 120) to develop various applications for the AR device 110. For example, the lens creation platform 114 enables the content creator client device 118 to select, identify, and place graphical user interface elements of the lens application 116. The content creator client device 118 can also define events triggered based on user interactions with the graphical user interface elements.
[0045] An ergonomics estimation system 122 simulates user interactions with the lens application 116 by applying simulated sensor data 202 from the data simulation platform 208. For example, the simulated sensor data 202 includes simulated sensor data of a user operating the AR device. The simulated sensor data includes simulated movements of a manipulated 3D human model to estimate the usability of the lens application 116 for different user groups. The simulated data is in the same format as the real data from the sensors of the AR device 110, such that the same ergonomics estimation system 122 can be used to infer upper body fatigue of the user over time. In another example, the ergonomics estimation system 122 can be calibrated when the estimated fatigue level does not match the actual usage time of the lens application 116.
[0046] The computer vision algorithm 204 operates on the output from the lens application 116 and the simulated sensor data 202 to identify the posture / movement of a potential user operating the lens application 116. For example, the simulated sensor data (images, IMU signals) are fed into the computer vision algorithm 204 (such as SLAM, HT, etc.) to identify joint angles (e.g., neck angle / trunk angle / wrist angle). In another example, when developing the lens application 116, the data simulation platform 206 simulates the use by animating a 3D manipulable human model and having it move with the left hand / right hand / both hands accordingly, so as to achieve the designed user interaction. The data simulation platform 206 uses 3D human models with different heights / arm lengths / genders for simulation to obtain various human movements. The ergonomics estimation system 122 will take the simulated data as input to infer the upper body fatigue of the user. In one example, the output from the computer vision algorithm 204 is fed into the ergonomics estimation system 122 to generate an ergonomics estimation or feedback. For example, the ergonomics feedback indicates the ergonomics risk level and provides guidance on which joint angles (which joint angles) are at risk and the recommended continuous usage duration (for a specific user 106 or for each user group).
[0047] Figure 3 is a block diagram showing the server 112 according to an example embodiment. The server 112 includes an ergonomics estimation system 122, a computer vision algorithm 204, a lens application 116, and real-time sensor data 302. The lens application 116 is uploaded to the AR device 110. In other words, the user 106 operates the lens application 116 using the AR device 110. When the user 106 interacts with the lens application 116, the IMU signals and images are captured by the AR device 110 and stored in the real-time sensor data 302. The real-time data (e.g., IMU signals, images from the AR device 110) are fed into the computer vision algorithm 204 (such as the SLAM / HT algorithm) to identify the head posture / gesture. The ergonomics estimation system 122 performs ergonomics estimation based on the user's head posture / gesture / various information (such as hand usage habits, height, etc.) to infer the upper body fatigue of the user 106. The ergonomics estimation system 122 provides an AR user ergonomics estimation to the user 106.
[0048] In other examples, before the user 106 downloads various AR applications from the server 112, the user 106 can obtain the ergonomics estimation of the AR application for reference. In addition, the ergonomics estimation system 122 can provide real-time reminders when they interact with the lens application 116.
[0049] The ergonomic estimation using real data plus the AR application usage time can be stored in the real-time sensor data 302 / simulated sensor data 202. The ergonomic estimation system 122 classifies users into groups and can perform further calibration based on the real-time sensor data 302.
[0050] Figure 4 FIG. is a block diagram showing modules (e.g., components) of an AR device 110 according to some example embodiments. The AR device 110 includes a sensor 402, a display 404, a processor 408, a graphics processing unit 418, a display controller 420, and a storage device 406. Examples of the AR device 110 include wearable computing devices, tablet computers, navigation devices, portable media devices, or smart phones.
[0051] The sensor 402 includes an optical sensor 414, an inertial sensor 416, and a depth sensor 426. The optical sensor 414 includes a combination of a color imaging device, a thermal imaging device, a depth sensor, and one or more grayscale, global shutter tracking imaging devices. The inertial sensor 416 includes a combination of a gyroscope, an accelerometer, and a magnetometer. The depth sensor 426 includes a combination of a structured light sensor, a time-of-flight sensor, a passive stereo sensor, and an ultrasonic device, a time-of-flight sensor. Other examples of the sensor 402 include a proximity sensor or a position sensor (e.g., near field communication, GPS, Bluetooth, Wifi), an audio sensor (e.g., a microphone), or any suitable combination thereof. Note that the sensor 402 described herein is for illustrative purposes, and thus the sensor 402 is not limited to the sensors described above.
[0052] The display 404 includes a screen or monitor configured to display images generated by the processor 408. In one example embodiment, the display 404 can be transparent or translucent such that the user 106 can view through the display 404 (in an AR use case). In another example, the display 404 (e.g., an LCOS display) presents each frame of the virtual content in multiple presentations.
[0053] The processor 408 includes an AR application 410, a 6DOF tracker 412, and an AR device ergonomics system 424. In one example, the AR application 410 includes a lens application 116. The AR application 410 uses computer vision to detect and identify a physical environment or physical object 108. The AR application 410 retrieves virtual objects (e.g., 3D object models) based on the identified physical object 108 or physical environment. The display 404 displays the virtual objects. The AR application 410 includes a local rendering engine that generates a visualization of the virtual objects that is superimposed on (e.g., overlaid on or otherwise displayed concurrently with) an image of the physical object 108 captured by the optical sensor 414. The visualization of the virtual objects can be manipulated by adjusting the positioning of the physical object 108 relative to the optical sensor 414 (e.g., its physical location, orientation, or both). Similarly, the visualization of the virtual objects can be manipulated by adjusting the pose of the AR device 110 relative to the physical object 108.
[0054] The 6DOF tracker 412 estimates the pose of the AR device 110. For example, the 6DOF tracker 412 uses image data and corresponding inertial data from the optical sensor 414 and the inertial sensor 416 to track the position and pose of the AR device 110 relative to a reference frame (e.g., the real-world environment 102). In one example, the 6DOF tracker 412 uses the sensor data to determine the three-dimensional pose of the AR device 110. The three-dimensional pose is the determined orientation and positioning of the AR device 110 relative to the user's real-world environment 102. For example, the AR device 110 can use an image of the user's real-world environment 102 and other sensor data to identify the relative positioning and orientation of the AR device 110 and physical objects in the real-world environment 102 surrounding the AR device 110. The 6DOF tracker 412 continuously collects and uses updated sensor data that describes the movement of the AR device 110 to determine the updated three-dimensional pose of the AR device 110, where the updated three-dimensional pose indicates a change in the relative positioning and orientation of the AR device 110 and physical objects in the real-world environment 102. The 6DOF tracker 412 provides the three-dimensional pose of the AR device 110 to the graphics processing unit 418.
[0055] The graphics processing unit 418 includes a rendering engine (not shown) configured to render frames of a 3D model of a virtual object based on virtual content provided by the AR application 410 and the pose of the AR device 110. In other words, the graphics processing unit 418 uses the three-dimensional pose of the AR device 110 to generate frames of virtual content to be presented on the display 404. For example, the graphics processing unit 418 uses the three-dimensional pose to render frames of virtual content such that the virtual content is presented in a certain orientation and position in the display 404 to appropriately enhance the user's sense of reality. As an example, the graphics processing unit 418 may use the three-dimensional pose data to render frames of virtual content such that when presented on the display 404, the virtual content overlaps with physical objects in the user's real-world environment 102. The graphics processing unit 418 generates updated frames of virtual content based on the updated three-dimensional pose of the AR device 110, which reflects changes in the user's positioning and orientation relative to physical objects in the user's real-world environment 102.
[0056] The graphics processing unit 418 transmits the rendered frames to the display controller 420. The display controller 420 is positioned as an intermediary between the graphics processing unit 418 and the display 404, receives image data (e.g., the rendered frames) from the graphics processing unit 418, reprojects the frames based on the latest pose of the AR device 110 (by performing a warping process), and provides the reprojected frames to the display 404.
[0057] The AR device ergonomics system 424 communicates with the ergonomics estimation system 122 of the server 112. In one example, the AR device ergonomics system 424 transmits data from the sensors 402 to the AR device ergonomics system 424. In another example, the AR device ergonomics system 424 retrieves ergonomics estimates / feedback for a specific AR application from the AR device ergonomics system 424. In other examples, the AR device ergonomics system 424 generates reminders for the user 106 based on the usage of the AR device 110, real-time data from the sensors 402, and ergonomics feedback.
[0058] The storage device 406 stores virtual object content 422 and ergonomics data 428. The virtual object content 422 includes, for example, a database of visual references (e.g., images, QR codes) and corresponding virtual content (e.g., 3D models of virtual objects). The ergonomics data 428 stores ergonomics feedback / ratings corresponding to the AR applications.
[0059] Other augmented data that may be stored within the storage device 406 includes augmented reality content items (e.g., corresponding to application lenses or augmented reality experiences). The augmented reality content items may be real-time special effects and sounds that can be added to images or videos.
[0060] As described above, enhanced data includes augmented reality content items, overlays, image transformations, AR images, and like terms referring to modifications that can be applied to image data (e.g., video or images). This includes real-time modifications that modify the image when it is captured using the device sensors of the AR device 110 (e.g., one or more camera devices), and the image is then displayed on the screen of the AR device 110 along with the modifications. This also includes modifications to stored content such as video clips in a library that can be modified. For example, in the AR device 110 accessing multiple augmented reality content items, a user can use a single video clip with multiple augmented reality content items to see how different augmented reality content items will modify the stored clip. For example, different pseudo-random movement models can be applied to the same content by selecting different augmented reality content items for the content. Similarly, real-time video capture can be used with the shown modifications to show how the video image currently being captured by the sensors of the AR device 110 will modify the captured data. Such data can be simply displayed on the screen rather than stored in memory, or the content captured by the device sensors can be recorded and stored in memory with or without modifications (or both). In some systems, a preview function can show how different augmented reality content items will look within different windows on the display. This can, for example, enable multiple windows with different pseudo-random animations to be viewed simultaneously on the display.
[0061] Thus, data and various systems using augmented reality content items or other such transformation systems that use the data to modify content can involve the detection of objects (e.g., faces, hands, bodies, cats, dogs, surfaces, objects, etc.) in video frames, the tracking of such objects as they leave the field of view, enter the field of view, and move around the field of view, and the modification or transformation of such objects when tracking them. In various examples, different methods for implementing such transformations can be used. Some examples can involve generating a three-dimensional mesh model of one or more objects and using the transformation and animated textures of the model within the video to implement the transformation. In other examples, points on the tracked object can be used to place an image or texture (which can be two-dimensional or three-dimensional) at the tracked location. In still other examples, neural network analysis of video frames can be used to place an image, model, or texture within the content (e.g., the image or frame of a video). Thus, augmented reality content items refer both to the images, models, and textures used to create transformations in content and to the additional modeling and analysis information required to implement such transformations using object detection, tracking, and placement.
[0062] Real-time video processing can be performed using any kind of video data (e.g., video stream, video file, etc.) stored in the memory of any kind of computerized system. For example, a user can load a video file and save it in the memory of the device, or can use the sensors of the device to generate a video stream. Additionally, any object such as a human face and human body parts, an animal, or a non-biological object such as a chair, a car, or other objects can be processed using a computer animation model.
[0063] In some examples, when a specific modification is selected along with the content to be transformed, the element to be transformed is identified by the computing device, and then if the element to be transformed exists in the video frame, the element to be transformed is detected and tracked. The elements of the object are modified according to the modification request, thus transforming the frame of the video stream. The transformation of the frame of the video stream can be performed by different methods for different kinds of transformations. For example, for a frame transformation mainly involving changing the form of the elements of an object, the feature points of each element of the object are calculated (e.g., using an active shape model (ASM) or other known methods). Then, a grid based on the feature points is generated for each of at least one element of the object. This grid is used for the subsequent stages of tracking the elements of the object in the video stream. During the tracking process, the mentioned grid for each element is aligned with the positioning of each element. Then, additional points are generated on the grid. A first set of first points is generated for each element based on the modification request, and a second set of second points is generated for each element based on the first set of points and the modification request. Then, the frame of the video stream can be transformed by modifying the elements of the object based on the first set of points, the second set of points, and the grid. In such a method, the background of the modified object can also be changed or deformed by tracking and modifying the background.
[0064] In some examples, a transformation of changing some regions of an object using the elements of the object can be performed by calculating the feature points of each element of the object and generating a grid based on the calculated feature points. Points are generated on the grid, and then various regions based on the points are generated. Then, the elements of the object are tracked by aligning the regions of each element with the positioning of each of at least one element, and the attributes of the regions can be modified based on the modification request, thus transforming the frame of the video stream. According to the specific modification request, the attributes of the mentioned regions can be transformed in different ways. Such modifications can involve: changing the color of the region; removing at least some parts of the region from the frame of the video stream; including one or more new objects into the region based on the modification request; and modifying or deforming the elements of the region or the object. In various examples, any combination of such modifications or other similar modifications can be used. For some models to be animated, some feature points can be selected as control points for determining the entire state space of the options for model animation.
[0065] In some examples of computer animation models that use face detection to transform image data, a specific face detection algorithm (e.g., Viola-Jones) is used to detect faces in an image. Then, an Active Shape Model (ASM) algorithm is applied to the face region of the image to detect facial feature reference points.
[0066] Other methods and algorithms applicable to face detection can be used. For example, in some examples, landmarks are used to locate features, where a landmark represents a distinguishable point that exists in most of the images under consideration. For example, for face landmarks, the position of the left eye pupil can be used. If the initial landmark is not recognizable (e.g., if a person has an eye patch), secondary landmarks can be used. Such a landmark recognition process can be used for any such object. In some examples, a set of landmarks forms a shape. The shape can be represented as a vector using the coordinates of the points in the shape. A similarity transformation (allowing translation, scaling, and rotation) that minimizes the average Euclidean distance between the shape points is used to align one shape with another. The average shape is the average of the aligned training shapes.
[0067] In some examples, the search for landmarks starts from an average shape that is aligned with the location and size of the face determined by the full-face detector. Then, such a search repeats the following steps until convergence: a tentative shape is proposed by adjusting the positions of the shape points through template matching of the image texture around each point, and then the tentative shape is made to conform to the global shape model. In some systems, individual template matching is unreliable, and the shape model pools the results of weak template matches to form a stronger overall classifier. The entire search is repeated at each level in the image pyramid from coarse resolution to fine resolution.
[0068] The transformation system can capture an image or video stream on a client device (e.g., AR device 110) and perform complex image manipulations locally on the AR device 110 while maintaining a suitable user experience, computation time, and power consumption. Complex image manipulations can include size and shape changes, emotion transfer (e.g., changing a face from a frown to a smile), state transfer (e.g., aging a subject, reducing apparent age, changing gender), style transfer, application of graphical elements, and any other suitable image or video manipulations implemented by a convolutional neural network that has been configured to execute efficiently on the AR device 110.
[0069] In some examples, a computer animation model for transforming image data can be used by a system in which a user can use an AR device 110 to capture an image or video stream of the user (e.g., a selfie), and the AR device 110 has a neural network operating as part of an AR application 410 operating on the AR device 110. A transformation system operating within the AR device 110 determines the presence of a face within the image or video stream and provides a modification icon associated with the computer animation model to transform the image data, or the computer animation model can be presented in association with the interfaces described herein. The modification icon includes changes that can be used to modify the basis of the user's face within the image or video stream as part of a modification operation. Once the modification icon is selected, the transformation system initiates the process of transforming the user's image to reflect the selected modification icon (e.g., generating a smiling face on the user). Once the image or video stream is captured and the specified modification is selected, the modified image or video stream can be presented in the graphical user interface displayed on the AR device 110. The transformation system can implement a complex convolutional neural network on a portion of the image or video stream to generate and apply the selected modification. That is, the user can capture an image or video stream, and once the modification icon is selected, the modification result is presented to the user in real time or near real time. Additionally, while a video stream is being captured, the modification can be persistent and the selected modification icon remains toggled. A machine-taught neural network can be used to implement such modifications.
[0070] A graphical user interface presenting the modifications performed by the transformation system can supply additional interaction options to the user. Such options can be based on the interface used to initiate content capture and selection for a particular computer animation model (e.g., initiated from a content creator user interface). In various examples, after an initial selection of a modification icon, the modification can be persistent. The user can toggle the modification on or off by tapping or otherwise selecting the face being modified by the transformation system and store it for later viewing or browsing to other areas of the imaging application. In the case where multiple faces are being modified by the transformation system, the user can globally toggle the modification on or off by tapping or selecting an individual face modified and displayed within the graphical user interface. In some examples, individual faces within a group of multiple faces can be modified separately, or such modifications can be toggled separately by tapping or selecting each individual face or a series of individual faces displayed within the graphical user interface.
[0071] Any one or more of the modules described herein can be implemented using hardware (e.g., a processor of a machine) or a combination of hardware and software. For example, any module described herein can configure a processor to perform the operations described herein for that module. Additionally, any two or more of these modules can be combined into a single module, and the functions described herein for a single module can be subdivided among multiple modules. Further, according to various example embodiments, modules described herein as being implemented within a single machine, database, or device can be distributed across multiple machines, databases, or devices.
[0072] Figure 5 is a block diagram showing the operation of an AR device ergonomics system 424 according to one example embodiment. The 6DOF tracker 412 accesses inertial sensor data from the inertial sensors 416 and optical sensor data from the optical sensors 414.
[0073] The 6DOF tracker 412 determines the pose (e.g., position, orientation, tilt) of the AR device 110 relative to a reference frame (e.g., the real-world environment 102). In one example embodiment, the 6DOF tracker 412 includes a VIO 502 and a SLAM 504. The 6DOF tracker 412 estimates the pose of the AR device 110 based on a 3D map of feature points from the images captured using the optical sensors 414 and the inertial sensor data captured using the inertial sensors 416.
[0074] The 6DOF tracker 412 provides pose data to the AR device ergonomics system 424. The optical sensors 414 provide image data (e.g., real-time stream images) to the AR device ergonomics system 424. The AR application provides user engagement data (e.g., the time the user 106 has operated the AR device 110 during the current session, the gestures and movements of the user 106, operations on the AR application 410) to the AR device ergonomics system 424. The AR device ergonomics system 424 provides real-time data (pose data, image data, and user engagement data) to the ergonomics estimation system 122. The AR device ergonomics system 424 receives ergonomics feedback from the ergonomics estimation system 122.
[0075] Figure 6 Shows an example method 600 for detecting changes in a scene. Although example method 600 depicts a particular order of operations, the order can be changed without departing from the scope of the present disclosure. For example, some of the depicted operations can be performed in parallel or in a different order that does not substantially affect the functionality of method 600. In other examples, different components of an example device or system implementing method 600 can perform functions substantially simultaneously or in a particular order.
[0076] The operations in method 600 can be performed by the ergonomics estimation system 122 using the components (e.g., modules, engines) described above with respect to Figure 2 Accordingly, method 600 is described by way of example with reference to the ergonomics estimation system 122. However, it should be understood that at least some of the operations in method 600 can be deployed on various other hardware configurations or performed by similar components residing elsewhere.
[0077] According to some examples, the method includes the ergonomics estimation system 122 accessing the lens application 116 from the lens creation platform 114 at block 602.
[0078] According to some examples, the method includes simulating user interactions of the lens application 116 with different human models at block 604.
[0079] According to some examples, the method includes identifying postures / motions based on applying computer vision algorithms to the simulation at block 606.
[0080] According to some examples, the method includes generating ergonomics feedback based on the identified postures and motions at block 608.
[0081] Figure 7 An example method 700 for providing ergonomics recommendations is shown. While example method 700 depicts a particular order of operations, the order can be changed without departing from the scope of the present disclosure. For example, some of the depicted operations can be performed in parallel or in a different order that does not substantially affect the functionality of method 700. In other examples, different components of an example device or system implementing method 700 can perform functions substantially simultaneously or in a particular order.
[0082] The operations in method 700 can be performed by the ergonomics estimation system 122 using the components (e.g., modules, engines) described above with respect to Figure 3 Accordingly, method 700 is described by way of example with reference to the ergonomics estimation system 122. However, it should be understood that at least some of the operations in method 700 can be deployed on various other hardware configurations or performed by similar components residing elsewhere.
[0083] According to some examples, the method includes receiving sensor data from the AR device 110 operating the lens application at block 702.
[0084] According to some examples, the method includes identifying postures / motions based on the sensor data at block 704.
[0085] According to some examples, the method includes providing ergonomic recommendations to the AR device 110 at block 706 based on the identified poses and movements.
[0086] System with a head-wearable device
[0087] Figure 8 Illustrates a network environment 800 in which a head-wearable device 802 can be implemented according to an example embodiment. Figure 8 Is a high-level functional block diagram of an example head-wearable device 802 that communicatively couples a mobile client device 838 and a server system 832 via various networks 840.
[0088] The head-wearable device 802 includes a camera device, such as at least one of a visible light camera device 812, an infrared emitter 814, and an infrared camera device 816. The client device 838 may be able to connect to the head-wearable device 802 using both communication 834 and communication 836. The client device 838 is connected to the server system 832 and the network 840. The network 840 can include any combination of wired and wireless connections.
[0089] The head-wearable device 802 also includes two image displays of an image display 804 of an optical component. The two image displays include one image display associated with the left side of the head-wearable device 802 and one image display associated with the right side of the head-wearable device 802. The head-wearable device 802 also includes an image display driver 808, an image processor 810, a low-power low-power circuit 826, and a high-speed circuit 818. The image display 804 of the optical component is used to present images and videos to the user of the head-wearable device 802, including images that may include a graphical user interface.
[0090] The image display driver 808 commands and controls the image displays in the image display 804 of the optical component. The image display driver 808 may deliver image data directly to the image displays in the image display 804 of the optical component for presentation, or may have to convert the image data into a signal or data format suitable for delivery to the image display device. For example, the image data may be video data formatted according to a compression format (e.g., H.264 (MPEG-4 Part 10), HEVC, Theora, Dirac, RealVideo RV40, VP8, VP9, etc.), and the still image data may be formatted according to a compression format (e.g., Portable Network Graphics (PNG), Joint Photographic Experts Group (JPEG), Tagged Image File Format (TIFF), or Exchangeable Image File Format (Exif), etc.).
[0091] As described above, the head-wearable device 802 includes a frame and a stem (or temple) extending from a side of the frame. The head-wearable device 802 also includes a user input device 806 (e.g., a touch sensor or a push button) including an input surface on the head-wearable device 802. The user input device 806 (e.g., a touch sensor or a push button) is configured to receive input selections from a user to manipulate a graphical user interface for presenting images.
[0092] Figure 8 The components for the head-wearable device 802 shown in are located on one or more circuit boards (e.g., a PCB or a flexible PCB) in a temple or a stem. Alternatively or additionally, the depicted components may be located in a block, a frame, a hinge, or a nose bridge of the head-wearable device 802. The left and right may include digital camera device elements such as complementary metal oxide semiconductor (CMOS) image sensors, charge-coupled devices, camera device lenses, or any other corresponding visible light or light-capturing elements that may be used to capture data, including images of scenes with unknown objects.
[0093] The head-wearable device 802 includes a memory 822 that stores instructions for performing a subset or all of the functions described herein. The memory 822 may also include a storage device.
[0094] As Figure 8 shown, the high-speed circuit 818 includes a high-speed processor 820, a memory 822, and a high-speed wireless circuit 824. In this example, an image display driver 808 is coupled to the high-speed circuit 818 and is operated by the high-speed processor 820 to drive a left image display and a right image display in the image display 804 of the optical assembly. The high-speed processor 820 may be any processor capable of managing high-speed communication and the operation of any general computing system required for the head-wearable device 802. The high-speed processor 820 includes processing resources required to manage high-speed data transmission on communication 836 to a wireless local area network (WLAN) using the high-speed wireless circuit 824. In certain examples, the high-speed processor 820 executes an operating system (e.g., a LINUX operating system or any other such operating system for the head-wearable device 802), and the operating system is stored in the memory 822 for execution. In addition to any other responsibilities, the high-speed processor 820 that executes the software architecture of the head-wearable device 802 manages data transmission with the high-speed wireless circuit 824. In certain examples, the high-speed wireless circuit 824 is configured to implement Institute of Electrical and Electronics Engineers (IEEE) 802.11 communication standards (also referred to herein as Wi-Fi). In other examples, other high-speed communication standards may be implemented by the high-speed wireless circuit 824.
[0095] The low-power wireless circuitry 830 and the high-speed wireless circuitry 824 of the head-wearable device 802 may include a short-range transceiver (Bluetooth TM ), and a wireless wide area network, local area network, or wide area network transceiver (e.g., cellular or WiFi). The client device 838 including the transceivers that communicate via communications 834 and 836 may be implemented using the details of the architecture of the head-wearable device 802, and so may other elements of the network 840.
[0096] The memory 822 includes any storage device capable of storing various data and applications, including the camera device data generated by the left and right, infrared camera devices 816 and the image processor 810, and the images for display generated by the image display driver 808 on the image display of the optical component 804, etc. Although the memory 822 is shown as integrated with the high-speed circuitry 818, in other examples, the memory 822 may be a separate stand-alone element of the head-wearable device 802. In some such examples, wire-by-wire circuitry may provide a connection from the image processor 810 or the low-power processor 828 to the memory 822 through the chip including the high-speed processor 820. In other examples, the high-speed processor 820 may manage the addressing of the memory 822 such that the low-power processor 828 will initiate the high-speed processor 820 at any time when a read or write operation involving the memory 822 is needed.
[0097] As Figure 8 shown, the low-power processor 828 or the high-speed processor 820 of the head-wearable device 802 may be coupled to a camera device (visible light camera device 812; infrared emitter 814 or infrared camera device 816), an image display driver 808, a user input device 806 (e.g., a touch sensor or a push button), and the memory 822.
[0098] The head-wearable device 802 is connected to a host computer. For example, the head-wearable device 802 pairs with the client device 838 via communication 836, or is connected to the server system 832 via the network 840. The server system 832 may be one or more computing devices that are part of a service or network computing system, e.g., including a processor, a memory, and a network communication interface to communicate with the client device 838 and the head-wearable device 802 via the network 840.
[0099] The client device 838 includes a processor and a network communication interface coupled to the processor. The network communication interface allows communication via the network 840, communication 834, or communication 836. The client device 838 may also store at least a portion of the instructions for generating binaural audio content in the memory of the client device 838 to implement the functions described herein.
[0100] The output components of the head-wearable device 802 include visual components such as a display (e.g., a liquid crystal display (LCD), a plasma display panel (PDP), a light-emitting diode (LED) display, a projector, or a waveguide). The image display of the optical component is driven by an image display driver 808. The output components of the head-wearable device 802 also include acoustic components (e.g., speakers), tactile components (e.g., vibration motors), other signal generators, etc. The input components (e.g., the user input device 806) of the head-wearable device 802, the client device 838, and the server system 832 may include alphanumeric input components (e.g., a keyboard, a touch screen configured to receive alphanumeric input, an optoelectronic keyboard, or other alphanumeric input components), pointing-based input components (e.g., a mouse, a touchpad, a trackball, a joystick, a motion sensor, or other pointing instruments), tactile input components (e.g., a physical button, a touch screen that provides the position and force of a touch or touch gesture, or other tactile input components), audio input components (e.g., a microphone), etc.
[0101] The head-wearable device 802 may optionally include additional peripheral device elements. Such peripheral device elements may include biometric sensors, additional sensors, or display elements integrated with the head-wearable device 802. For example, the peripheral device elements may include any I / O components, including output components, motion components, positioning components, or any other such elements described herein.
[0102] For example, biometric components include components for detecting expressions (e.g., hand expressions, facial expressions, voice expressions, body postures, or eye tracking), measuring biosignals (e.g., blood pressure, heart rate, body temperature, sweating, or brain waves), identifying people (e.g., voice recognition, retina recognition, facial recognition, fingerprint recognition, or electroencephalogram-based recognition), etc. Motion components include acceleration sensor components (e.g., accelerometers), gravity sensor components, rotational sensor components (e.g., gyroscopes), etc. Positioning components include position sensor components that generate position coordinates (e.g., a global positioning system (GPS) receiver component), WiFi or Bluetooth TM transceivers, altitude sensor components (e.g., an altimeter or barometer that detects air pressure, from which altitude can be obtained), orientation sensor components (e.g., magnetometers), etc. Such positioning system coordinates may also be received from the client device 838 via communication 836 through a low-power wireless circuit 830 or a high-speed wireless circuit 824.
[0103] When using phrases such as "at least one of A, B, or C", "at least one of A, B, and C", "one or more of A, B, or C", or "one or more of A, B, and C", it is intended that the phrase be interpreted to mean that A can exist alone in an embodiment, B can exist alone in an embodiment, C can exist alone in an embodiment, or any combination of elements A, B, and C can exist in a single embodiment; for example, A and B, A and C, B and C, or A, B, and C.
[0104] Without departing from the scope of the present disclosure, changes and modifications can be made to the disclosed embodiments. These and other changes or modifications are intended to be included within the scope of the present disclosure as expressed in the appended claims.
[0105] Figure 9 is a block diagram 900 showing a software architecture 904 that can be installed on any one or more of the devices described herein. The software architecture 904 is supported by hardware such as a machine 902 that includes a processor 920, a memory 926, and I / O components 938. In this example, the software architecture 904 can be conceptualized as a stack of layers, where each layer provides a specific function. The software architecture 904 includes layers such as an operating system 912, libraries 910, frameworks 908, and applications 906. In operation, the application 906 makes API calls 950 through the software stack and receives messages 952 in response to the API calls 950.
[0106] The operating system 912 manages hardware resources and provides common services. The operating system 912 includes, for example, a kernel 914, services 916, and drivers 922. The kernel 914 serves as an abstraction layer between the hardware and other software layers. For example, the kernel 914 provides memory management, processor management (e.g., scheduling), component management, networking, and security settings, as well as other functions. The services 916 can provide other common services to other software layers. The drivers 922 are responsible for controlling or interfacing with the underlying hardware. For example, the drivers 922 can include a display driver, a camera device driver, or a low-power driver, a flash driver, a serial communication driver (e.g., a Universal Serial Bus (USB) driver), drivers, an audio driver, a power management driver, etc.
[0107] The library 910 provides low-level common infrastructure used by the application 906. The library 910 may include a system library 918 (e.g., the C standard library), which provides functions such as memory allocation functions, string manipulation functions, mathematical functions, etc. Additionally, the library 910 may include an API library 924, such as a media library (e.g., a library for supporting the presentation and manipulation of various media formats, such as Moving Picture Experts Group-4 (MPEG4), High Efficiency Video Coding (H.264 or AVC), Moving Picture Experts Group Layer-3 (MP3), Advanced Audio Coding (AAC), Adaptive Multi-Rate (AMR) audio codec, Joint Photographic Experts Group (JPEG or JPG), or Portable Network Graphics (PNG)), a graphics library (e.g., the OpenGL framework for 2D and 3D rendering in graphical content on a display), a database library (e.g., SQLite that provides various relational database functions), a web library (e.g., WebKit that provides web browsing functions), etc. The library 910 may also include various other libraries 928 to provide many other APIs to the application 906.
[0108] The framework 908 provides high-level common infrastructure used by the application 906. For example, the framework 908 provides various Graphical User Interface (GUI) functions, advanced resource management, and advanced location services. The framework 908 may provide a wide range of other APIs that can be used by the application 906, some of which may be specific to a particular operating system or platform.
[0109] In an example embodiment, the application 906 may include a home application 936, a contacts application 930, a browser application 932, a book reader application 934, a location application 942, a media application 944, a messaging application 946, a gaming application 948, and various other applications such as third-party applications 940. The application 906 is a program that executes functions defined in the program. One or more of the applications in the application 906 can be created using various programming languages, such as object-oriented programming languages (e.g., Objective-C, Java, or C++) or procedural programming languages (e.g., the C language or assembly language). In a particular example, the third-party application 940 (e.g., an application developed using an ANDROID TM or IOS TM Software Development Kit (SDK)) can be an application on platforms such as IOS TM 、ANDROID TM 、 Mobile software running on the mobile operating system of a Phone or another mobile operating system. In this example, the third-party application 940 can call the API call 950 provided by the operating system 912 to facilitate the functions described herein.
[0110] Figure 10 is a graphical representation of a machine 1000 within which instructions 1008 (e.g., software, program, application, applet, app, or other executable code) can be executed to cause the machine 1000 to perform any one or more of the methods discussed herein. For example, the instructions 1008 can cause the machine 1000 to perform any one or more of the methods described herein. The instructions 1008 transform the general unprogrammed machine 1000 into a particular machine 1000 programmed to perform the described and shown functions in the described manner. The machine 1000 can operate as a stand-alone device or can be coupled (e.g., networked) to other machines. In a networked deployment, the machine 1000 can operate as a server machine or a client machine in a server-client network environment, or as a peer machine in a peer-to-peer (or distributed) network environment. The machine 1000 can include, but is not limited to: server computers, client computers, personal computers (PCs), tablet computers, laptop computers, netbooks, set-top boxes (STBs), PDAs, entertainment media systems, cellular phones, smartphones, mobile devices, wearable devices (e.g., smartwatches), smart home devices (e.g., smart appliances), other smart devices, web devices, network routers, network switches, network bridges, or any machine capable of sequentially or otherwise executing the instructions 1008 specifying the actions to be taken by the machine 1000. Moreover, although only a single machine 1000 is shown, the term "machine" should also be regarded as including a collection of machines that individually or jointly execute the instructions 1008 to perform any one or more of the methods discussed herein.
[0111] The machine 1000 can include a processor 1002, a memory 1004, and I / O components 1042 configured to communicate with each other via a bus 1044. In an example embodiment, the processor 1002 (e.g., a central processing unit (CPU), a reduced instruction set computing (RISC) processor, a complex instruction set computing (CISC) processor, a graphics processing unit (GPU), a digital signal processor (DSP), an application specific integrated circuit (ASIC), a radio frequency integrated circuit (RFIC), other processors, or any suitable combination thereof) can include, for example, a processor 1006 and a processor 1010 that execute the instructions 1008. The term "processor" is intended to include multi-core processors, which can include two or more independent processors (sometimes referred to as "cores") that can execute instructions simultaneously. AlthoughFigure 10 Multiple processors 1002 are shown, but machine 1000 can include a single processor with a single core, a single processor with multiple cores (e.g., a multi-core processor), multiple processors with a single core, multiple processors with multiple cores, or any combination thereof.
[0112] Memory 1004 includes main memory 1012, static memory 1014, and storage unit 1016, all of which are accessible by processor 1002 via bus 1044. Main memory 1004, static memory 1014, and storage unit 1016 store instructions 1008 embodying any one or more of the methods or functions described herein. The instructions 1008 may also reside, completely or partially, within main memory 1012, within static memory 1014, within machine-readable medium 1018 within storage unit 1016, within at least one of the processors 1002 (e.g., within a cache memory of the processor), or within any suitable combination thereof, during execution by machine 1000.
[0113] The I / O components 1042 can include a variety of components that receive input, provide output, generate output, transmit information, exchange information, capture measurement results, etc. The specific I / O components 1042 included in a particular machine will depend on the type of the machine. For example, a portable machine such as a mobile phone may include a touch input device or other such input mechanism, while a headless server machine is less likely to include such a touch input device. It should be understood that the I / O components 1042 can include many other components not shown in Figure 10 In various example embodiments, the I / O components 1042 can include output components 1028 and input components 1030. The output components 1028 can include visual components (e.g., a display such as a plasma display panel (PDP), a light-emitting diode (LED) display, a liquid crystal display (LCD), a projector, or a cathode ray tube (CRT)), acoustic components (e.g., speakers), haptic components (e.g., a vibration motor, a resistance mechanism), other signal generators, etc. The input components 1030 can include alphanumeric input components (e.g., a keyboard, a touch screen configured to receive alphanumeric input, an optical keyboard, or other alphanumeric input components), pointing-based input components (e.g., a mouse, a touchpad, a trackball, a joystick, a motion sensor, or other pointing instrument), haptic input components (e.g., a physical button, a touch screen that provides the location and / or force of a touch or touch gesture, or other haptic input components), audio input components (e.g., a microphone), etc.
[0114] In other example embodiments, the I / O component 1042 can include a biometric component 1032, a motion component 1034, an environmental component 1036, or a positioning component 1038, as well as various other components. For example, the biometric component 1032 includes components for detecting expressions (e.g., hand expressions, facial expressions, vocal expressions, body postures, or eye tracking), measuring biometric signals (e.g., blood pressure, heart rate, body temperature, sweating, or brain waves), identifying people (e.g., voice recognition, retina recognition, facial recognition, fingerprint recognition, or electroencephalogram-based recognition), etc. The motion component 1034 includes an acceleration sensor component (e.g., an accelerometer), a gravity sensor component, a rotation sensor component (e.g., a gyroscope), etc. The environmental component 1036 includes, for example, a lighting sensor component (e.g., a photometer), a temperature sensor component (e.g., one or more thermometers for detecting ambient temperature), a humidity sensor component, a pressure sensor component (e.g., a barometer), an auditory sensor component (e.g., one or more microphones for detecting background noise), a proximity sensor component (e.g., an infrared sensor for detecting nearby objects), a gas sensor (e.g., a gas detection sensor for detecting the concentration of hazardous gases to ensure safety or measuring pollutants in the atmosphere), or other components that can provide an indication, measurement, or signal corresponding to the surrounding physical environment. The positioning component 1038 includes a position sensor component (e.g., a GPS receiver component), an altitude sensor component (e.g., an altimeter or barometer that detects air pressure and from which altitude can be obtained based on the air pressure), an orientation sensor component (e.g., a magnetometer), etc.
[0115] A variety of techniques can be used to implement communication. The I / O component 1042 also includes a communication component 1040, which is operable to couple the machine 1000 to the network 1020 or the device 1022 via the couplings 1024 and 1026, respectively. For example, the communication component 1040 can include a network interface component or another suitable device to interface with the network 1020. In other examples, the communication component 1040 can include a wired communication component, a wireless communication component, a cellular communication component, a near field communication (NFC) component, components (e.g., low power), components, and other communication components for providing communication via other modalities. The device 1022 can be another machine or any of a variety of peripheral devices (e.g., a peripheral device coupled via USB).
[0116] In addition, communication component 1040 can detect an identifier or include components operable to detect an identifier. For example, communication component 1040 can include a radio frequency identification (RFID) tag reader component, an NFC smart tag detection component, an optical reader component (e.g., for detecting one-dimensional barcodes such as Universal Product Code (UPC) barcodes, multi-dimensional barcodes such as Quick Response (QR) codes, Aztec codes, Data Matrix, Dataglyph, MaxiCode, PDF417, Ultra Code, UCC RSS-2D barcodes, and other optical codes), or an acoustic detection component (e.g., a microphone for identifying an audio signal of a tag). Additionally, various information can be obtained via communication component 1040, such as a location via Internet Protocol (IP) geolocation, a location via signal triangulation, a location via detecting an NFC beacon signal that can indicate a specific location, and so on.
[0117] Various memories (e.g., memory 1004, main memory 1012, static memory 1014, and / or the memory of processor 1002) and / or storage unit 1016 can store a set or more sets of instructions and data structures (e.g., software) that embody any one or more of the methods or functions described herein or are used by any one or more of the methods or functions described herein. These instructions (e.g., instruction 1008), when executed by processor 1002, cause various operations to implement the disclosed embodiments.
[0118] Instruction 1008 can be sent or received via a network interface device (e.g., a network interface component included in communication component 1040) using a transmission medium and using any one of a plurality of well-known transmission protocols (e.g., Hypertext Transfer Protocol (HTTP)) over network 1020. Similarly, instruction 1008 can be sent or received to device 1022 via coupling 1026 (e.g., a peer-to-peer coupling) using a transmission medium.
[0119] As used herein, the terms "machine storage medium", "device storage medium", and "computer storage medium" mean the same thing and may be used interchangeably in this disclosure. These terms refer to a single or multiple storage devices and / or media that store executable instructions and / or data (e.g., a centralized or distributed database, and / or associated caches and servers). Thus, these terms should be regarded as including, without limitation, solid-state memory as well as optical and magnetic media, including memory internal or external to a processor. Specific examples of machine storage media, computer storage media, and / or device storage media include: non-volatile memory, including, for example, semiconductor memory devices such as erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), field-programmable gate arrays (FPGA), and flash memory devices; magnetic disks, such as internal hard disks and removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks. The terms "machine storage medium", "computer storage medium", and "device storage medium" expressly exclude carrier waves, modulated data signals, and other such media, at least some of which are covered by the term "signal medium" discussed below.
[0120] The terms "transmission medium" and "signal medium" mean the same thing and may be used interchangeably in this disclosure. The terms "transmission medium" and "signal medium" should be understood to include any non-transitory medium that is capable of storing, encoding, or carrying instructions 1416 for execution by machine 1400, and includes digital or analog communication signals or other non-transitory media that facilitate the communication of such software. Thus, the terms "transmission medium" and "signal medium" should be regarded as including any form of modulated data signal, carrier wave, etc. The term "modulated data signal" means a signal that has one or more of its characteristics set or changed in such a manner as to encode information in the signal.
[0121] The terms "machine-readable medium", "computer-readable medium", and "device-readable medium" mean the same thing and may be used interchangeably in this disclosure. These terms are defined to include both machine storage media and transmission media. Thus, these terms include both storage devices / media and carrier waves / modulated data signals.
[0122] Although the embodiments have been described with reference to specific example embodiments, it will be apparent that various modifications and changes can be made to these embodiments without departing from the broader scope of the disclosure. Accordingly, the specification and drawings are to be regarded in an illustrative rather than a restrictive sense. The accompanying drawings, which form a part of this invention, illustrate specific embodiments in which the subject matter can be practiced by way of illustration and not by way of limitation. The illustrated embodiments are described in sufficient detail to enable those skilled in the art to practice the teachings disclosed herein. Other embodiments can be utilized and logical substitutions and changes can be made therefrom without departing from the scope of the disclosure. Accordingly, this detailed description should not be taken in a limiting sense, and the scope of the various embodiments is defined only by the appended claims and the full scope of equivalents to which such claims are entitled.
[0123] Such embodiments of the subject matter of the invention may be referred to herein individually and / or collectively by the term "invention" merely for convenience and are not intended to voluntarily limit the scope of this application to any single invention or inventive concept if more than one invention or inventive concept is in fact disclosed. Accordingly, although specific embodiments have been shown and described herein, it should be understood that any arrangement calculated to achieve the same purpose may be substituted for the specific embodiments shown. The disclosure is intended to cover any and all changes or variations of the various embodiments. After reviewing the above description, combinations of the above embodiments and other embodiments not specifically described herein will be apparent to those skilled in the art.
[0124] A summary of the disclosure is provided to enable the reader to quickly ascertain the nature of the technical disclosure. It is submitted with the understanding that it will not be used to interpret or limit the scope or meaning of the claims. Additionally, in the foregoing detailed description, it can be seen that for purposes of simplifying the disclosure, various features are grouped together in a single embodiment. This method of the disclosure should not be interpreted as reflecting an intention that the claimed embodiments require more features than are expressly recited in each claim. Rather, as reflected by the appended claims, the subject matter of the invention lies in less than all of the features of a single disclosed embodiment. Thus, the claims are hereby incorporated into the detailed description, with each claim standing on its own as a separate embodiment.
[0125] Examples
[0126] Implementations of the described subject matter can include, individually or in combination, one or more features as illustrated below by way of example.
[0127] Example 1 is a method that includes: accessing user interface elements of an augmented reality application; accessing multiple user models and corresponding simulated augmented reality device sensor data for the multiple user models; identifying, based on the multiple user models and the corresponding simulated augmented reality device sensor data for the multiple user models, simulated user interactions with the simulated augmented reality device that operates the augmented reality application; applying a computer vision algorithm to the simulated user interactions; identifying a user pose and user movement based on the simulated user interactions; and generating first ergonomic feedback based on the user pose and user movement.
[0128] Example 2 includes the method of Example 1, wherein the first ergonomic feedback includes an ergonomic estimate of the user interface elements of the augmented reality application.
[0129] Example 3 includes the method of Example 2, wherein the ergonomic estimate indicates a recommended adjustment to the user interface elements, and the recommended adjustment corresponds to a target user group of the augmented reality application.
[0130] Example 4 includes the method of Example 3, wherein the recommended adjustment includes a combination of user interface element scale, hand usage configuration, and virtual object size.
[0131] Example 5 includes the method of Example 3, wherein the ergonomic feedback includes an estimated change in user time spent on the augmented reality application based on the recommended adjustment to the user interface elements.
[0132] Example 6 includes the method of Example 1, and further includes: accessing sensor data from a user's augmented reality device operated by the user, the user's augmented reality device operating the augmented reality application, the sensor data including images captured by an image sensor of the user's augmented reality device and inertial motion unit signals from an inertial motion unit device of the user's augmented reality device; applying a computer vision algorithm to the sensor data to identify the user's pose and movement; and generating second ergonomic feedback to the user based on the user's pose and movement, the second ergonomic feedback indicating a recommended adjustment to the pose when the user operates the user's augmented reality device.
[0133] Example 7 includes the method of Example 6, and further includes: recording the sensor data; and updating the multiple user models and the corresponding simulated augmented reality device sensor data for the multiple user models using the sensor data.
[0134] Example 8 includes the method of Example 6, and further includes: identifying a user group corresponding to the user based on the sensor data of the user's augmented reality device and the user's profile; and calibrating the multiple user models and the corresponding simulated augmented reality device sensor data for the multiple user models based on the user group, the sensor data of the user's augmented reality device, and the user's profile.
[0135] Example 9 includes the method of Example 8, further comprising: receiving an augmented reality application query request from a user's augmented reality device; in response to receiving the augmented reality application query request, identifying at least one augmented reality application compatible with a user group corresponding to the user; and presenting an ergonomic estimate of user interface elements of the at least one augmented reality application at the user's augmented reality device.
[0136] Example 10 includes the method of Example 1, wherein each of the plurality of user models indicates a physical size range of a body part of the user.
[0137] Example 11 is a computing device, comprising: a processor; and a memory storing instructions that, when executed by the processor, configure the device to: access user interface elements of an augmented reality application; access a plurality of user models and corresponding simulated augmented reality device sensor data for the plurality of user models; based on the plurality of user models and the corresponding simulated augmented reality device sensor data for the plurality of user models, identify a simulated user interaction with a simulated augmented reality device operating the augmented reality application from the plurality of user models; apply a computer vision algorithm to the simulated user interaction; identify a user pose and user movement based on the simulated user interaction; and generate first ergonomic feedback based on the user pose and user movement.
[0138] Example 12 includes the computing device of Example 11, wherein the first ergonomic feedback includes an ergonomic estimate of user interface elements of the augmented reality application.
[0139] Example 13 includes the computing device of Example 12, wherein the ergonomic estimate indicates a suggested adjustment to the user interface element, and the suggested adjustment corresponds to a target user group of the augmented reality application.
[0140] Example 14 includes the computing device of Example 13, wherein the suggested adjustment includes a combination of user interface element scale, hand usage configuration, and virtual object size.
[0141] Example 15 includes the computing device of Example 13, wherein the ergonomic feedback includes an estimated change in user time spent on the augmented reality application based on the suggested adjustment to the user interface element.
[0142] Example 16 includes the computing device of Example 11, wherein the instructions further configure the device to: access sensor data from a user augmented reality device operated by a user, the user augmented reality device operating an augmented reality application, the sensor data including images captured by an image sensor of the user augmented reality device and inertial motion unit signals from an inertial motion unit device of the user augmented reality device; apply a computer vision algorithm to the sensor data to identify the user's pose and motion; and generate second ergonomic feedback to the user based on the user's pose and motion, the second ergonomic feedback indicating a recommended adjustment to the pose when the user operates the user augmented reality device.
[0143] Example 17 includes the computing device of Example 16, wherein the instructions further configure the device to: record the sensor data; and update a plurality of user models and corresponding simulated augmented reality device sensor data for the plurality of user models using the sensor data.
[0144] Example 18 includes the computing device of Example 16, wherein the instructions further configure the device to: identify a user group corresponding to the user based on the sensor data of the user augmented reality device and the user's profile; and calibrate a plurality of user models and corresponding simulated augmented reality device sensor data for the plurality of user models based on the user group, the sensor data of the user augmented reality device, and the user's profile.
[0145] Example 19 includes the computing device of Example 18, wherein the instructions further configure the device to: receive an augmented reality application query request from the user augmented reality device; in response to receiving the augmented reality application query request, identify at least one augmented reality application compatible with the user group corresponding to the user; and present an ergonomic estimate of user interface elements of the at least one augmented reality application at the user augmented reality device.
[0146] Example 20 is a non-transitory computer-readable storage medium including instructions that, when executed by a computer, cause the computer to: access user interface elements of an augmented reality application; access a plurality of user models and corresponding simulated augmented reality device sensor data for the plurality of user models; based on the plurality of user models and the corresponding simulated augmented reality device sensor data for the plurality of user models, identify simulated user interactions with a simulated augmented reality device operating the augmented reality application from the plurality of user models; apply a computer vision algorithm to the simulated user interactions; identify a user pose and user motion based on the simulated user interactions; and generate first ergonomic feedback based on the user pose and user motion.
Claims
1. A method, comprising: Accessing user interface elements of an augmented reality application; Accessing a plurality of user models and corresponding simulated augmented reality device sensor data for the plurality of user models; Based on the plurality of user models and the corresponding simulated augmented reality device sensor data for the plurality of user models, identifying simulated user interactions with a simulated augmented reality device for operating the augmented reality application from the plurality of user models; Applying a computer vision algorithm to the simulated user interactions; Identifying a user's pose and user movement based on the simulated user interactions; And Generating first ergonomic feedback based on the user's pose and the user movement.
2. The method according to claim 1, wherein The first ergonomic feedback includes an ergonomic estimate of the user interface elements of the augmented reality application.
3. The method according to claim 2, wherein The ergonomic estimate indicates a recommended adjustment to the user interface elements, the recommended adjustment corresponding to a target user group of the augmented reality application.
4. The method according to claim 3, wherein, The recommended adjustment includes a combination of user interface element scale, hand usage configuration, and virtual object size.
5. The method according to claim 3, wherein: The ergonomic feedback includes an estimated change in user time spent on the augmented reality application based on the recommended adjustment to the user interface elements.
6. The method according to claim 1, further comprising: Accessing sensor data from a user's augmented reality device operated by the user, the user's augmented reality device operating the augmented reality application, the sensor data including images captured by an image sensor of the user's augmented reality device and inertial motion unit signals from an inertial motion unit device of the user's augmented reality device; Applying the computer vision algorithm to the sensor data to identify the user's pose and movement; And Generating second ergonomic feedback to the user based on the user's pose and movement, the second ergonomic feedback indicating a recommended adjustment to the pose of the user when operating the user's augmented reality device.
7. The method according to claim 6, further comprising: Recording the sensor data; And Using the sensor data to update the plurality of user models and the corresponding simulated augmented reality device sensor data for the plurality of user models.
8. The method according to claim 6, further comprising: Identifying a user group corresponding to the user based on the sensor data of the user's augmented reality device and the user's profile; And Calibrating the plurality of user models and the corresponding simulated augmented reality device sensor data for the plurality of user models based on the user group, the sensor data of the user's augmented reality device, and the user's profile.
9. The method according to claim 8, further comprising: Receiving an augmented reality application query request from the user's augmented reality device; In response to receiving the augmented reality application query request, identifying at least one augmented reality application compatible with the user group corresponding to the user; And Presenting an ergonomic estimate of the user interface elements of the at least one augmented reality application at the user's augmented reality device.
10. The method according to claim 1, wherein: Each user model in the plurality of user models indicates a physical size range of a user's body part.
11. A computing device, comprising: a processor; and a memory storing instructions that, when executed by the processor, configure the device to: access user interface elements of an augmented reality application; access a plurality of user models and corresponding simulated augmented reality device sensor data for the plurality of user models; identify, based on the plurality of user models and the corresponding simulated augmented reality device sensor data for the plurality of user models, simulated user interactions with a simulated augmented reality device operating the augmented reality application from the plurality of user models; apply a computer vision algorithm to the simulated user interactions; identify a user pose and user movement based on the simulated user interactions; and generate first ergonomic feedback based on the user pose and the user movement.
12. The computing device according to claim 11, wherein, The first ergonomic feedback includes an ergonomic estimate of the user interface elements of the augmented reality application.
13. The computing device according to claim 12, wherein: The ergonomic estimate indicates a recommended adjustment to the user interface elements, the recommended adjustment corresponding to a target user group of the augmented reality application.
14. The computing device according to claim 13, wherein, The recommended adjustment includes a combination of user interface element scale, hand usage configuration, and virtual object size.
15. The computing device according to claim 13, wherein, The ergonomic feedback includes an estimated change in user time spent on the augmented reality application based on the recommended adjustment to the user interface elements.
16. The computing device according to claim 11, wherein, The instructions further configure the device to: access sensor data from a user augmented reality device operated by a user, the user augmented reality device operating the augmented reality application, the sensor data including images captured by an image sensor of the user augmented reality device and inertial motion unit signals from an inertial motion unit device of the user augmented reality device; apply the computer vision algorithm to the sensor data to identify the user's pose and movement; and generate second ergonomic feedback to the user based on the user's pose and movement, the second ergonomic feedback indicating a recommended adjustment to the pose of the user when operating the user augmented reality device.
17. The computing device according to claim 16, wherein, The instructions further configure the device to: record the sensor data; and update the plurality of user models and the corresponding simulated augmented reality device sensor data for the plurality of user models using the sensor data.
18. The computing device according to claim 16, wherein, The instructions further configure the device to: identify a user group corresponding to the user based on the sensor data of the user augmented reality device and the user's profile; and calibrate the plurality of user models and the corresponding simulated augmented reality device sensor data for the plurality of user models based on the user group, the sensor data of the user augmented reality device, and the user's profile.
19. The computing device of claim 18, wherein: The instructions further configure the device to: receive an augmented reality application query request from the user augmented reality device; in response to receiving the augmented reality application query request, identify at least one augmented reality application compatible with the user group corresponding to the user; and present an ergonomic estimate of the user interface elements of the at least one augmented reality application at the user augmented reality device.
20. A non-transitory computer-readable storage medium, the computer-readable storage medium comprising instructions that, when executed by a computer, cause the computer to: Access user interface elements of an augmented reality application; Access a plurality of user models and corresponding simulated augmented reality device sensor data for the plurality of user models; Based on the plurality of user models and the corresponding simulated augmented reality device sensor data for the plurality of user models, identify, from the plurality of user models, simulated user interactions with a simulated augmented reality device that operates the augmented reality application; Apply a computer vision algorithm to the simulated user interactions; Identify user poses and user movements based on the simulated user interactions; And Generate first ergonomic feedback based on the user poses and the user movements.