User guidance system based on augmented reality and / or gesture detection technology
By combining AR/VR technology and human posture detection, real-time screen and voice guidance is provided, solving the problem of lack of guidance for users during 3D scanning and improving the accuracy and efficiency of scanning.
Patent Information
- Application Number
- CN202080027336.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-09-12
- Filing Date
- 2020-04-01
- Publication Date
- 2025-11-21
- Estimated Expiration
- 2040-04-01
AI Technical Summary
In existing technologies, users lack effective user guidance when using smart devices for 3D scanning, resulting in a complex scanning process that is difficult to complete accurately.
By combining augmented reality (AR) and/or virtual reality (VR) technologies with human posture detection, on-screen and voice guidance are generated through real-time analysis of sensor data to help users adjust their posture and position to complete 3D scanning.
It achieves a user-friendly 3D scanning process, improves scanning accuracy and efficiency, and reduces the complexity of user operation.
Smart Images

Figure CN113994396B_ABST
Abstract
Description
[0001] BACKGROUND
[0002] The following discussion of the background to the application can reflect a hindsight bias from the disclosure of the application; and these features are not necessarily considered prior art.
[0003] Increasingly, sensors and display devices have been integrated into smartphones and similar smartphone-like devices, including phones, iPads, tablets, or any similar type of device that is mass produced or custom made for a particular market (i.e., that has a display screen, a processor, and one or more of the sensors listed below). For the purposes of this discussion, we will generalize these devices as “smart devices.” For example, these sensors can include:
[0004] a. a camera;
[0005] b. a camera array;
[0006] c. a depth sensor;
[0007] d. a GPS, location sensor;
[0008] e. a gyroscope sensor;
[0009] f. an acceleration sensor;
[0010] g. an orientation sensor;
[0011] h. a laser;
[0012] i. other types of sensors; and
[0013] j. augmented reality / virtual reality sensors.
[0014] These display devices can include, for example:
[0015] a. a general-purpose smartphone display screen; and
[0016] b. an augmented reality / virtual reality-specific display device.
[0017] Of all the above sensors, there is a class of sensor suites that combine all of these sensors and display technologies to produce augmented reality (AR) and / or virtual reality (VR). These augmented reality-based (AR-based) and virtual reality-based (VR-based) software packages, such as AR Core, AR Kit, VR kit, etc., are increasingly being adapted for consumer applications.
[0018] SUMMARY
[0019] Systems and methods for augmented reality (AR) user guidance are described herein, where multiple embodiments of the systems and methods can include some or all of the elements, features, and steps described below.
[0020] The systems and methods described herein can provide accurate and simple user guidance instructions for a human user during a scanning process. The discussion below further illustrates AR-based guidance techniques.
[0021] In a method of the present disclosure, a three-dimensional model of a target object can be generated using a camera and a display screen with augmented reality on-screen guidance. An augmented reality component is computer-generated to guide a camera operator to position a camera in a particular position relative to the target object with a particular tilt orientation relative to the target object to capture an image that includes a region of the target object. The computer-generated augmented reality component is displayed on the display screen, where the camera operator can use the computer-generated augmented reality component to position the camera in the particular position with the particular tilt orientation in order to then capture the image. Then, the image is received from the camera (e.g., in a computer-readable memory of the camera).
[0022] Another method of the present disclosure can be used to generate a three-dimensional model of at least one body part of a person using a camera and a display screen with on-screen guidance based on human body pose detection. In this method, an image (taken by a camera operator) is received from a camera. The camera image includes at least one body part of a person to be modeled. The image of the at least one body part of the person is analyzed to identify a body pose of the at least one body part and a position and orientation of the at least one body part in the received image. On-screen guidance is provided on a display screen to guide positioning of the camera and positioning of the body part.
[0023] Also, in accordance with the present disclosure, a computer-readable storage medium storing program code that, when executed on a computer system, performs a method for generating a three-dimensional model by the methods described herein. BRIEF DESCRIPTION OF DRAWINGS
[0025] Figure 1 The flow of communication between components in an example guidance system is schematically illustrated, indicating that the subject 14 or the scanner 12 performs an adjustment in accordance with feedback given by the example guidance system.
[0026] Figure 2 A 3D scan of the subject 14 by the scanner 12 using a camera in the electronic device 16 is shown.
[0027] Figure 3Self-scanning is shown, in which the subject 14 scans himself by placing an electronic device 16 (such as a smartphone) with a camera 18 in a stationary position and rotating his body in front of the camera 18.
[0028] Figure 4 A screenshot of an application with multiple on-screen indicators 20 is shown. Figure 24 The multiple screen indicators 20 are generated on the screen of the smartphone 16 and aligned with the representation of the human figure 22 to provide feedback on whether the person being scanned 14 has achieved the desired pose.
[0029] Figure 5 It shows Figure 4 The subject 14 is scanned and an indicator 20 is displayed on the screen, wherein five body parts aligned with the indicator 20 are in the desired positions, and wherein the indicator 20 is colored (e.g., green) to indicate compliance.
[0030] Figure 6 It shows Figure 4 and Figure 5 On the screen, indicator 20 indicates that the feet of the person being scanned 14 are too close together, and the indicator 20” closer to the feet is a different color (red) than the indicators 20’ of other body parts, which are correctly positioned (their indicators are green).
[0031] Figure 7 It shows Figures 4-6 On the screen, indicator 20 indicates that the right arm of the person being scanned 14 is in the wrong position (as indicated by the coloring of the nearby indicator 20″), while the coloring of other indicators 20′ indicates that other body parts are in the desired position.
[0032] Figure 8 Another self-scanning scenario is shown, in which the subject 14 is using the rear camera of smartphone 16 to take 2D images of her feet in order to reconstruct a 3D foot model.
[0033] Figure 9 The image shows the use of an AR floor mat 26 and an augmented reality foot outline 28 when scanning the scanner's feet.
[0034] Figure 10 It shows the Figure 9 The use of AR floor mat 26, in which the defective scan portion 28 needs to be rescanned.
[0035] Figure 11 and Figure 12 The image shows a screenshot of a smartphone display running an application that requests preliminary data to be used when executing the user-bootstrapping method described herein.
[0036] Figure 13 、 Figure 14 and Figure 15 shows screenshots in the execution of an exemplary 3D foot scan process driven by an application executing the user guidance method described herein.
[0037] Figures 16-18 shows additional screenshots of the application in the execution of an exemplary 3D foot scan process, where the application provides instructions to the user to scan the floor.
[0038] Figure 19 and Figure 20 shows additional screenshots of the application in the execution of an exemplary 3D foot scan process, where the application provides additional instructions 32 to the user for size and scaling calibration.
[0039] Figure 21 is an illustration based on a real picture of a room, while Figure 22 and Figure 23 are screenshots of the application, which are captured when taking the camera image from different positions, where a thicker border is shown and the deployed AR components are shown, including virtual pillars 34 for phone positioning and orientation, and AR floor mat 26 and AR foot outline 28.
[0040] Figure 24 and Figure 25 shows additional screenshots of the application employed in the scan process using the exemplary 3D foot shape scan application, while Figure 26 shows an illustration based on a photo of a real floor, where a piece of paper is used as a reference object 36 for the scan.
[0041] Figure 27 、 Figure 28 and Figure 29 shows additional screenshots of the application employed in the scan process using the exemplary 3D foot shape scan application, where virtual pillars 34 extending from the floor are generated as AR components of the image.
[0042] Figure 30 、 Figure 31 and Figure 32 shows additional screenshots of the application employed in the scan process using the exemplary 3D foot shape scan application, which shows how the AR components 26, 28 and 34 guide the user to accomplish the task of moving his / her real smartphone to the desired position.
[0043] Figures 33-38Additional screenshots of the application are shown that were employed during the scanning process using the example 3D foot shape scanning application, which illustrate how the AR components 26, 28, 38, 40, 42, 44, 46, 48, and 50 guide the user through the task of tilting his / her smartphone to position it at the desired angle.
[0044] Figure 39 and Figure 40 A screenshot of the application is shown that displays some task completion indicators, while Figure 41 A screenshot of a transition indicator is shown.
[0045] Figure 42 and Figure 43 A screenshot of the application is shown that displays the scanning process using the example 3D body / foot scanning application, with the AR components 26, 28, and 34 turned on and off.
[0046] In the drawings, like reference numerals refer to same or similar functionalities throughout the several views; and primes are used to distinguish multiple instances of the same item or different embodiments of an item sharing the same reference numeral. The drawings are not necessarily to scale; instead, emphasis is placed on illustrating the particular principles in the examples discussed below. For any drawing that includes text (words, reference characters, and / or numbers), an alternate version of the drawing without the text will be understood to be part of the present disclosure; and thus, an official alternate drawing without such text can be substituted.
[0047] DETAILED DESCRIPTION
[0048] The foregoing and other features and advantages of various aspects of the present application will become apparent upon reading the following more particular description of the various concepts and specific embodiments, the purposes of which are to illustrate key principles underlying the application. Various aspects of the subject matter introduced above and discussed in greater detail below can be implemented in any of numerous ways, as the subject matter is not limited to any particular manner of implementation. Examples of specific implementations and applications are provided primarily for illustrative purposes.
[0049] Unless otherwise defined, used or characterized herein, terms used herein (including technical and scientific terms) will be interpreted as having a meaning that is consistent with their accepted meaning in the context of the relevant art and not in an idealized or overly formal sense, unless expressly defined otherwise herein. For example, if a particular shape is referenced, the shape is intended to include imperfect differences from the ideal shape, e.g., due to manufacturing tolerances. Unless otherwise noted, the following processes, procedures, and phenomena can occur at ambient pressure (e.g., about 50-120 kiloPascals — e.g., about 90-110 kiloPascals) and temperature (e.g., -20 to 50 degrees Celsius — e.g., about 10-35 degrees Celsius).
[0050] Although the terms first, second, third, etc. can be used herein to describe various elements, components, regions, layers and / or sections, these elements should not be limited by these terms. These terms are only used to distinguish one element from another. Thus, a first element discussed below could be termed a second element without departing from the teachings of the example embodiments.
[0051] Spatially relative terms, such as "up", "down", "left", "right", "front", "back", and the like, can be used herein for ease of description to describe one element's or feature's relationship to another element(s) or feature(s) as illustrated in the figures. It will be understood that the spatially relative terms are intended to encompass different orientations of the device in use or operation in addition to the orientations depicted in the figures. For example, if a device is inverted in the figure, a spatially relative term such as "below" or "under" can be interpreted to mean "above" or "over" in view of the device's inverted orientation. The exemplary term "above" can encompass both the above and below orientations. The device can be otherwise oriented (e.g., rotated 90 degrees or at other orientations) and the spatially relative descriptors used herein interpreted accordingly. The term "about" can refer to a range of ±10% of the stated value. In addition, where a range of values is provided, every sub-range and every individual value within the range is intended to be contemplated, and thus is disclosed.
[0052] Still further, in the disclosure, when an element is referred to as being "on", "connected to", "coupled to", "contacted to", etc., another element, it can be directly on, connected to, coupled to, or contacted to the other element or intervening elements can be present. Unless otherwise specified, the terms "connected" and "coupled" are used in the broadest context to encompass both direct and indirect connections and couplings.
[0053] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting of example embodiments. As used herein, the singular forms "a", "an" and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. Additionally, the terms "includes", "including", "comprises" and "comprising" are intended to be inclusive and mean that there can be additional elements or steps, which are not specifically listed, but which are present in addition to those specifically listed.
[0054] Smart devices (e.g., smart phones) can be used to generate a three-dimensional model (3D model) of a target object, such as a human user. The smart device first acquires three-dimensional (3D) information of the human user using sensors attached to the smart device. The 3D information can include the background of the scene around the human user, 2D images of the human user, and other useful 3D information of the human user. The smart device can then upload this information to an external device, or use the smart device's own computing power, i.e., itself, to reconstruct a 3D model of the human user. In this disclosure, from here on, this entire process is referred to as "3D scanning" or "3D scanning process", and the human user is referred to as "user".
[0055] If the smart device is a smart phone, the 3D scanning can be done by a smart phone application (app) in the form of processor-executable software code stored in the memory of the smart phone.
[0056] At some times, the 3D scanning needs to be performed by two users, where the first user holds the smart device to scan the second user. The first user is further distinguished as the scanner, while the second user is denoted as the scanned. At some times, the 3D scanning can be done by a single user; thus, it is referred to as self-scanning, and the user is referred to as both the scanner and the scanned. The camera operator is defined as the user who operates the smart device (e.g., a smart phone with a camera) and conducts the 3D scanning of the scanned. In the case of two users performing the scanning, the camera operator is the scanner, while in the case of self-scanning, the camera operator is the user himself / herself (both the scanner and the scanned). The smart device will collect 3D information of the scanned for use in constructing a 3D body model of the scanned.
[0057] During the 3D scanning process, a user guidance system is usually needed to help the user (i.e., the scanned and / or the scanner) to collect 3D information (e.g., pictures / videos of the scanned) of the scanned. In the case of the smart device being a smart phone, this user guidance system is usually part of the smart phone application. The user guidance system in the smart phone application tells the scanned and / or the scanner what should be done during the 3D scanning process in order to 3D scan the body / body parts of the scanned, e.g., where to stand relative to the camera position of the smart phone, how to pose in front of the smart phone camera, how to position the smart phone, how to tilt the smart phone to the correct shooting angle, how long to hold a certain pose, etc.
[0058] In this disclosure, we propose a user guidance system that utilizes AR / VR technology and / or human pose detection technology to read real-time data from sensors, dynamically compute the data from the sensors, and generate real-time guidance from a smart device to the user to guide the user step-by-step to perform a 3D scan. The user guidance system can include on-screen guidance, voice guidance, haptic feedback guidance, flash light guidance, etc.
[0059] As shown in Figure 1 high level, an example user guidance system includes the following components. During a 3D scanning process, the user guidance system takes the output data from any or all of the sensors 11 (e.g., cameras, depth cameras, gyroscope sensors, GPS sensors, motion sensors, etc.), analyzes the data and determines whether the user (the subject 14 or the scanner 12 or both) is performing the expected tasks (e.g., the subject has the correct pose, is standing at the correct position relative to the smart device, the background is correct, etc.) via a computing algorithm 13 (e.g., on a smartphone) stored on a computer readable medium in communication with a computer processor that processes the algorithm. The processor then issues instructions to generate appropriate real-time feedback 15 that is communicated to the user from the smart device via on-screen guidance, voice guidance, haptic guidance, etc. The user then adjusts his / her pose, moves himself / herself to the correct position relative to the smart device, moves, tilts or rotates the smart device to the correct position at the correct camera angle, cleans up the background environment, or moves to a better background, etc. in response to these prompts. The system functions as a feedback system that can generate clear instructions in real-time and communicate the clear instructions to the user 14 and / or the user 12 so that the user can easily follow the instructions to meet the requirements of the 3D scanning process.
[0060] A 3D scan can be performed by a scanner 12 operating a smart device 16 with one type of sensor (e.g., a smartphone with a camera or a depth sensor or a light detection and ranging “lidar” sensor) to capture sensor data (e.g., images, video, depth data, lidar data, etc.) of a subject 14 as shown in Figure 2 or by a subject 14 performing a self-scan as shown in Figure 3 The sensor data can include preliminary images surveying the scene prior to generating the guidance component, at least one image captured from the guidance component via the guidance, and additional images captured from the same or different perspective to collect additional data. Other types of sensors such as depth sensors or lidar sensors can be used in other examples to collect data that helps to assess the position, shape, orientation, etc. of the target.
[0061] Example One:
[0062] Figure 3A self-scanning scenario is shown in which the scanner 14 is using the front-facing camera 18 of a smartphone 16 and is facing the screen of the smartphone 16. As Figure 4 shown, the smartphone application first generates and communicates a command to produce a human outline on the screen as an on-screen guide to visually instruct the scanner 14 how to pose in a specific standing pose in front of the smartphone camera 18. As can be seen, this standing pose is a simple standing pose with the scanner's feet shoulder-width apart and arms hanging at his / her sides. Voice guidance via voice prompts can also deliver this instruction to the scanner 14.
[0063] The user guidance system then (a) uses the image / video stream captured by the camera 18 as real-time sensor input data, (b) processes the input data quickly, (c) conducts image analysis, and (d) uses algorithms for real-time human pose detection and extracts the scanner's 14 pose and his / her relative position from the smartphone. Note that suitable techniques and algorithms for determining user pose using computer vision are described in, for example, "Human posture recognition and classification", O. Khalifa et al., International Conference on Computing, Electrical and Electronics Engineering 40-43 (2013); "Human posture recognition based on images captured by the Kinect sensor", W. Wang et al., 13 International Journal of Advanced Robotic Systems 1-16 (2016); and "Human posture recognition using human skeleton provided by Kinect", T. Le et al., International Conference on Computing, Management and Telecommunications 340-45 (2013). This algorithm can detect in real time where the scanner's hands, feet, head, etc. are located.
[0064] Based on the human pose detection results, the guidance system generates in real time a plurality of on-screen guide indicators 20 to indicate whether the scanner's pose is in compliance with the requirements of the 3D scanning system. As Figure 4As shown, five colored circles (with color variations) next to the head, left hand, right hand, left foot, and right foot of a human-shaped silhouette are displayed on the screen of smartphone 16 and serve as on-screen guidance indicators 20 for compliance. These circles 20 change color to provide the scanned person with real-time feedback on whether they have correctly adopted the expected posture and / or are at the expected distance from the camera and / or are standing in the expected position in the room. In addition to on-screen guidance, voice guidance can also be used. Therefore, this interactive user guidance system can help the scanned person 14 correct their posture in real time.
[0065] like Figure 5 As shown, if each of the five parts of the subject's body to be scanned is in its corresponding predetermined position, all the colored circles 20 will be green.
[0066] like Figure 6 As shown, if the two feet of the person being scanned are too close together as required, the circle 20 next to the feet turns red, while the other circles 20 are green.
[0067] like Figure 7 As shown, in another case, if the feet and left arm of the person being scanned 14 are in the correct position and posture, but the right arm of the person being scanned is not in the required posture, then only the circle 20 next to the right arm turns red.
[0068] With this interactive, real-time, step-by-step on-screen and voice guidance, users can easily complete the scanning process based on their human posture detection results. In 3D scanning, especially in the downstream process where algorithms reconstruct the 3D model of the scanned subject contain any assumptions about the subject's posture, the subject's compliance with matching the specified posture is particularly important.
[0069] The example above illustrates only one scenario where the sensor input is an image / video generated by a camera. Other types of sensor input can be utilized and processed to generate real-time on-screen and voice guidance, thus guiding the scanned subject through the image / video / user data acquisition process for 3D model reconstruction.
[0070] Furthermore, in a two-user scanning scenario where the scanner holds a smartphone to perform a 3D scan of the subject, the scanner may also need to follow real-time screen and voice guidance to perform specific tasks, such as adjusting the distance between the smartphone and the subject and / or changing the camera's shooting angle.
[0071] Finally, on-screen and voice guidance can also be used to ensure that the background environment meets the requirements during the 3D scanning process.
[0072] Example Two:
[0073] Figure 8 Another self-scanning scenario is shown in which the scanner 14 is using the rear-facing camera of the smartphone 16 to take a 2D image and / or video of his / her feet (both feet) and is performing the 3D scanning of the feet by himself / herself while looking at the screen of the smartphone 16 during the 3D scanning process.
[0074] The smartphone 16 will first survey the background environment around the scanner and generate an AR floor mat 26 at a specific location in the room (typically the center of the room or a place where the background environment changes less, such as Figure 9 as shown in the screen shot of the smartphone display) as shown.
[0075] This AR floor mat 26 (as well as all other AR components described in this disclosure) does not exist in reality; rather, it is a computer-generated object that can only be observed in the perspective of the smartphone screen. Note that this is not just a drawing pasted on the display screen of the smartphone at a specific location; rather, it is an AR component that exhibits its proper optical perspective when the smartphone changes its orientation and camera angle of view and maintains its relative 3D position with respect to the rest of the real physical background.
[0076] Within the AR floor mat shown in Figure 9 as shown. These AR foot outlines 28 indicate to the scanner where his / her feet should be positioned during the 3D scanning of the feet. Figure 9
[0077] The on-screen guidance using the AR floor mat and AR foot outlines will also be accompanied by voice guidance and / or additional on-screen indicators to provide clear instructions to the scanner that the scanner will comply with.
[0078] For this particular example, after the scanner follows the instructions of the user guidance system to stand on the two AR foot outlines 28, he / she will perform a sweeping motion of his / her smartphone from his / her body left side to his / her body right side to take a video stream (or a series of picture snapshots) of his / her feet from his / her body left side to his / her body right side. During this sweeping motion, the edges of the AR floor mat 26 can turn into progress bars to indicate the progress (completion) of the 3D scanning process. As Figure 10 shown, for example, if the scanner misses a portion during the sweeping motion (e.g., possibly because the sweep was too fast in that portion), the color of the progress bar for that portion 30 turns to “red” and indicates to the scanner that the video stream taken from that particular portion needs to be recaptured.
[0079] Example Three:
[0080] In this example, we further present a complete step-by-step description of a user guidance system that employs AR technology and uses AR components as components of a real-time user guidance system to guide a user to complete a 3D scan in an efficient, interactive, and user-friendly manner. This example is also a self-scan of the foot (both feet), however the system and method can be applied to self-scan and multi-person scan scenarios, as well as 3D scans of any body part.
[0081] There are multiple figures associated with this example, some of which are illustrations of screenshots of the smartphone application when the user guidance system is activated; others are illustrations based on real-life scenario photos. To let the reader distinguish between these two different types of figures, we use the following convention from now on. For screenshots of the smartphone application, from Figure 11 onwards, the figures have thick black borders as shown in Figure 11 . Meanwhile, for illustrations based on real-life scenario photos, the figures do not have thick black borders as shown in Figure 21 .
[0082] The following provides a step-by-step explanation of how a user can use this smartphone application to self-scan his / her feet as instructed by the user guidance system.
[0083] Step 1: Information Intake
[0084] Figure 11 and Figure 12 are screenshots of an exemplary 3D foot scan application. Upon starting the execution of the application, the application generates screen prompts (e.g. printed text) that ask the user some questions to obtain some preliminary data. This data can be used to better construct the 3D model of the user’s body / body part (e.g. feet). These data can also be used to provide a better user experience.
[0085] In this exemplary application, it asks the user’s height (in metric units or in imperial units), and also asks what type of scaling reference object they are using (e.g. an A4 paper or a letter-sized paper).
[0086] Step 2: Norms and Tutorial
[0087] Figures 13-15 are screenshots of an exemplary 3D foot scan application. Upon completion of Step 1 (Information Intake), the application generates a quick tutorial of the 3D scanning process as well as some prompts and requirements on the display. The tutorial of the 3D scanning process can be in the format of pictures, videos, audio, and / or text. The prompt / requirement steps can also be interactive.
[0088] The screenshots presented in this step only show an example, which is in picture format with text explanation; however, the system and method can be employed to communicate with the user by using any of a number of other formats.
[0089] This exemplary application shows the user how much open space is needed to perform a 3D scan (in Figure 13 ); what type of floor tone is favorable for a more accurate 3D scan (in Figure 14 ); and where to place the scale reference object on the floor (in Figure 15 ).
[0090] Step 3: AR Environment Setup
[0091] Figures 16-18 are screenshots of an exemplary 3D foot scan application. After step 2 (specification and tutorial) is completed, the user-guided system asks the user to perform an AR environment setup; this setup can be performed by the user sweeping the smartphone around the user's real environment. Of course, the user can also perform this AR environment setup process by other means (e.g., by automatic displacement of a camera- incorporated device or by another person-generated motion), as the present invention encompasses the more general concept of setting up the AR environment as one of the steps in an AR-based user-guided system; therefore, this setup step is not limited to sweeping motion only.
[0092] However, as shown in Figures 16-18 , this exemplary application generates an indicator asking the user to perform a sweeping motion to scan the floor to complete the AR environment setup.
[0093] Step 4: Scale / Position Reference Scan
[0094] Figure 19 and Figure 20 are screenshots of an exemplary 3D foot scan application when it generates on-screen guidance 32 for scale calibration. The application asks the user to perform scale calibration using a standard-sized object or shape as a reference object to obtain actual dimensional information of the environment. In this exemplary application, the application uses an A4 or letter-sized paper. Note that there are other types of standard-sized reference objects that can be used for scale calibration, such as a credit card, a library card, a driver's license, a general ID card, a coin, a soda can, the smartphone itself, etc. Generally, this method can use a standard-sized object to calibrate the scale factor of the AR environment and / or the scanning system, and is not limited to an A4 or letter-sized paper.
[0095] Note that two or more standard-sized reference objects can be used together in this scale calibration process, and this scenario is also included in the more general concept of the method.
[0096] With this exemplary application, we demonstrate the process by using a single standard-sized reference object for the AR system and / or 3D scanning system's scale calibration. The application generates on-screen guide indicators 32, as shown, prompting the user to aim at the standard-sized reference object and, as shown, hold the camera still to complete the task. Figure 19 Figure 20
[0097] Step 5: AR Deployment
[0098] Figure 21 are illustrations based on real-life photos of the room where the 3D scanning is being performed, while Figure 22 and Figure 23 are screen shots of an exemplary 3D foot scanning application. After Step 4 (Calibration), the application deploys AR components 26, 28, and 34 to the on-screen guide, along with other user guide components, as shown in Figure 22 and Figure 23
[0099] The motivation for implementing these AR components into the user guide system is described as follows. It is difficult to perform accurate 3D body scanning using a smartphone, especially when it requires the user to take photos / videos of the user's body / body part from the correct distance and at the correct angle. In most cases, the accuracy of the 3D model generated from these pictures / videos depends on the proper position and angle of the camera relative to the scanned body / body part. Picture, voice, and / or text guidance would help; however, by incorporating AR components into the user guide system, it would be much easier for the user to accurately understand what needs to be done in the actual physical space to comply with the system requirements.
[0100] AR components can also transform the somewhat boring 3D body scanning process into a gamified enjoyable experience, which provides more motivation for the user to perform this task.
[0101] Figure 21 are illustrations based on real-life photos of the room where the 3D scanning is being performed, while Figure 22 Figure 23 The same set of deployed AR components 26, 28, 34, and 35 are shown from different perspectives. As you can see, even though there is only one sheet of paper (as reference object 36) in reality on an empty floor - in the AR environment generated by the app on the smartphone display, there are multiple computer-simulated objects (AR components) 26, 28, 34, and 35. As the smartphone changes perspective between Figure 22 and Figure 23 , these AR components exhibit proper optical perspective and maintain their relative 3D positions with respect to the rest of the physical background environment. AR components 26, 28, 34, and 35 will be used in the next few steps to guide the user to complete the 3D scan.
[0102] There are several different types of AR components in this example, including the following:
[0103] • multiple virtual pillars 34 extending from the surface of the real floor, with a virtual carton smartphone 35 at the top of each virtual pillar 34;
[0104] • a virtual floor mat 26; for this example, the virtual floor mat 26 is divided into four sections, with each section facing a virtual pillar 34; and
[0105] • a pair of virtual footprints 28 that provide guidance to the user about where to stand.
[0106] The detailed functionality of each of these AR components 26, 28, and 34 will be detailed in the next few steps.
[0107] Step 6: AR-based guidance - user positioning
[0108] Figure 24 and Figure 25 are screen shots of an example 3D foot scan app. The user guidance system of this app displays some AR components 26, 28, and 34 on the app screen as on-screen guidance to guide the user to stand in the correct position on the floor. Figure 24 and Figure 25 Two of the three AR components shown in Figures 6A and 6B are being used in this step: the virtual floor mat 26 and the pair of virtual footprint outlines 28.
[0109] The app uses on-screen guidance and voice commands to guide the user to stand on the virtual footprint outlines 28. The edges of the virtual floor mat 26 will be used as a progress indicator 20 for the user as he / she takes multiple photos of his / her feet from four different angles to perform the 3D scan. The virtual floor mat 26 can also be used to display trademarks, advertisements, slogans, and other commercial uses on the screen through instructions provided by the app.
[0110] As Figure 26 shown, a representation based on a photo of the actual floor (in reality) is provided; as we can see, it is just an empty floor with a piece of paper on it (used as a reference object 36).
[0111] Step 7: AR-based guidance - camera distance and shooting angle
[0112] As Figures 27-29 shown, additional screenshots of the exemplary 3D foot scanning application are provided. These screenshots show the virtual pillar 34 on the floor, with the virtual smartphone 35 on top of the virtual pillar 34. Figures 27-29 Three screenshots are shown, which illustrate what the scanned person sees on his / her smartphone display when he / she moves the smartphone from the left side of his / her left leg to the right side of his / her left leg. One can see that the characteristics and behavior of these AR components are very realistic.
[0113] As Figures 30-32 shown, another set of screenshots of the exemplary 3D foot scanning application are provided. These screenshots show how the AR components 26, 28, 34 and 35 will guide the user to accomplish the task of moving his / her real smartphone (and thus, the camera of his / her smartphone) to the specified location in reality.
[0114] When no user action is taken (i.e., the user keeps the smartphone still in real life), the virtual pillar 34 remains stationary, while the virtual paper-box smartphone 35 swings up and down. The swinging motion of the virtual paper-box smartphone 35, together with the real-time voice guidance and text guidance from the application, instructs the user to move the smartphone in real life to the location of the virtual paper-box smartphone 35. Figures 30-32 The swinging motion of the virtual paper-box smartphone 35 is shown when the screenshots are taken at different times while the user keeps the smartphone still in real life. One can see that the swinging motion of the virtual paper-box smartphone 35 is a function of time.
[0115] Once the user has followed the on-screen guidance and voice guidance and has moved the real smartphone to the location of the virtual paper-box smartphone 35, the virtual pillar 34 and the virtual paper-box smartphone 35 then disappear to indicate that the user has completed this operation.
[0116] Note that if the user moves the smartphone in real life away from this required location, the disappeared virtual pillar 34 and the virtual paper-box smartphone 35 will reappear to indicate that the location of the user's smartphone no longer meets the requirements of the user guidance system; and then, the user needs to accomplish this step again.
[0117] Once the user has completed this step, the user smartphone camera in real life x-y-z axis position is fixed; in the next step, the user will follow the user guidance system to tilt the smartphone camera to the specified angle, aiming at the user's foot in real life, for image / video capturing.
[0118] As Figures 33-38 shown, screen shots of an exemplary 3D foot scanning application are provided. These screen shots show how the AR component is used to guide the person being scanned 14 to complete the task of tilting his / her smartphone (and thus, the camera of his / her smartphone) at a specified angle in real life. Note that the foot shown on the screen is an image of the real foot of the person being scanned 14 generated by the application using the camera of the smartphone, not the AR component.
[0119] At the top of the application screen in Figures 33-35 there is a virtual paper box smartphone 35 and a text message next to it. The angle of tilt of the virtual paper box smartphone 35 reflects the actual angle of tilt of the real smartphone of the user in real life. The text message generated by the application on the screen next to the virtual paper box smartphone 35 reads "Tilt your phone to match".
[0120] Also as Figures 33-35 shown, the application generates a gray virtual circular ring structure 40, a white virtual double arrow 56 attached to the gray virtual circular ring structure 40, and a red virtual line segment 38 with a rounded end in the middle part of the screen, however the choice of color or pattern used for each component is arbitrary and non-limiting in order to distinguish the different components from each other. The gray virtual circular ring structure 40 indicates the point on the floor that the real smartphone camera of the user is currently pointing at; the white virtual double arrow 56 indicates the direction in which the user guidance system instructs the user to change the pointing direction of the smartphone camera in real life. The length of the red virtual line segment 38 with a rounded end indicates the distance between the point on the floor that the real smartphone camera of the user is currently pointing at and the point on the floor that the user guidance system instructs the smartphone camera to point at.
[0121] The rounded end in the red virtual line segment 38 indicates the final desired pointing point of the smartphone camera, which is the upper part of the virtual footprint outline 28. As can be seen, the guidance and voice guidance instructions on the screen instruct the person being scanned 14 to point his / her real smartphone camera at his / her foot.
[0122] The screen guides the user with its AR components and its non-AR components, such as actual video streams and text and voice guidance, to guide the scanned person to complete the following task: to tilt his / her real smartphone (and thus its camera) to a specified angle to aim at the virtual footprint outline 28 (and thus at his / her right foot, as the user's foot has been positioned on the virtual footprint outline 28 in the previous step), so that the real smartphone camera can take a photo / video of the scanned person's 14 foot from the specified shooting angle.
[0123] In the previous step, the x-y-z position of the real-life user's smartphone camera was locked, and now the tilt angle of the real-life user's smartphone camera is locked in this step. Thus, all degrees of freedom of the smartphone (and thus the camera) are locked by the user guidance system.
[0124] Figures 33-35 A series of screenshots are shown when the user gradually tilts his / her real smartphone from an initial tilt angle that is parallel to the floor (in Figure 33 ) to a tilt angle that is about 45 degrees relative to the floor (in Figure 35 ). During this process, the gray virtual circular ring structure 40 moves with the user's real smartphone's motion and gets closer and closer to the final desired pointing point on the virtual footprint outline 28 (representing the scanned person's 14 real foot); at the same time, the red virtual line segment 38 gets shorter and shorter.
[0125] Figure 36 When the pointing of the real smartphone's camera is very close to the required pointing point, the red virtual line segment 38 is replaced by a beige circular target 44, and the gray virtual circular ring structure 40 is replaced by a beige rifle scope 42. This instantaneous AR component change is generated to further help the user (the scanned person) to control his / her tilt motion more finely, so that the user can carefully tilt the real-life smartphone camera to point at the point represented by the beige circular target 44.
[0126] Figure 37 When the pointing of the user's real smartphone's camera is exactly on the required pointing point, the colors of the "beige circular target 44" and "beige rifle scope 42" both change to green (48 and 46); a large font text that reads "hold still" is generated in the center of the screen. This generation of text occurs when the user's smartphone camera is at the correct x-y-z position and at the correct tilt angle. This guidance message instructs the user to hold still when applying the photo(s) and / or video(s) of the user's foot. These pictures and / or videos will be used for the 3D reconstruction of the user's foot.
[0127] Figure 38It is shown that when the application is such that pictures / videos are taken, a progress bar 50 is displayed for the user.
[0128] This step 7 is repeated several times around the scanned person's foot at different angles of shooting to capture images / videos of his / her foot from different angles. An algorithm will be used to reconstruct a 3D model of the user's foot based on these pictures / videos. One example of a system that builds a 3D model of a target object through its images is a computer vision system. Two general references for computer vision systems are: "Computer Vision: A Modern Approach" by David A. Forsyth and Jean Ponce, Prentice Hall, 2002; and "Multiple View Geometry in Computer Vision" 2nd Edition by Richard Hartley and Andrew Zisserman, Cambridge University Press, 2004.
[0129] A similar process is repeated for the other foot of the scanned person to complete the 3D reconstruction of the scanned person's feet.
[0130] The above three examples illustrate the details of the novel user guidance system that utilizes AR components and / or human pose detection technology and conveys instructions to the user through on-screen guidance and voice guidance to help the user with the 3D scanning. Some additional features not mentioned above include the following:
[0131] Feature A: Task completion indicators and transition indicators
[0132] Figures 39-41 Some screen shots of task completion indicators are provided in the Appendix; they show screen shots of task transition indicators with and without AR components. AR components can be used as part of the task completion indicators and task transition indicators.
[0133] Feature B: Stepwise release of guidance components
[0134] Figure 42 and Figure 43 Screen shots of the exemplary 3D scanning application in Example Three are provided. These screen shots show that not all AR components need to be turned on at the same time. Depending on the task that the user guidance system instructs the user to complete, it can be a good practice to turn on a specific set of AR components during a specific task and turn off all other AR components; this way, the user can focus on the current task and irrelevant AR components will not distract the user. Figure 42 It is shown that all the virtual supports 34 that the application will deploy throughout the process of the 3D scanning of the user's feet. Figure 43It is shown that when alignment instructions are provided to take a photo from a specific position and angle, the app only deploys one virtual prop 34 for that specific position and angle of taking the photo, thereby enabling the user to focus on that task at that moment, while all other virtual props 34 are hidden via the instructions generated by the app.
[0135] Feature C: Non-AR components in the guidance
[0136] In addition to the AR components of the user guidance system, other sensor inputs and / or smartphone outputs can be used to generate instructions as part of the user guidance system. These sensor inputs can come from the smartphone’s camera, GPS, wifi readings, 4G signal readings, gyroscope sensors, magnetic sensors, geographic sensors, accelerometers, depth cameras, IR cameras, camera arrays, lidar sensors / scanners, etc. These smartphone outputs can include voice messages, images, sounds, text, the smartphone’s camera flash, the smartphone’s vibration, or even a voice call from a customer service department directly to the smartphone to issue direct verbal guidance instructions to the user completing a task using the app.
[0137] Feature D: Gamification of the 3D body scan
[0138] The systems and methods described herein can make the 3D scanning task fun and interactive for the user by implementing AR components and / or human pose detection technology in the user guidance system.
[0139] The app can employ a one-step: increase a reward system, where once the user completes a specific task in the 3D body scan process, he / she can earn points and benefits.
[0140] A reward system can also be employed to encourage the user to use his / her 3D body after the scan is complete; for example, when the user uses his / her body to purchase certain brands of clothing through the app’s own e-commerce system, the app can reward the user with points or discount coupons.
[0141] Computer implementation:
[0142] While most of the discussion above is exemplified through a smartphone (a form of computing device) running an app, the systems and methods can equally be implemented / executed using any of a variety of computing devices, including a camera and an output device (e.g., a display screen), or in communication with a camera and an output device.
[0143] In any event, the computing device operates as a system controller and can include a logic device, such as a microprocessor, microcontroller, programmable logic device, or other suitable digital circuitry for executing control algorithms; and the systems and methods of the present disclosure can be implemented in a computing system environment. Examples of well-known computing system environments that can be suitable for use with the systems and methods include, but are not limited to, personal computers, hand-held or laptop devices, tablet devices, smart phones, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, distributed computing environments that include any of the above systems or devices, or the like. Typical computing system environments and their operations and components are described in numerous existing patents, such as U.S. Patent No. 7,191,467 to Microsoft Corporation.
[0144] The methods can be performed via non-transitory computer-executable instructions, such as program modules, for example, in the form of applications. Generally, program modules include routines, programs, objects, components, and data structures that perform particular tasks or implement particular types of data. The methods can also be practiced in a distributed computing environment where tasks are performed by remote processing devices that are linked through a communications network. In a distributed computing environment, program modules can be located in both local and remote computer storage media including memory storage devices.
[0145] The processes and functions described herein can be stored non-transitorily in the form of software instructions (e.g., an application as described above) in a computing device. Components of the computing device can include, but are not limited to, a computer processor; a computer storage medium that functions as memory; and a system bus that couples the various system components, including the memory, to the computer processor. The system bus can be any of several types of bus structures including a memory bus or memory controller, a peripheral bus, and a local bus using any of a variety of bus architectures.
[0146] The computing device can include one or more of various computer-readable media, which include volatile and nonvolatile media, and removable and non-removable media. By way of example, computer-readable media can include computer storage media and communication media.
[0147] Computer storage media can store software and data that when executed by processors, can cause the processors to perform various actions. The above-described apparatuses and methods of embodiments of the present disclosure can be directed to computer storage media with computer executable instructions stored thereon. For example, computer readable storage media can bear computer readable instructions thereon which, when executed by a computer processor(s), can cause the processor(s) to carry out the operations described herein. These computer readable instructions can represent a computer program, procedure, or process. Computer storage media can be non-transitory, in that it can be a tangible medium which can store programming for access by a processor. The programming stored by the computer storage media can describe the fabric or structure of the programming and be stored for a long duration. In some examples, the computer storage media can be a non-transitory medium because it can store programming for an extended period of time. In some examples, the computer storage media can be non-transitory because it can store programming persistently, over time. In some examples, the computer storage media can be non-transitory because it can employ program instructions that can be stored permanently.
[0148] Memory includes computer storage media in the form of volatile and / or nonvolatile memory, such as read only memory (ROM) and random access memory (RAM). A basic input / output system (BIOS) typically contains the basic routines that help to transfer information between elements within the computer, such as during start-up. The RAM typically contains data and / or program modules that are immediately accessible to and / or presently being operated on by the processor.
[0149] The computing device also can include other removable / non-removable, volatile / non-volatile computer storage media, such as for example (a) a storage card reader / writer, (b) a hard disk drive, to read from and / or write to non-removable, nonvolatile magnetic media; (c) a magnetic disk drive, to read from and / or write to a removable, non-volatile magnetic disk; and (d) an optical disk drive, to read from and / or write to a removable, non-volatile optical disk such as a CD ROM or other optical media. The computer storage media can be coupled to the system bus by a storage interface. Storage interface can include, for example, an interface to transfer digital or optical signals. Other removable / non-removable, volatile / non-volatile computer storage media that can be used with the exemplary operating environment include, but are not limited to, magnetic tape cassettes, flash memory cards, digital versatile disks, digital video tape, solid state RAM, solid state ROM, and the like. The drives and their associated computer readable media provide storage of computer executable instructions, data structures, program modules, and other data for the computer. For example, a hard disk drive can store an operating system, application programs, and program data.
[0150] The drives and their associated computer readable media, such as the hard disk drive and computer readable media described above, provide storage of computer executable instructions, data structures, program modules, and other data for the computer. In example, a computer internal or external hard disk drive can store an operating system, application programs, and program data.
[0151] The computing device can also include a network interface controller in communication with the processor and with input / output devices that communicate with external devices such as printers and wireless network routers, to which external devices are also in wireless communication.
[0152] Further examples consistent with the present teachings are set forth in the following numbered clauses:
[0153] 1. A method for generating a three-dimensional model of a target object using a camera and a display screen having augmented reality on-screen guidance, comprising:
[0154] using augmented reality, computer-generated augmented reality components to guide a camera operator to position the camera in a particular position relative to the target object with a particular tilt orientation relative to the target object to capture an image that includes a region of the target object;
[0155] displaying the computer-generated augmented reality components on the display screen, wherein the camera operator can use the computer-generated augmented reality components to position the camera in the particular position with the particular tilt orientation to capture the image; and
[0156] receiving the image from the camera (e.g., in a computer-readable memory in the camera).
[0157] 2. The method of clause 1, further comprising generating a three-dimensional model of the target object from the received camera image.
[0158] 3. The method of clause 1 or 2, further comprising:
[0159] using augmented reality, computer-generated additional augmented reality components to guide the camera operator to position the camera in at least a second particular position relative to the target object with at least a second particular tilt orientation relative to the target object to capture an additional image that includes a second region of the target object;
[0160] displaying the additional augmented reality components on the display screen to guide the camera operator to position the camera in at least the second particular position with at least the second particular tilt orientation to capture the additional image; and
[0161] receiving the additional image from the camera.
[0162] 3.1 The method of clause 3, further comprising iteratively removing or changing the appearance of augmented reality components from the display screen as the camera operator completes capturing images of one region of the target object and begins capturing images of another region of the target object.
[0163] 4. The method of clauses 1-3.1, wherein the target object comprises at least one body part of a living being.
[0164] 5. The method of clause 4, further comprising:
[0165] repeating the method of claim 3 at selected intervals over an extended period of time; and
[0166] comparing the three-dimensional models generated from camera images taken at different times to assess changes in the living being’s adaptability or growth of the living being over time.
[0167] 6. The method of clauses 4 or 5, wherein the living being is a human.
[0168] 7. The method of clause 6, wherein the human is the camera operator.
[0169] 8. The method of clause 7, further comprising presenting commands on the screen to instruct the human to move the at least one body part to a favorable position and posture for image capture.
[0170] 9. A method for generating a three-dimensional model of at least one body part of a human using a camera and a display screen with on-screen guidance based on human posture detection, comprising:
[0171] receiving images taken by a camera operator from a camera, wherein the camera images include the at least one body part of the human to be modeled;
[0172] analyzing the images of the at least one body part to identify a body posture of the at least one body part and a position and orientation of the at least one body part in the received images;
[0173] using the identified body posture, position, and orientation of the at least one body part, providing real-time on-screen guidance to the camera operator to position and orient the camera at a specified position and orientation for the camera to capture at least a second additional image of the at least one body part; and
[0174] using the identified body pose, position, and orientation, providing real-time on-screen guidance for the person to assume another intended pose in front of the camera so that the camera captures at least a fourth image of the at least one body part in the intended pose.
[0175] 10. The method of clause 9, further comprising:
[0176] using the identified body pose, position, and orientation of the at least one body part, providing real-time on-screen guidance to the camera operator to position and orient the camera at another specified position and orientation so that the camera captures at least a third image of the at least one body part; and
[0177] using the identified body pose, position, and orientation, providing real-time on-screen guidance for the person to assume another intended pose in front of the camera so that the camera captures at least a fourth image of the at least one body part in the intended pose.
[0178] 11. The method of clause 10, further comprising generating a three-dimensional model of the target object from the received camera images.
[0179] 12. The method of any one of clauses 9-11, wherein the camera operator is the person being imaged and modeled.
[0180] 13. The method of any one of clauses 9-11, wherein the camera operator is a person other than the person having the at least one body part being imaged and modeled.
[0181] 14. The method of any one of clauses 9-13, further comprising providing voice commands that accompany the on-screen guidance.
[0182] 15. The method of any one of clauses 9-14, further comprising producing an augmented reality component that provides a visual representation indicating a position and tilt angle for camera placement to generate the camera images as part of the on-screen guidance.
[0183] 16. The method of any one of clauses 6-15, further comprising manufacturing a wearable garment sized and profiled to fit the at least one body part or selecting from pre-manufactured garments of different sizes to fit the at least one body part.
[0184] 17. The method of any one of clauses 6-16, further comprising analyzing the images and identifying body parts of the person from the images.
[0185] 18. The method of any of clauses 6-17, further comprising generating a body part outline on the display screen using augmented reality to indicate where the person positions and orients the at least one body part.
[0186] 19. The method of any of clauses 1-18, wherein the camera and the display screen are incorporated into a smartphone.
[0187] 20. The method of clause 19, further comprising generating a three-dimensional model of the target object from data from at least one sensor.
[0188] 21. The method of clause 15, wherein the sensor is the camera, and wherein the data comprises at least one camera image.
[0189] 22. The method of clause 15, wherein the sensor is selected from a depth sensor and a lidar sensor.
[0190] 23. The method of any of clauses 1-22, further comprising:
[0191] generating a visual representation of a scaling reference object on the screen using augmented reality to indicate where the camera operator is to place the scaling reference object;
[0192] capturing at least one image of the scaling reference object after placement by the camera operator; and
[0193] determining dimensions of the target object using information from the image of the scaling reference object and known dimensions of the scaling reference object, or by positioning the camera at a fixed distance from the scaling reference object to provide explicit dimensional scaling information by providing instructions on the screen to the camera operator.
[0194] 23.1. The method of any of clauses 1-23, further comprising using at least one of sound, voice, haptic feedback, and a flash as an additional part of the guidance to the camera operator.
[0195] 24. A computer-readable storage medium storing program code that, when executed on a computer system, performs a method for generating a three-dimensional model of a target object using a camera and a display screen with augmented reality on-screen guidance, the method comprising:
[0196] generating instructions to display a computer-generated augmented reality component on the display screen, wherein the computer-generated augmented reality component is usable by an operator of the camera to position the camera in the particular position with the particular tilt orientation to capture the image.
[0197] generating instructions to display a computer-generated augmented reality component on the display screen, wherein the computer-generated augmented reality component is usable by an operator of the camera to position the camera in the particular position with the particular tilt orientation to capture the image.
[0198] 25. A computer-readable storage medium storing program code that, when executed on a computer system, performs a method for generating a three-dimensional model of at least one body part of a person using a camera and a display screen having augmented reality on-screen guidance, the method comprising:
[0199] receiving an image taken by a camera operator from a camera, wherein the camera image includes the at least one body part of the person to be modeled;
[0200] analyzing the image of the at least one body part of the person taken with the camera to identify a body pose of the at least one body part and a location and orientation of the at least one body part in the received image; and
[0201] using the identified body pose, location, and orientation, generating instructions to provide real-time on-screen guidance for the person to assume an intended pose in front of the camera so that the camera captures at least a second image of the at least one body part assuming the intended pose.
[0202] In describing the embodiments of the application specific terminology is used for the sake of clarity. The specific terminology is intended to indicate that the technology and the functional equivalents thereof operate in a similar manner and achieve similar results in accordance with the description of the embodiments. Furthermore, in some instances in which a particular embodiment of the application includes a number of system elements or method steps, those elements or steps can be replaced by a single element or step. Similarly, a single element or step can be replaced by a number of elements or steps serving the same purpose. Moreover, where parameters or other values are specified herein for the embodiments of the application, such parameters or values can be adjusted up or down by 1 / 100, 1 / 50, 1 / 20, 1 / 10, 1 / 5, 1 / 3, 1 / 2, 2 / 3, 3 / 4, 4 / 5, 9 / 10, 19 / 20, 49 / 50, 99 / 100 (or rounded to the nearest such increment), unless otherwise indicated. Furthermore, although the application has been described with reference to specific embodiments thereof, it will be apparent to one of ordinary skill in the art that various alterations and modifications to the embodiments described herein can be made without departing from the scope of the application. Furthermore, other aspects, features, and advantages of the application are apparent from the scope of the application; and it is further intended that all such additional aspects, features, and advantages be within the scope of the application. Moreover, steps, elements, and features of the embodiments described herein can be used in combination with each other, unless otherwise indicated. The contents of all references, including references to literature and patents, journal articles, and patent applications cited throughout this application are hereby expressly incorporated by reference in their entirety for all purposes; and all such references and the embodiments, features, properties, and methods described therein can be associated with the embodiments of the application. Furthermore, components and steps identified in the Background section as being known in the art are incorporated by reference in the disclosure and can be used in connection with the embodiments of the application or can be replaced by other components and steps that are within the scope of the disclosure. In method claims that are presented in a phase-ordered format, the phases are not to be interpreted as being temporally ordered unless otherwise indicated or implied by the terminology and context.
Claims
1. A method for generating a three-dimensional model of a target object using a camera and a display screen guided on an augmented reality screen, comprising: Using augmented reality, computer-generated and displayed augmented reality components, the augmented reality components including a virtual cardboard box at a location on the display screen to guide a camera operator to align the camera with the location of the virtual cardboard box on the display screen, wherein the location of the virtual cardboard box corresponds to a specific location relative to the target object, at which the camera is capable of capturing an image including a target area of the target object; Adjust at least one of the augmented reality components to provide an indication of a specific tilt orientation of the camera on the display screen, at which the camera is capable of capturing an image of the target region including the target object; and After the camera operator uses the camera to capture the image, the image is received from the camera, wherein the camera is aligned with the position of the virtual cardboard box and has a tilt orientation indicated by at least one augmented reality component on the display screen.
2. The method of claim 1 further includes generating a three-dimensional model of the target object from the received camera image.
3. The method according to claim 1, further comprising: Using augmented reality, computer-generated and displayed additional augmented reality components, the additional augmented reality components including the virtual cardboard box at an updated position on the display screen to guide the camera operator to position the camera so as to align the camera with the updated position of the virtual cardboard box on the display screen, wherein the updated position of the virtual cardboard box corresponds to at least a second specific position relative to the target object, at the updated position, the camera captures an additional image including a second target area of the target object; Adjust at least one of the augmented reality components to provide an indication of an updated tilt orientation of the camera on the display screen, at which the camera is capable of capturing the additional image including the second target region of the target object; and After the camera operator uses the camera to capture an image, the additional image is received from the camera, wherein the camera is located at the updated position corresponding to the virtual cardboard box and has the updated tilt orientation indicated by the at least one augmented reality component on the display screen.
4. The method of claim 3, further comprising iteratively removing augmented reality components from the display screen or changing the appearance of augmented reality components as the camera operator completes capturing an image of one area of the target object and begins capturing an image of another area of the target object.
5. The method of claim 1, further comprising using at least one of sound, voice, haptic feedback, and flash as an additional component for guiding the camera operator.
6. The method according to claim 1, wherein, The target object includes at least one body part of an organism.
7. The method according to claim 6, further comprising: Repeat the method according to claim 3 at selected intervals within the extended time period; Each time the method according to claim 3 is repeated, a three-dimensional model of the body part is generated from the received camera images; and The three-dimensional model generated from camera images taken at different times is compared to assess the organism's adaptability or changes in the organism's growth over time.
8. The method according to claim 6, wherein, The organism in question is a human being.
9. The method according to claim 8, wherein, The person in question is the camera operator.
10. The method of claim 9, further comprising presenting a command on a screen to instruct the person to move the at least one body part to an advantageous position and posture for image capture.
11. The method of claim 8, further comprising manufacturing wearable garments whose size and profile are tailored to fit the at least one body part, or selecting from pre-made garments of different sizes to fit the at least one body part.
12. The method of claim 8, further comprising analyzing the image and identifying body parts of the person from the image.
13. The method of claim 8, further comprising using augmented reality to generate body part outlines on the display screen to indicate where the person locates and orients the at least one body part.
14. The method according to claim 1, wherein, The camera and the display screen are integrated into the smartphone.
15. The method of claim 14, further comprising generating a three-dimensional model of the target object from data from at least one sensor, wherein, The sensor: (a) is the camera, and the data includes at least one camera image, or (b) is selected from a depth sensor and a lidar sensor.
16. The method according to claim 1, further comprising: Augmented reality is used to generate a visual representation of a scaled reference object on the screen to instruct the camera operator where to place the scaled reference object. After placement by the camera operator, at least one image of the scaled reference object is captured; as well as The size of the target object is determined using information from an image of the scaling reference object and the known size of the scaling reference object, or by providing on-screen instructions to the camera operator to position the camera at a fixed distance from the scaling reference object to provide explicit size scaling information.
17. The method of claim 1, further comprising: A preliminary image is received from the camera, wherein... The preliminary image was captured by the camera and includes the target object; and The augmented reality component is generated using the initial image.
18. The method according to claim 1, wherein, The virtual cardboard box has the shape of a smartphone, and the augmented reality component further includes a virtual support, on which the virtual cardboard box is mounted to indicate the height at which the camera is to be positioned.
19. A computer-readable storage medium storing program code, which, when executed on a computer system, performs a method for generating a three-dimensional model of a target object using a camera and a display screen guided on an augmented reality screen, the method comprising: Instructions are provided to generate computer-generated augmented reality components, including a virtual cardboard box displayed at a location on the display screen, to guide an operator of the camera to align the camera with the location of the virtual cardboard box on the display screen, wherein the location of the virtual cardboard box corresponds to a specific location relative to the target object, at which the camera is capable of capturing an image including a target area of the target object; and Instructions are generated to adjust at least one of the augmented reality components to provide an indication of a specific tilt orientation of the camera on the display screen, at which the camera is able to capture an image of the target region including the target object.
20. The computer-readable storage medium according to claim 19, wherein, The virtual cardboard box has the shape of a smartphone, and the augmented reality component further includes a virtual support, on which the virtual cardboard box is mounted to indicate the height at which the camera is to be positioned.
Citation Information
Patent Citations
Automated three dimensional model generation
US20160284123A1
Foot measuring and sizing application
US20180160777A1
Augmented reality for three-dimensional model reconstruction
US20190028637A1
Image processing methods and systems
US20210321035A1