Method, system and storage medium for posture guidance

By using the context of the camera viewfinder flow in the imaging system to determine the sample posture image and generate flexible and accurate posture guidance, the problems of low efficiency and inaccurate posture guidance in the existing system are solved, and efficient and accurate posture guidance is achieved.

CN114816041BActive Publication Date: 2025-05-13ADOBE INC
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202111265211.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2021-01-19
Filing Date
2021-10-28
Publication Date
2025-05-13
Estimated Expiration
2041-10-28

AI Technical Summary

Technical Problem

Existing imaging systems are inefficient when providing posture guidance, and users need to frequently switch applications to imitate professional postures, resulting in wasted system resources and inaccurate guidance.

Method used

By determining the context in the camera viewfinder stream, a sample pose image is provided for the user and a flexible and accurate pose guidance is generated based on the selected sample pose image.

Benefits of technology

It realizes efficient provision of posture guidance in a single interface, reduces the waste of system resources, and improves the flexibility and accuracy of posture guidance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114816041B_ABST
    Figure CN114816041B_ABST
Patent Text Reader

Abstract

The present disclosure relates to providing contextual augmented reality photo pose guidance. The present disclosure relates to systems, non-transitory computer-readable media, and methods for generating and providing context-tailored pose guidance for a camera viewfinder stream. Specifically, in one or more embodiments, the disclosed system determines the context of the camera viewfinder stream and provides a sample pose image corresponding to the determined context. In response to a selection of the sample pose image, the disclosed system generates and displays pose guidance customized to the proportions of an object depicted in the camera viewfinder stream. The disclosed system also iteratively modifies portions of the generated pose guidance to indicate that an object depicted in the camera viewfinder stream is being aligned with the generated pose guidance. When the object is fully aligned with the generated pose guidance, the disclosed system automatically captures a digital image.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0001] In recent years, there have been significant improvements in imaging systems. For example, conventional imaging systems provide a vivid camera viewfinder display and capture colorful and detailed digital images via a mobile device. Specifically, the imaging experience provided by conventional imaging systems includes using a mobile device camera viewfinder to position a mobile device's camera, and then capturing a digital image in response to a user interaction with a shutter function of the mobile device.

[0002] Often, users want to imitate (or instruct others to imitate) interesting and engaging poses they see in professional photography or other media. To do this, conventional imaging systems typically require users to access web browsers and other applications on their mobile devices to find professional images that include poses for imitation. Thereafter, conventional imaging systems also require users to navigate and interact between these web browsers and other applications to view and imitate (or instruct others to imitate) the displayed poses. Therefore, conventional imaging systems are particularly inefficient when operating in conjunction with client computing devices (such as smartphones or tablets) with small screens, where it is difficult to interact with and switch between multiple applications to find and imitate professional photography poses.

[0003] Moreover, by forcing the user to estimate and guess with respect to mimicking gestures between applications, conventional imaging systems result in various system-level inefficiencies. For example, in forcing the user to switch back and forth between applications in order to mimic professional and attractive gestures, conventional imaging systems result in excessive use and ultimate waste of system resources associated with generating graphical displays, storing user selections, maintaining application data, and capturing digital images. Additionally, given the guesswork involved in attempting to mimic gestures between applications, conventional imaging systems waste additional system resources in capturing and deleting digital images that fail to render in the manner the user intended.

[0004] Even though conventional imaging systems provide some degree of in-application gesture guidance, such conventional imaging systems are generally inflexible and inaccurate. For example, to provide some degree of gesture guidance, conventional imaging systems are limited to static, silhouette-based overlays. To illustrate, conventional imaging systems may provide gesture guidance as a generic humanoid gesture silhouette overlaid onto a camera viewfinder of a client computing device.

[0005] This degree of posing guidance provided by conventional imaging systems is inflexible. For example, as discussed, conventional imaging systems provide a one-size-fits-all posing guidance that is not constrained by the proportions, characteristics, and attributes of the person who is posing. Thus, conventionally provided posing guidance is too rigid to adapt to the body of any particular poser.

[0006] Furthermore, such conventionally provided pose guidance is extremely inaccurate. For example, the pose guidance provided by conventional imaging systems is non-specific with respect to the context and position of the posing user within the camera viewfinder. As a result, conventional imaging systems often inaccurately capture digital images in which the posing user is in a pose that is different from the pose indicated by the pose guidance and / or the posing user is in a pose that is inappropriate relative to the context of the posing user.

[0007] These and additional problems and issues exist with conventional imaging systems. Summary of the invention

[0008] The present disclosure describes one or more embodiments of systems, non-transitory computer-readable media, and methods that solve one or more of the foregoing or other problems in the art. Specifically, the disclosed system determines and provides sample pose images tailored for the context of a user's camera viewfinder stream. For example, the disclosed system determines the context of the camera viewfinder stream based on objects, backgrounds, clothing, and other characteristics depicted in the camera viewfinder stream. The disclosed system then identifies sample pose images corresponding to the determined context. The disclosed system provides the identified sample pose images as optional display elements overlaid on the camera viewfinder, so that the user can select a specific sample pose image to imitate without having to switch to a different application.

[0009] In addition to providing context-specific sample pose images via a camera viewfinder, the disclosed system also generates and provides pose guidance based on the selected sample pose images. For example, the disclosed system extracts an object body frame representing an object (e.g., a person) depicted in a camera viewfinder stream. The disclosed system also extracts a reference body frame representing the person depicted in the selected sample pose image. For example, in order to generate pose guidance, the disclosed system redirects the reference body frame based on the proportions of the object body frame. The disclosed system then overlays the redirected reference body frame onto the camera viewfinder by aligning the redirected reference body frame with landmarks relative to the object. When the object moves a body part to align with the pose indicated by the redirected reference body frame, the disclosed system modifies the display characteristics of the redirected reference body frame overlaid on the camera viewfinder to indicate alignment. In response to determining the complete alignment between the object and the redirected reference body frame, the disclosed system optionally automatically captures a digital image from the camera viewfinder stream.

[0010] Additional features and advantages of one or more embodiments of the present disclosure are outlined in the description which follows, and in part will be obvious from the description, or may be learned by practicing such example embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] The detailed description provides one or more embodiments with additional specificity and detail through the use of the accompanying drawings, as briefly described below:

[0012] Figure 1 Illustrated is a diagram of an environment in which an AR gesture system operates in accordance with one or more embodiments.

[0013] Figures 2A to 2G An AR gesture system that provides sample gesture images and interactive gesture guidance via a camera viewfinder of a client computing device is illustrated in accordance with one or more embodiments.

[0014] Figure 3A An overview of an AR gesture system that determines a context associated with a camera viewfinder stream of a client computing device and provides sample gesture images corresponding to the determined context is illustrated in accordance with one or more embodiments.

[0015] Figure 3B An AR gesture system is illustrated that determines different subsets of a set of sample gesture images in accordance with one or more embodiments.

[0016] Figure 4A An AR gesture system that provides gesture guidance via a camera viewfinder of a client computing device is illustrated in accordance with one or more embodiments.

[0017] Figure 4B An AR gesture system utilizing a gesture neural network to extract a body frame from a digital image is illustrated in accordance with one or more embodiments.

[0018] Figure 4C and Figure 4D An AR gesture system that redirects a reference body frame based on the scale of an object depicted in a camera viewfinder stream of a client computing device is illustrated in accordance with one or more embodiments.

[0019] Figure 4E and Figure 4F An AR gesture system is illustrated that aligns a redirected reference body frame with a target object in accordance with one or more embodiments.

[0020] Figure 5A An AR gesture system is illustrated in accordance with one or more embodiments, whereby alignment between gesture guidance and objects depicted in a camera viewfinder stream is determined, and display characteristics of the gesture guidance are modified to iteratively and continuously provide gesture guidance to a user of a client computing device.

[0021] Figure 5B Illustrated are additional details regarding how an AR gesture system determines alignment between a redirected portion of a reference body frame and an object depicted in a camera viewfinder of a client computing device in accordance with one or more embodiments.

[0022] Figure 6 A schematic diagram of an AR gesture system is illustrated in accordance with one or more embodiments.

[0023] Figure 7 Illustrated is a flow diagram of a series of acts for generating an interactive overlay of context-tailored sample gesture images in accordance with one or more embodiments.

[0024] Figure 8 A flow diagram illustrating a series of actions for generating augmented reality gesture guidance in accordance with one or more embodiments is illustrated.

[0025] Fig. 9 Illustrated is a flow diagram of another series of actions for automatically capturing a digital image in response to determining perfect alignment between an object depicted in a camera viewfinder stream and a gesture guide in accordance with one or more embodiments.

[0026] Fig.10 A block diagram of an example computing device for implementing one or more embodiments of the present disclosure is illustrated. DETAILED DESCRIPTION

[0027] The present disclosure describes one or more embodiments of an augmented reality (AR) gesture system that provides interactive augmented reality gesture guidance via a camera viewfinder based on context-related sample gesture images. For example, the AR gesture system determines a context associated with a camera viewfinder stream of a client computing device and identifies a set of sample gesture images corresponding to the determined context. In response to detecting a user selection of one of the sample gesture images, the AR gesture system generates AR gesture guidance based on a body frame extracted from the camera viewfinder stream and the selected sample gesture image. The AR gesture system also aligns the AR gesture guidance with an object in the camera viewfinder, and iteratively determines the alignment between a portion of the AR gesture guidance and a corresponding body part of the object in the camera viewfinder. In response to determining that all AR gesture guidance portions and corresponding body parts are aligned, the AR gesture system captures a digital image from the camera viewfinder stream.

[0028] In more detail, the AR gesture system optionally determines a context associated with a camera viewfinder stream of a client computing device based on an analysis of a digital image from the camera viewfinder stream. For example, the AR gesture system extracts a digital image (e.g., an image frame) from the camera viewfinder stream of the client computing device. The AR gesture system also analyzes the digital image to determine an object (e.g., a person) within the digital image. The AR gesture system performs additional analysis on the digital image to determine object tags, gender tags, and clothing tags associated with the digital image. In one or more embodiments, the AR gesture system determines the context of the digital image based on the determined tags associated with the identified object.

[0029] In response to determining the context of a digital image from a camera viewfinder stream of a client computing device, the AR gesture system generates a collection of sample gesture images corresponding to the determined context. For example, in one embodiment, the AR gesture system generates the collection by querying one or more sample gesture image repositories and search engines using a search query based on one or more context tags associated with the digital image. To illustrate, in one or more embodiments, the AR gesture system generates a search query that includes one or more context tags associated with an object depicted in the digital image, a scene depicted in the digital image, and other objects depicted in the digital image. The AR gesture system also utilizes the generated search query in conjunction with one or more sample gesture image repositories, which include but are not limited to: a local sample gesture image repository, a general search engine, and other third-party applications.

[0030] In one or more embodiments, the AR gesture system optimizes the limited amount of display space shared by the client computing devices by providing different subsets of the set of sample gesture images. For example, in one embodiment, the AR gesture system utilizes one or more clustering techniques to group similar sample gesture images from the identified set of sample gesture images together. The AR gesture system also provides different subsets of sample gesture images by identifying and providing sample gesture images from each group or cluster.

[0031] In one or more embodiments, the AR gesture system provides different subsets of the set of sample gesture images via a camera viewfinder of a client computing device. For example, the AR gesture system generates an interactive overlay including different subsets of the sample gesture images. The AR gesture system also positions the interactive overlay on the camera viewfinder of the client computing device. In one or more alternative embodiments, the AR gesture system retrieves a plurality of commonly selected gesture images, determines a plurality of popular gesture images, or otherwise determines a set of gesture images to provide without determining the context of a camera viewfinder.

[0032] In response to a selection of a sample pose image detected from an interactive overlay, the AR pose system generates and provides AR pose guidance via a camera viewfinder. For example, in at least one embodiment, the AR pose system generates AR pose guidance that indicates how an object depicted in a camera viewfinder stream should position one or more body parts to simulate a pose depicted in a selected sample pose image. In one or more embodiments, the AR pose system generates the AR pose guidance by extracting an object body frame representing the object from a camera viewfinder stream of a client computing device. The AR pose system then extracts a reference body frame representing the pose from the selected sample pose image. Finally, the AR pose system generates the AR pose guidance by redirecting the reference body frame based on the proportions of the object body frame.

[0033] The AR gesture system provides the redirected reference body frame as an AR gesture guide via a camera viewfinder of the client computing device. In one or more embodiments, for example, the AR gesture system provides the AR gesture guidance by generating a visualization of the redirected reference body frame. The AR gesture system then anchors at least one predetermined point of the visualization to at least one landmark of the object depicted in the camera viewfinder stream. Thus, the AR gesture system provides the AR gesture guidance via the camera viewfinder so that the user of the client computing device can see how the object's body is aligned with the gesture indicated by the AR gesture guidance.

[0034] The AR gesture system iteratively determines alignment between portions of the redirected reference body frame and portions of the object depicted in the camera viewfinder stream. For example, in at least one embodiment, the AR gesture system aligns both the object body frame and the redirected reference body frame with one or more regions (e.g., hip region, chest region) of the object depicted in the camera viewfinder. The AR gesture system then iteratively determines alignment of one or more segments of the object body frame with corresponding segments of the redirected reference body frame.

[0035] In one or more embodiments, for each determined segment alignment, the AR gesture system modifies the display characteristics (e.g., color, line width) of the aligned segment of the redirected reference body frame. Thus, the AR gesture system provides a simple visual cue to the user of the client computing device indicating whether the object in the camera viewfinder correctly simulates the gesture from the selected sample gesture image. In response to determining that all segments of the object's body frame are aligned with corresponding segments of the redirected reference body frame, the AR gesture system captures a digital image from the camera viewfinder stream. For example, the AR gesture system optionally captures the digital image automatically in response to determining the alignment between the gesture guide and the object. In other embodiments, when the AR gesture system determines the alignment between the gesture guide and the object, the AR gesture system captures the digital image in response to the user's selection of a shutter button selection.

[0036] Although the embodiments discussed herein focus on a single object in a camera viewfinder stream and a single object in a selected sample pose image, the AR pose system is not limited thereto, and in other embodiments generates and provides pose guidance for multiple objects depicted in the camera viewfinder stream. For example, the AR pose system extracts body frames for multiple objects depicted in the camera viewfinder stream. The AR pose system also extracts body frames for multiple pose objects depicted in the selected sample pose image. The AR pose system then redirects and anchors the pose guidance to each object depicted in the camera viewfinder stream, and iteratively determines the alignment between the pose guidance and the multiple objects.

[0037] As mentioned above, the AR gesture system provides many advantages and benefits over conventional imaging systems. For example, rather than requiring the user to access and switch between multiple applications to find sample gesture images, the AR gesture system provides sample gesture images in an interactive overlay located on the camera viewfinder of the client computing device. Thus, the AR gesture system provides an efficient single-interface method for providing gesture guidance in conjunction with a camera viewfinder.

[0038] Additionally, the AR gesture system overcomes and improves upon various system-level inefficiencies common to conventional imaging systems. To illustrate, by avoiding application and interface switching common to conventional imaging systems, the AR gesture system efficiently utilizes system resources to generate a single interactive overlay including sample gesture images, and positions the overlay on the camera viewfinder of the client computing device. Thus, the AR gesture system avoids the use and ultimate waste of system resources associated with generating, maintaining, and otherwise persisting additional user interfaces and applications.

[0039] Moreover, the AR gesture system also improves the efficiency of conventional imaging systems by providing sample gesture images that are context-specific to the scene depicted in the camera viewfinder stream. For example, in situations where a conventional imaging system is unable to provide specific gesture guidance, the AR gesture system identifies and provides sample gesture images that are specific to the objects and scenes depicted in the camera viewfinder stream. Thus, the AR gesture system avoids the waste of system resources involved in multiple user searches for sample gesture images that are specific to the objects and scenes depicted in the camera viewfinder stream.

[0040] The tailored gesture guidance method provided by the AR gesture system is also flexible and accurate. For example, where some conventional imaging systems provide a generic, contour-based overlay to try to assist users in simulating various gestures, the AR gesture system generates and provides a specific reference body frame that is tailored to the proportions of the object depicted in the camera viewfinder stream. Thus, the AR gesture system provides gesture guidance that is specific to the object's body. Moreover, the AR gesture system anchors the gesture guidance to the object within the camera viewfinder, so that if the object moves within the camera viewfinder stream, the gesture guidance moves with the object.

[0041] As indicated by the foregoing discussion, the present disclosure utilizes various terms to describe the features and advantages of the disclosed AR gesture system. Additional details are now provided regarding the meaning of such terms. For example, as used herein, the term "digital image" refers to a collection of digital information representing an image. More specifically, a digital image is composed of pixels, each including a digital representation of color and / or grayscale. Pixels are arranged two-dimensionally in a digital image, where each pixel has a spatial coordinate including an x ​​value and a y value. In at least one embodiment, a "target digital image" refers to a digital image to which editing can or will be applied. In one or more embodiments, a digital image is stored as a file (e.g., a ".jpeg" file, a ".tiff" file, a ".bmp" file, a ".pdf" file).

[0042] As used herein, the term "pose" refers to the configuration of an object. Specifically, a pose includes an arrangement of joints (e.g., of a human figure) and / or segments connecting the joints. In some embodiments, a pose includes a visible depiction of the joints and segments, while in other cases, a pose includes a computerized representation of the joint locations and / or segment locations. In some cases, a pose includes an abstract representation of the joint locations and / or segment locations represented using vectors or other features (e.g., depth features) in a pose feature space or a pose prior space.

[0043] Relatedly, "joint" refers to the endpoints of segments that join a depicted figure or virtual human model. For example, a joint refers to a location where two or more segments connect. In some embodiments, a joint includes a location where segments rotate, pivot, or otherwise move with respect to each other. In some cases, a joint includes a computerized or abstract vector representation corresponding to the location of the joint of the depicted figure or virtual human model.

[0044] Along these lines, a "segment" refers to a representation or depiction of a length or portion of a portrait or virtual human model. In some embodiments, a segment refers to a line or other connector between joints of the depicted portrait or virtual human model. For example, a segment represents an upper arm between a shoulder joint and an elbow joint, a forearm between an elbow joint and a wrist joint, or a thigh between a hip joint and a knee joint. In some cases, a segment includes a computerized or abstract vector representation of a line or connecting component between two joint locations of the depicted portrait or virtual human model.

[0045] As used herein, an "object" refers to a likeness, depiction, or description of a human or humanoid figure within a digital image. For example, an object includes a captured depiction of a real person within a digital image, a drawing of a human figure in a digital image, a cartoon depiction of a character in a digital image, or some other humanoid figure in a digital image, such as a humanoid machine, creature, stick figure, or other likeness. In some cases, an object includes one or more arms, one or more legs, a torso, and a head. Although many of the example embodiments described herein include human figures, the gesture system is not limited thereto, and in other embodiments, the gesture search system operates with respect to other human figures such as animals, animated characters, and the like.

[0046] As used herein, a "sample pose image" refers to a digital image depicting one or more objects in a pose. For example, the AR pose system determines and provides one or more sample pose images including poses that are contextually relevant to the objects depicted in the camera viewfinder of the client computing device. In one or more embodiments, the sample pose images also include backgrounds, additional objects, clothing, and / or metadata describing the content of the sample pose images. In at least one embodiment, the AR pose system accesses the sample pose images from a private image repository, a public image repository, an additional application, and / or a search engine.

[0047] As used herein, a "body frame" refers to a representation of the joints and segments of an object. For example, a body frame representing a human object includes joint representations associated with the subject's hips, knees, shoulders, elbows, etc. The body frame also includes segment representations associated with the upper and lower arms, thighs and calves, etc. In at least one embodiment, the body frame also includes a circular representation of the subject's head. As used herein, a "reference body frame" refers to a body frame representing an object depicted in a sample pose image. As used herein, a "subject body frame" refers to a body frame representing an object depicted in a camera viewfinder of a client computing device.

[0048] The term "neural network" refers to a machine learning model that is trained and / or tuned based on inputs to determine a classification or approximate an unknown function. For example, the term neural network includes a model of interconnected artificial neurons (e.g., organized in layers) that communicate and learn to approximate complex functions and generate outputs based on multiple inputs provided to the neural network. In some cases, a neural network refers to an algorithm (or a collection of algorithms) that implements deep learning techniques to model high-level abstractions in data. For example, a neural network includes a convolutional neural network, a recursive neural network (e.g., an LSTM neural network), a graph neural network, or a generative neural network.

[0049] As used herein, the term "gesture neural network" refers to a neural network that is trained or tuned to identify gestures. For example, a gesture neural network determines the pose of a digital image by processing a digital image to identify the location and arrangement of joints and segments of a person depicted in the digital image. As another example, a gesture neural network determines the pose of a virtual human model by processing a virtual human model to identify the location of joints and segments of the virtual human model. Additional details regarding the architecture of a gesture neural network are provided in more detail below.

[0050] Additional details about the AR gesture system will now be provided with reference to the accompanying drawings. For example, Figure 1 A schematic diagram of an example system environment for implementing AR gesture system 102 according to one or more embodiments is illustrated. An overview of AR gesture system 102 is provided with reference to FIG. Figure 1 Thereafter, a more detailed description of the components and processes of AR gesture system 102 is provided with respect to subsequent figures.

[0051] As shown, the environment includes (multiple) servers 106, client computing devices 108, sample gesture image repository 112, (multiple) one or more third-party systems 116, and network 114. Each of the components of the environment communicates via network 114, and network 114 is any suitable network through which computing devices communicate. Fig.10 The example network is discussed in more detail.

[0052] As mentioned, the environment includes a client computing device 108. The client computing device 108 includes one of a variety of computing devices, including a smartphone, a tablet, a smart TV, a desktop computer, a laptop computer, a virtual reality device, an augmented reality device, or a computer system related to a computer. Fig.10 Another computing device described. Figure 1A single client computing device 108 is illustrated, but in some embodiments, the environment includes multiple different client devices, each associated with a different user (e.g., a digital image editor). The client computing device 108 communicates with the server(s) 106 via a network 114. For example, the client computing device 108 receives user input from a user interacting with the client computing device 108 (e.g., via the image capture application 110), such as to select a sample pose image. The AR pose system 102 receives information or instructions to generate a set of sample pose images from one or more sample pose image repositories 112 or (multiple) third-party systems 116.

[0053] As shown, the client computing device 108 includes an image capture application 110. Specifically, the image capture application 110 is a web application, a native application (e.g., a mobile application, a desktop application, etc.) installed on the client computing device 108, or a cloud-based application, wherein all or part of the functionality is performed by (multiple) servers 106. The image capture application 110 presents or displays information to the user, including a camera viewfinder (including a camera viewfinder stream), an interactive overlay including one or more sample pose images, pose guidance including a redirected reference body frame, and / or additional information associated with a determined context of the camera viewfinder stream. The user interacts with the image capture application 110 to provide user input to perform the above-mentioned operations, such as selecting a sample pose image.

[0054] like Figure 1 As shown, the environment includes (multiple) servers 106. (Multiple) servers 106 generate, track, store, process, receive and send electronic data, such as digital images, search queries, sample pose images and pose guidance. For example, (multiple) servers 106 receive data from client computing devices 108 in the form of digital images from a camera viewfinder stream to identify sample pose images corresponding to the context of the digital image. In addition, (multiple) servers 106 send data to client computing devices 108 to provide sample pose images corresponding to the context of the camera viewfinder stream and one or more pose guidance for display via the camera viewfinder. In fact, (multiple) servers 106 communicate with client computing devices 108 to send and / or receive data via network 114. In some embodiments, (multiple) servers 106 include distributed servers, wherein (multiple) servers 106 include multiple server devices distributed on network 114 and located at different physical locations. (Multiple) servers 106 include content servers, application servers, communication servers, web hosting servers, multidimensional servers or machine learning servers.

[0055] The image capture system 104 communicates with the client computing device 108 to perform various functions associated with the image capture application 110, such as storing and managing a repository of digital images, determining or accessing tags of digital content depicted within digital images, and retrieving digital images based on one or more search queries. For example, the AR gesture system 102 communicates with a sample gesture image repository to access sample gesture images. In practice, as Figure 1 As further shown, the environment includes a sample pose image repository 112. Specifically, the sample pose image repository 112 stores information such as a repository of digital images depicting objects in various poses within various scenes and various neural networks including pose neural networks. The AR pose system 102 also communicates with (multiple) third-party systems 116 to access additional sample pose images.

[0056] like Figure 1 As shown, a particular arrangement of the environment is illustrated, but in some embodiments, the environment has a different arrangement of components and / or may have a completely different number or set of components. For example, Figure 1 It is shown that the client computing device 108, the server(s) 106, or both may implement the AR gesture system 102. In some embodiments, the AR gesture system 102 is implemented by (e.g., entirely located at) the client computing device 108. In some embodiments, the client computing device 108 may download the AR gesture system 102 from the server(s) 106. Alternatively, the AR gesture system 102 is implemented by the server(s) 106, and the client computing device 108 accesses the AR gesture system 102 through a web-hosted application or website. Still further, the AR gesture system 102 is implemented on both the client computing device 108 and the server(s) 106, and some functions are performed on the client computing device 108 while other functions are performed on the server(s) 106. Additionally, in one or more embodiments, the client computing device 108 communicates directly with the AR gesture system 102, bypassing the network 114. Further, in some embodiments, sample pose image repository 112 is located external to server(s) 106 (eg, communicated via network 114 ), or is located on server(s) 106 and / or client computing device 108 .

[0057] In one or more embodiments, the AR gesture system 102 receives a digital image from a camera viewfinder stream of a client computing device 108 and utilizes various image analysis techniques to determine a context of the digital image. The AR gesture system 102 then identifies and provides at least one sample gesture image corresponding to the determined context of the digital image. In response to a detected selection of the provided sample gesture image, the AR gesture system 102 generates and provides a gesture guide so that a user of the client computing device 108 can easily see how the body of the object depicted in the camera viewfinder stream is aligned with the gesture represented in the selected sample gesture image. The AR gesture system 102 iteratively determines that various body parts of the object are aligned with the gesture guide, and updates one or more display characteristics of the gesture guide to indicate to the user of the client computing device 108 that the object is correctly simulating the gesture depicted in the selected sample gesture image. In response to determining that the object depicted in the camera viewfinder stream is aligned with the gesture guide overlaid onto the camera viewfinder, the AR gesture system 102 automatically captures a digital image from the camera viewfinder stream without any additional input from the user of the client computing device 108. In an alternative implementation, AR gesture system 102 captures a digital image in response to a user selection of a shutter button.

[0058] Figures 2A to 2G The AR gesture system 102 is illustrated as providing sample gesture images and interactive gesture guidance via a camera viewfinder of the client computing device 108. For example, Figure 2A 2 shows a camera viewfinder 202 of a client computing device 108. In one or more embodiments, the camera viewfinder 202 displays a camera viewfinder stream of digital images or image frames depicting an object 204 within a scene 206. Figure 2A As further shown, objects 204 include clothing (e.g., clothes, jewelry, accessories), gender, and other attributes. In addition, scenes 206 include buildings, plants, backgrounds, vehicles, sky areas, other objects, animals, and the like.

[0059] Figure 2BThe AR gesture system 102 is shown providing an entry point option 210 associated with the AR gesture system 102 in conjunction with a camera viewfinder stream of the camera viewfinder 202. In one or more embodiments, in response to detecting initialization of the image capture application 110 on the client computing device 108, the AR gesture system 102 provides a message box 208 including the entry point option 210. Additionally or alternatively, the AR gesture system 102 provides the message box 208 including the entry point option 210 in response to determining that the camera viewfinder stream in the camera viewfinder 202 depicts: one or more objects, one or more object poses (e.g., not moving for more than a threshold amount of time), and / or a scene including at least one object (e.g., a human in front of a background). For example, the AR gesture system 102 utilizes one or more neural networks to make any of these determinations. Additionally or alternatively, in response to determining that the object 204 is posing for a photo, the AR gesture system 102 provides the message box 208 including the entry point option 210. For example, in response to determining that object 204 has not moved within camera viewfinder 202 for a threshold amount of time, AR gesture system 102 optionally provides message box 208 .

[0060] In response to detecting the selection of entry point option 210, AR gesture system 102 determines the context of the camera viewfinder stream. Figure 3A As discussed in more detail, the AR gesture system 102 utilizes one or more machine learning models, neural networks, and other algorithms to generate various tags or identifiers associated with digital images from the camera viewfinder stream of the client computing device 108. To illustrate, the AR gesture system 102 utilizes object detectors, object detectors, and other detectors to generate one or more object tags, gender tags, clothing tags, and / or other tags associated with the digital image. For example, the AR gesture system 102 utilizes these detectors to generate tags indicating the content of the scene depicted in the digital image (e.g., buildings, plants, vehicles, animals, cars) and the attributes of the objects within the scene (e.g., clothing type, hair type, facial hair). In one or more embodiments, the AR gesture system 102 determines the context associated with the digital image based on one or more determined tags.

[0061] AR gesture system 102 also provides a set of sample gesture images corresponding to the determined context of the digital image. For example, in one embodiment, AR gesture system 102 generates a search query based on the determined context, and utilizes the search query in conjunction with sample gesture image repository 112 and one or more of third-party systems 116 to generate a set of sample gesture images. To illustrate, in response to determining that the context of the digital image is a bride and groom at a wedding, AR gesture system 102 generates a set of sample gesture images including images of other brides and grooms wearing wedding dresses in a range of poses (e.g., including professional models, popular images, celebrities).

[0062] In at least one embodiment, AR gesture system 102 also identifies different subsets of the set of sample gesture images. For example, AR gesture system 102 avoids providing multiple sample gesture images that depict the same or similar gestures. Therefore, in one or more embodiments, AR gesture system 102 clusters the sample gesture images in the set of sample gesture images based on similarity. AR gesture system 102 also identifies different objects of the sample gesture images by selecting the sample gesture images from each cluster in the clusters. In at least one embodiment, AR gesture system 102 utilizes k-means clustering to identify different objects of the set of sample gesture images.

[0063] like Figure 2C As shown, AR gesture system 102 generates interactive overlay 212 including different subsets of sample gesture images, including sample gesture images 214a, 214b, and 214c. In one or more embodiments, AR gesture system 102 generates interactive overlay 212 including a maximum threshold number of sample gesture images 214a through 214c. Additionally or alternatively, AR gesture system 102 generates interactive overlay 212 having sample gesture images 214a through 214c in a horizontally slidable portion such that AR gesture system 102 displays additional sample gesture images in response to a detected horizontal sliding touch gesture.

[0064] like Figure 2CAs further shown, the AR gesture system 102 positions the interactive overlay 212 on the camera viewfinder 202. For example, the AR gesture system 102 positions the interactive overlay 212 in a lower portion (e.g., horizontal lower half, horizontal lower third) of the camera viewfinder 202 such that a majority of the camera viewfinder 202 is unobstructed. In an alternative embodiment, the AR gesture system 102 positions the interactive overlay 212 in a vertical portion (e.g., vertical half) or upper portion of the camera viewfinder 202. In an alternative embodiment, the AR gesture system 102 selectively positions the interactive overlay 212 based on the content of the camera viewfinder stream. For example, in response to determining that an object is in the lower third of the camera viewfinder stream, the AR gesture system 102 positions the interactive overlay 212 in the upper third of the camera viewfinder.

[0065] like Figure 2C As further shown, in at least one embodiment, the AR gesture system 102 generates an interactive overlay 212 that includes an indication 216 of a context of the digital image. For example, the AR gesture system 102 provides an indication of a context (e.g., "man standing in front of a building") including one or more identified objects and other labels associated with a digital image acquired from a camera viewfinder stream.

[0066] Additionally, if Figure 2C As shown, AR gesture system 102 provides a search button 218 within interactive overlay 212. In one or more embodiments, in response to detecting a selection of search button 218, AR gesture system 102 provides a search interface in which a user of client computing device 108 enters a search query for additional sample gesture images. In addition, using the search query provided by the user, AR gesture system 102 identifies additional sample gesture images (e.g., from sample gesture image repository 112 and / or (multiple) third-party systems 116). In at least one embodiment, AR gesture system 102 displays the additional identified sample gesture images by expanding interactive overlay 212. Alternatively, AR gesture system 102 generates and provides an additional user interface that includes the additional identified sample gesture images, and shifts display focus from camera viewfinder 202 to the additional user interface.

[0067] In one or more embodiments, the AR gesture system 102 generates and provides gesture guidance corresponding to the selected sample gesture image. For example, in response to detecting the selection of the sample gesture image 214a, the AR gesture system 102 generates and provides gesture guidance 220, such as Figure 2D In at least one embodiment, and as will be described below with respect to Figures 4A to 4FAs discussed in more detail, the AR gesture system 102 generates the gesture guidance 220 by extracting an object body frame from a digital image acquired from a camera viewfinder stream that represents the proportions of the object 204. The AR gesture system 102 also extracts a reference body frame from the selected sample gesture image 214a that represents the proportions and gestures of the person in the selected sample gesture image 214a. The AR gesture system 102 also generates the gesture guidance 220 by reorienting the reference body frame based on the proportions of the object body frame. Figure 2D As shown, the resulting pose guidance 220 retains the pose shown in the selected sample pose image 214 a , but also has the proportions of the object 204 .

[0068] The AR gesture system 102 aligns the gesture guide 220 with the object 204 by overlaying the gesture guide 220 based on one or more landmarks of the object 204. For example, the AR gesture system 102 anchors the gesture guide 220 to at least one landmark of the object 204, such as a hip region of the object 204. With the gesture guide 220 so anchored, the AR gesture system 102 maintains the positioning of the gesture guide 220 relative to the object 204 even when the object 204 moves within the camera viewfinder 202. Moreover, the AR gesture system 102 anchors the gesture guide 220 to additional regions of the object 204, such as a chest region of the object 204. With this additional anchoring, the AR gesture system 102 maintains the position of the gesture guide 220 relative to the object 204 even when the object 204 rotates toward or away from the client computing device 108.

[0069] like Figure 2E As further shown, AR gesture system 102 iteratively determines alignment between portions of gesture guide 220 and portions of an object depicted in camera viewfinder 202. For example, in response to determining the alignment, AR gesture system 102 modifies one or more display characteristics of the aligned portion of gesture guide 220. For illustration, Figure 2E As shown, in response to determining that segments 222a, 222b, 222c, 222d, 222e, and 222f of gesture guide 220 are aligned with corresponding portions of object 204, AR gesture system 102 modifies the display colors of segments 222a through 222f of gesture guide 220. By iteratively modifying the display characteristics of the segments of gesture guide 220, AR gesture system 102 provides effective guidance to the user of client computing device 108 when positioning object 204 into the gesture depicted by selected sample gesture image 214a.

[0070] In one or more embodiments, AR gesture system 102 continues to iteratively determine the alignment between gesture guide 220 and object 204. For example, Figure 2FAs shown, AR gesture system 102 has modified the display color of the additional segment of gesture guide 220 to indicate that object 204 has moved the corresponding body part to align with gesture guide 220. In at least one embodiment, AR gesture system 102 also modifies the display characteristics to indicate that the body part of object 204 is no longer aligned with the corresponding segment of gesture guide 220. For example, AR gesture system 102 changes the color of the segment of gesture guide 220 back to the original color to indicate that the corresponding body part is no longer aligned.

[0071] like Figure 2G 204, and in response to determining that all portions of the gesture guide 220 are aligned with corresponding portions of the object 204, the AR gesture system 102 automatically captures a digital image from the camera viewfinder stream shown in the camera viewfinder 202 of the client computing device 108. Additionally or alternatively, the AR gesture system 102 automatically captures the digital image in response to determining that all portions of the gesture guide 220 are aligned with corresponding portions of the object 204 for a threshold amount of time (e.g., 2 seconds). Additionally or alternatively, the AR gesture system 102 automatically captures the digital image in response to determining that a predetermined number or percentage of portions of the gesture guide 220 are aligned with a corresponding number or percentage of portions of the object 204. In other embodiments, the AR gesture system 102 captures the digital image in response to a user selection of a shutter button.

[0072] In one or more embodiments, AR gesture system 102 saves the automatically captured digital image in local storage on client computing device 108. Additionally or alternatively, AR gesture system 102 saves the automatically captured digital image in sample gesture image repository 112, enabling AR gesture system 102 to use the automatically captured digital image as a sample gesture image for the same or additional users of AR gesture system 102. Additionally or alternatively, AR gesture system 102 also automatically uploads the automatically captured digital image to one or more social media accounts associated with the user of client computing device 108.

[0073] Figure 3AThe AR gesture system 102 is illustrated as determining a context associated with a camera viewfinder stream of a client computing device 108 and providing an overview of sample gesture images corresponding to the determined context. For example, the AR gesture system 102 determines the context of an image from the camera viewfinder stream of the client computing device 108 based on the content of the digital image. The AR gesture system 102 then generates a collection of sample gesture images by generating a search query based on the determined context and utilizing the generated search query in conjunction with the sample gesture image repository 112 and / or additional search engines available via (multiple) third-party systems 116. The AR gesture system 102 also identifies and provides different subsets of the collection of sample gesture images via the camera viewfinder of the client computing device 108.

[0074] In more detail, the AR gesture system 102 performs an act 302 of determining a context of a digital image from a camera viewfinder stream of the client computing device 108. For example, by utilizing one or more machine learning models, neural networks, and algorithms in conjunction with the digital image, the AR gesture system 102 determines the context of the digital image. More specifically, the AR gesture system 102 utilizes one or more machine learning models, neural networks, and algorithms to identify characteristics and attributes of objects and scenes depicted in the digital image.

[0075] In one or more embodiments, the AR gesture system 102 utilizes an object detector neural network to generate one or more object labels associated with the digital image. For example, the AR gesture system 102 utilizes an object detector neural network to generate an object label indicating that the digital image depicts one or more of an object, an animal, a car, a plant, a building, etc. In at least one embodiment, the object detector neural network generates an object label including a string identifying a corresponding object (e.g., "man," "dog," "building"), a location of the corresponding object (e.g., corner coordinates of a bounding box surrounding the corresponding object), and a confidence score.

[0076] In more detail, the AR gesture system 102 detects one or more objects in a digital image using a Faster-RCNN model (e.g., ResNet-101) trained to detect objects across multiple categories and categories. Additionally or alternatively, the AR gesture system 102 detects one or more objects using a different neural network, such as ImageNet or DenseNet. Additionally or alternatively, the AR gesture system 102 detects one or more objects using an algorithmic approach, such as the You Only Look Once (YOLO) algorithm. In one or more embodiments, the AR gesture system 102 detects one or more objects by generating an object identifier (e.g., an object label) and an object position / location (e.g., an object bounding box) within a digital image. In one or more embodiments, the AR gesture system 102 utilizes an automatic annotation neural network to generate tags, such as those described in U.S. Patent No. 9,767,386, filed on June 23, 2015, “Training A Classifier Algorithm Used For Automatically Generating Tags To Be Applied To Images”; and U.S. Patent No. 10,235,623, filed on April 8, 2016, “Accurate Tag Relevance Prediction For Image Search”, the entire contents of both patents are incorporated herein by reference.

[0077] The AR gesture system 102 utilizes additional neural networks to generate other tags associated with the digital image. For example, the AR gesture system 102 utilizes a gender neural network to generate one or more gender tags associated with the digital image. More specifically, the AR gesture system 102 utilizes a gender neural network to perform gender recognition and generates a gender tag associated with each object depicted in the digital image. For example, the AR gesture system 102 can utilize a facial detection model to determine the gender of any object in the digital image, such as described in Face Detection and Recognition using Open CV Based on FisherFaces Algorithm, published by J. Manikandan et al. in International Journal of Latest Technology and Engineering, Vol. 8, No. 5, January 2020, the entire contents of which are incorporated herein by reference in their entirety. In further embodiments, the AR gesture system 102 can utilize a deep cognitive attribution neural network to determine the gender of an object in a digital image, such as described in U.S. patent application Ser. No. 16 / 564,831, filed on Sept. 9, 2019, and entitled “Identifying Digital Attributes From Multiple Attribute Groups Within Target Digital Images Utilizing A Deep Cognitive Attribution Neural Network,” the entire contents of which are incorporated herein by reference.

[0078] The AR gesture system 102 also optionally utilizes a clothing neural network to generate one or more clothing tags associated with the digital image. For example, the AR gesture system 102 utilizes a clothing neural network to generate clothing tags indicating the items and types of clothing worn by the object depicted in the digital image. For illustration, the AR gesture system 102 utilizes a clothing neural network to generate clothing tags indicating that the object depicted in the digital image is wearing formal wear, casual wear, sportswear, wedding dresses, etc. For example, the AR gesture system 102 utilizes a trained convolutional neural network to generate clothing tags and other determinations. In one or more embodiments, the clothing neural network includes an object expert network, such as a clothing expert detection neural network. Additional details about utilizing a dedicated object detection neural network are found in U.S. Patent Application No. 16 / 518,880, entitled “Utilizing Object Attribute Detection Models To Automatically Select Instances Of Detected Objects In Images” filed on July 19, 2019, which is incorporated herein by reference in its entirety.

[0079] In one or more embodiments, the AR gesture system 102 determines the context of the digital image based on the generated tags. For example, the AR gesture system 102 determines the context by identifying all or a subset of the generated tags that are relevant to the gesture-based search query. To illustrate, the AR gesture system 102 identifies tags that are specific to the object depicted in the digital image (e.g., a gender tag, one or more clothing tags). The AR gesture system 102 also identifies scene-based tags that further provide information about the digital image. For example, the AR gesture system 102 identifies scene-based tags that indicate that the object is located in a city, at a party, at a park, etc. In at least one embodiment, the AR gesture system 102 avoids identifying duplicate tags so that the identified tags make the body unique.

[0080] AR gesture system 102 also performs an action 304 of generating a set of sample gesture images corresponding to the determined context. For example, AR gesture system 102 generates the set of sample gesture images by first generating a search query based on the identified tags. For illustration, AR gesture system 102 generates the search query by adapting some or all of the identified tags into a logical order using natural language processing. Additionally or alternatively, AR gesture system 102 generates the search query including the identified tags in any order.

[0081] The AR gesture system 102 utilizes the generated search query to retrieve one or more sample gesture images from the sample gesture image repository 112 and / or (multiple) third-party systems 116. For example, the AR gesture system 102 utilizes the search query to identify one or more corresponding sample gesture images from the sample gesture image repository 112. Additionally or alternatively, the AR gesture system 102 utilizes the search query in conjunction with the (multiple) third-party systems 116. For example, the AR gesture system 102 provides the search query to one or more third-party search engines. Additionally or alternatively, the AR gesture system 102 provides the search query to one or more third-party applications that are capable of searching for and providing sample gesture images.

[0082] In response to generating the set of sample pose images, AR pose system 102 performs act 306 of providing different subsets of the set of sample pose images via a camera viewfinder of client computing device 108. For example, in one or more embodiments, AR pose system 102 avoids providing similar or repeated sample pose images via a camera viewfinder of client computing device 108. Thus, AR pose system 102 identifies different subsets of the set of sample pose images such that the object includes unique and diverse sample pose images.

[0083] In at least one embodiment, AR gesture system 102 identifies different subsets by clustering sample gesture images in the sample gesture image collection based on similarity. For example, AR gesture system 102 groups visually or semantically similar sample gesture images together using one or more clustering techniques. AR gesture system 102 then selects sample gestures from each cluster to provide different subsets of the sample gesture image collection.

[0084] The AR gesture system 102 also provides different subsets of the set of sample gesture images via the camera viewfinder of the client computing device 108. For example, the AR gesture system 102 generates an interactive overlay including different subsets of the set of sample gesture images, and overlays the interactive overlay onto the camera viewfinder. In one or more embodiments, the AR gesture system 102 generates an interactive overlay including a horizontal slider that includes different subsets of the sample gesture images. The AR gesture system 102 also generates an interactive overlay such that each of the different subsets of sample gesture images is selectable. In at least one embodiment, the AR gesture system 102 generates an interactive overlay including an indicator of a determined context of a digital image acquired from the camera viewfinder stream and a search button, whereby the AR gesture system 102 receives additional input contextual search terms.

[0085] Figure 3BThe AR gesture system 102 is illustrated as determining different subsets of the sample gesture image set. As mentioned above, the AR gesture system 102 utilizes one or more clustering techniques to identify and provide unique and non-repetitive sample gesture images. For example, Figure 3B As shown, AR gesture system 102 performs action 304 of generating a set of sample gesture images corresponding to the context of the digital image acquired from the camera viewfinder stream. Figure 3A As discussed, AR gesture system 102 generates a set of sample gesture images in response to receiving sample gesture images from sample gesture image repository 112 and / or third party system(s) 116 .

[0086] In response to generating a set of sample pose images, the AR pose system 102 performs an action 308 of extracting a feature vector from one of the sample pose images from the generated set. In one or more embodiments, the AR pose system 102 extracts a feature vector from the sample pose image by generating one or more numerical values ​​representing the characteristics and attributes of the sample pose image. Specifically, the AR pose system 102 generates a feature vector including encoded information describing the characteristics of the sample pose image. For example, the AR pose system 102 generates a feature vector including a set of values ​​corresponding to potential and / or proprietary attributes and characteristics of the sample pose image. In one or more embodiments, the AR pose system 102 generates a feature vector as a multidimensional dataset representing or characterizing the sample pose image. In one or more embodiments, the extracted feature vector includes a set of digital metrics learned by a machine learning algorithm (such as a neural network). In at least one embodiment, the AR pose system 102 utilizes one or more algorithms (e.g., SciKit learning library) to extract feature vectors from the sample pose image.

[0087] Next, AR gesture system 102 performs an action 310 of determining whether there are more sample gesture images in the set of sample gesture images that do not have corresponding feature vectors. If there are additional sample gesture images (e.g., "yes" in action 310), AR gesture system 102 performs an action 312 of identifying the next sample gesture image in the set of sample gesture images. AR gesture system 102 then repeats action 310 in conjunction with the next sample gesture image. AR gesture system 102 continues to extract feature vectors from sample gesture images in the set of sample gesture images until all feature vectors have been extracted (e.g., "no" in action 310).

[0088] After all feature vectors are extracted, AR gesture system 102 performs an action 314 of mapping the extracted feature vectors into a vector space. For example, AR gesture system 102 maps the extracted feature vectors into points in the vector space. In one or more embodiments, the vector space is in an n-dimensional vector space, where n is the number of features represented in each vector.

[0089] Next, the AR gesture system 102 performs an act 316 of determining clusters of feature vectors in the vector space. For example, the AR gesture system 102 determines clusters of feature vectors by grouping each feature vector with its nearest neighbors. In one embodiment, the AR gesture system 102 performs k-means clustering to cluster the feature vectors. For example, the AR gesture system 102 utilizes k-means clustering by partitioning the vector space so that each feature vector belongs to a cluster with the nearest mean (e.g., a cluster center). In alternative embodiments, the AR gesture system 102 utilizes other clustering algorithms.

[0090] As part of partitioning the vector space, the AR gesture system 102 determines the distances between feature vectors. For example, the AR gesture system 102 calculates the distances between feature vectors to determine appropriate clusters for grouping the feature vectors. The AR gesture system 102 can use various methods to calculate the distances between feature vectors. In one embodiment, the AR gesture system 102 determines the Euclidean distance between feature vectors. In another embodiment, the AR gesture system 102 calculates the distances between feature vectors using the Minkowski method.

[0091] To further identify different subsets of the set of sample pose images, the AR pose system 102 performs an action 318 of identifying sample pose images from each cluster. For example, the AR pose system 102 identifies feature vectors from each cluster and then provides different subsets as sample pose images corresponding to the identified feature vectors. In one or more embodiments, the AR pose system 102 identifies feature vectors from a particular cluster by randomly selecting feature vectors from the cluster. Additionally or alternatively, the AR pose system 102 identifies feature vectors from a cluster by selecting feature vectors that are closest to the center of the cluster.

[0092] The AR gesture system 102 also performs an action 320 of providing the identified sample gesture image via an interactive overlay located on the camera viewfinder of the client computing device 108. For example, the AR gesture system 102 generates an interactive overlay including a predetermined number of different subsets of the set of sample gesture images. Additionally or alternatively, the AR gesture system 102 generates an interactive overlay including all of the different subsets of the set of sample gesture images in a horizontal slider. The AR gesture system 102 also positions the generated interactive overlay on a portion of the camera viewfinder of the client computing device 108.

[0093] Figure 4A The AR gesture system 102 is illustrated providing gesture guidance via a camera viewfinder of a client computing device 108. As mentioned above, the AR gesture system 102 provides gesture guidance, which is a body frame extracted from a selected sample gesture image and redirected based on the scale of the object. Using the redirected reference body frame as a gesture guide, the AR gesture system 102 provides visual guidance to enable objects in the camera viewfinder stream to pose their body parts to simulate the gestures shown in the selected sample gesture image.

[0094] like Figure 4A As shown, AR gesture system 102 performs action 402 of detecting a selection of a sample gesture image. As discussed above, AR gesture system 102 generates an interactive overlay including a selectable sample gesture image. Thus, AR gesture system 102 detects a selection of a sample gesture image (e.g., a tap touch gesture, a mouse click) via a camera viewfinder of client computing device 108.

[0095] In response to the detected selection of the sample pose image, the AR pose system 102 performs an action 404 of extracting a reference body frame from the selected sample pose image. In one or more embodiments, the AR pose system 102 extracts the reference body frame from the selected sample pose image by identifying the locations of joints and segments of the object (e.g., a portrait) depicted in the selected sample pose image. For example, the AR pose system 102 utilizes full body tracking via a pose neural network to identify the locations of joints and segments of the object in the selected sample pose image. In at least one embodiment, the AR pose system 102 also utilizes the pose neural network to generate a reference body frame (e.g., a digital skeleton) that includes the joints and segments in the determined locations.

[0096] Furthermore, in response to detecting the selection of the sample pose image, the AR pose system 102 performs an act 406 of extracting a body frame of the object from the camera viewfinder stream of the client computing device 108. For example, the AR pose system 102 utilizes a pose neural network in conjunction with a digital image of the camera viewfinder stream from the client computing device 108 to identify locations of joints and segments of the object depicted in the camera viewfinder stream. The AR pose system 102 also utilizes the pose neural network to generate a body frame of the object (e.g., a digital skeleton) that includes the joints and segments in the determined locations. Although Figure 4A While AR gestural system 102 is shown performing actions 404 and 406 in parallel, in additional or alternative embodiments, AR gestural system 102 performs actions 404 and 406 sequentially (eg, performs action 404 and then performs action 406 ).

[0097] The AR gesture system 102 also performs an action 408 of redirecting the reference body frame based on the object body frame. For example, the AR gesture system 102 redirects the reference body frame by first determining the lengths of the segments between the joints of the object body frame (e.g., indicating the scale of the object depicted in the digital image from the camera viewfinder stream). The AR gesture system 102 also redirects the reference body frame by modifying the lengths of the segments between the joints in the reference body frame to match the lengths of the segments between the corresponding joints of the object body frame. Thus, the redirected reference body frame retains the original gesture indicated by the selected sample gesture image, but has the scale of the object depicted in the camera viewfinder stream.

[0098] The AR gesture system 102 next performs an action 410 of providing a redirected reference body frame aligned with an object in the camera viewfinder stream. For example, the AR gesture system 102 provides the redirected reference body frame as a gesture guide overlaid onto the camera viewfinder based on one or more landmarks relative to the object depicted in the camera viewfinder stream. To illustrate, the AR gesture system 102 identifies one or more landmarks of the object depicted in the camera viewfinder stream, such as, but not limited to, a hip region and a chest region. The AR gesture system 102 then generates a visualization of the redirected reference body frame, including joints and segments of the redirected reference body frame, and anchors the visualization to the camera viewfinder at the identified landmarks of the object.

[0099] In one or more embodiments, AR gesture system 102 performs actions 402 to 410 in conjunction with the additional object depicted in the camera viewfinder stream and the selected sample gesture image. For example, if the camera viewfinder stream depicts two objects and the selected sample gesture image also depicts two gesture objects, AR gesture system 102 repeats actions 402 to 410 for the second object. In at least one embodiment, AR gesture system 102 determines which object in the sample gesture image corresponds to which object in the camera viewfinder stream in response to user input. If the camera viewfinder stream depicts two objects and the selected sample gesture image depicts one object, AR gesture system 102 performs actions 402 to 410 in response to user input indicating which of the two objects AR gesture system 102 will utilize in conjunction with generating the gesture guidance.

[0100] Figure 4BThe AR gesture system 102 is illustrated utilizing the gesture neural network 412 to extract a body frame (e.g., a reference body frame, an object body frame) from a digital image. As mentioned, in some embodiments, the AR gesture system 102 determines or extracts a body frame of an object depicted within a digital image (e.g., a digital image or a sample gesture image from a camera viewfinder stream of the client computing device 108). Specifically, the AR gesture system 102 utilizes the gesture neural network to determine locations and labels of joints and segments of an object (e.g., a portrait) depicted in the digital image.

[0101] like Figure 4B As illustrated, the AR gesture system 102 identifies a digital image 414. Specifically, the AR gesture system 102 receives an extracted digital image 414 or a selection or upload of a digital image 414 as a sample gesture image from a camera viewfinder stream of the client computing device 108. In addition, the AR gesture system 102 determines a two-dimensional body frame 418 (e.g., an object body frame or a reference body frame) associated with the digital image 414. For example, the AR gesture system 102 identifies an object 416 depicted within the digital image 414 using a gesture neural network 412, and determines the locations of joints and segments of the object 416. In some embodiments, the AR gesture system 102 utilizes the gesture neural network 412 to identify all two-dimensional joints from the digital image 414. For example, the AR gesture system 102 utilizes the gesture neural network 412 to extract a two-dimensional joint vector or segment 422a, 422b, 422c associated with each joint 420a, 420b, 420c, 420d, 420e corresponding to the object 416 in the digital image 414.

[0102] In one or more embodiments, the AR gesture system 102 utilizes a gesture neural network 412 in the form of a convolutional neural network to jointly predict a confidence map for body part detection and a part affinity field for learning associated body parts of the object from an input digital image 414 of an object 416 (e.g., a portrait). For example, to identify a body part, the AR gesture system 102 generates a confidence map that includes a two-dimensional representation of a confidence measure that a particular body part (e.g., a head or torso) is located at any given pixel. To identify limbs that connect body parts, the AR gesture system 102 also generates a part affinity field that includes a two-dimensional vector field for each limb, including location and orientation information across the limb support area. The AR gesture system 102 generates a part affinity field for each type of limb that joins two associated body parts. In addition, the AR gesture system 102 utilizes the gesture neural network 412 to parse the digital image 414 into parts for two-part matching of associated body part candidates. For example, the AR pose system 102 utilizes a pose neural network 412, such as the pose neural network described in OpenPose: Realtime Multi-Person 2D Pose Estimation using Part Affinity Fields, published by Zhe Cao, Gines Hidalgo, Tomas Simon, Shih-en Wei, and Yaser Seikh in 2018 at arXiv:1812.08008, which is incorporated herein by reference in its entirety. In some cases, the pose neural network 412 is a hybrid neural network based on a combination of GoogleNet and OpenPose. The AR pose system 102 utilizes a variety of neural network architectures to determine poses.

[0103] like Figure 4B As further illustrated, the AR gesture system 102 generates a three-dimensional body frame 425 based on the two-dimensional body frame 418. Specifically, the AR gesture system 102 utilizes the 2D to 3D neural network 423 to process the two-dimensional body frame 418 and generates the three-dimensional body frame 425. For example, the AR gesture system 102 utilizes the 2D to 3D neural network 423 to generate three-dimensional joint features for the joints identified in the two-dimensional body frame 418. In practice, the AR gesture system 102 utilizes the 2D to 3D neural network 423 to project the two-dimensional joint features onto a unit sphere, thereby generating the three-dimensional joint features.

[0104] In some embodiments, the AR gesture system 102 utilizes a 2D to 3D neural network 423 that estimates body joint locations in a three-dimensional space (e.g., a three-dimensional body frame 425) from a two-dimensional input (e.g., a two-dimensional body frame 418). For example, the AR gesture system 102 utilizes a 2D to 3D neural network 423 in the form of a deep feed-forward neural network that generates a series of points in a three-dimensional space from a series of two-dimensional points. Specifically, the AR gesture system 102 utilizes the 2D to 3D neural network 423 to learn a function that reduces or minimizes the prediction error of a predicted three-dimensional point by projecting the two-dimensional point onto a fixed global space (relative to a root joint) over a dataset of a specific number of gesture objects and corresponding body frames. For example, the AR gesture system 102 utilizes a 2D to 3D neural network 423, such as the 2D to 3D neural network described in A Simple Yet Effective Baseline for 3D Human Pose Estimation, published by Julieta Martinez, Rayat Hossain, Javier Romero, and James J. Little in 2017 at arXiv:1705.03098, which is incorporated herein by reference in its entirety. The AR gesture system 102 utilizes various machine learning models (e.g., neural networks) to project two-dimensional joint features and generate three-dimensional joint features.

[0105] In some embodiments, the AR gesture system diagram 102 also generates a visualization of the three-dimensional body frame 425. More specifically, as Figure 4B As shown, AR gesture system 102 arranges 3D virtual joints and 3D virtual segments overlaid at locations corresponding to joints and segments of object 416 depicted in digital image 414. In one or more embodiments, after reorienting 3D body frame 425 based on the proportions of the object depicted in the camera viewfinder stream of client computing device 108, AR gesture system 102 utilizes the visualization of 3D body frame 425 as a gesture guide.

[0106] Figure 4C and Figure 4D The AR gesture system 102 is illustrated as reorienting the reference body frame based on the scale of the object depicted in the camera viewfinder stream of the client computing device 108. For example, Figure 4C As shown, AR gesture system 102 utilizes gesture neural network 412 in conjunction with 2D to 3D neural network 423 to extract reference body frame 426 from reference object 424 (e.g., depicted in a selected sample gesture image) and extract object body frame 430 from target object 428 (e.g., depicted in a camera viewfinder stream of client computing device 108).

[0107] like Figure 4C As shown, the scale of reference object 424 is different from the scale of target object 428. This scale difference is common between sample pose images that often depict professional models and objects in the camera viewfinder stream that are not usually professional models. If AR gesture system 102 uses reference body frame 426 as a pose guide without redirection, the result will be as follows Figure 4D For example, because target object 428 has a different scale than reference body frame 426 and is unconstrained, a person represented by target object 428 may have difficulty imitating a pose dictated by reference body frame 426 .

[0108] Thus, to provide customized gestural guidance relative to the person represented by target object 428, AR gestural system 102 reorients reference body frame 426. In one or more embodiments, AR gestural system 102 reorients reference body frame 426 by first determining the lengths of segments between joints of subject body frame 430 (e.g., Figure 4C 426. For example, the AR gesture system 102 determines the pixel length of each segment in the object body frame 430 and the position of each segment relative to its adjacent joints.

[0109] Next, the AR gesture system 102 redirects the reference body frame 426 by modifying the lengths of the segments between the joints of the reference body frame 426 to match the determined lengths of the segments between the corresponding joints of the subject body frame 430. For example, the AR gesture system 102 lengthens or shortens the segments of the reference body frame 426 to match the lengths of the corresponding segments in the subject body frame 430. By maintaining the relative positions of the segments and surrounding joints between the reference body frame 426 and the subject body frame 430, the AR gesture system 102 determines that the segments in the reference body frame 426 correspond to the segments in the subject body frame 430. Thus, the AR gesture system 102 determines, for example, that the segment between the knee and ankle joints in the reference body frame 426 corresponds to the segment between the knee and ankle joints in the subject body frame 430.

[0110] Therefore, if Figure 4D As further shown, AR gesture system 102 generates redirected reference body frame 426'. In one or more embodiments, AR gesture system 102 generates redirected reference body frame 426' that includes the proportions of target object 426 by lengthening or shortening segments of reference body frame 426. Thus, a person represented by target object 426 can more easily simulate a gesture indicated by redirected reference body frame 426' (e.g., Figure 4D 426' in the target object).

[0111] In additional or alternative embodiments, the AR gesture system 102 utilizes motion redirection to generate the redirected reference body frame 426'. For example, the AR gesture system 102 utilizes motion redirection in conjunction with the object body frame 430 by arranging segments of the object body frame 430 to match a gesture indicated by the reference body frame 426. To illustrate, the AR gesture system 102 determines the relative positions and angles between consecutive segments of the reference body frame 426. The AR gesture system 102 then manipulates the corresponding segments of the object body frame 430 to match the determined positions and angles.

[0112] Figure 4E and Figure 4F The AR gesture system 102 is illustrated aligning the redirected reference body frame 426' with the target object 428. In one or more embodiments, the AR gesture system 102 aligns the redirected reference body frame 426' based on at least one landmark 432a of the target object 428. For example, the AR gesture system 102 determines the landmark 432a (e.g., a hip region) based on a label generated by the gesture neural network 412. To illustrate, the AR gesture system 102 determines that the landmark 432a corresponds to the location of the joint having the label "hip" in the redirected reference body frame 426'.

[0113] AR gesture system 102 also determines corresponding landmarks for target object 428. For example, AR gesture system 102 determines corresponding landmarks for target object 428 by generating an updated object body frame (e.g., including locations and labels of segments and joints) for target object 428, and aligning redirected reference body frame 426a with the updated object body frame at segments and / or joints having labels corresponding to landmark 432a (e.g., “hip joint”). AR gesture system 102 performs this alignment in conjunction with the object body frame (e.g., not visible in the camera viewfinder), the redirected reference body frame (e.g., not visible in the camera viewfinder), and / or a visualization of the redirected reference body frame (e.g., visible in the camera viewfinder).

[0114] Figure 4F The AR gesture system 102 is illustrated as aligning the redirected reference body frame 426' based on more than one landmark of the target object 428. In one or more embodiments, the AR gesture system 102 aligns the redirected reference body frame 426a' across multiple landmarks 432a, 432b (e.g., hip area and torso area) to account for possible rotation of the target object 428 toward or away from the camera of the client computing device 108.

[0115] For example, Figure 4FAs shown, when target object 428 is rotated relative to the camera of client computing device 108 and AR gesture system 102 anchors redirected reference body frame 426' to only one landmark 432a, the result is that the person represented by target object 428 will find it impossible to properly align their limbs with the limbs represented by redirected reference body frame 426'. In contrast, when AR gesture system 102 anchors redirected reference body frame 426' to at least two landmarks, the shoulders, hips, legs, and arms represented by redirected reference body frame 426' are correctly positioned relative to corresponding areas of target object 428.

[0116] Figure 5A The AR gesture system 102 is illustrated as determining an alignment between a gesture guide and an object depicted in a camera viewfinder stream, and modifying display characteristics of the gesture guide to iteratively and continuously provide gesture guidance to a user of the client computing device 108. For example, as will be discussed in more detail below, the AR gesture system 102 overlays a visualization of a redirected reference body frame (e.g., a gesture guide) on a camera viewfinder. The AR gesture system 102 iteratively determines an alignment between portions of the redirected reference body frame and corresponding portions of an object depicted in the camera viewfinder. For each determined alignment, the AR gesture system 102 modifies display characteristics of a corresponding portion or segment of the redirected reference body visualization to indicate the alignment. In response to determining that all segments of the redirected reference body frame are aligned with corresponding segments of the object, the AR gesture system 102 automatically captures a digital image from the camera viewfinder stream of the client computing device 108.

[0117] In more detail, AR gesture system 102 performs an action 502 of overlaying a visualization of the redirected reference body frame on a camera viewfinder of client computing device 108. As discussed above, AR gesture system 102 generates the redirected reference body frame by modifying the scale of segments of the reference body frame extracted from the selected sample gesture image based on the scale of the object's body frame extracted from the digital image depicting the object. AR gesture system 102 also generates a visualization of the redirected reference body frame including segment lines having a color, pattern, animation, etc. and joints represented by dots or other shapes having the same or different color, pattern, animation, etc. as the segment lines.

[0118] Also as discussed above, AR gesture system 102 overlays the redirected visualization of the reference body frame onto the camera viewfinder of client computing device 108 by anchoring the redirected visualization of the reference body frame to the object depicted in the camera viewfinder. For example, AR gesture system 102 determines one or more landmarks of the object (e.g., hip region, torso region) and anchors corresponding points of the redirected visualization of the reference body frame to those landmarks.

[0119] AR gesture system 102 also performs an action 504 of determining an alignment between a portion of the redirected reference body frame and an object depicted in the camera viewfinder. For example, and as will be described below with reference to Figure 5B Discussed in more detail, based on an updated object body frame simultaneously anchored to the object at the same one or more landmarks as the redirected reference body frame, the AR gesture system 102 determines that a portion of the redirected reference body frame (e.g., one or more joints and segments) is aligned with a corresponding portion of the object. More specifically, when one or more segments and / or joints of the redirected reference body frame overlap with corresponding one or more segments and / or joints of the updated object body frame, the AR gesture system 102 determines that a portion of the redirected reference body frame is aligned with a corresponding portion of the object when both body frames are anchored to the object at the same landmark(s).

[0120] In response to determining the alignment between the redirected portion of the reference body frame and the object, the AR gesture system 102 performs an action 506 of modifying the display characteristics of the visualization of the reference body frame based on the alignment. For example, in response to determining the alignment between the arm portion of the redirected reference body frame and the object, the AR gesture system 102 modifies the display characteristics of the corresponding arm portion of the visualization of the redirected reference body frame in the camera viewfinder. In one or more embodiments, the AR gesture system 102 modifies the display characteristics of the visualization, including but not limited to modifying the display color of the aligned portion of the visualization, modifying the line type of the aligned portion of the visualization (e.g., from a solid line to a dashed line), and modifying the line width of the aligned portion of the visualization (e.g., from a thin line to a thick line). In additional or alternative embodiments, the AR gesture system 102 modifies the display characteristics of the visualization by adding animation or highlights to the visualized portion to indicate the alignment.

[0121] AR gesture system 102 also performs determining whether there are additional misaligned portions of the visualization of the redirected reference body frame 508. For example, AR gesture system 102 determines that there are additional misaligned portions of the visualization in response to determining that there is at least one portion of the visualization that exhibits original or unmodified display characteristics.

[0122] In response to determining that there is an additional misaligned portion of the visualization (e.g., "yes" in action 508), AR gesture system 102 repeats actions 504 and 506 of determining the alignment between a portion of the visualization and the object and modifying display characteristics of the portion. In at least one embodiment, AR gesture system 102 performs action 504 in conjunction with an updated object body frame representing an updated pose of the object depicted in the camera viewfinder. For example, and to account for additional movement of the object as the object attempts to simulate the pose represented by the redirected reference body frame, AR gesture system 102 utilizes gesture neural network 412 in conjunction with 2D to 3D neural network 423 to generate an updated object body frame corresponding to the object.

[0123] Thus, each time alignment is determined, the AR gesture system 102 generates an updated object body frame associated with the object. Additionally, the AR gesture system 102 generates an updated object body frame at regular intervals. For example, the AR gesture system 102 generates an updated object body frame every predetermined number of camera viewfinder stream frames (e.g., every 30 frames). In another example, the AR gesture system 102 generates an updated object body frame after a predetermined amount of time (e.g., every 5 seconds). Prior to the next iteration of actions 504 and 506, the AR gesture system 102 also anchors the updated object body frame to the same landmark of the object as the redirected reference body frame.

[0124] In one or more embodiments, AR gesture system 102 continues to iteratively perform actions 504, 506, and 508 until AR gesture system 102 determines that there are no additional misaligned portions of the redirected visualization of the reference body frame (e.g., "No" in action 508). In response to determining that there are no additional misaligned portions of the redirected visualization of the reference body frame, AR gesture system 102 performs action 510 of automatically capturing a digital image from a camera viewfinder stream of client computing device 108. For example, AR gesture system 102 stores the captured digital image locally (e.g., in a camera roll of client computing device 108). AR gesture system 102 also provides the captured digital image to image capture system 104 along with information associated with the redirected reference body frame, the object body frame, and / or the selected sample gesture image. In response to automatically capturing the digital image, AR gesture system 102 also removes the redirected visualization of the reference body frame from the camera viewfinder of client computing device 108. In one or more embodiments, AR gesture system 102 performs actions 502 - 510 concurrently in conjunction with multiple objects depicted in the camera viewfinder.

[0125] Figure 5BThe diagram illustrates additional details regarding how the AR gesture system 102 determines alignment between a portion of the redirected reference body frame and an object depicted in the camera viewfinder of the client computing device 108. For example, as mentioned above with respect to action 504 performed by the AR gesture system 102, the AR gesture system 102 determines the alignment based on a continuously updated object body frame that is simultaneously anchored to the object in the camera viewfinder.

[0126] In more detail, the AR gesture system 102 performs an action 512 of generating a redirected reference body frame and an object body frame. As discussed above, in at least one embodiment, the AR gesture system 102 generates these body frames using a gesture neural network 412 that identifies the relative positions of joints and segments of an object displayed in a digital image and a 2D to 3D neural network 423 that generates a three-dimensional body frame. In one or more embodiments, the AR gesture system 102 iteratively and continuously generates an updated object body frame to account for movement of the object within the camera viewfinder stream.

[0127] For each iteration, the AR gesture system 102 performs an action 514 of anchoring the redirected reference body frame and the object body frame (e.g., whether the original object body frame or the updated object body frame in a subsequent iteration) to the object by one or more regions. For example, the AR gesture system 102 anchors both body frames to the object by at least a hip region of the object. It is noted that while the AR gesture system 102 may display a visualization of the redirected reference body frame anchored to the object within the camera viewfinder, the AR gesture system 102 may not display a visualization of the object body frame anchored to the object within the camera viewfinder. Therefore, the simultaneously anchored object body frame may not be viewable even if it exists.

[0128] The AR gesture system 102 also performs an act of determining that at least a portion of the redirected reference body frame overlaps with a corresponding portion of the subject's body frame 516. In one or more embodiments, the AR gesture system 102 determines that a portion of the redirected reference body frame overlaps with a corresponding portion of the subject's body frame by determining 1) whether any portion of the redirected reference body frame overlaps with any portion of the subject's body frame, and 2) whether the overlapping portions of the two body frames correspond (e.g., represent the same one or more body parts).

[0129] In more detail, the AR gesture system 102 determines whether any portion of the redirected reference body frame (e.g., one or more segments and / or joints) overlaps with any portion of the object body frame in various ways. For example, in response to determining that the joints at both ends of the redirected reference body frame segment are located at the same position (e.g., position coordinates) as the joints at both ends of the object body frame segment, the AR gesture system 102 determines that the segment of the redirected reference body frame overlaps with the segment of the object body frame. Additionally or alternatively, the AR gesture system 102 generates a vector representing each segment of the redirected reference body frame in a vector space. The AR gesture system 102 also generates a vector representing each segment of the object body frame in the same vector space. The AR gesture system 102 then determines whether any vectors between the two body frames occupy the same location in the vector space.

[0130] In one embodiment, the AR gesture system 102 determines that a portion of the redirected reference body frame overlaps a portion of the object body frame based on the total coverage of the corresponding portions. For example, if two portions have the same starting and ending coordinate points, the AR gesture system 102 determines that the corresponding portions overlap—which means that the two portions have the same length and position relative to the object in the camera viewfinder. Additionally or alternatively, the AR gesture system 102 determines that a portion of the redirected reference body frame overlaps a portion of the object body frame based on a threshold amount of coverage of the corresponding portions. For example, if two segments have the same starting coordinates and the end points of both segments are within a threshold angle (e.g., ten degrees), the AR gesture system 102 determines that a segment of the redirected reference body frame overlaps a segment of the object body frame. Similarly, if two segments have the same starting coordinates and the end points of both segments are within a threshold distance (e.g., ten pixels), the AR gesture system 102 determines that a segment of the redirected reference body frame overlaps a segment of the object body frame.

[0131] Next, in response to determining that a portion of the redirected reference body frame overlaps a portion of the object body frame, the AR gesture system 102 determines whether the overlapping portions correspond. In one embodiment, the AR gesture system 102 determines that the overlapping portions correspond based on segment labels associated with each body frame. For example, as mentioned above, the AR gesture system 102 generates a body frame using a gesture neural network, which outputs segments and joints of the body frame and labels identifying the segments and joints (e.g., "femur segment," "hip joint," "tibia segment," "ankle joint"). Therefore, when labels associated with one or more segments and / or joints in the body frame portion match, the AR gesture system 102 determines that the body frame portions correspond.

[0132] Finally, in response to determining that at least a portion of the redirected reference body frame overlays a corresponding portion of the subject body frame, the AR gesture system 102 performs an action 518 of modifying the display characteristics of the determined portion of the redirected reference body frame. For example, as discussed above, the AR gesture system 102 modifies the display characteristics of the determined portion by modifying one or more of: a display color of the portion, a line type of the portion, or a line width of the portion. Additionally or alternatively, the AR gesture system 102 flashes the portion, and / or depicts another type of animation to indicate alignment. It is to be understood that the AR gesture system 102 modifies the display characteristics of the determined portion of the visualization of the redirected reference body frame overlaid on the camera viewfinder of the client computing device 108 so that the user of the client computing device 108 understands that the corresponding portion of the subject's body is aligned with the gesture guidance indicated by the visualization of the redirected reference body frame.

[0133] Figure 6 A detailed schematic diagram of an embodiment of an AR gesture system 102 operating on a computing device 600 according to one or more embodiments is illustrated. As discussed above, the AR gesture system 102 is operable on a variety of computing devices. Thus, for example, the computing device 600 is optionally a (multiple) server 106 and / or a client computing device 108. In one or more embodiments, the AR gesture system 102 includes a communication manager 602, a context detector 604, a sample gesture image manager 606, a body frame generator 608, an alignment manager 610, and a gesture neural network 412. Additionally, the AR gesture system 102 corresponds to and / or is connected to a sample gesture image repository 112.

[0134] As mentioned above, and as Figure 6 As shown, AR gesture system 102 includes communication manager 602. In one or more embodiments, communication manager 602 handles communication between AR gesture system 102 and Figure 1 The communications manager 602 handles communications between the AR gesture system 102 and other devices and components within the illustrated environment 100. For example, when the AR gesture system 102 resides on the server(s) 106, the communications manager 602 handles communications between the AR gesture system 102 and the client computing device 108. To illustrate, the communications manager 602 receives data from the client computing device 108 including an indication of a user interaction (e.g., a user selection of a sample gesture image) and one or more digital images from a camera viewfinder stream of the client computing device 108. The communications manager 602 also provides data to the client computing device 108, including the sample gesture images and an interactive overlay of the AR gesture guidance.

[0135] As mentioned above, and as Figure 6As shown, AR gesture system 102 includes context detector 604. In one or more embodiments, context detector 604 utilizes one or more detectors (e.g., object detector, clothing detector), including one or more neural networks, algorithms, and / or machine learning models, to generate tags associated with a digital image from a camera viewfinder stream of client computing device 108. Context detector 604 also determines the context of the digital image based on the generated tags, thereby utilizing natural language processing to arrange the generated tags in a logical order. In at least one embodiment, context detector 604 selects unique and / or relevant tags to determine the context of the digital image.

[0136] As mentioned above, and as Figure 6 As shown, the AR gesture system 102 includes a sample gesture image manager 606. In one or more embodiments, the sample gesture image manager 606 utilizes a determined context of a digital image extracted from a camera viewfinder stream to generate a set of sample gesture images tailored to the determined context. For example, the sample gesture image manager 606 utilizes the determined context to generate a search query. The sample gesture image manager 606 also provides the search query to one or more search engines and / or repositories to receive the set of sample gesture images. In at least one embodiment, the sample gesture image manager 606 utilizes one or more clustering techniques (such as, k-means clustering) to generate different subsets of the set of sample gesture images.

[0137] As mentioned above, and as Figure 6 As shown, AR gesture system 102 includes body frame generator 608. In one or more embodiments, body frame generator 608 utilizes gesture neural network 412 in conjunction with 2D to 3D neural network 423 to generate one or more body frames. For example, body frame generator 608 utilizes gesture neural network 412 in conjunction with 2D to 3D neural network 423 to generate a reference body frame based on an object depicted in a selected sample gesture image. Body frame generator 608 also utilizes gesture neural network 412 in conjunction with 2D to 3D neural network 423 to generate an object body frame based on an object depicted in a camera viewfinder stream of client computing device 108. Body frame generator 608 also redirects the reference body frame based on a scale of an object depicted in the camera viewfinder stream (e.g., represented by a corresponding object body frame), such that the redirected reference body frame retains its original pose, but has the scale of the object depicted in the camera viewfinder stream.

[0138] As mentioned above, and as Figure 6As shown, AR gesture system 102 includes alignment manager 610. In one or more embodiments, alignment manager 610 determines an alignment between a portion of the redirected reference body frame and an object depicted in a camera viewfinder stream of client computing device 108. Alignment manager 610 also modifies display characteristics of a visualization of the redirected reference body frame in response to the determined alignment. In response to determining a complete alignment between the redirected reference body frame and the object, alignment manager 610 automatically captures a digital image from the camera viewfinder stream of client computing device 108.

[0139] Each of the components 412, 423, 602 to 610 of the AR gesture system 102 includes software, hardware, or both. For example, the components 412, 423, 602 to 610 include one or more instructions stored on a computer-readable storage medium and executable by a processor of one or more computing devices such as a client device or a server device. When executed by one or more processors, the computer-executable instructions of the AR gesture system 102 cause the (multiple) computing devices to perform the methods described herein. Alternatively, the components 412, 423, 602 to 610 include hardware, such as a dedicated processing device that performs a specific function or group of functions. Alternatively, the components 412, 423, 602 to 610 of the AR gesture system 102 include a combination of computer-executable instructions and hardware.

[0140] In addition, the components 412, 423, 602 to 610 of the AR gesture system 102 can be implemented, for example, as one or more operating systems, one or more independent applications, one or more modules of an application, one or more plug-ins, one or more library functions, or functions that can be called by other applications and / or a cloud computing model. Therefore, components 412, 423, 602 to 610 can be implemented as independent applications, such as desktop computers or mobile applications. In addition, components 412, 423, 602 to 610 can be implemented as one or more web-based applications hosted on a remote server. Components 412, 423, 602 to 610 can also be implemented in a set of mobile device applications or "apps". For illustration, components 412, 423, 602 to 610 can be implemented in applications, including but not limited to ADOBE CREATIVE CLOUD, such as ADOBE PHOTOSHOP or ADOBE PHOTOSHOP CAMERA. “ADOBE”, “CREATIVE CLOUD”, “PHOTOSHOP” and “PHOTOSHOP CAMERA” are registered trademarks or trademarks of Adobe Systems Incorporated in the United States and / or other countries.

[0141] Figures 1 to 6, corresponding text and examples provide a number of different methods, systems, devices, and non-transitory computer-readable media for the AR gesture system 102. In addition to the foregoing, one or more embodiments may also be described in terms of a flowchart including actions for achieving a particular result, such as Figure 7 , Figure 8 and Fig. 9 shown. Figure 7 , Figure 8 and Fig. 9 More or fewer actions may be used to perform. Further, these actions may be performed in different orders. Additionally, the actions described herein may be repeated or performed in parallel with each other or with different instances of the same or similar actions.

[0142] As mentioned, Figure 7 A flow diagram illustrating a series of actions 700 for generating an interactive overlay of context-tailored sample gesture images in accordance with one or more embodiments. Figure 7 The actions according to one embodiment are illustrated, but alternative embodiments may omit, add, reorder and / or modify Figure 7 Any action shown. Figure 7 The actions of may be performed as part of a method. Alternatively, the non-transitory computer readable medium may include a computer program that, when executed by one or more processors, causes a computing device to perform Figure 7 In some embodiments, the system may execute Figure 7 action.

[0143] like Figure 7 As shown, a series of actions 700 includes an action 710 of receiving a digital image from a computing device. For example, action 710 involves receiving a digital image from a camera viewfinder stream of a computing device. In one or more embodiments, receiving a digital image from the camera viewfinder stream is responsive to detecting a selection of an entry point option displayed in conjunction with the camera viewfinder stream on the computing device.

[0144] like Figure 7 As shown, a series of actions 700 includes an action 720 of determining a context of a digital image. For example, action 720 involves determining the context of the digital image based on objects and scenes depicted within the digital image. In one or more embodiments, determining the context of the digital image includes: generating one or more object tags associated with the digital image; generating a gender tag for each person depicted in the digital image; generating one or more clothing tags for each person depicted in the digital image; and determining the context of the digital image based on the one or more object tags, the gender tag for each person depicted in the digital image, and the one or more clothing tags for each person depicted in the digital image.

[0145] like Figure 7 As shown, a series of actions 700 includes an action 730 of generating a set of sample pose images based on the context. For example, action 730 involves generating a set of sample pose images corresponding to the context of the digital image. In one or more embodiments, generating the set of sample pose images includes: generating a search query based on one or more object tags, a gender tag of each person depicted in the digital image, and one or more clothing tags of each person depicted in the digital image; and receiving the set of sample pose images in response to providing the search query to one or more search engines.

[0146] In one or more embodiments, a series of actions 700 includes actions of determining different subsets of the sample posture image set by: generating feature vectors for sample posture images in the sample posture image set; clustering the feature vectors to determine one or more categories of the sample posture images; selecting a feature vector from each of the one or more categories of the sample posture images; and determining different subsets of the sample posture image set as sample posture images corresponding to the selected feature vectors.

[0147] like Figure 7 As shown, a series of actions 700 includes an action 740 of providing different subsets of a set of sample pose images. For example, action 740 involves providing different subsets of the set of sample pose images via an interactive overlay located on a camera viewfinder of a computing device. In one or more embodiments, providing different subsets of the set of sample pose images for display includes: generating an interactive overlay including a scrollable display of different subsets of the set of sample pose images, wherein each of the sample pose images in the scrollable display is selectable; and positioning the interactive overlay to a lower portion of the camera viewfinder of the computing device. In at least one embodiment, generating the interactive overlay also includes providing an indication of the context of the digital image within the interactive overlay.

[0148] As mentioned, Figure 8 A flow chart of a series of actions 800 for generating augmented reality gesture guidance is illustrated in accordance with one or more embodiments. Figure 8 The actions according to one embodiment are illustrated, but alternative embodiments may omit, add, reorder and / or modify Figure 8 Any action shown. Figure 8 The actions of may be performed as part of a method. Alternatively, the non-transitory computer readable medium may include a computer program that, when executed by one or more processors, causes a computing device to perform Figure 8 In some embodiments, the system may execute Figure 8 action.

[0149] like Figure 8As shown, sequence of acts 800 includes an act of determining a selection of a sample gesture image 810. For example, act 810 involves detecting a selection of a sample gesture image from an interactive overlay located on a camera viewfinder of a client computing device.

[0150] like Figure 8 As shown, a series of actions 800 includes an action 820 of extracting a reference body frame from a selected sample pose image and extracting a subject body frame from a digital image from a camera viewfinder stream. For example, action 820 involves utilizing a pose neural network to: extract a first reference body frame from the selected sample pose image, and extract a first subject body frame from a digital image from a camera viewfinder stream from a client computing device.

[0151] like Figure 8 As shown, a series of actions 800 includes an action 830 of redirecting a reference body frame based on a proportion of a subject body frame. For example, action 830 involves redirecting a first reference body frame to include a proportion of the first subject body frame in a pose of the first reference body frame. In one or more embodiments, redirecting the first reference body frame to include a proportion of the first subject body frame in a pose of the first reference body frame includes: determining a length of a segment between joints of the first subject body frame; and modifying the length of the segment between the joints of the first reference body frame to match the length of the segment between corresponding joints of the first subject body frame.

[0152] like Figure 8 As shown, a series of actions 800 includes an action 840 of providing a redirected reference body frame via a camera viewfinder stream. For example, action 840 involves providing a redirected first reference body frame aligned with a first object in the camera viewfinder stream. In one or more embodiments, providing the redirected first reference body frame aligned with the first object in the camera viewfinder stream is based on at least one landmark relative to the first object. In at least one embodiment, providing the redirected first reference body frame aligned with the first object in the camera viewfinder stream based on at least one landmark relative to the first object includes: determining at least one landmark of the first object in the camera viewfinder stream; generating a visualization of the redirected first reference body frame; and providing a visualization of the redirected first reference body frame on the camera viewfinder stream by anchoring at least one predetermined point of the visualization of the redirected first reference body frame to at least one landmark of the first object. For example, determining at least one landmark of the first object includes determining at least one of the following: a hip area of ​​the first object, and a torso area of ​​the first object.

[0153] In at least one embodiment, the sequence of actions 800 includes the actions of utilizing a pose neural network to: extract a second reference body frame from the selected sample pose image, and extract a second object body frame from a digital image from a camera viewfinder stream of a client computing device. The sequence of actions 800 also includes the actions of: reorienting the second reference body frame to include a proportion of the second object body frame in the pose of the second reference body frame; and providing the reorientated second reference body frame aligned with the second object in the camera viewfinder stream based on at least one landmark relative to the second object.

[0154] As mentioned, Fig. 9 900 is a flow diagram illustrating a series of actions for automatically capturing a digital image in response to determining perfect alignment between an object depicted in a camera viewfinder stream and a gesture guide in accordance with one or more embodiments. Fig. 9 The actions according to one embodiment are illustrated, but alternative embodiments may omit, add, reorder and / or modify Fig. 9 Alternatively, the non-transitory computer readable medium may include a computer program that, when executed by one or more processors, causes a computing device to perform Fig. 9 In some embodiments, the system may execute Fig. 9 action.

[0155] like Fig. 9 As shown, a series of actions 900 includes an action 910 of generating an object body frame associated with an object in a camera viewfinder stream. For example, action 910 involves generating an object body frame associated with an object depicted in a camera viewfinder stream of a client computing device. In one or more embodiments, the series of actions 900 also includes an action of providing a sample pose image in response to determining that the sample pose image corresponds to a context of the camera viewfinder stream; and wherein generating the object body frame is in response to detecting a selection of the sample pose image.

[0156] like Fig. 9 As shown, a series of actions 900 includes an action 920 of reorienting a reference body frame based on a subject body frame. For example, action 920 involves reorienting a reference body frame extracted from a sample pose image based on a subject body frame.

[0157] like Fig. 9As shown, a series of actions 900 includes an action 930 of overlaying the redirected reference body frame onto the camera viewfinder stream. For example, action 930 involves overlaying the redirected reference body frame onto the camera viewfinder stream by anchoring the redirected reference body frame to the object. In at least one embodiment, action 900 also includes: overlaying the object body frame onto the camera viewfinder stream by anchoring the object body frame to the object. For example, anchoring the redirected reference body frame to the object includes anchoring at least the hip region of the redirected reference body frame to the hip region of the object; and anchoring the object body frame to the object includes anchoring at least the hip region of the object body frame to the hip region of the object.

[0158] like Fig. 9 As shown, a series of actions 900 includes an action 940 of iteratively determining an alignment between a redirected reference body frame and an object. For example, action 940 involves iteratively determining an alignment between a portion of the redirected reference body frame and a portion of an object depicted in a camera viewfinder stream. In one or more embodiments, iteratively determining an alignment between a portion of the redirected reference body frame and a portion of an object depicted in a camera viewfinder stream includes iteratively determining that at least one segment of the redirected reference body frame that is anchored to the object overlaps with a corresponding at least one segment of the object body frame that is anchored to the object. In addition, the series of actions 900 may also include the following action: determining that a portion of the redirected reference body frame and a portion of the object are aligned by determining that each portion of the redirected reference body frame is aligned with a corresponding portion of the object body frame that is anchored to the object.

[0159] like Fig. 9 As shown, a series of actions 900 includes an action 950 of modifying a display characteristic of the redirected reference body frame based on the alignment. For example, action 950 involves, for each determined alignment, modifying a display characteristic of the redirected reference body frame to indicate the alignment. In one or more embodiments, modifying the display characteristic of the redirected reference body frame includes modifying at least one of the following: a display color of a portion of the redirected reference body frame that is aligned with a corresponding portion of the object, a line type of a portion of the redirected reference body frame that is aligned with a corresponding portion of the object, and a line width of a portion of the redirected reference body frame that is aligned with a corresponding portion of the object.

[0160] like Fig. 9 As shown, a series of actions 900 includes an action 960 of automatically capturing a digital image from a camera viewfinder stream based on a complete alignment between the redirected reference body frame and the object. For example, action 960 involves automatically capturing a digital image from a camera viewfinder stream in response to determining that a portion of the redirected reference body frame and a portion of the object are aligned.

[0161] Embodiments of the present disclosure may include or use a special-purpose or general-purpose computer, including computer hardware, such as, for example, one or more processors and system memory, as discussed in more detail below. Embodiments within the scope of the present disclosure also include physical and other computer-readable media for carrying or storing computer-executable instructions and / or data structures. Specifically, one or more processes described herein may be implemented at least in part as instructions implemented in a non-transitory computer-readable medium and executable by one or more computing devices (e.g., any media content access device described herein). Typically, a processor (e.g., a microprocessor) receives instructions from a non-transitory computer-readable medium (e.g., a memory) and executes those instructions, thereby performing one or more processes, including one or more processes described herein.

[0162] Computer readable media can be any available media that is accessible by a general or special purpose computer system. A computer readable medium that stores computer executable instructions is a non-transitory computer readable storage medium (device). A computer readable medium that carries computer executable instructions is a transmission medium. Therefore, by way of example and not limitation, embodiments of the present disclosure may include at least two distinct computer readable media: a non-transitory computer readable storage medium (device) and a transmission medium.

[0163] Non-transitory computer-readable storage media (devices) include RAM, ROM, EEPROM, CD-ROM, solid-state drives (“SSD”) (e.g., RAM-based), flash memory, phase change memory (“PCM”), other types of memory, other optical disk storage devices, magnetic disk storage devices or other magnetic storage devices, or any other media that can be used to store desired program code components in the form of computer-executable instructions or data structures and accessed by a general or special purpose computer.

[0164] "Network" is defined as one or more data links capable of conveying electronic data between computer systems and / or modules and / or other electronic devices. When information is transmitted or provided to a computer via a network or another communication connection (hardwired, wireless, or a combination of hardwired or wireless), the computer correctly views the connection as a transmission medium. Transmission media include networks and / or data links, which can be used to carry desired program code components in the form of computer executable instructions or data structures and are accessed by general or special computers. The above combinations should also be included within the scope of computer-readable media.

[0165] Further, upon reaching various computer system components, program code components in the form of computer executable instructions or data structures can be automatically transferred from the transmission medium to the non-transient computer readable storage medium (device) (and vice versa). For example, computer executable instructions or data structures received via a network or data link can be buffered in RAM within a network interface module (e.g., "NIC") and then ultimately transferred to the computer system RAM and / or a less volatile computer storage medium (device) at the computer system. Therefore, it should be understood that non-transient computer readable storage media (devices) can be included in computer system components that also (or even primarily) use transmission media.

[0166] For example, computer executable instructions include instructions and data that make a general-purpose computer, a special-purpose computer or a special-purpose processing device perform a specific function or function group when executed by a processor. In some embodiments, computer executable instructions are executed by a general-purpose computer to convert a general-purpose computer into a special-purpose computer for implementing the elements of the present disclosure. Computer executable instructions can be, for example, binary, intermediate format instructions (such as, assembly language) or even source code. Although the subject matter has been described in a language specific to structural features and / or method actions, it is to be understood that the subject matter defined in the appended claims is not necessarily limited to the features or actions described above. On the contrary, the described features and actions are disclosed as example forms for implementing the claims.

[0167] Those skilled in the art will appreciate that the present disclosure can be practiced in a network computing environment using many types of computer system configurations, including personal computers, desktop computers, laptop computers, message processors, handheld devices, multiprocessor systems, microprocessor-based or programmable consumer electronics, network PCs, minicomputers, mainframe computers, mobile phones, PDAs, tablet computers, pagers, routers, switches, etc. The present disclosure can also be practiced in a distributed system environment, where local and remote computer systems linked by a network (by a hardwired data link, a wireless data link, or by a combination of hardwired and wireless data links) all perform tasks. In a distributed system environment, program modules can be located in local and remote memory storage devices.

[0168] Embodiments of the present disclosure may also be implemented in a cloud computing environment. As used herein, the term "cloud computing" refers to a model for implementing on-demand network access to a shared pool of configurable computing resources. For example, cloud computing may be used in a marketplace to provide universal and convenient on-demand access to a shared pool of configurable computing resources. A shared pool of configurable computing resources may be rapidly provisioned via virtualization and published with low management effort or service provider interaction, and then scaled accordingly.

[0169] The cloud computing model can consist of various characteristics, such as, for example, on-demand self-service, broad network access, resource pooling, rapid elasticity, measured services, etc. The cloud computing model can also expose various service models, such as, for example, software as a service ("SaaS"), platform as a service ("PaaS"), and infrastructure as a service ("IaaS"). The cloud computing model can also be deployed using different deployment models, such as private cloud, community cloud, public cloud, hybrid cloud, etc. In addition, as used herein, the term "cloud computing environment" refers to an environment that employs cloud computing.

[0170] Fig.10 A block diagram of an example computing device 1000 is illustrated, which can be configured to perform one or more of the above-described processes. It is to be understood that one or more computing devices such as computing device 1000 can represent the above-described computing devices (e.g., (multiple) servers 106, client computing devices 108). In one or more embodiments, computing device 1000 can be a mobile device (e.g., a mobile phone, a smart phone, a PDA, a tablet computer, a laptop computer, a camera, a tracker, a watch, a wearable device, etc.). In some embodiments, computing device 1000 can be a non-mobile device (e.g., a desktop computer or another type of client computing device). Further, computing device 1000 can be a server device that includes cloud-based processing and storage capabilities.

[0171] like Fig.10 As shown, computing device 1000 includes one or more processors 1002, memory 1004, storage device 1006, input / output interface 1008 (or, "I / O interface 1008"), and communication interface 1010, which may be communicatively coupled via a communication infrastructure (e.g., bus 1012). Fig.10 However, Fig.10 The components illustrated are not intended to be limiting. Additional or alternative components may be used in other embodiments. Furthermore, in some embodiments, the computing device 1000 includes more than Fig.10 Fewer components than those shown. Fig.10 The components of the illustrated computing device 1000 will now be described in additional detail.

[0172] In certain embodiments, processor(s) 1002 include hardware for executing instructions, such as instructions that make up a computer program. As an example and not by way of limitation, to execute instructions, processor(s) 1002 may retrieve (or fetch) instructions from internal registers, internal caches, memory 1004, or storage device 1006, and decode and execute them.

[0173] The computing device 1000 includes a memory 1004 coupled to the processor(s) 1002. The memory 1004 may be used to store data, metadata, and programs for execution by the processor(s). The memory 1004 may include one or more of volatile and non-volatile memories, such as random access memory ("RAM"), read-only memory ("ROM"), solid state disk ("SSD"), flash memory, phase change memory ("PCM"), or other types of data storage. The memory 1004 may be internal or distributed memory.

[0174] The computing device 1000 includes a storage device 1006, including a storage device for storing data or instructions. As an example and not by way of limitation, the storage device 1006 includes the non-transitory storage medium described above. The storage device 1006 may include a hard disk drive (HDD), a flash memory, a universal serial bus (USB) drive, or a combination of these or other storage devices.

[0175] As shown, computing device 1000 includes one or more I / O interfaces 1008, which are provided to allow a user to provide input (such as user stylus) to computing device 1000, receive output from computing device 1000, and otherwise transfer data between computing devices 1000. These I / O interfaces 1008 may include a mouse, keypad or keyboard, touch screen, camera, optical scanner, network interface, modem, other known I / O devices, or a combination of such I / O interfaces 1008. The touch screen may be activated using a stylus or a finger.

[0176] The I / O interface 1008 may include one or more devices for presenting output to a user, including, but not limited to, a graphics engine, a display (e.g., a display screen), one or more output drivers (e.g., a display driver), one or more audio speakers, and one or more audio drivers. In some embodiments, the I / O interface 1008 is configured to provide graphical data to the display for presentation to the user. The graphical data may represent one or more graphical user interfaces and / or any other graphical content, which may serve a particular implementation.

[0177] The computing device 1000 may also include a communication interface 1010. The communication interface 1010 includes hardware, software, or both. The communication interface 1010 provides one or more interfaces for communicating between the computing device and one or more other computing devices or one or more networks (such as, for example, packet-based communications). As an example and not by way of limitation, the communication interface 1010 may include a network interface controller (NIC) or a network adapter for communicating with an Ethernet or other wired-based network, or a wireless NIC (WNIC) or a wireless adapter for communicating with a wireless network (such as, WI-FI). The computing device 1000 may also include a bus 1012. The bus 1012 includes hardware, software, or both, which connects the components of the computing device 1000 to each other.

[0178] In the foregoing description, the present invention has been described with reference to specific example embodiments thereof. Various embodiments and aspects of the present invention(s) are described with reference to the details discussed herein, and the accompanying drawings illustrate various embodiments. The above description and the accompanying drawings illustrate the present invention and are not to be construed as limiting the present invention. Many specific details are described to provide a thorough understanding of the various embodiments of the present invention.

[0179] The present invention may be specifically implemented in other specific forms without departing from its spirit or essential characteristics. The described embodiments are considered to be illustrative in all respects only, and not restrictive. For example, the method described herein may be performed with fewer or more steps / actions, or these steps / actions may be performed in different orders. Additionally, the steps / actions described herein may be repeated or performed in parallel with each other or with different instances of the same or similar steps / actions. Therefore, the scope of the present invention is indicated by the appended claims, rather than by the foregoing description. All changes that fall within the equivalent meaning and scope of the claims will be included within their scope.

Claims

1. A computer-implemented method comprising: generating a subject body frame associated with a subject depicted in a camera viewfinder stream of a client computing device; reorienting a reference body frame extracted from a sample pose image based on the subject body frame; overlaying the redirected reference body frame onto the camera viewfinder stream by anchoring the redirected reference body frame to the object; iteratively determining an alignment between the redirected portion of the reference body frame and the portion of the object depicted in the camera viewfinder stream; for each determined alignment, modifying a display characteristic of the redirected reference body frame to indicate the alignment; as well as In response to determining that the redirected portion of the reference body frame is aligned with the portion of the object, a digital image is automatically captured from the camera viewfinder stream.

2. The computer-implemented method of claim 1 , further comprising: The subject body frame is overlaid onto the camera viewfinder stream by anchoring the subject body frame to the subject.

3. The computer-implemented method of claim 2, wherein: Anchoring the redirected reference body frame to the subject includes: anchoring at least a hip region of the redirected reference body frame to a hip region of the subject; and Anchoring the subject body frame to the subject includes anchoring at least a hip region of the subject body frame to the hip region of the subject.

4. The computer-implemented method of claim 3, wherein iteratively determining an alignment between the redirected portion of the reference body frame and the portion of the object depicted in the camera viewfinder stream comprises: It is iteratively determined that at least one segment of the redirected reference body frame anchored to the subject covers a corresponding at least one segment of the subject body frame anchored to the subject.

5. The computer-implemented method of claim 4, further comprising: By determining that each portion of the redirected reference body frame is aligned with a corresponding portion of the subject body frame anchored to the subject, it is determined that the portions of the redirected reference body frame are aligned with portions of the subject.

6. A computer-implemented method according to claim 1, wherein modifying the display characteristics of the redirected reference body frame includes modifying at least one of the following: the display color of the portion of the redirected reference body frame aligned with the corresponding portion of the object, the line type of the portion of the redirected reference body frame aligned with the corresponding portion of the object, and the line width of the portion of the redirected reference body frame aligned with the corresponding portion of the object.

7. The computer-implemented method of claim 1 , further comprising: In response to determining that the sample gesture image corresponds to a context of the camera viewfinder stream, providing the sample gesture image; and Wherein generating the subject body frame is responsive to a detected selection of the sample pose image.

8. A system for posture guidance, comprising: at least one computer memory device including a plurality of sample pose images and a pose neural network; as well as One or more servers, configured to enable the system to: detecting a selection of a sample pose image from an interactive overlay located on a camera viewfinder of a client computing device; The posture neural network is used to: extracting a first reference body frame from the selected sample pose image, and extracting a first subject body frame from a digital image streamed from a camera viewfinder of the client computing device; reorienting the first reference body frame to include a proportion of the first subject body frame in a pose of the first reference body frame; as well as The redirected first reference body frame is provided aligned with a first object in the camera viewfinder stream.

9. The system of claim 8, wherein the one or more servers are further configured to cause the system to reorient the first reference body frame to include the proportions of the first subject body frame in the pose of the first reference body frame by: determining the length of the segments between the joints of the first subject's body frame; and The lengths of segments between joints of the first reference body frame are modified to match the lengths of the segments between corresponding joints of the first subject body frame.

10. The system of claim 8, wherein the one or more servers are further configured to cause the system to: provide the redirected first reference body frame aligned with the first object in the camera viewfinder stream based on at least one landmark relative to the first object.

11. The system of claim 10, wherein the one or more servers are further configured to cause the system to provide the redirected first reference body frame aligned with the first object in the camera viewfinder stream based on at least one landmark relative to the first object by: determining the at least one landmark of the first object in the camera viewfinder stream; generating a visualization of the redirected first reference body frame; as well as The visualization of the redirected first reference frame on the camera viewfinder stream is provided by anchoring at least one predetermined point of the visualization of the redirected first reference frame to the at least one landmark of the first object. 12 . The system of claim 11 , wherein determining the at least one landmark of the first subject comprises determining at least one of: a hip region of the first subject, and a torso region of the first subject.

13. The system of claim 8, wherein the one or more servers are further configured to cause the system to: The posture neural network is used to: extracting a second reference body frame from the selected sample pose image, and extracting a second subject body frame from a digital image of the camera viewfinder stream from the client computing device; reorienting the second reference body frame to include a proportion of the second subject body frame in a pose of the second reference body frame; as well as The redirected second reference body frame is provided aligned with the second object in the camera viewfinder stream based on at least one landmark relative to the second object.

14. A non-transitory computer-readable storage medium comprising instructions that, when executed by at least one processor, cause a computing device to: receiving a digital image from a camera viewfinder stream of the computing device; determining a context of the digital image based on objects and scenes depicted within the digital image; generating a set of sample pose images corresponding to the context of the digital image; as well as Different subsets of the set of sample gesture images are provided for display via an interactive overlay located on a camera viewfinder of the computing device.

15. The non-transitory computer-readable storage medium of claim 14, further comprising instructions that, when executed by the at least one processor, cause the computing device to determine the context of the digital image by: generating one or more object tags associated with the digital image; generating a gender label for each person depicted in the digital image; generating one or more clothing tags for each person depicted in the digital image; as well as The context of the digital image is determined based on the one or more object tags, the gender tag for each person depicted in the digital image, and the one or more clothing tags for each person depicted in the digital image.

16. The non-transitory computer-readable storage medium of claim 15, further comprising instructions that, when executed by the at least one processor, cause the computing device to generate the set of sample pose images by: generating a search query based on the one or more object tags, the gender tag for each person depicted in the digital image, and the one or more clothing tags for each person depicted in the digital image; and In response to providing the search query to one or more search engines, the set of sample gesture images is received.

17. The non-transitory computer-readable storage medium of claim 14, further comprising instructions that, when executed by the at least one processor, cause the computing device to determine the different subsets of the set of sample gesture images by: generating a feature vector for the sample posture image in the sample posture image set; Clustering the feature vectors to determine one or more categories of the sample posture image; selecting a feature vector from each of the one or more categories of sample gesture images; as well as The different subset of the set of sample pose images is determined as the sample pose images corresponding to the selected feature vector.

18. The non-transitory computer-readable storage medium of claim 14, further comprising instructions that, when executed by the at least one processor, cause the computing device to provide the different subset of the set of sample gesture images for display by: generating the interactive overlay, the interactive overlay comprising a scrollable display of the different subsets of the set of sample gesture images, wherein each of the sample gesture images in the scrollable display is selectable; and The interactive overlay is positioned over a lower portion of the camera viewfinder of the computing device.

19. The non-transitory computer-readable storage medium of claim 14, wherein generating the interactive overlay further comprises: An indication of the context of the digital image is provided within the interactive overlay.

20. The non-transitory computer-readable storage medium of claim 14, further comprising instructions that, when executed by the at least one processor, cause the computing device to: receive the digital image from the camera viewfinder stream in response to detecting a selection of an entry point option displayed in conjunction with the camera viewfinder stream on the computing device.

Citation Information

Patent Citations

  • Accurate tag relevance prediction for image search

    US10235623B2

  • Utilizing object attribute detection models to automatically select instances of detected objects in images

    US11107219B2

  • Identifying digital attributes from multiple attribute groups within target digital images utilizing a deep cognitive attribution neural network

    US20210073267A1

  • Training a classifier algorithm used for automatically generating tags to be applied to images

    US9767386B2

  • Auxiliary photographing method, mobile terminal and storage medium

    CN109743504A