Method, apparatus, and computer system for estimating three-dimensional gestures in an image

By receiving the data corresponding to the two hand images, calculating the optical flow value, generating a heat map, estimating the hand grid map and determining the gesture, the problem of inaccurate estimation of three-dimensional gestures from RGB images in the prior art is solved, and efficient and accurate gesture recognition is achieved.

CN115004293BActive Publication Date: 2025-06-24TENCENT AMERICA LLC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202180009241.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-02-13
Filing Date
2021-02-08
Publication Date
2025-06-24
Estimated Expiration
2041-02-08

AI Technical Summary

Technical Problem

The prior art is difficult to effectively estimate three-dimensional gestures from RGB images, especially when the three-dimensional gesture data samples are limited, the determination of gestures is not accurate enough.

Method used

By receiving data corresponding to the two hand images, the optical flow value is calculated, the heat map is generated, the hand grid map is estimated, and the gesture is determined based on the hand grid map. This method uses only two-dimensional spatial information, using self-supervised learning models such as TASSN, to estimate 3D gestures and mesh from video annotated with 2D keyframe locations through time consistency constraints.

Benefits of technology

The gesture recognition of three-dimensional gestures is realized based on two-dimensional spatial information, which improves the convenience and accuracy of gesture recognition, especially when the three-dimensional gesture data samples are limited.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115004293B_ABST
    Figure CN115004293B_ABST
Patent Text Reader

Abstract

A method, a computer program, and a computer system for estimating three-dimensional gestures in an image are provided. Data corresponding to two hand images is received, and an optical flow value corresponding to a change in hand pose in the received hand image data is calculated. A heat map is generated based on the calculated optical flow value, and a hand mesh map is estimated based on the generated heat map. A gesture present in the hand image is determined based on the estimated hand mesh map.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross - Reference to Related Applications

[0002] This application claims priority to U.S. Application No. 16 / 789,507, filed on February 13, 2020, which is hereby incorporated by reference in its entirety into this application. Technical Field

[0003] The present disclosure generally relates to the field of computing, and more particularly to methods, apparatuses, and computer systems for estimating three - dimensional gestures in images. Background Art

[0004] Gesture estimation is the task of finding the joints of a hand from an image or a set of video frames. Estimating three - dimensional (3D) gestures in Red - Green - Blue (RGB) color images is essential for a wide range of potential applications such as computer vision, virtual reality, augmented reality, and other forms of human - computer interaction. Due to the ease of capturing RGB images through webcameras, Internet of Thing (IoT) cameras, and smartphones, gesture estimation in RGB images has become significantly more popular. Summary of the Invention

[0005] Embodiments relate to methods, systems, and computer - readable media for estimating 3D gestures. According to one aspect, a method for estimating 3D gestures is provided. The method may include receiving data corresponding to two hand images. Calculating optical flow values corresponding to hand - pose changes in the received data and generating a heatmap based on the calculated optical flow values. Estimating a hand mesh map based on the generated heatmap and determining the gestures present in the two hand images based on the estimated hand mesh map.

[0006] According to another aspect, a computer system for estimating 3D gestures is provided. The computer system may include one or more processors, one or more computer - readable memories, one or more computer - readable tangible storage devices, and program instructions stored on at least one of the one or more storage devices, the program instructions being for execution by at least one of the one or more processors via at least one of the one or more memories, whereby the computer system is capable of performing the method described above.

[0007] According to yet another aspect, a computer - readable medium for estimating 3D gestures is provided. The computer - readable medium may include one or more computer - readable storage devices and program instructions stored on at least one of the one or more tangible storage devices, the program instructions being executable by a processor. The program instructions are executable by the processor to perform the method described above.

[0008] In the solution of the present application, first, data corresponding to two hand images is received. Then, an optical flow value corresponding to the change in hand posture in the received data is calculated, and a heat map is generated based on the calculated optical flow value. After that, a hand mesh map is estimated based on the generated heat map, and a gesture existing in the two hand images is determined based on the estimated hand mesh map. In this way, a three-dimensional gesture can be determined only using two-dimensional spatial information, making gesture recognition more convenient, and further making the determination of the gesture more accurate in the case where the existing three-dimensional gesture data samples are limited. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] These and other objects, features, and advantages will become apparent from the following detailed description of the illustrative embodiments to be read in conjunction with the accompanying drawings. The various features of the drawings are not to scale as the illustrations are for clarity to facilitate understanding by those skilled in the art in conjunction with the detailed description. In the drawings:

[0010] Figure 1 shows a networked computer environment according to at least one embodiment;

[0011] Figure 2 is a block diagram of a program for estimating 3D gestures according to at least one embodiment;

[0012] Figure 3 is an operational flowchart showing the steps performed by a program for estimating 3D gestures according to at least one embodiment;

[0013] Figure 4 is according to at least one embodiment Figure 1 a block diagram of the internal and external components of the computer and server depicted in

[0014] Figure 5 is according to at least one embodiment including Figure 1 a block diagram of an illustrative cloud computing environment of the computer system depicted in

[0015] Figure 6 is according to at least one embodiment Figure 5 a block diagram of the functional layers of an illustrative cloud computing environment of DETAILED DESCRIPTION OF THE EMBODIMENTS

[0016] This disclosure describes in detail embodiments of the claimed structures and methods; however, it is to be understood that the disclosed embodiments are merely illustrative of the claimed structures and methods that may be implemented in various forms. Those structures and methods may, however, be embodied in many different forms and should not be construed as limited to the exemplary embodiments set forth herein. Rather, these exemplary embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope to those skilled in the art. In the description, details of well-known features and techniques may be omitted to avoid unnecessarily obscuring the presented embodiments.

[0017] Embodiments generally relate to the field of computing, and more particularly, to machine learning. The exemplary embodiments described below provide systems, methods, and program products for estimating 3D gestures and the like. Thus, some embodiments have the ability to improve the field of computing by allowing the use of deep neural networks to allow a computer to determine 3D gestures using only 2D spatial information.

[0018] As previously mentioned, gesture estimation is the task of finding the joints of a hand in an image or a set of video frames. Estimating three-dimensional (3D) gestures in red-green-blue (RGB) color images is essential for a wide range of potential applications such as computer vision, virtual reality, augmented reality, and other forms of human-computer interaction. Due to the ease of capturing RGB images via webcameras, Internet of Things (IoT) cameras, and smartphones, estimating gestures in RGB images has become significantly more popular. Directly estimating 3D gestures from RGB images is a challenging task, but progress has been made in training deep models using annotated 3D gestures. However, annotating 3D gestures can be difficult, and thus, only a few 3D gesture datasets may be available with limited sample sizes. Therefore, it is advantageous to train a 3D gesture estimation model using machine learning and neural networks based on RGB images without explicit 3D annotation (i.e., training using only 2D information).

[0019] Aspects are described herein with reference to flowcharts and / or block diagrams of methods, apparatus (systems), and computer-readable media according to various embodiments. It should be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable program instructions.

[0020] The exemplary embodiments described below provide a system, method, and computer-readable medium for estimating 3D gestures. According to this embodiment, a self-supervised learning model called the Temporal-Aware Self-Supervised Network (TASSN) can be used to estimate 3D gestures and meshes from videos annotated only with 2D keyframe positions by implementing temporal consistency constraints.

[0021] Now referring to Figure 1 , a functional block diagram of a networked computer environment is shown, which shows a gesture estimation system 100 (hereinafter referred to as the "system") for improved estimation of 3D gestures. It should be understood that Figure 1 only an illustration of one implementation is provided, and it does not imply any limitations on the environment in which different embodiments can be implemented. Many modifications can be made to the depicted environment based on design and implementation requirements.

[0022] System 100 may include a computer 102 and a server computer 114. Computer 102 may communicate with server computer 114 via a communication network 110 (hereinafter referred to as the "network"). Computer 102 may include a processor 104 and a software program 108 stored on a data storage device 106. Computer 102 is capable of interfacing with a user and communicating with server computer 114. As will be discussed below with reference to Figure 4 , computer 102 may include internal components 800A and external components 900A respectively, and server computer 114 may include internal components 800B and external components 900B respectively. Computer 102 may be, for example, a mobile device, a phone, a personal digital assistant, a netbook, a laptop computer, a tablet computer, a desktop computer, or any type of computing device capable of running programs, accessing the network, and accessing a database.

[0023] As discussed below with respect to Figure 5 and Figure 6 , server computer 114 may also operate in a cloud computing service model such as Software as a Service (SaaS), Platform as a Service (PaaS), or Infrastructure as a Service (IaaS). Server computer 114 may also be located in a cloud computing deployment model such as a private cloud, a community cloud, a public cloud, or a hybrid cloud.

[0024] Server computer 114, which can be used to estimate three-dimensional gestures in an image, is capable of running a gesture estimation program 116 (hereinafter referred to as the "program") that can interact with a database 112. In the following with respect toFigure 3 Describe the gesture estimation program method in more detail. In one embodiment, the computer 102 can operate as an input device including a user interface, while the program 116 can run mainly on the server computer 114. In an alternative embodiment, the program 116 can run mainly on one or more computers 102, and the server computer 114 can be used to process and store the data used by the program 116. It should be noted that the program 116 can be an independent program or can be integrated into a larger gesture estimation program.

[0025] However, it should be noted that in some cases, the processing for the program 116 can be shared between the computer 102 and the server computer 114 at any ratio. In another embodiment, the program 116 can operate on more than one computer, server computer, or some combination of computers and server computers, such as multiple computers 102 communicating with a single server computer 114 across the network 110. In another embodiment, for example, the program 116 can operate on multiple server computers 114 communicating with multiple client computers across the network 110. Alternatively, the program can operate on a network server communicating with servers and multiple client computers across the network.

[0026] The network 110 can include a wired connection, a wireless connection, a fiber optic connection, or some combination thereof. Generally, the network 110 can be any combination of connections and protocols that support communication between the computer 102 and the server computer 114. The network 110 can include various types of networks, such as, for example, a local area network (LAN), a wide area network (WAN) such as the Internet, a telecommunications network such as a public switched telephone network (PSTN), a wireless network, a public switched network, a satellite network, a cellular network (e.g., a fifth generation (5G) network, a long-term evolution (LTE) network, a third generation (3G) network, a code division multiple access (CDMA) network, etc.), a public land mobile network (PLMN), a metropolitan area network (MAN), a private network, an ad hoc network, an intranet, a fiber optic-based network, etc., and / or a combination of these or other types of networks.

[0027] Provide Figure 1The number and arrangement of devices and networks shown are examples. In practice, there may be Figure 1 The devices and / or networks shown may include additional devices and / or networks, fewer devices and / or networks, different devices and / or networks, or differently arranged devices and / or networks. Figure 1 Two or more of the devices shown may be implemented in a single device, or Figure 1 The single device shown may be implemented as multiple distributed devices. Additionally or alternatively, one or more devices of a set of devices (eg, one or more devices) of the system 100 may perform one or more functions described as being performed by another set of devices of the system 100.

[0028] Reference Figure 2 , depicting Figure 1 Block diagram 200 of the gesture estimation program 116 . Figure 2 Can be helped by Figure 1 . Accordingly, the gesture estimation program 116 may include a flow estimation module 202, a heat map estimation module 204, a graph convolutional network (GCN) hand mesh estimation module 206, and a 3D gesture estimation module 208. The gesture estimation program 116 may be configured to receive hand image data 210, create feature data 212 for one or more features F, and output gesture data 214. The gesture estimation program 116 may also generate a heat map loss value L h and the grid loss value L m According to one embodiment, the gesture estimation program 116 may be located on the computer 102 ( Figure 1 According to an alternative embodiment, the gesture estimation program 116 may be located on the server computer 114 ( Figure 1 )superior.

[0029] The gesture image data 210 may include an RGB hand motion video x having N frames, etc., where x={I t ,…,I t+n}, can be the t-th frame, and W and H can be the frame width and height, respectively. 3D gesture at frame t It can be represented by a set of 3D keypoint coordinates of the hand, where K can be the number of keypoints. Using the temporal consistency properties of the video, the predicted gestures and meshes in the forward and backward inference processes can perform mutual supervision. With such an approach, self-supervised learning can be used to train the model, and 3D hand keypoint annotations may no longer be required. Using the hand mesh to train the gesture estimator improves performance because the hand mesh can be used as an intermediate guide for gesture prediction.

[0030] The optical flow estimation module 202 can estimate the optical flow between two consecutive frames using forward and backward inference. The heatmap estimation module 204 can calculate 2D hand key points and generate features and mesh estimators for 3D gestures. The estimated 2D key point heatmap can be represented by , where K can represent the number of key points. Two stacked hourglass networks can be used to infer the hand key point heatmap H and calculate the feature data 212.

[0031] The heatmap estimation module 204 can connect I t+1 , o t+1 and H t as inputs to the stacked hourglass network, which can generate the heatmap H t+1 . The estimated H t+1 can include K heatmaps where can represent the confidence map of the position of the k-th key point. The ground truth heatmap can be the Gaussian blur of the Dirac-δ distribution centered at the ground truth position of the k-th key point. The heatmap loss L h at frame t can be defined by .

[0032] The graph convolutional network (GCN) hand mesh estimation module 206 can take the hand feature data 212 as input and can infer the 3D hand mesh. The output hand mesh can be represented by a set of 3D mesh vertices, where C can be the number of vertices in the hand mesh. The hand mesh can be constructed in a coarse-to-fine manner, and a multi-level clustering algorithm can be used to coarsen the graph. The graph can be stored at each level, and the mapping between graph nodes can be stored for every two consecutive levels. In forward inference, the GCN can upsample the node features according to the stored mapping and graph and can perform graph convolutional operations. To avoid a folded mesh, the mesh loss value Lm can be used to calculate the difference between the contour s t of the predicted hand mesh at frame t and the ground truth contour . The contour loss can be defined by . To obtain the hand contour can be estimated from the training images, and the contour of the predicted hand mesh m t can be obtained by using a neural rendering method.

[0033] The 3D gesture estimation module 208 can directly infer the 3D hand key points p t from the predicted hand mesh m t. The 3D gesture estimation module 208 may include networks such as two stacked GCNs. Pooling layers may be added to each GCN to extract gesture features from the grid, and the gesture features may be fed into two fully connected layers to regress the 3D gesture p t for regression.

[0034] Now referring to Figure 3 , an operation flowchart 300 showing the steps performed by a program for estimating 3D gestures is depicted. Figure 3 It can be described with the aid of Figure 1 and Figure 2 . As previously mentioned, the gesture estimation program 116 ( Figure 1 ) can quickly and effectively determine 3D gestures based on 2D data.

[0035] At 302, the computer receives data corresponding to two hand images. The hand image data may be two-dimensional data. The hand image data may be two discrete images, or may be consecutive frames extracted from a video source. In operation, the gesture estimation program ( Figure 1 ) may receive the hand image data 210 ( Figure 1 ) through the communication network 110 ( Figure 2 ).

[0036] At 304, the computer calculates the optical flow value corresponding to the change in the hand pose in the received hand image data. By determining the change in the hand position between images, temporal data can be used to help determine the gestures present in the image. In operation, the flow estimation module 202 ( Figure 2 ) may use forward and backward inference to determine the change between two consecutive images from the hand image data 210 ( Figure 2 ).

[0037] At 306, the computer generates a heatmap based on the calculated optical flow value. Two-dimensional hand key points can be used to calculate the heatmap, which can allow the determination of the features present in the three-dimensional gesture and the generation of a mesh estimator. In operation, the heatmap estimation module 204 ( Figure 2 ) may receive the optical flow value from the flow estimation module 202 ( Figure 2 ), and may take the hand image data 210 ( Figure 2 ) as input. The heatmap estimation module 204 can generate a heatmap and identify the feature data 212 ( Figure 2 ) that may correspond to the features of the hand image data 210.

[0038] At 308, the computer estimates a hand mesh map based on the generated heat map. The hand mesh map can implement a coarse-to-fine construction of a gesture by applying operations of layers of a convolutional neural network. In operation, the GCN hand mesh estimation module 206 can receive feature data 212 and the heat map from the heat map estimation module 204. The GCN hand mesh estimation module 206 can estimate a three-dimensional mesh corresponding to the gesture present in the hand image data 210.

[0039] At 310, the computer determines the gesture present in the hand image based on the estimated hand mesh map. A gesture can be generated without using three-dimensional annotations and only using two-dimensional image data. In operation, the 3D gesture estimation module 208 can receive the hand mesh from the GCN hand mesh estimation module and can determine the gesture associated with the hand based on the 3D hand mesh.

[0040] It can be understood that Figure 3 only an illustration of one implementation is provided and does not imply any limitation on how different implementations can be achieved. Many modifications can be made to the depicted environment based on design and implementation requirements.

[0041] Figure 4 is according to an illustrative embodiment Figure 1 is a block diagram 400 of the internal and external components of the computer depicted in Figure 4 only an illustration of one implementation is provided and does not imply any limitation on the environment in which different implementations can be achieved. Many modifications can be made to the depicted environment based on design and implementation requirements.

[0042] The computer 102 ( Figure 1 ) and the server computer 114 ( Figure 1 ) can include Figure 4 the corresponding groups of internal components 800A, 800B and external components 900A, 900B as shown. Each group of internal components 800 includes one or more processors 820 on one or more buses 826, one or more computer-readable RAMs 822, and one or more computer-readable ROMs 824, one or more operating systems 828, and one or more computer-readable tangible storage devices 830.

[0043] Processor 820 is implemented in hardware, firmware, or a combination of hardware and software. Processor 820 is a central processing unit (CPU), a graphics processing unit (GPU), an accelerated processing unit (APU), a microprocessor, a microcontroller, a digital signal processor (DSP), a field-programmable gate array (FPGA), an application-specific integrated circuit (ASIC), or another type of processing component. In some implementations, processor 820 includes one or more processors that can be programmed to perform functions. Bus 826 includes components that permit communication between internal components 800A, 800B.

[0044] One or more operating systems 828, software programs 108 ( Figure 1 ), and gesture estimation programs 116 ( Figure 1 ) on server computer 114 ( Figure 1 ) are stored on one or more respective computer-readable tangible storage devices 830 for execution by one or more respective processors 820 via one or more respective RAMs 822 (which typically include cache memory). In Figure 4 the illustrated implementation, each of computer-readable tangible storage devices 830 is a magnetic disk storage device of an internal hard disk drive. Alternatively, each of computer-readable tangible storage devices 830 is a semiconductor storage device such as ROM 824, EPROM, flash memory, optical disk, magneto-optical disk, solid state disk, compact disc (CD), digital versatile disk (DVD), floppy disk, cassette tape, magnetic tape, and / or another type of non-transitory computer-readable tangible storage device that can store computer programs and digital information.

[0045] Each set of internal components 800A, 800B also includes an R / W drive or interface 832 to read from and write to one or more portable computer-readable tangible storage devices 936 such as CD-ROMs, DVDs, memory sticks, magnetic tapes, magnetic disks, optical disks, or semiconductor storage devices. Such as software programs 108 ( Figure 1 ) and gesture estimation programs 116 ( Figure 1The software program ( ) can be stored on one or more corresponding portable computer-readable tangible storage devices 936, read via the corresponding R / W drive or interface 832, and loaded into the corresponding hard disk drive 830.

[0046] Each set of internal components 800A, 800B also includes a network adapter or interface 836, such as a TCP / IP adapter card; a wireless Wi-Fi interface card; or a 3G, 4G, or 5G wireless interface card or other wired or wireless communication link. The software program 108 ( Figure 1 ) and the gesture estimation program 116 ( Figure 1 ) on the server computer 114 ( Figure 1 ) can be downloaded from an external computer to the computer 102 ( Figure 1 ) and the server computer 114 via a network (such as the Internet, a local area network, or other wide area network) and the corresponding network adapter or interface 836. The software program 108 and the gesture estimation program 116 on the server computer 114 are loaded from the network adapter or interface 836 into the corresponding hard disk drive 830. The network can include copper wire, optical fiber, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers.

[0047] Each set of external components 900A, 900B can include a computer monitor 920, a keyboard 930, and a computer mouse 934. The external components 900A, 900B can also include a touch screen, a virtual keyboard, a touchpad, a pointing device, and other human-machine interface devices. Each set of internal components 800A, 800B also includes a device driver 840 connected to the computer monitor 920, the keyboard 930, and the computer mouse 934. The device driver 840, the R / W drive or interface 832, and the network adapter or interface 836 include hardware and software (stored in the storage device 830 and / or ROM 824).

[0048] It is understood in advance that although the present disclosure includes a detailed description of cloud computing, the implementation of the teachings described herein is not limited to a cloud computing environment. Instead, some embodiments can be implemented in conjunction with any other type of computing environment known now or developed later.

[0049] Cloud computing is a service delivery model for enabling convenient, on-demand network access to a shared pool of configurable computing resources (such as networks, network bandwidth, servers, processing, memory, storage devices, applications, virtual machines, and services), which can be rapidly provisioned and released with minimal management effort or interaction with the service provider. The cloud model can include at least five characteristics, at least three service models, and at least four deployment models.

[0050] The characteristics are as follows:

[0051] On-demand self-service: Cloud consumers can unilaterally and automatically provision computing capabilities such as server time and network storage as needed, without human interaction with the service provider.

[0052] Broad network access: Capabilities are available over a network and accessed through standard mechanisms that promote the use of heterogeneous thin or thick client platforms (such as mobile phones, laptops, and PDAs).

[0053] Resource pooling: A provider's computing resources are pooled to serve multiple consumers using a multi-tenant model, where different physical and virtual resources are dynamically assigned and reassigned according to demand. There is a sense of location independence in that consumers generally have no control or knowledge of the exact location of the provided resources, but can be able to specify location at a higher level of abstraction (such as country, state, or data center).

[0054] Rapid elasticity: Capabilities can be provisioned quickly and elastically (automatically in some cases) to scale out rapidly and released to scale in rapidly. To the consumer, the capabilities available for provisioning generally appear to be unlimited and can be purchased in any quantity at any time.

[0055] Measured service: Cloud systems automatically control and optimize resource use by leveraging metering capabilities at some level of abstraction appropriate to the type of service (such as storage, processing, bandwidth, and active user accounts). Resource use can be monitored, controlled, and reported, providing transparency for both the provider and consumer of the utilized services.

[0056] The service models are as follows:

[0057] Software as a Service (SaaS): The capabilities provided to the consumer are to use the provider's applications running on a cloud infrastructure. The applications can be accessed from various client devices through a thin client interface such as a web browser (such as web-based email). The consumer does not manage or control the underlying cloud infrastructure, including the network, servers, operating systems, storage devices, or even individual application capabilities, with the possible exception of limited user-specific application configuration settings.

[0058] Platform as a Service (PaaS): The capabilities provided to the consumer are to deploy consumer-created or acquired applications that use the programming languages and tools supported by the provider onto the cloud infrastructure. The consumer does not manage or control the underlying cloud infrastructure, including the network, servers, operating systems, or storage devices, but has control over the deployed applications and possible application-level configuration of the hosting environment.

[0059] Infrastructure as a Service (IaaS): The ability provided to the consumer is to provide processing, storage, networking, and other fundamental computing resources that the consumer can deploy and run arbitrary software, which may include operating systems and applications. The consumer does not manage or control the underlying cloud infrastructure, but has control over the operating system, storage, deployed applications, and possibly limited control over selected networking components (e.g., the main firewall).

[0060] The deployment models are as follows:

[0061] Private cloud: The cloud infrastructure is for the exclusive use of an organization. The cloud infrastructure can be managed by the organization or a third party and can exist on-premises or off-premises.

[0062] Community cloud: The cloud infrastructure is shared by several organizations and supports a specific community with common concerns (e.g., missions, security requirements, policies, and compliance considerations). The cloud infrastructure can be managed by the organization or a third party and can exist on-premises or off-premises.

[0063] Public cloud: The cloud infrastructure is available for general public or large industry groups and is owned by the organization selling the cloud services.

[0064] Hybrid cloud: The cloud infrastructure is a combination of two or more clouds (private cloud, community cloud, or public cloud), and the two or more clouds remain distinct entities but are bound together through standardized or proprietary technologies that enable data and application portability (e.g., cloud bursting for load balancing between clouds).

[0065] The cloud computing environment is service-oriented, focusing on statelessness, low coupling, modularity, and semantic interoperability. The core of cloud computing is the infrastructure that includes a network of interconnected nodes.

[0066] Refer to Figure 5 , which depicts an illustrative cloud computing environment 500. As shown, the cloud computing environment 500 includes one or more cloud computing nodes 10, and local computing devices used by cloud consumers such as a personal digital assistant (PDA) or cellular phone 54A, desktop computer 54B, laptop computer 54C, and / or in-vehicle computer system 54N can communicate with one or more cloud computing nodes 10. The cloud computing nodes 10 can communicate with each other. The cloud computing nodes 10 can be physically or virtually grouped (not shown) in one or more networks such as the private cloud, community cloud, public cloud, or hybrid cloud or a combination thereof as described above. This allows the cloud computing environment 500 to provide infrastructure, platform, and / or software as a service, and cloud consumers do not need to maintain resources on local computing devices for this service. It should be understood thatFigure 5 The types of computing devices 54A through 54N shown are only illustrative, and the cloud computing node 10 and cloud computing environment 500 can communicate with any type of computerized device via any type of network and / or network addressable connection (e.g., using a web browser).

[0067] Referring to Figure 6 , a set of functional abstraction layers 600 provided by the cloud computing environment 500 ( Figure 5 ) is shown. It should be understood upfront that Figure 6 the components, layers, and functions shown are only illustrative, and the implementation is not limited thereto. As depicted, the following layers and corresponding functions are provided:

[0068] The hardware and software layer 60 includes hardware components and software components. Examples of hardware components include: mainframes 61; servers 62 based on RISC (Reduced Instruction Set Computer) architecture; servers 63; blade servers 64; storage devices 65; and network and networking components 66. In some implementations, the software components include network application server software 67 and database software 68.

[0069] The virtualization layer 70 provides an abstraction layer from which the following examples of virtual entities can be provided: virtual servers 71; virtual storage devices 72; virtual networks 73 including virtual private networks; virtual applications and operating systems 74; and virtual clients 75.

[0070] In one example, the management layer 80 can provide the following functions. Resource provisioning 81 provides the dynamic acquisition of computing resources and other resources for performing tasks within the cloud computing environment. Metering and pricing 82 provides cost detection when resources are utilized within the cloud computing environment, as well as billing or invoicing for the consumption of these resources. In one example, these resources can include application software licenses. Security provides authentication for cloud consumers and tasks, as well as protection for data and other resources. The user portal 83 provides access to the cloud computing environment for consumers and system administrators. Service level management 84 provides cloud computing resource allocation and management such that the required service levels are met. Service level agreement (SLA) planning and fulfillment 85 provides the pre-arrangement and acquisition of cloud computing resources, anticipating future demands for the cloud computing resources according to the SLA.

[0071] The workload layer 90 provides examples of functions that can utilize the capabilities of a cloud computing environment. Examples of workloads and functions that can be provided from this layer include: drawing and navigation 91; software development and lifecycle management 92; virtual classroom teaching delivery 93; data analysis processing 94; transaction processing 95; and gesture estimation 96. Gesture estimation 96 can estimate 3D gestures based on 2D data.

[0072] Some embodiments may relate to systems, methods, and / or computer-readable media at any possible level of integration technical detail. The computer-readable media may include one or more computer-readable non-transitory storage media having computer-readable program instructions thereon for causing a processor to perform operations.

[0073] A computer-readable storage medium may be a tangible device that can retain and store instructions for use by an instruction execution device. For example, a computer-readable storage medium may be, but is not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. A non-exhaustive list of more specific examples of computer-readable storage media includes the following: a portable computer floppy disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disc (DVD), a memory stick, a floppy disk, a mechanically encoded device such as a punched card or raised structures in a groove having instructions recorded thereon, and any suitable combination of the foregoing. As used herein, a computer-readable storage medium should not be construed as being a transitory signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., an optical pulse transmitted through an optical fiber cable), or an electrical signal transmitted through a wire.

[0074] The computer-readable program instructions described herein can be downloaded to a corresponding computing / processing device from a computer-readable storage medium or downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network can include copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium within the corresponding computing / processing device.

[0075] The computer-readable program code / instructions for performing the operations can be assembly instructions, instruction-set-architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state-setting data, configuration data for integrated circuits, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages (e.g., Smalltalk, C++, etc.) and procedural programming languages (e.g., the "C" programming language or similar programming languages). The computer-readable program instructions can be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the latter case, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (e.g., through the Internet using an Internet service provider). In some embodiments, an electronic circuit, including, for example, a programmable logic circuit, a field-programmable gate array (FPGA), or a programmable logic array (PLA), can execute the computer-readable program instructions by utilizing the state information of the computer-readable program instructions to specialize the electronic circuit for performing various aspects or operations.

[0076] These computer-readable program instructions may be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions executed via the processor of the computer or other programmable data processing apparatus create means for implementing the functions / acts specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions may also be stored in a computer-readable storage medium that can direct a computer, a programmable data processing apparatus, and / or other devices to function in a particular manner, such that the computer-readable storage medium in which the instructions are stored comprises an article of manufacture including instructions implementing various aspects of the function / act specified in one or more blocks of the flowchart and / or block diagram.

[0077] The computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus, or other device to produce a computer-implemented process, such that the instructions executed on the computer, other programmable apparatus, or other device implement the functions / acts specified in one or more blocks of the flowchart and / or block diagram.

[0078] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods and computer-readable media according to various embodiments. In this regard, each block in the flowchart or block diagram may represent a module, segment, or portion of instructions, which comprises one or more executable instructions for implementing the specified logical function. The method, computer system and computer-readable media may include additional blocks, fewer blocks, different blocks, or differently arranged blocks compared to those depicted in the figures. In some alternative implementations, the functions noted in the blocks may occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently, or the blocks may sometimes be executed in the reverse order, depending upon the functionality involved. It should also be noted that each block of the block diagrams and / or flowchart illustrations, and combinations of blocks in the block diagrams and / or flowchart illustrations, can be implemented by special purpose hardware-based systems that perform the specified functions or acts, or combinations of special purpose hardware and computer instructions.

[0079] It will be apparent that the systems and / or methods described herein may be implemented in different forms of hardware, firmware, or a combination of hardware and software. The actual special control hardware or software code used to implement these systems and / or methods does not limit the implementation. Thus, the operations and behaviors of the systems and / or methods are described herein without reference to specific software code - it being understood that software and hardware can be designed to implement the systems and / or methods based on the description herein.

[0080] Unless expressly described as such, no element, act, or instruction used herein shall be construed as critical or essential. Additionally, as used herein, the article "a" is intended to include one or more items and may be used interchangeably with "one or more." Further, as used herein, the term "group" is intended to include one or more items (e.g., related items, unrelated items, combinations of related and unrelated items, etc.) and may be used interchangeably with "one or more." Where only one item is intended, the term "one" or similar language is used. Additionally, as used herein, the terms "has," "have," "having," etc. are intended to be open-ended terms. Further, unless expressly stated otherwise, the phrase "based on" is intended to mean "at least partially based on."

[0081] The description of the various aspects and embodiments has been presented for purposes of illustration, but is not intended to be exhaustive or limited to the disclosed embodiments. Even where combinations of features are recited in the claims and / or disclosed in the specification, such combinations are not intended to limit the disclosure of possible implementations. Indeed, many of these features may be combined in ways not specifically recited in the claims and / or not disclosed in the specification. Although each of the listed dependent claims may directly refer to only one claim, the disclosure of possible implementations includes combinations of each dependent claim with every other claim in the claim set. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope of the described embodiments. The terms used herein were chosen to best explain the principles of the embodiments, practical application, or technical improvement of technologies found in the marketplace, or to enable other ordinary skilled artisans in the art to understand the embodiments disclosed herein.

Claims

1. A method for estimating three-dimensional (3D) gestures in an image, characterized in that, The method includes: Receiving data corresponding to two hand images; Calculating an optical flow value corresponding to a hand pose change in the received data; Generating a heatmap based on the calculated optical flow value; Estimating a hand mesh map based on the generated heatmap; and Determining a gesture present in the two hand images based on the estimated hand mesh map; Wherein, calculating the optical flow value includes: Inferring a first optical flow between the two hand images in the forward direction; Inferring a second optical flow between the two hand images in the reverse direction; and Determining the optical flow value based on the inferred first optical flow and second optical flow.

2. The method according to claim 1, wherein Generating the heatmap includes: Calculating one or more two-dimensional hand key points; and Generating one or more features based on the calculated two-dimensional hand key points.

3. The method according to claim 2, wherein Calculating the two-dimensional hand key points includes: Inferring one or more features from the hand image data; Generating a confidence map corresponding to each inferred feature; and Calculating a Gaussian blur of the Dirac-δ distribution corresponding to the generated confidence map.

4. The method according to claim 1, wherein Estimating the hand mesh map includes: Inferring a three-dimensional hand mesh based on one or more features in the hand image data; and Comparing the contour corresponding to the inferred three-dimensional hand mesh with the contour corresponding to the predicted hand mesh associated with the training image.

5. The method according to claim 1, characterized in that, Determining the gesture includes: Applying one or more graph convolutional networks to the estimated hand mesh map; Extracting one or more gesture features based on applying a pooling layer to the one or more graph convolutional networks; and Generating the gesture in response to applying one or more fully connected layers to the extracted features.

6. The method according to claim 1, characterized in that, Determining the gesture using only two-dimensional data without using three-dimensional annotations.

7. The method according to claim 1, characterized in that, Further includes generating a heatmap loss value and a mesh loss value.

8. The method according to claim 7, characterized in that, Further includes training gesture estimation by minimizing the heatmap loss value and the mesh loss value.

9. The method according to claim 1, characterized in that, The hand image data includes two consecutive frames of a video.

10. A computer system for estimating three-dimensional (3D) gestures in an image, characterized in that, The computer system includes: One or more computer-readable non-transitory storage media configured to store computer program code; and One or more computer processors configured to access the computer program code and perform the method according to any one of claims 1-9 as indicated by the computer program code.

11. A non-transitory computer-readable medium storing a computer program for estimating three-dimensional (3D) gestures in an image, characterized in that, The computer program is configured to cause one or more computer processors to execute the method according to any one of claims 1-9.

12. An apparatus for estimating a three-dimensional (3D) gesture in an image, characterized in that, The apparatus includes: A receiving module for receiving data corresponding to two hand images; A calculating module for calculating an optical flow value corresponding to a hand pose change in the received data; A generating module for generating a heatmap based on the calculated optical flow value; An estimating module for estimating a hand mesh map based on the generated heatmap; and A determining module for determining a gesture present in the two hand images based on the estimated hand mesh map; Wherein, calculating the optical flow value includes: Inferring a first optical flow between the two hand images in the forward direction; Infer a second optical flow between the two hand images in the reverse direction; and Determine the optical flow value based on the inferred first optical flow and second optical flow.

Citation Information

Patent Citations

  • System and method for real-time large image homography processing

    US20190147621A1

  • Activity recognition method using videotubes

    US20190213406A1