Action recognition method and system using bidirectional space-time transformer
A bidirectional space-time transformer neural network enhances action recognition by predicting skeletal joint coordinates through self-supervised learning, addressing accuracy issues in noisy environments and improving semantic feature tagging for diverse applications.
Patent Information
- Application Number
- JP2023561248
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-05-09
- Filing Date
- 2022-04-01
- Publication Date
- 2025-12-16
- Estimated Expiration
- 2042-04-01
AI Technical Summary
Existing technologies face challenges in accurately recognizing human actions and body poses, particularly in the presence of background noise, and improving semantic feature tagging for robust detection.
A bidirectional space-time transformer neural network is employed to process skeletal joint coordinates, utilizing self-supervised adversarial learning to enhance the prediction of joint coordinates, spatial and temporal ordering, and semantic classification, thereby improving action recognition accuracy.
The solution provides robust and accurate detection of body poses and actions, even in noisy environments, enabling applications in home appliances, vehicle control, virtual reality, sign language translation, and neuromuscular health assessment.
Smart Images

Figure 0007786856000020 
Figure 0007786856000021 
Figure 0007786856000022
Abstract
Description
[Technical Field]
[0001] The present invention relates to electrical, electronic, and computer technologies, and more particularly to computer vision. Summary of the Invention
[0002] The present principles provide techniques for skeletal-based action recognition using a bidirectional space-time transformer. In one aspect, an exemplary method includes instantiating a bidirectional space-time transformer neural network, obtaining a plurality of frames including coordinates of skeletal joints and coordinates of other joints, generating spatially masked frames from the eigenframes by masking on the original coordinates of the skeletal joints, providing at least one or more of the eigenframes, the spatially masked frames, and the plurality of frames to a coordinate prediction head of the bidirectional space-time transformer network, obtaining predictions of coordinates for the skeletal joints in the spatially masked frames from the coordinate prediction head, and training the bidirectional space-time transformer neural network to predict the original coordinates of the skeletal joints in the eigenframes through relative relationships between the skeletal joints and other joints and with states of the skeletal joints in other frames by adjusting parameters of the bidirectional space-time transformer neural network until a mean square error between the predictions of the coordinates for the skeletal joints and the original coordinates of the skeletal joints converges.
[0003] One or more embodiments of the present invention, or elements thereof, can be implemented in the form of a computer program product including a computer-readable storage medium having computer-usable program code for facilitating the indicated method steps. Furthermore, one or more embodiments of the present invention, or elements thereof, can be implemented in the form of a system (or apparatus) including a memory embodying computer-executable instructions and at least one processor, coupled to the memory and operable by the instructions, to facilitate the exemplary method steps. Furthermore, in another aspect, one or more embodiments of the present invention, or elements thereof, can be implemented in the form of a means for performing one or more of the method steps described herein, which means can include (i) hardware modules, (ii) software modules stored on a tangible computer-readable storage medium(s) and executed on a hardware processor, or (iii) a combination of (i) and (ii), any of (i)-(iii) performing the specific techniques described herein.
[0004] As used herein, "facilitating" an action includes performing an action, making an action easy, assisting in performing an action, or causing an action to be performed. Thus, by way of example and not limitation, instructions executing on one processor may facilitate an action to be performed by instructions executing on a remote processor by sending appropriate data or commands to cause the action to be performed or assist in the action being performed. For the avoidance of doubt, where an actor facilitates an action by other than performing the action, the action is still performed by some entity or combination of entities.
[0005] In view of the foregoing, the techniques of the present invention can provide many beneficial technical effects. For example, one or more embodiments may provide one or more of the following:
[0006] To improve the technical process of machine body pose detection by providing robust detection of body pose even in the presence of background image noise.
[0007] Accurate tagging of body poses with advanced semantic features (e.g., emotions, actions).
[0008] Some embodiments may not have these potential advantages, and these potential advantages are not necessarily required for all embodiments. These and other features and advantages of the present invention will become apparent from the following detailed description of illustrative embodiments thereof, which is to be read in conjunction with the accompanying drawings. [Brief explanation of the drawings]
[0009] [Figure 1] 1 is a diagram of a cloud computing environment in accordance with an embodiment of the present invention. [Figure 2] FIG. 2 is a diagram of abstraction model layers according to an embodiment of the present invention. [Figure 3] 1 is a diagram of the overall architecture of a skeleton-based action recognition system incorporating a bidirectional space-time transformer in accordance with an exemplary embodiment; [Figure 4] FIG. 4 is a diagram of details of an embodiment of the architecture of FIG. 3. [Figure 5] FIG. 4 is a diagram of details of another embodiment of the architecture of FIG. 3. [Figure 6] FIG. 1 is a diagram of a vector converter architecture in accordance with an example embodiment. [Figure 7] FIG. 1 is a diagram of a spatial transformer architecture in accordance with an exemplary embodiment. [Figure 8] FIG. 2 is a diagram of the architecture of a time converter according to an exemplary embodiment; [Figure 9] FIG. 10 is a diagram of forward and backward temporal masking of skeletal frames for bidirectional self-attention. [Figure 10] Figure 7 shows the data flow through the architecture. [Figure 11] FIG. 1 is a diagram of an architecture of a skeleton-based action recognition system incorporating multiple bidirectional space-time transformers in accordance with an exemplary embodiment. [Figure 12] FIG. 1 is a diagram of a home appliance using a skeleton-based action recognition system to interpret user behavior. [Figure 13] FIG. 1 is a diagram of a virtual reality system using a skeleton-based action recognition system to interpret a user's actions. [Figure 14] FIG. 1 is a diagram of a real-time translation system using a skeleton-based action recognition system to interpret user behavior. [Figure 15] FIG. 1 is a diagram of a safety monitoring system using a skeleton-based action recognition system to interpret user behavior. [Figure 16] FIG. 1 illustrates the use of a neuromuscular health assessment system using a skeleton-based action recognition system to interpret user behavior. [Figure 17] 1 is a diagram of method steps for interpreting results of a skeleton-based action recognition system. [Figure 18] FIG. 1 is a diagram of method steps for training and running a skeleton-based action recognition system. [Figure 19] FIG. 10 illustrates further steps in training a skeleton-based action recognition system. [Figure 20] FIG. 10 illustrates further steps in training a skeleton-based action recognition system. [Figure 21] FIG. 10 illustrates further steps in training a skeleton-based action recognition system. [Figure 22] FIG. 1 is a diagram of a computer system that may be useful in implementing one or more aspects and / or elements of the present invention, also representing a cloud computing node according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0010] Although this disclosure includes detailed descriptions of cloud computing, it should be understood that implementations of the teachings recited herein are not limited to cloud computing environments. Rather, embodiments of the present invention are capable of being executed in conjunction with any other type of computing environment now known or later developed.
[0011] Cloud computing is a service delivery model for enabling convenient, on-demand network access to a shared pool of configurable computing resources (e.g., networks, network bandwidth, servers, processing, memory, storage, applications, virtual machines, and services) that can be rapidly provisioned and released with minimal management effort or interaction by the service provider. This cloud model may include at least five characteristics, at least three service models, and at least four deployment models.
[0012] The characteristics are as follows:
[0013] On-Demand Self-Service: Cloud customers can unilaterally provide computing capacity, such as server time and network storage, automatically as needed without the need for human interaction with the service provider.
[0014] Broad Network Access: Capabilities are available over the network and accessed through standard mechanisms that facilitate use by heterogeneous thin or thick client platforms (e.g., mobile phones, laptops, and PDAs).
[0015] Resource Pooling: A provider's computing resources are pooled to serve multiple consumers using a multi-tenant model, where various physical and virtual resources are dynamically allocated and reallocated according to demand. Location independence is meant in that consumers generally have no control or knowledge of the exact location of the resources provided, but may specify location at a higher level of abstraction (e.g., country, state, or data center).
[0016] Rapid Elasticity: Capacity can be rapidly and elastically provisioned, sometimes automatically, to quickly scale out, and rapidly released to quickly scale in. To the consumer, the capacity available for provisioning often appears unlimited, and can be purchased at any time and in any quantity.
[0017] Measured Services: Cloud systems automatically control and optimize resource usage by leveraging metering capabilities at several levels of abstraction appropriate to the type of service (e.g., storage, processing, bandwidth, and active user accounts). Resource utilization can be monitored, controlled, and reported, providing transparency to both providers and consumers of the services used.
[0018] The service model is as follows:
[0019] Software as a Service (SaaS): The consumer is provided with the ability to use the provider's applications running on a cloud infrastructure. The applications are accessible from a variety of client devices through a thin-client interface, such as a web browser (e.g., web-based email). The consumer does not manage or control the underlying cloud infrastructure, including the network, servers, operating systems, storage, or even individual application capabilities, with the possible exception of limited user-specific application configuration settings.
[0020] Platform as a Service (PaaS): Provides a consumer with the ability to deploy consumer-created or acquired applications, created using programming languages and tools supported by the provider, onto a cloud infrastructure. The consumer does not manage or control the underlying cloud infrastructure, including the network, servers, operating systems, or storage, but rather exercises control over the deployed applications and, in some cases, the application hosting environment configuration.
[0021] Infrastructure as a Service (IaaS): The capability provided to a customer is the provision of processing, storage, networking, and other basic computing resources on which the customer can deploy and run any software, which may include operating systems and applications. The customer does not manage or control the underlying cloud infrastructure, but rather exercises control over the operating systems, storage, deployed applications, and in some cases, limited control over selected networking components (e.g., host firewalls).
[0022] The deployment model is as follows:
[0023] Private Cloud: The cloud infrastructure is operated solely for the organization. The cloud infrastructure may be managed by the organization or a third party and may be on-site or off-site.
[0024] Community Cloud: Cloud infrastructure is shared by several organizations and supports a unique community with shared concerns (e.g., mission, security requirements, policies, and compliance considerations). The cloud infrastructure may be managed by the organization or a third party and may be on-site or off-site.
[0025] Public Cloud: Cloud infrastructure is made available to the general public or large industry groups and is owned by organizations that sell cloud services.
[0026] Hybrid Cloud: A cloud infrastructure is a composite of two or more clouds (private, community, or public) that remain unique entities but are tied together by standard or proprietary technologies that enable data and application portability (e.g., cloud bursting for load balancing between clouds).
[0027] A cloud computing environment is service-oriented, with a focus on statelessness, loose coupling, modularity, and semantic interoperability. At the heart of cloud computing is an infrastructure that includes a network of interconnected nodes.
[0028] Referring now to FIG. 1, an illustrative cloud computing environment 50 is depicted. As shown, the cloud computing environment 50 includes one or more cloud computing nodes 10 with which local computing devices used by cloud users, such as, for example, a personal digital assistant (PDA) or cellular phone 54A, a desktop computer 54B, a laptop computer 54C, or an automobile computer system 54N, or combinations thereof, may communicate. The nodes 10 may also communicate with each other. The nodes 10 may be physically or virtually grouped in one or more networks, such as private, community, public, or hybrid clouds, or combinations thereof, as described above (not shown). This enables the cloud computing environment 50 to provide infrastructure, platform, and / or software as a service without the need for cloud users to maintain resources on their local computing devices. The types of computing devices 54A-N shown in FIG. 1 are intended to be illustrative only, and it is understood that computing node 10 and cloud computing environment 50 can communicate with any type of computerized device via any type of network and / or network-addressable connection (e.g., using a web browser).
[0029] Referring now to Figure 2, a set of functional abstraction layers provided by cloud computing environment 50 (Figure 1) is shown. It should be understood in advance that the components, layers, and functions shown in Figure 2 are intended to be illustrative only, and embodiments of the present invention are not limited thereto. As depicted, the following layers and corresponding functions are provided:
[0030] Hardware and software layer 60 includes hardware and software components. Examples of hardware components include mainframe 61, RISC (reduced instruction set computer) architecture-based servers 62, servers 63, blade servers 64, storage devices 65, and network and networking components 66. In some embodiments, software components include network application server software 67 and database software 68.
[0031] The virtualization layer 70 provides an abstraction layer in which examples of virtual entities are provided, such as virtual servers 71, virtual storage 72, virtual networks including virtual private networks 73, virtual applications and operating systems 74, and virtual clients 75.
[0032] In one example, management layer 80 may provide the functions described below. Resource provisioning 81 dynamically procures computing and other resources utilized to perform tasks within the cloud computing environment. Metering and pricing 82 tracks costs as resources are utilized within the cloud computing environment and bills or invoices for the usage of these resources. In one example, these resources may include application software licenses. Security validates cloud subscribers and tasks and protects data and other resources. User portal 83 provides subscribers and system administrators with access to the cloud computing environment. Service level management 84 allocates and manages cloud computing resources to meet required service levels. Service level agreement (SLA) planning and fulfillment 85 pre-provisions and procures cloud computing resources to anticipate future requirements according to SLAs.
[0033] The workload layer 90 provides examples of functions for which cloud computing environments are utilized. Examples of workloads and functions provided from this layer include mapping and navigation 91, software development and lifecycle management 92, virtual classroom instruction delivery 93, data analytics processing 94, transaction processing 95, and a skeleton-based action recognition system incorporating a bidirectional space-time transformer 96 according to an example embodiment.
[0034] Referring to FIG. 3, the action recognition system 96 incorporates a bi-directional spatial-temporal transformer (BDSTT) 100. The BDSTT 100 includes two modules (shown in FIGS. 7 and 8): a spatial transformer block (STB) 112 (shown in FIG. 7) and a temporal transformer block (TTB) 114 (shown in FIG. 8). During normal operation (i.e., inference), the system 96 receives a skeletal action frame 106 at an embedding module 102. The skeletal action frame 106 represents the positions of the joints of a human skeleton moving through space and time. The embedding module 102 encodes the frame 106 as a joint matrix X (described further below with reference to FIG. 6). During normal operation, the embedding module 102 passes the joint matrix 106 to the BDSTT 100, which produces features that are delivered to the action classifier 104, which performs inference using averaging layers, multilayer perceptrons, and softmax optimizers to generate predictions. During training, the data generation module 108 produces the joint matrix X that is passed through the BDSTT 100 to a self-supervised adversarial learning task head (SSL head) 110. The SSL head 110 adjusts the weights and / or parameters of the BDSTT 100 to improve the accuracy of the features produced by the BDSTT 100.
[0035] Referring to FIG. 4, in an embodiment of system 96, embedding unit 102 receives frame 106 and produces spatial matrix 111 and temporal matrix 113. Embedding unit 102 implements an algorithm for mapping spatial features of an image within frame 106, familiar to those skilled in the art of computer vision programming. BDSTT 100 includes STB 112 and TTB 114, arranged sequentially such that STB processes spatial matrix 111 and TTB processes temporal matrix 113. SSL head 110 includes coordinate prediction (reconstruction) head 116, joint classification (semantic) head 118, spatial ordering head 120, and temporal ordering head 122. Data generator 108 receives original frame 106 and includes spatial reordering unit 124, temporal reordering unit 126, semantic token mask 128, and coordinate token mask 130. The spatial reordering unit 124 implements an algorithm for spatially segmenting the frames 106 and randomly or pseudo-randomly rearranging them (shuffling them in space). The temporal reordering unit 126 implements an algorithm for shuffling the frames 106 in time. The semantic token mask 128 returns zeros for some elements of the one-hot semantic vector that identifies joints by type. The coordinate token mask 130 returns zeros for some coordinates in the spatial matrix 111.
[0036] Figure 5 depicts another embodiment in which the STB 112 and TTB 114 are arranged in parallel so that they process the spatial matrix 111 and the temporal matrix 113 simultaneously and then combine the results. The remaining elements of Figure 5 are similar to those of Figure 4, as described elsewhere herein.
[0037] In general, Figure 6 depicts a general transducer block 400, which utilizes a self-attention mechanism 402 to capture the relationships between all input elements 404.
number
number
number
number
[0038] In one or more embodiments, X is a square matrix, i.e., N = C. For some embodiments, it may be intended that C > N to provide hidden channels or hidden dimensions that allow for additional intermediate features.
[0039] Based on the attention map A, a hidden representation H of all elements is generated as follows: H=LayerNorm(ψ(AX)+X) where ψ is a linear transformation and LayerNorm is a layer normalization function. Thus, in one or more embodiments, shortcut connections are applied to improve the stability of the model.
[0040] In Figure 6, the self-attention mechanism 402 uses a multi-head strategy, applying several respective self-attention operations and combining all of the outputs as a single aggregated multi-head output. Given an input X 404, we update X in the transformer block 400 according to: X=LayerNorm(H+FF(H)) where FF is an arbitrary row-wise feed-forward (fully connected) layer of the neural network (i.e., FF processes the features of each element separately and identically). In other words, X can be updated as follows: X=F(X,A) The entire converter block can be expressed as: TB(X)=F(X,Att(X))
[0041] Neural networks generally include multiple computer processors configured to work together to execute one or more machine learning algorithms. Execution can be synchronous or asynchronous. In a neural network, the processors simulate thousands or millions of neurons connected by axons and synapses. Each connection can be permissive, inhibitory, or neutral in its influence on the activation state of the connected neuronal units. Each individual neuronal unit has a summation function that combines the values of all of the neuronal unit's inputs together. In some implementations, there are threshold or limiting functions for at least some connections and / or at least some neuronal units, so that signals must exceed a limit before being propagated to other neurons. Neural networks can perform supervised, unsupervised, or semi-supervised machine learning. A common approach to training a neural network is to calculate the "loss," or difference, between the "true" value and the value predicted by the network, then iteratively adjust the weights of the neurons in the network and recalculate the loss until the loss converges, i.e., the loss remains within the threshold for multiple iterations. For example, loss may be said to converge if it changes by no more than 5% over three iterations, or if it changes by no more than 1% over five iterations, or if it changes by no more than 2% over two iterations. Selecting the convergence condition is a matter of design choice within the skill of those in the art given the teachings herein.
[0042] Figure 7 depicts the spatial transformer block (STB) 112, a dedicated transformer block that models the skeletal pose for each frame by calculating the relationships between joints. STB 112 uses the multi-head attention setup shown in Figure 6. In STB 112, matrix X is a joint-by-channel matrix, and there is a sequence of T matrices, and each matrix X in the sequence tcorresponds to frame t, where 1 ≤ t ≤ T, where T is the number of frames. It is useful to have the model distinguish between types of joints, i.e., "neck" or "elbow". For this, a one-hot classification vector is added to the joint matrix, e.g.,
number
number
[0043] According to the attention map equation, the attention map A for frame t is t teeth,
number
[0044] The STB 112 therefore incorporates a multi-head self-attention unit 132 that implements the self-attention mechanism 402 described with reference to Figure 6. The STB 112 also incorporates two ADD & Normalization blocks 134, 138, and a feed-forward (or multi-layer perceptron) layer 136.
[0045] Figure 8 depicts the Temporal Transformer Block (TTB) 114, a dedicated transformer block capable of capturing posture changes over time and modeling entire sequences of posture changes. The difference between the TTB 114 and the STB 112 is that the TTB incorporates a bidirectional multi-head self-attention unit 142, which implements various self-attention mechanisms 402 (described further below) previously described with reference to Figure 6. The TTB 114 also incorporates two ADD & Normalization blocks 144, 148 and a feed-forward (or multi-layer perceptron) layer 146.
[0046] Figure 9 shows the operation of unit 142. The sequence of the sth joint in every frame is
number
number
[0047] Then, the position encoding P T is each input frame X S will be added.
number
[0048] Nevertheless, the cyclic symmetry of trigonometric functions can disrupt the ordering of frames. Therefore, the TTB 114 is trained to unravel the sequence of frames by applying a directional mask strategy to the training data to force the model to recognize the time order. A single self-attention operation, such as that in the transformer block 400, is replaced by one time-forward self-attention operation and one time-backward self-attention operation, as shown in FIG. 9. For forward self-attention, a triangular matrix of 1s is used so that, for each joint, each frame is correlated only to its previous frame in time.
number
number
number
[0049] An aspect of the present invention is self-supervised training of the BDSTT 100. Self-supervised learning aims to learn feature representations from large amounts of unlabeled data. For example, self-supervised learning helps the BDSTT 100 learn high-level semantic information in actions represented by skeletal sequences. Four self-supervised adversarial learning tasks address various exceptional situations in skeleton-based action recognition. Referring again to FIGS. 4 and 5, these self-supervised adversarial learning tasks include a coordinate prediction task performed by the coordinate prediction head 116, a joint-type (semantic) prediction task performed by the semantic head 118, a spatial order prediction task performed by the spatial ordering head 120, and a temporal order prediction task performed by the temporal ordering head 122. In equations expressing these tasks, the BDSTT 100 is represented as an encoder f(·).
[0050] In one or more embodiments, the skeleton-based model perceives the laws of posture change. Therefore, it is advantageous for a skeleton-based model such as the BDSTT 100 to predict the coordinates of a joint in an intrinsic frame through its relative relationship to other joints and the state of the joint in other frames. To train the BDSTT 100 for such predictions, the data generator 108 masks the original coordinates of some joints in some frames by setting the masked coordinates to zero according to a specific ratio, while keeping the remaining joints unchanged. The appropriate ratio can be heuristically selected by one skilled in the art depending on the details of a given application. For example, setting approximately 15% of the coordinates to zero can enhance the training of the BDSTT 100 and especially the STB 112 (other embodiments can use other values). Randomly masking the data X masked The encoder f(·) 100 reads the input sequence, extracts a representation from the input, and then generates a coordinate prediction head h C The coordinate prediction head h (·) 116 receives the learned representation and generates a sequence for reproducing or predicting the coordinates of all joints in the input sequence. C(·) is the neural network. The mean squared error (MSE) loss function L PC Using h C The network parameters for (·) are estimated as follows:
number
[0051] Thus, the BDSTT 100 learns like an autoencoder, but the input is a token sequence with data set to 1 or 0 for the eigenvalues. The reconstruction head 116 is a neural network trained to predict the original coordinates of masked joints from the coordinates of unmasked joints. In other words, the reconstruction head 116 is trained on an artificially incomplete set of perfectly known coordinates so that it can predict missing coordinates in another set of less perfectly known coordinates.
[0052] The semantic tokens help the model BDSTT100 distinguish between joint types and appropriate motion envelopes. Fortunately, a good skeleton-based model has the ability to infer the type of joint from its motion history and its position relative to other joints. To train the BDSTT100 on this inferential ability, the data generator 108 removes a few percent of the semantic tokens for joints in the encoder and uses the following cross-entropy classification loss function L: PJ Joint type (semantic) prediction head h trained by J (·) Pass the masked matrix X to 118.
number
[0053] To reduce the impact of incorrect spatial ordering of joints, BDSTT100 uses the matrix X ShfS It can be trained to predict the correct spatial ordering of spatially shuffled joints.
number
number
[0054] This is useful if the skeleton-based model can ascertain the correct temporal order for a potentially time-shuffled sequence of frames. To enhance the ability of the BDSTT 100 to do this, the data generator 108 may generate a time-reordered permutation for the shuffled sequence X ShfT Within each segment
number
number
[0055] Figure 10 depicts the data flow through an architecture such as that shown in Figure 7. From frames 106, embedding unit 102 (shown in Figures 3, 4, and 5, but not shown in Figure 10) produces frame-wise (spatial) matrices 111 and joint-wise (temporal) matrices 113. STB 112 receives spatial matrices 111 and produces matrices representing skeletal pose features within each frame. TTB 114 receives temporal matrices 113 and produces matrices representing the evolution of skeletal pose features across a sequence of frames.
[0056] In one or more embodiments, as shown in FIG. 11 , the action recognition system 96 includes eight (8) stacked instances of the BDSTT 100, with the instances having 64, 64, 128, 128, 256, 256, 256, and 256 output channels, respectively. Each BDSTT 100 includes a spatial transformer block 112 and a temporal transformer block 114, along with three (3) heads of a multi-head attention network. In one or more embodiments, the system 96 also includes a preprocessor that randomly and uniformly samples all skeletal sequences to 150 frames and then randomly and centrally crops the sampled sequences to 128 frames for the training / test split. In one or more embodiments, the system 96 is trained for 120 epochs with a batch size of 32 using stochastic gradient descent with a Nesterov momentum of 0.9. In one or more embodiments, the initial learning rate (LR) is set to 0.1 at epochs 60 and 90 with a step LR decay (factor 0.1).
[0057] Embodiments of the present invention have numerous practical applications. One non-limiting exemplary application for embodiments of the present invention is in the area of home appliance control. Another non-limiting exemplary application is in the area of vehicle control. In both areas, it may be useful for a user to be able to adjust the operation of a device (e.g., a refrigerator, a car, a wheelchair) without touching the device's controls. For example, a camera associated with the device can capture images (frames) of the user's movements, and system 96 can process the frames to detect skeletal joint movement sequences by applying a trained bidirectional space-time transformer network 100 to the frames. In response to the detected skeletal joint movement sequences, system 96 can transmit control signals to the device. For example, system 96 can be implemented in a computer system / server 12 (or in a microprocessor without peripherals such as a monitor and keyboard) as shown in FIG. 22 , communicatively connected with an external electromechanical, electro-optical, or electronic device 14 via input / output interface 22. The system 96 can transmit control signals to the device 14 via the input / output interface 22 in response to detecting a skeletal joint movement sequence through operation of the trained bidirectional space-time transformer network 100 .
[0058] FIG. 12 depicts a home appliance (robot) 1200 that uses a skeleton-based action recognition system (not shown in FIG. 12) to interpret user actions. Based on the action recognition, the robot can provide unique services. As an example, in family services, the robot can understand people's actions and intentions. In FIG. 12, the robot is facing the person and is ready to collect or clean when it recognizes that the person is drinking water.
[0059] 13 depicts a virtual reality system 1300 that uses a skeletal-based action recognition system to interpret user behavior, allowing people to easily interact with the VR system 1300 using actions / gestures rather than using bulky controllers. In FIG. 13, the VR system 1300 rotates a virtual object 1302 in a user's virtual hand 1304 in response to detecting a skeletal articulation sequence of the user's actual hand 1306.
[0060] Figure 14 depicts a real-time translation system using a skeletal-based action recognition system to interpret user actions. Skeletal-based action recognition can function in a sign language translation system to help deaf people / people who cannot communicate by speaking with hearing people. In Figure 14, a handheld electronic device 1402 uses its camera 1404 to capture gestures of a user's hand 1406 and applies skeletal-based action recognition to look up the gestures in a gesture database 1408 for translation into utterances.
[0061] Another practical application is for remote monitoring and care of vulnerable groups. For vulnerable groups such as the elderly and children, abnormal behavior detection using a skeleton-based action recognition scheme is highly beneficial. FIG. 15 depicts an example in which a skeletal joint movement sequence 1502 is detected and compared with a "normal" sequence 1504, and the system 96 (not shown in FIG. 15) can alert parents / caregivers / hospital staff when abnormal movements are detected. FIG. 16 depicts a medical robot 1602 implementing a neuromuscular health assessment system that uses a skeleton-based action recognition system to interpret user behavior.
[0062] FIG. 17 depicts the steps of a basic method 1700 for implementing any of the practical applications of FIGS. 12 through 16. At 1702, an image of a person in motion is captured. At 1704, the image is converted into a skeletal articulation sequence. At 1706, a skeletal-based action recognition system is applied to the skeletal articulation sequence. At 1708, the actions recognized by the skeletal-based action recognition system are compared to a database of potentially related actions. At 1710, the device is actuated depending on whether the recognized action matches or does not match any action from the database. For example, in FIG. 12, when a person places a glass of water (matching a database action), the robot 1200 is actuated to get a glass of water. In FIG. 15, instead of opening a door (not matching any database action), an alert buzzer (not shown) is actuated to fetch staff when a person falls.
[0063] In view of the foregoing discussion, and with particular reference to the accompanying FIGS. 18-21, it will be appreciated that an exemplary method 1800 according to one embodiment of the present invention generally includes instantiating, at 1802, a bidirectional space-time transformer (BDSTT) neural network. At 1804, the bidirectional space-time transformer neural network is trained to predict the original coordinates of skeletal joints in the eigenframe through the relative relationships of the skeletal joints to other joints and to the state of the skeletal joints in other frames. Those skilled in the art will appreciate that prediction of known original coordinates is performed for training purposes so that the predictions can be compared to ground truth and that parameters of the BDSTT can be adjusted until the predictions approach the original coordinates within satisfactory limits. Training includes several steps: At 1806, a plurality of frames including the coordinates of the skeletal joints and the coordinates of other joints are obtained; and At 1808, a spatially masked frame is generated from the eigenframe by masking the original coordinates of the skeletal joints. At 1810, at least one or more of the intrinsic frame, the spatially masked frame, and the plurality of frames are provided to a coordinate prediction head of a bidirectional space-time transformer network. At 1812, predictions of coordinates for skeletal joints in the spatially masked frame are obtained from the coordinate prediction head. At 1814, parameters of the bidirectional space-time transformer neural network are adjusted until a mean square error between the predictions of coordinates for skeletal joints and the original coordinates of the skeletal joints converges. In one or more embodiments, method 1800 may also include training a BDSTT to predict the correct temporal order at 1900, training a BDSTT to predict the correct spatial arrangement at 2000, or training a BDSTT to predict the correct semantic coding at 2100, or a combination thereof.
[0064] 19 in detail, training the BDSTT to predict the correct time order includes several steps in the illustrated example. At 1902, time-shuffling a plurality of frames to produce a plurality of time-shuffled frames. At 1904, providing the plurality of time-shuffled frames along with the plurality of frames to a time classification head. At 1906, obtaining a prediction of the correct time order for the plurality of time-shuffled frames from the time classification head. At 1908, adjusting parameters of a bidirectional space-time transformer neural network until the cross-entropy loss between the prediction of the correct time order and the plurality of frames converges within a threshold.
[0065] Referring specifically to FIG. 20 , training the BDSTT to predict correct spatial arrangements includes several steps in the illustrated example. At 2002, a plurality of spatially shuffled frames are generated by spatially rearranging a plurality of joints in one or more of the frames. At 2004, the plurality of spatially shuffled frames are provided to a spatial classification head along with the plurality of frames. At 2006, predictions of correct spatial arrangements for the plurality of joints in the plurality of spatially shuffled frames are obtained from a temporal classification head. At 2008, parameters of a bidirectional space-time transformer neural network are adjusted until the cross-entropy loss between the prediction of correct spatial arrangements and the plurality of frames converges within a threshold.
[0066] Referring now to FIG. 21 , training the BDSTT to predict correct semantic coding includes several steps in the illustrated example. At 2102, a semantically masked frame is generated from the eigenframe by masking at least a portion of a matrix of one-hot vectors corresponding to multiple joints in the eigenframe. At 2104, the semantically masked frame and the eigenframe are provided to a semantic prediction head of a bidirectional space-time transformer network. At 2106, a predicted matrix of one-hot vectors for the semantically masked frame is obtained from the temporal classification head. At 2108, parameters of the bidirectional space-time transformer neural network are adjusted until a cross-entropy classification loss between the predicted matrix of one-hot vectors and the matrix of one-hot vectors corresponding to multiple joints converges within a threshold.
[0067] 18, it will be appreciated that in one or more embodiments, aspects of the present invention may include controlling a device in response to skeletal movement. For example, at 1816, a skeletal articulation movement sequence is detected by applying a trained bidirectional space-time transformer neural network to the sequence of frames. Then, at 1818, control signals are transmitted to at least one of an electromechanical device, an electro-optical device, and an electronic device in response to the detected skeletal articulation movement sequence (e.g., via I / O interface 22 and / or network adapter 20, discussed below).
[0068] One or more embodiments of the present invention, or elements thereof, may be implemented in the form of an apparatus including a memory and at least one processor coupled to the memory and operable to perform exemplary method steps, or in the form of a non-transitory computer-readable medium embodying computer-executable instructions that, when executed by a computer, cause the computer to perform exemplary method steps. FIG. 22 depicts a computer system that may be useful in implementing one or more aspects and / or elements of the present invention, which also represents a cloud computing node according to embodiments of the present invention. Referring now to FIG. 22, cloud computing node 10 is only one example of a suitable cloud computing node and is not intended to suggest any limitation on the scope of use or functionality of embodiments of the present invention described herein. Nevertheless, cloud computing node 10 is capable of performing and / or implementing any of the functions described above.
[0069] Cloud computing node 10 includes computer system / server 12, which is operable in numerous other general-purpose or special-purpose computing system environments or configurations. Examples of well-known computing systems, environments, or configurations, or combinations thereof, that may be suitable for use with computer system / server 12 include, but are not limited to, personal computer systems, server computer systems, thin clients, thick clients, handheld or laptop devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputer systems, mainframe computer systems, and distributed cloud computing environments that include any of the above systems or devices, and the like.
[0070] Computer system / server 12 may be described in the general context of computer system-executable instructions, such as program modules, executed by a computer system. Generally, program modules may include routines, programs, objects, components, logic, data structures, etc. that perform particular tasks or implement particular abstract data types. Computer system / server 12 may also be practiced in distributed cloud computing environments where tasks are performed by remote processing devices linked through a communications network. In a distributed cloud computing environment, program modules may be located in both local and remote computer system storage media, including memory storage devices.
[0071] 22, computer system / server 12 in cloud computing node 10 is shown in the form of a general-purpose computing device. Components of computer system / server 12 may include, but are not limited to, one or more processors or processing units 16, system memory 28, and a bus 18 coupling various system components, including system memory 28, to processor 16.
[0072] Bus 18 represents one or more of any of several types of bus structures, including a memory bus or memory controller, a peripheral bus, an accelerated graphics port, and a processor or local bus using any of a variety of bus architectures. By way of example and without limitation, such architectures include an Industry Standard Architecture (ISA) bus, a Micro Channel Architecture (MCA) bus, an Enhanced ISA (EISA) bus, a Video Electronics Standards Association (VESA) local bus, and a Peripheral Component Interconnect (PCI) bus.
[0073] Computer system / server 12 typically includes a variety of computer system-readable media, which may be any available media that can be accessed by computer system / server 12 and includes both volatile and nonvolatile media, removable and non-removable media.
[0074] System memory 28 may include computer system-readable media in the form of volatile memory, such as random access memory (RAM) 30 and / or cache memory 32. Computer system / server 12 may also include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, storage system 34 may be provided for reading from and writing to non-removable, non-volatile magnetic media (not shown, typically referred to as a "hard drive"). Although not shown, a magnetic disk drive may be provided for reading from and writing to removable, non-volatile magnetic disks (e.g., "floppy disks"), and an optical disk drive may be provided for reading from and writing to removable, non-volatile optical disks, such as CD-ROMs, DVD-ROMs, or other optical media. In such cases, each may be connected to bus 18 by one or more data media interfaces. As further depicted and explained below, memory 28 may include at least one program product having a set (e.g., at least one) of program modules configured to perform the functions of embodiments of the present invention.
[0075] By way of example and not limitation, a program / utility 40 having a set (at least one) of program modules 42 may be stored in memory 28, as well as an operating system, one or more application programs, other program modules, and program data. Each of the operating system, one or more application programs, other program modules, and program data, or some combination thereof, may comprise an implementation of a networking environment. The program modules 42 generally perform the functions and / or methods of embodiments of the present invention as described herein.
[0076] The computer system / server 12 may also communicate with one or more external devices 14, such as a keyboard, pointing device, display 24, one or more devices that allow a user to interact with the computer system / server 12, or any device (e.g., a network card, modem, etc.) that allows the computer system / server 12 to communicate with one or more other computing devices, or a combination thereof. Such communication may occur via an input / output (I / O) interface 22. Additionally, the computer system / server 12 may communicate with one or more networks, such as a local area network (LAN), a general wide area network (WAN), or a public network (e.g., the Internet), or a combination thereof, via a network adapter 20. As depicted, the network adapter 20 communicates with other components of the computer system / server 12 via a bus 18. It should be understood that other hardware and / or software components, not shown, may be used with the computer system / server 12. Examples include, but are not limited to, microcode, device drivers, redundant processing units, and external disk drive arrays, RAID systems, tape drives, and data archive storage systems.
[0077] Accordingly, one or more embodiments may employ software running on a general-purpose computer or workstation. Referring to FIG. 22, such an implementation may employ, for example, a processor 16, memory 28, a display 24, and an input / output interface 22 to external devices 14 (such as a keyboard, pointing device, or the like). The term "processor" as used herein is intended to include any processing device, such as, for example, one that includes a CPU (central processing unit) and / or other forms of processing circuitry. Furthermore, the term "processor" may refer to two or more individual processors. The term "memory" is intended to include memory associated with a processor or CPU, such as, for example, RAM (random access memory) 30, ROM (read-only memory), fixed memory devices (e.g., hard drives 34), removable memory devices (e.g., diskettes), flash memory, and the like. Additionally, the phrase "input / output interface" as used herein is intended to contemplate an interface to, for example, one or more mechanisms for inputting data into a processing unit (e.g., a mouse) and one or more mechanisms for providing results associated with the processing unit (e.g., a printer). The processor 16, memory 28, and input / output interface 22 may be interconnected, for example, via a bus 18 as part of the data processing unit 12. Suitable interconnection, for example, via the bus 18, may also be provided to a network interface 20, such as a network card, which may provide for interfacing with a computer network and to a media interface, such as a diskette or CD-ROM drive, which may provide for interfacing with suitable media.
[0078] Thus, computer software containing instructions or code for implementing the methods of the present invention as described herein may be stored in one or more associated memory devices (e.g., ROM, fixed or removable memory) and, when ready for use, loaded partially or entirely into (e.g., RAM) and executed by a CPU. Such software may include, but is not limited to, firmware, resident software, microcode, and the like.
[0079] A data processing system suitable for storing and / or executing program code will include at least one processor 16 coupled directly or indirectly to memory elements 28 through a system bus 18. The memory elements may include local memory, mass storage, and cache memory 32 employed during the actual implementation of the program code, with cache memory 32 providing temporary storage of at least some of the program code to reduce the number of times the code must be retrieved from mass storage during execution.
[0080] Input / output or I / O devices (including but not limited to keyboards, displays, pointing devices, and the like) can be coupled to the system either directly or through intervening I / O controllers.
[0081] A network adapter 20 may also be coupled to the system to enable the data processing system to become coupled to other data processing systems or remote printers or storage devices through intervening private or public networks. Modems, cable modems, and Ethernet cards are just a few of the currently available types of network adapters.
[0082] As used herein, including in the claims, a "server" includes a physical data processing system (e.g., system 12 as shown in FIG. 22) that runs a server program. It will be understood that such a physical server may or may not include a display and keyboard.
[0083] By way of example and not limitation, one or more embodiments may be implemented at least in part in the context of a cloud or virtual machine environment. Reference is again made to Figures 1-2 and the accompanying text.
[0084] Any of the methods described herein may include the additional step of providing a system comprising separate software modules embodied on a computer-readable storage medium, where the modules may include any or all of the appropriate elements, e.g., depicted in block diagrams and / or described herein, and it should be noted that, by way of example and not limitation, any one, some, or all of the modules / blocks and / or sub-modules / sub-blocks are described. The method steps may be performed using the separate software modules and / or sub-modules of the system, as described above, which then execute on one or more hardware processors, such as 16. Additionally, a computer program product may include a computer-readable storage medium having code adapted to be implemented to perform one or more method steps described herein, including providing the separate software modules to the system.
[0085] Exemplary Systems and Product Details The present invention may be a system, method, or computer program product, or combination thereof, at any possible level of technical detail of integration. The computer program product may include a computer-readable storage medium (or media) having computer-readable program instructions for causing a processor to carry out aspects of the present invention.
[0086] A computer-readable storage medium can be any tangible device capable of retaining and storing instructions for use by an instruction-execution device. A computer-readable storage medium may be, for example, but is not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. A non-exhaustive list of more specific examples of computer-readable storage media includes portable computer diskettes, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disk read-only memory (CD-ROM), digital versatile disk (DVD), memory sticks, floppy disks, mechanically encoded devices such as punch cards or ridge-in-groove structures having instructions recorded thereon, and any suitable combination of the foregoing. Computer-readable storage media as used herein should not be construed as signals that are transitory in nature, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through a waveguide or other transmission medium (e.g., light pulses passing through fiber optic cable), or electrical signals transmitted through wires.
[0087] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to each computing / processing device or to an external computer or external storage device over a network, such as the Internet, a local area network, a wide area network, or a wireless network, or a combination thereof. The network may comprise copper transmission cables, optical fiber transmissions, wireless transmissions, routers, firewalls, switches, gateway computers, or edge servers, or a combination thereof. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage on a computer-readable storage medium within the respective computing / processing device.
[0088] Computer-readable program instructions for carrying out operations of the present invention may be source or object code written in any combination of one or more programming languages, including assembler instructions, instruction set architecture (ISA) instructions, machine language instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, configuration data for integrated circuit devices, or object-oriented programming languages such as Smalltalk®, C++, or the like, and procedural programming languages such as the "C" programming language or similar programming languages. The computer-readable program instructions may execute entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the latter scenario, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection may be to an external computer (e.g., through the Internet using an Internet Service Provider). In some embodiments, electronic circuitry including, for example, a programmable logic circuit (PLC), a field programmable gate array (FPGA), or a programmable logic array (PLA) may execute computer-readable program instructions by utilizing state information of the computer-readable program instructions to individualize the electronic circuitry to implement aspects of the present invention.
[0089] Aspects of the present invention are described herein with reference to flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable program instructions.
[0090] These computer-readable program instructions may be provided to a processor of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus to produce a machine, the instructions of which execute on the processor of the computer or other programmable data processing apparatus to produce means for performing the functions / acts specified in one or more blocks of the flowcharts and / or block diagrams. These computer-readable program instructions may also be stored on a computer-readable storage medium such that the computer-readable storage medium comprises an article of manufacture containing instructions for performing aspects of the functions / acts specified in one or more blocks of the flowcharts and / or block diagrams, and may direct a computer, programmable data processing apparatus, or other device, or combination thereof, to function in a particular manner.
[0091] The computer-readable program instructions may also be loaded into a computer, other programmable data processing apparatus, or other device to cause the computer, other programmable apparatus, or other device to perform a series of operational steps to produce a computer-executed process, the instructions executing on the computer, other programmable apparatus, or other device to perform the functions / acts specified in one or more blocks of the flowcharts and / or block diagrams.
[0092] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowcharts or block diagrams may represent a module, segment, or portion of instructions, including one or more executable instructions for implementing specified logical functions. In some alternative implementations, the functions noted in the blocks may occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently, or the blocks may sometimes be executed in the reverse order, depending on the functionality involved. It is also noted that each block of the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented by a dedicated hardware-based system that performs the specified functions or acts or executes a combination of dedicated hardware and computer instructions.
[0093] The description of various embodiments of the present invention has been presented for purposes of illustration and is not intended to be exhaustive or limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope of the described embodiments. The terminology used herein has been chosen to best explain the principles of the embodiments, practical applications, or technical improvements over technology found in the marketplace, or to enable those skilled in the art to understand the embodiments disclosed herein.
Claims
1. A computer-implemented method for processing information on a computer, comprising: instantiating a bidirectional space-time transformer neural network; acquiring a plurality of frames including coordinates of skeletal joints and coordinates of other joints; generating a spatially masked frame from the eigenframe by masking to the original coordinates of the skeletal joints; providing the intrinsic frame and the spatially masked frame, together with at least one of the plurality of frames, to a coordinate prediction head of the bidirectional space-time transformer network; obtaining predictions of coordinates for the skeletal joints in the spatially masked frames from the coordinate prediction head; and adjusting parameters of the bidirectional space-time transformer neural network until a mean square error between the prediction of coordinates for the skeletal joints and the original coordinates of the skeletal joints converges. training the bidirectional space-time transformer neural network to predict the original coordinate of the skeletal joint in the intrinsic frame through a relative relationship between the skeletal joint and a state of another joint in the same frame, and a relative relationship between the skeletal joint and a state of the skeletal joint in another frame, by 20. A computer-implemented method comprising:
2. time-shuffling the plurality of frames to produce a plurality of time-shuffled frames; providing the plurality of time-shuffled frames together with the plurality of frames to a time sorting head; obtaining a prediction of a correct time order for the plurality of time-shuffled frames from the time classification head; adjusting parameters of the bidirectional space-time transformer neural network until a cross-entropy loss between the prediction and the plurality of frames in the correct time order converges; training the bidirectional space-time transformer neural network to predict the correct time sequence of successive coordinates of the skeletal joints by The method of claim 1 further comprising:
3. detecting a skeletal joint movement sequence by applying the trained bidirectional space-time transformer neural network to a sequence of frames; transmitting a control signal to at least one of an electromechanical device, an electro-optical device, and an electronic device in response to the detected skeletal articulation movement sequence; The method of claim 2 further comprising:
4. generating a plurality of spatially shuffled frames by spatially rearranging a plurality of joints in one or more of the frames; providing the plurality of spatially shuffled frames together with the plurality of frames to a spatial sorting head; obtaining predictions of correct spatial locations for the joints in the spatially shuffled frames from the spatial classification head; adjusting parameters of the bidirectional space-time transformer neural network until a cross-entropy loss between the prediction of correct spatial arrangement and the plurality of frames converges; training the bidirectional space-time transformer neural network to predict the correct spatial arrangement of coordinates of a plurality of skeletal joints by The method of claim 1 further comprising:
5. detecting a skeletal joint movement sequence by applying the trained bidirectional space-time transformer neural network to a sequence of frames; transmitting a control signal to at least one of an electromechanical device, an electro-optical device, and an electronic device in response to the detected skeletal articulation movement sequence; The method of claim 4 further comprising:
6. generating a semantically masked frame from the eigenframe by masking at least a portion of a matrix of one-hot vectors corresponding to a plurality of joints included in the frame; providing the semantically masked frames and the eigenframes to a semantic prediction head of the bidirectional space-time transformer network; obtaining a prediction matrix of one-hot vectors for the semantically masked frame from the semantic prediction head; adjusting parameters of the bidirectional space-time transformer network until a cross-entropy classification loss between the prediction matrix of one-hot vectors and the matrix of one-hot vectors corresponding to the plurality of joints converges; training the bidirectional space-time transformer neural network to predict correct semantic coding of multiple skeletal joints by The method of claim 1 further comprising:
7. detecting a skeletal joint movement sequence by applying the trained bidirectional space-time transformer neural network to a sequence of frames; transmitting a control signal to at least one of an electromechanical device, an electro-optical device, and an electronic device in response to the detected skeletal articulation movement sequence; The method of claim 6 further comprising:
8. A computer program product that, by processing information on a computer, instantiating a bidirectional space-time transformer neural network; acquiring a plurality of frames including coordinates of skeletal joints and coordinates of other joints; generating a spatially masked frame from the eigenframe by masking to the original coordinates of the skeletal joints; providing the intrinsic frame and the spatially masked frame, together with at least one of the plurality of frames, to a coordinate prediction head of the bidirectional space-time transformer network; obtaining predictions of coordinates for the skeletal joints in the spatially masked frames from the coordinate prediction head; and adjusting parameters of the bidirectional space-time transformer neural network until a mean square error between the prediction of coordinates for the skeletal joints and the original coordinates of the skeletal joints converges. training the bidirectional space-time transformer neural network to predict the original coordinate of the skeletal joint in the intrinsic frame through a relative relationship between the skeletal joint and a state of another joint in the same frame, and a relative relationship between the skeletal joint and a state of the skeletal joint in another frame, by 1. A computer program product comprising one or more computer-readable storage media embodying computer-executable instructions that cause said computer to perform a method comprising:
9. The method comprises: time-shuffling the plurality of frames to produce a plurality of time-shuffled frames; providing the plurality of time-shuffled frames together with the plurality of frames to a time sorting head; obtaining a prediction of a correct time order for the plurality of time-shuffled frames from the time classification head; adjusting parameters of the bidirectional space-time transformer neural network until a cross-entropy loss between the prediction and the plurality of frames in the correct time order converges; training the bidirectional space-time transformer neural network to predict the correct time sequence of successive coordinates of the skeletal joints by 9. The computer program product of claim 8, further comprising:
10. The method comprises: detecting a skeletal joint movement sequence by applying the trained bidirectional space-time transformer neural network to a sequence of frames; transmitting a control signal to at least one of an electromechanical device, an electro-optical device, and an electronic device in response to the detected skeletal articulation movement sequence; 10. The computer program product of claim 9, further comprising:
11. The method comprises: generating a plurality of spatially shuffled frames by spatially rearranging a plurality of joints included in the frame; providing the plurality of spatially shuffled frames together with the plurality of frames to a spatial sorting head; obtaining predictions of correct spatial locations for the joints in the spatially shuffled frames from the spatial classification head; adjusting parameters of the bidirectional space-time transformer neural network until a cross-entropy loss between the prediction of correct spatial arrangement and the plurality of frames converges; training the bidirectional space-time transformer neural network to predict the correct spatial arrangement of coordinates of a plurality of skeletal joints by 9. The computer program product of claim 8, further comprising:
12. The method comprises: detecting a skeletal joint movement sequence by applying the trained bidirectional space-time transformer neural network to a sequence of frames; transmitting a control signal to at least one of an electromechanical device, an electro-optical device, and an electronic device in response to the detected skeletal articulation movement sequence; 12. The computer program product of claim 11, further comprising:
13. The method comprises: generating a semantically masked frame from the eigenframe by masking at least a portion of a matrix of one-hot vectors corresponding to a plurality of joints included in the frame; providing the semantically masked frames and the eigenframes to a semantic prediction head of the bidirectional space-time transformer network; obtaining a prediction matrix of one-hot vectors for the semantically masked frame from the semantic prediction head; adjusting parameters of the bidirectional space-time transformer network until a cross-entropy classification loss between the prediction matrix of one-hot vectors and the matrix of one-hot vectors corresponding to the plurality of joints converges; training the bidirectional space-time transformer neural network to predict correct semantic coding of multiple skeletal joints by 9. The computer program product of claim 8, further comprising:
14. The method comprises: detecting a skeletal joint movement sequence by applying the trained bidirectional space-time transformer neural network to a sequence of frames; transmitting a control signal to at least one of an electromechanical device, an electro-optical device, and an electronic device in response to the detected skeletal articulation movement sequence; 14. The computer program product of claim 13, further comprising:
15. 1. An apparatus comprising: a memory embodying computer-executable instructions; at least one processor coupled to the memory; instantiating a bidirectional space-time transformer neural network; acquiring a plurality of frames including coordinates of skeletal joints and coordinates of other joints; generating a spatially masked frame from the eigenframe by masking to the original coordinates of the skeletal joints; providing the intrinsic frame and the spatially masked frame, together with at least one of the plurality of frames, to a coordinate prediction head of the bidirectional space-time transformer network; obtaining predictions of coordinates for the skeletal joints in the spatially masked frames from the coordinate prediction head; and adjusting parameters of the bidirectional space-time transformer neural network until a mean square error between the prediction of coordinates for the skeletal joints and the original coordinates of the skeletal joints converges. training the bidirectional space-time transformer neural network to predict the original coordinate of the skeletal joint in the intrinsic frame through a relative relationship between the skeletal joint and a state of another joint in the same frame and a relative relationship between the skeletal joint and a state of the skeletal joint in another frame by at least one processor operable by said computer-executable instructions to perform a method comprising: An apparatus comprising:
16. The method comprises: time-shuffling the plurality of frames to produce a plurality of time-shuffled frames; providing the plurality of time-shuffled frames together with the plurality of frames to a time sorting head; obtaining a prediction of a correct time order for the plurality of time-shuffled frames from the time classification head; adjusting parameters of the bidirectional space-time transformer neural network until a cross-entropy loss between the prediction and the plurality of frames in the correct time order converges; training the bidirectional space-time transformer neural network to predict the correct time sequence of successive coordinates of the skeletal joints by 16. The apparatus of claim 15, further comprising:
17. The method comprises: detecting a skeletal joint movement sequence by applying the trained bidirectional space-time transformer neural network to a sequence of frames; transmitting a control signal to at least one of an electromechanical device, an electro-optical device, and an electronic device in response to the detected skeletal articulation movement sequence; 17. The apparatus of claim 16, further comprising:
18. The method comprises: generating a plurality of spatially shuffled frames by spatially rearranging a plurality of joints included in the frame; providing the plurality of spatially shuffled frames together with the plurality of frames to a spatial sorting head; obtaining predictions of correct spatial locations for the joints in the spatially shuffled frames from the spatial classification head; adjusting parameters of the bidirectional space-time transformer neural network until a cross-entropy loss between the prediction of correct spatial arrangement and the plurality of frames converges; training the bidirectional space-time transformer neural network to predict the correct spatial arrangement of coordinates of a plurality of skeletal joints by 16. The apparatus of claim 15, further comprising:
19. The method comprises: detecting a skeletal joint movement sequence by applying the trained bidirectional space-time transformer neural network to a sequence of frames; transmitting a control signal to at least one of an electromechanical device, an electro-optical device, and an electronic device in response to the detected skeletal articulation movement sequence; 20. The apparatus of claim 18, further comprising:
20. The method comprises: generating a semantically masked frame from the eigenframe by masking at least a portion of a matrix of one-hot vectors corresponding to a plurality of joints included in the frame; providing the semantically masked frames and the eigenframes to a semantic prediction head of the bidirectional space-time transformer network; obtaining a prediction matrix of one-hot vectors for the semantically masked frame from the semantic prediction head; adjusting parameters of the bidirectional space-time transformer network until a cross-entropy classification loss between the prediction matrix of one-hot vectors and the matrix of one-hot vectors corresponding to the plurality of joints converges; training the bidirectional space-time transformer neural network to predict correct semantic coding of multiple skeletal joints by 16. The apparatus of claim 15, further comprising:
21. The method comprises: detecting a skeletal joint movement sequence by applying the trained bidirectional space-time transformer neural network to a sequence of frames; transmitting a control signal to at least one of an electromechanical device, an electro-optical device, and an electronic device in response to the detected skeletal articulation movement sequence; 21. The apparatus of claim 20, further comprising:
Citation Information
Patent Citations
Information processing device, program, and information processing method
WO2021005745A1