Action recognition using implicit pose expressions
By employing implicit pose representations and a two-neural-network architecture, the system effectively addresses the challenges of human action recognition, achieving robust and efficient action classification even in complex contexts.
Patent Information
- Application Number
- JP2020151934
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-06-17
- Filing Date
- 2020-09-10
- Publication Date
- 2025-06-23
- Estimated Expiration
- 2040-09-10
AI Technical Summary
Existing methods for human action recognition face challenges such as high computational cost, complexity in training, and inability to generate robust predictions due to factors like camera motion, noise, and partial or absent contextual information.
The proposed system and method utilize implicit pose representations to classify human actions in images, bypassing the need for explicit pose estimation and reducing ambiguity in representation. This is achieved by generating feature vectors from a first neural network trained for pose determination and classifying these vectors using a second neural network trained for action recognition.
The approach provides an efficient and robust method for human action recognition, capable of accurately identifying actions even in contexts where explicit pose information is incomplete or ambiguous, while reducing computational complexity.
Smart Images

Figure 0007696703000005 
Figure 0007696703000006 
Figure 0007696703000007
Abstract
Description
Technical Field
[0001] The present disclosure relates to action recognition, and more particularly, to a system and method for classifying one or more images according to the actions of individuals shown in the images.
Background Art
[0002] The description of the background art provided herein is for generally presenting the context of the present disclosure. Within the scope described in this background art section, the work of the currently described inventors and the embodiments of the specification that do not qualify as prior art as of the filing date are not expressly or implicitly admitted as prior art to the present disclosure.
[0003] Neural networks including convolutional neural networks (CNNs) can be used, for example, in computer vision applications. Human action recognition is a difficult task in computer vision because of the large number and variability of possible actions and the importance of various cues (e.g., appearance, motion, pose, objects, etc.) for different actions. Also, actions can span a variety of time scales and can be performed at different speeds.
[0004] Some methods can be constructed based on two-stream architectures, RNNs (Recurrent Neural Networks), or spatio-temporal 3D convolutions. In a two-stream architecture based on a deep spatiotemporal convolutional architecture, a method for combining appearance and motion information processes RGB (Red Green Blue) video frames and optical flow separately. This method can accurately determine human actions for specific datasets such as the Kinetics dataset.
[0005] Some drawbacks of these methods correspond to the high computational cost and complexity for training. Also, the method may not be able to generate robust human action predictions for some recorded attributes of the image such as camera motion and noise. Therefore, what some datasets and convolutional neural networks (CNNs) learn can be biased towards scenes and objects. Information in such situations can be useful for predicting human actions, but it is not sufficient to accurately determine human actions and accurately understand what is happening in the scene or image.
[0006] Some methods for action recognition can utilize contexts such as scenes or objects instead of focusing on human actions. Therefore, such methods are prone to errors when the context (e.g., scene or object) is only partially available, absent, or there is room for misinterpretation. Examples of when the context is partial, absent, or has room for misinterpretation include the case of indoor surfing where the correct object is present but not in its usual location, the case of a mime artist imitating a person playing the violin where there are no objects or scenes, and the case of a soccer player mimicking a bowling scene in a soccer stadium with a soccer ball where the context is open to misinterpretation. 3D CNNs and other types of action recognition systems may not be able to recognize such actions. Figure 4 includes an example of an image with room for misinterpretation.
[0007] Humans can understand actions more perfectly and recognize actions without a context, object, or scene. One particularly challenging example is human action recognition for mimed actions. A mime artist can use only expressions, gestures, and movements to suggest emotions and actions to the audience without words or context.
[0008] For example, the information conveyed by a person's pose can be used to understand the mimed action. Action recognition methods used as input 3D skeleton data can be independent of situational information. However, such methods can be used in a very restricted environment where accurate data is collected through a motion capture system or a range sensor.
[0009] In real-world scenarios, several problems that make it difficult to extract reliable full-body 3D human poses need to be addressed. For example, such methods can be sensitive to noise, complex camera motion, image blur, and lighting artifacts, as well as occlusion and truncation of the human body. Such problems can further complicate the estimation of 3D human poses and, consequently, make human action recognition more difficult and provide inaccurate results.
[0010] An explicit human pose can be defined by 3D coordinates (xi, yi, zi) at which there are i key points (e.g., corresponding to joints of the human body) or body key point coordinates such as skeletal data.
[0011] Methods based on explicit poses can require that the person be fully visible and can be limited to the recognition of one person's actions. Some of such methods rely only on 2D poses for action recognition tasks in actual videos. Compared with 2D poses, 3D poses have the advantage of being less ambiguous and better representing motion dynamics. However, 3D human poses can only be obtained in motion capture systems, Kinect sensors, or multiple camera setups and cannot be obtained in actual videos.
[0012] Designing a neural network for capturing spatio-temporal information is not easy. A neural network can be configured to generate a specific output such as a classification result based on an input. However, without knowledge of the specific structure of the neural network and the specific filters or weight values, it may be impossible to determine or evaluate on what features a specific classification is based. Therefore, a combination of two neural networks utilized for a specific purpose based on different ideas respectively, or a part of both of these neural networks, may not guarantee the desired effect which can only be assumed based on different ideas. For example, when the specific features of the input underlying the result in a neural network are not clear, a combination with another neural network or a part of another neural network may not be able to generate the expected or desired result. Determining whether a combination of two different approaches or networks provides a specific result can be quite costly in terms of evaluation.
[0013] Techniques for dealing with the problem of human action recognition are needed. It is desirable to provide an improved method for human action recognition that overcomes the above drawbacks. In particular, it is desirable to provide a method for accurately identifying one or more human actions based on one or more images.
Prior Art Documents
Patent Documents
[0014]
Patent Document 1
Summary of the Invention
Means for Solving the Problems
[0015] To overcome the above-mentioned drawbacks and solve the problem of biased image classification, a system and method for accurately determining human behavior are disclosed. The present disclosure provides a method and system for improved action recognition that directly utilizes implicit pose representation. Example embodiments of the summary are described in relation to human action recognition, but hereinafter can be used for individual action recognition, where an individual is used as including inanimate objects such as humans, animals, and robots. An individual can be a living object and / or a moving object that can belong to a particular class, species, or part of a collection.
[0016] In one embodiment, a computer-implemented method for human action recognition includes obtaining one or more images showing at least a portion of one or more humans, generating one or more implicit representations of one or more human poses based at least in part on the one or more images, and determining at least one human action by classifying one or more of the implicit representations of the one or more human poses. By directly classifying one or more of the implicit representations of the one or more human poses, an improved method and system for robust action recognition related to the context shown in the one or more images are provided. At least one human action can be determined by classifying one or more of the implicit representations of the one or more human poses instead of classifying explicit representations of the one or more human poses. By implicitly maintaining the representation, there is no need to resolve the ambiguity of the representation, which provides efficient processing of the implicit representation compared to the explicit representation.
[0017] In one embodiment, a computer-implemented method for training a neural network configured to recognize actions performed by an individual in an image includes obtaining, by one or more processors, a first image and a first label corresponding to an action performed by the individual in the first image, respectively; generating, by one or more processors, a feature vector by the first neural network by inputting a part of the first image into the first neural network configured to determine an action performed by the individual based on the input image; and training, by one or more processors, a second neural network configured to recognize an action performed by the individual in the image based on the input image and the feature vector corresponding to the first label corresponding thereto.
[0018] In one embodiment, a system includes one or more processors and a memory including code that, when executed by the one or more processors, performs functions including obtaining an image including at least a portion of an individual, generating an implicit representation of a pose of the individual in the image based on the image, and determining an action performed by the individual and captured in the image by classifying the implicit representation of the pose of the individual.
Brief Description of the Drawings
[0019] The present disclosure will be more fully understood from the detailed description and the accompanying drawings.
[0020]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8A
Figure 8B
Figure 8C
Figure 8D
Figure 8E
Figure 8F
Figure 9
[0021] In the drawings, reference numerals can be reused to identify similar and / or identical elements.
Best Mode for Carrying Out the Invention
[0022] <Summary of the Invention> In one embodiment, one or more implicit representations of one or more human poses include one or more feature vectors generated by at least a portion of a first neural network. The first neural network can be initially trained to determine two-dimensional (2D) and / or three-dimensional (3D) poses of a person. Alternatively or additionally, at least one human action can be determined by classifying one or more implicit representations of one or more human poses, including feature vectors, by a second neural network that can be different from the first neural network. The second neural network can be trained to classify one or more feature vectors according to the actions of a person. By classifying the feature vectors generated by the neural network for pose determination, implicit representations of human poses can be generated in an efficient manner. For example, a feature vector can be extracted in a first neural network initially trained to detect the pose of a person.
[0023] In an additional embodiment, a method of obtaining one or more images includes obtaining a sequence of images, such as a sequence of video and / or RGB (Red-Green-Blue) images, generating one or more implicit representations including generating a sequence of implicit representations of human poses, concatenating the implicit representations of the sequence of human pose implicit representations, and determining at least one human action including performing the same convolution as a one-dimensional (1D) temporal convolution on the concatenated implicit representations. Concatenating the implicit representations can include concatenating all of the implicit representations of the sequence of implicit representations. By performing a convolution on the concatenated implicit pose representations, spatio-temporal information is processed in a single operation (e.g., with the support of a simple 1D convolution), so that human actions can be determined in an efficient manner. As a result, a less complex architecture can implement this embodiment compared to architectures having complex processing steps or having multiple layers for encoding and processing temporal information.
[0024] In an additional embodiment, the method further includes determining one or more candidate boxes around each person or one or more persons partially or fully shown in at least one of the one or more images. One or more implicit representations of one or more person poses can be generated based on the one or more candidate boxes and the one or more images. If the one or more images include a sequence of images such as a video sequence, the method can further include extracting tubes based on the candidate boxes along the sequence of images. Further, the step of generating one or more implicit representations of one or more person poses can include generating a sequence of implicit representations of one or more person poses along the tubes. A 1D temporal convolution can be performed or applied to a sequence of implicit representations that can be a sequence of feature vectors such as stacked feature vectors to determine at least one person action. By generating implicit representations based on the one or more candidate boxes and the one or more images, the actions of one or more persons can be determined in a more accurate manner. For example, RoI (Region-of-Interest) pooling can be performed based on the candidate boxes to further enhance the robustness to the context of the images.
[0025] In one embodiment, a computer-implemented method for training at least a portion of a neural network for human action recognition or classification includes obtaining a first set of one or more images, such as a sequence of videos or images, and a first label that describes an action displayed in the one or more images and that can correspond to the first set; generating one or more sets of feature vectors by inputting at least one image of each set of the first set of one or more images into a first neural network trained to determine a human pose; and training at least a portion of a second neural network, such as one or more layers of the second neural network, based on the feature vectors of at least one set of feature vectors for classification of human actions. The feature vectors of the at least one set of feature vectors correspond to the corresponding label among the input images and the first label, and the feature vectors are generated by at least a portion of the first neural network. By training at least a portion of a neural network for classification of human actions based on the feature vectors generated by a neural network for human pose determination, a neural network for classification or recognition of human actions can be generated quickly and efficiently. For example, a neural network for classification or recognition of human actions can be combined with an existing neural network for human pose determination or a portion of such a neural network. Further, by directly training a neural network with respect to the pose features or feature vectors of a neural network trained to determine an initial human pose, an improved neural network for robust action recognition can also be generated with respect to the context displayed in the image.
[0026] In an additional embodiment, each of the first set of one or more images includes a sequence of images, and the feature vectors of at least one set of feature vectors include a sequence of feature vectors that at least partially corresponds to the sequence of images. A plurality of the feature vectors in the sequence of feature vectors are concatenated or stacked, and a temporal convolutional layer, such as a 1D temporal convolutional layer of a second neural network, is trained with the concatenated or stacked plurality of feature vectors. By training a temporal convolutional layer of a neural network for action recognition with the stacked features generated by a neural network for pose detection, an architecture is provided or trained that processes the temporal and spatial information of the images in a simplified manner. Thereby, the training cost of the neural network for action recognition can be reduced.
[0027] In an additional embodiment, the first neural network can be trained based on a second set of labels corresponding to a second set prior to the second neural network and a second set of one or more images. The weights of the first neural network can be fixed during the training of the second neural network. By fixing the weights of the first neural network during training, the training of the second neural network can be performed efficiently and effectively. For example, more processing resources can be used in training a part of the neural network for action classification. Here, when a sequence of feature vectors is used for training, steps can be included that allow for a larger temporal window. Alternatively, at least a part of the first neural network is trained together with the second neural network based on a first set of images and corresponding first labels. This can increase the accuracy of the combination of the first neural network and the second neural network.
[0028] In an additional embodiment, a computer-readable storage medium storing computer-executable instructions is provided. When executed by one or more processors, the computer-executable instructions perform a method for human action recognition or a method for training at least a part of the neural network for human action recognition described herein.
[0029] In an additional embodiment, an apparatus including a processing circuit is provided. The processing circuit is configured to perform a method for human action recognition or a method for training at least a part of the neural network for human action recognition described herein.
[0030] In one embodiment, a computer-implemented method for recognizing an action performed by an individual includes obtaining, by one or more processors, an image including at least a portion of the individual; generating, by one or more processors, based on the image, an implicit representation of a pose of the individual in the image; and determining, by one or more processors, an action performed by the individual and captured in the image by classifying the implicit representation of the pose of the individual.
[0031] In additional embodiments, the implicit representation includes a feature vector from which keypoints having three-dimensional (3D) coordinates corresponding to the pose of the individual's skeleton can be derived.
[0032] In additional embodiments, an implicit representation that does not include keypoints having 3D coordinates is used to derive keypoints corresponding to the pose of the individual's skeleton.
[0033] In additional embodiments, the implicit representation does not include two-dimensional (2D) or three-dimensional (3D) skeleton data of the individual.
[0034] In additional embodiments, generating the implicit representation includes generating, using a first neural network, an implicit representation of a pose of the individual in the image.
[0035] In additional embodiments, the method further includes training the first neural network to determine at least one of a two-dimensional (2D) and a three-dimensional (3D) pose of the individual in the image.
[0036] In additional embodiments, determining the action performed by the individual includes determining the action by classifying the implicit representation of the pose using a second neural network.
[0037] In additional embodiments, the implicit representation includes a feature vector.
[0038] In an additional embodiment, the method further includes training a second neural network to classify feature vectors based on the actions of an individual.
[0039] In an additional embodiment, the individual includes a person, the pose includes the pose of the person, and the action includes the actions performed by the person.
[0040] In an additional embodiment, the image includes a sequence of images from a video.
[0041] In an additional embodiment, the step of generating the implicit representation includes generating an implicit representation based on each image. The method further includes concatenating the implicit representations to generate a concatenated implicit representation. The step of determining the action includes performing a convolution on the concatenated implicit representation to generate a final implicit representation and determining the action performed by the individual by classifying the final implicit representation.
[0042] In an additional embodiment, the convolution corresponds to a one-dimensional (1D) convolution.
[0043] In an additional embodiment, the convolution corresponds to a one-dimensional (1D) temporal convolution.
[0044] In an additional embodiment, the method further includes determining, by one or more processors, candidate boxes around the individual captured in the image, and the step of generating the implicit representation includes generating an implicit representation based on the candidate boxes.
[0045] In an additional embodiment, the method further includes extracting, by one or more processors, a tube from the image based on the candidate boxes.
[0046] In an additional embodiment, the step of determining the candidate box includes the step of determining the candidate box using a regional proposal network (RPN) module.
[0047] In an additional embodiment, the first image includes a sequence of images of an individual performing an action.
[0048] In an additional embodiment, the method further includes the step of concatenating the feature vectors by one or more processors, and training includes training a one-dimensional (1D) temporal convolutional layer of a second neural network using the concatenated feature vectors.
[0049] In an additional embodiment, the method further includes the step of training a first neural network by one or more processors based on a second image and a second label corresponding to the second image, and training the second neural network includes training the second neural network after training the first neural network, and the weight values of the first neural network are maintained in a fixed state during the training of the second neural network.
[0050] In an additional embodiment, the method further includes the step of training the first neural network by one or more processors based on the first image and the first label corresponding to the first image together with the training of the second neural network.
[0051] Additional application areas of the present disclosure will become apparent from the detailed description, claims, and drawings. The detailed description and specific examples are for illustrative purposes only and do not limit the scope of the present disclosure.
[0052] <Detailed Description of the Invention> This specification describes a system and method for human action recognition. For the purpose of the description, numerous exemplifications and specific details are presented to provide a complete understanding of the described embodiments. The embodiments defined by the claims can include some or all of the features of these examples alone or in combination with other features described below, and can further include modifications and equivalents of the features and concepts described herein. Exemplary embodiments are described with reference to drawings in which elements and structures are denoted by reference numerals. Also, in the case of a method of an embodiment, the method steps and elements can be combined in parallel or sequential execution. Unless inconsistent, all embodiments described hereinafter can be combined with each other. The methods described herein can be performed by one or more processors.
[0053] FIG. 1 is a process flowchart of an exemplary method (100) for human action recognition that directly utilizes one or more implicit representations of a human pose, such as human pose features, according to one embodiment. Specifically, the method (100) can be used to capture from one or more images or predict actions such as human actions (actions performed by a human) occurring in a video. The video is generated from a sequence of images captured over time.
[0054] At reference number 110, for example, one or more images are obtained by one or more processors of a computing device. The one or more images can be captured, for example, using a camera of the computing device, obtained from the memory of the computing device, received from another computing device (e.g., via a network), or obtained by other suitable means. The one or more images can include a sequence of images or frames, such as a trimmed video. A trimmed video means that the video persists for only a few seconds (e.g., 2 to 10 seconds) for a particular task performed by a person. The one or more obtained images can display at least a portion of one or more persons. In one example, the one or more images include videos such as real-world videos, videos of websites (e.g., YouTube, TikTok, etc.), television images, movies, camera recordings, or videos captured on a mobile platform. The image or video can include a sequence of multiple RGB (red, green, blue) frames / images.
[0055] At reference number 120, based on at least a portion in one or more acquired images, one or more implicit representations for one or more human poses are generated. Examples of implicit representations of human poses can be feature vectors or features generated by at least a portion of a first neural network such as a pose detector module. The feature vectors or features can include deep or mid-level 3D pose features. One or more implicit representations of one or more human poses can indicate information related to one or more human poses or by which one or more human poses can be determined. Implicit representations of one or more human poses can include ambiguity regarding one or more human poses. Conversely, explicit representations of human poses are clearly defined in relation to human poses. For example, an explicit representation of a human pose can include 13 keypoints along with each keypoint having known three-dimensional (3D) coordinates (xi, yi, zi). FIG. 7 includes an exemplary stick figure including 39 keypoints (13 keypoints per frame / image * 3 = 39). Conversely, an implicit representation of a pose can be a feature vector from which the 39 keypoints can be derived.
[0056] By implicitly maintaining the representation, the ambiguity in pose detection can be made unnecessary to resolve, and thus the processing for implicit representations can be improved compared to explicit representations. For example, the information content of implicit representations can be increased compared to explicit representations.
[0057] The first neural network can be a neural network trained or initially trained to recognize or determine human poses such as two-dimensional (2D) and / or three-dimensional (3D) poses of a person. One or more implicit representations of one or more human poses can be generated based on a single frame (or image) of a video. For example, one or more implicit representations can be generated for one frame, two or more frames, or each frame of a video.
[0058] At reference number 130, at least one human action is determined by classifying one or more implicit expressions of one or more human poses. In various embodiments, the step of determining at least one human action by classifying one or more implicit expressions of one or more human poses is performed by a second neural network. The second neural network can be trained to classify one or more feature vectors generated by a first neural network based on human actions. The method (100) can be performed by a neural network including at least a portion of the first neural network and the second neural network. The first neural network and the second neural network can be executed on a single computing device or server, or on distributed computing devices and / or servers.
[0059] FIG. 2 is a process flowchart of a method (200) for generating an implicit human pose expression. The method (200) or one or more portions of the method (200) can be used to generate one or more implicit expressions of reference number 120 in FIG. 1.
[0060] The method (200) can be started at reference number 210 including the step of processing the acquired image. The step of processing can include the step of inputting the acquired image into a first neural network or the step of inputting the acquired image into one or more convolutional layers of the first neural network to generate a feature vector. The first neural network can be a deep convolutional neural network (CNN) configured to determine a human pose (e.g., an implicit expression of a human pose).
[0061] At reference number 220, one or more candidate boxes or bounding boxes are determined around a person shown in at least one of the one or more images. An exemplary bounding box is illustrated around the person in FIG. 7. In an embodiment, the RPN (Region Proposal Network) module is used to determine candidate boxes or bounding boxes around each person shown in the one or more images. The candidate boxes can be determined based on the generated feature vectors or the processed images. In various embodiments, the candidate boxes correspond to the detected person poses or the person keypoint coordinates.
[0062] At reference number 230, one or more implicit representations of one or more person poses are generated based on the one or more candidate boxes and the processed image. The generation of one or more implicit representations of one or more person poses can include modifying the feature vectors generated by performing region of interest (RoI) pooling based on the determined candidate boxes. In this example, one or more implicit representations of one or more person poses include the one or more modified and generated feature vectors.
[0063] FIG. 3 is a process flowchart of a method (300) for human action recognition according to at least one embodiment. One or more parts of the method (300) can correspond to one or more parts of the method (100) of FIG. 1 and / or the method (200) of FIG. 2.
[0064] At reference number 310, a sequence of images or frames is acquired. The sequence of images can be a video, for example, a video recorded by a camera or a video acquired from a computer-readable medium. At reference number 320, the sequence of images is processed. The processing step can include inputting at least two images of the sequence of images into a first neural network or one or more convolutional layers of the first neural network to obtain a series of feature vectors. In various embodiments, each image or frame of the sequence of images or frames is processed to obtain a sequence of feature vectors.
[0065] At reference number 330, candidate boxes around the people shown in at least two images are determined, for example, using an RPN module. The candidate boxes can be determined based on the generated feature vectors and / or the people detected from the at least two images. As described above, the feature vectors generated by a person pose detector or a neural network for person pose detection can be an implicit representation of the person pose.
[0066] To determine the evolution of poses over time in the sequence of images, each person can be tracked over time. This can be achieved by linking the candidate boxes around the person along the sequence of images.
[0067] At reference number 340, the tube is extracted based on candidate boxes determined along a sequence of at least two images. The tube can be a temporal series of spatial candidates or a bounding box, i.e., a spatio-temporal video tube. For example, candidate boxes of consecutive images can be concatenated to generate a tube. Starting from the candidate box with the highest person detection score, the candidate boxes can be matched with detections in subsequent frames, for example, based on the IoU (Intersection-over-Union) between candidate boxes. In one exemplary case, two candidate boxes in subsequent images of a sequence of images are concatenated when the IoU is greater than a predetermined value such as 0.3. Otherwise, the candidate box is matched with candidate boxes of images after the subsequent image, and linear interpolation can be performed for the missing candidate box of the image or the candidate box of the subsequent image. For example, if there is no match among 10 consecutive images (e.g., IoU less than a predetermined value), the tube can be aborted. This procedure can be executed forward (video in the forward time direction) and / or backward (video in the reverse time direction) to obtain tubes containing persons. Thereafter, all detections of the first link can be deleted, and the procedure can be repeated for the remaining detections.
[0068] In summary, all sets of person detections in all frames are considered first. Then, the first link is created as described above. Next, a second link of bounding boxes that do not include the bounding boxes of the previous link is created. For this, the same process considering all bounding boxes of all frames except the previously linked boxes can be repeated.
[0069] At reference number 350, the sequence of implicit pose representations is generated based on the sequence of processed images and the extracted tubes. For example, the sequence of implicit pose representations can include the generated feature vectors by performing RoI-pooling on the processed images among the sequence of images.
[0070] At reference number 360, the implicit pose representations of the sequence of implicit pose representations are stacked or concatenated end to end. At reference number 370, convolution can be applied to the concatenated or stacked representations to determine an action or multiple action scores. The convolution can be a 1D temporal convolution or another suitable type of convolutional layer. Based on the result of the convolution, an action score is determined and can be used to classify the representations concatenated or stacked by the action. By using the feature vectors as implicit representations, there is no need to resolve the ambiguity of pose detection, and the temporal convolution can be made in an efficient manner to generate results.
[0071] In various embodiments, instead of explicitly estimating a sequence of poses, such as 3D skeleton data, and operating based on such an explicit sequence of poses, an action recognition framework is disclosed that uses an implicit deep pose representation learned by a 3D pose detector. The implicit deep pose representation has the advantage of encoding pose ambiguity and has the advantage of not requiring tracking or temporal inference that a pure pose-based method may need to resolve complex cases. In at least one embodiment, the first neural network for pose determination includes a pose detector such as an LCR-Net++ pose detector, an LCR-Net pose detector, or other suitable type of pose detector. The LCR-Net++ pose detector is described in "LCR-Net++: Multi-person 2D and 3D pose detection in natural images" by G. Rogez, P. Weinzaepfel, and C. Schmid, IEEE trans, PAMI, 2019. The LCR-Net++ pose detector can be robust even in difficult cases such as occlusion and truncation by the image boundaries. LCR-Net++ is configured to estimate full-body 2D and 3D poses for one or more persons shown in an image. The human-level intermediate pose feature representation of LCR-Net++ can be stacked, and 1D temporal convolution can be applied to the stacked representation.
[0072] FIG. 4 is a schematic diagram illustrating a functional block diagram (400) for determining human actions from an image. FIG. 4 illustrates three images (410a, 410b, and 410c) that form a sequence of images (410) over time. However, the sequence of images can include two or more images. In this illustration, the images (410a, 410b, 410c) show a person dressed in the uniform of an American football player playing basketball (not American football) in an American football stadium. Thus, the images (410a, 410b, 410c) include the context related to American football, but show actions performed by two people that are not related to the context of the image or are subject to misinterpretation, namely, the actions of a person throwing or pretending to throw a basketball.
[0073] The image can be processed by at least a part of a first neural network for human pose detection. A first neural network such as an RPN module or at least a part of the first neural network can be configured to determine candidate boxes (412a and 412b) around the person displayed in the image. A human track can be defined by a tube generated by concatenating candidate boxes around the person for each image in a sequence of images. The human track or tube is displayed in the same grayscale / color for each person along the sequence of images (410a - 410c). Also, the first neural network or at least a part of the first neural network can be configured to generate a sequence of feature vectors (420) including feature vectors (420a, 420b, 420c). In the illustration of FIG. 4, it is displayed only for one of the human tracks / tubes for mobility. The feature vectors (420a, 420b, 420c) correspond to the images (410a, 410b, 410c). The feature vectors (420a, 420b, 420c) of the sequence of feature vectors can be mid-level pose features of the first neural network along the human track.
[0074] After generating the sequence of feature vectors (420), the feature vectors (420a, 420b, 420c) are stacked / concatenated to utilize temporal information. The order of the stacked feature vectors (430) can correspond to the order of the feature vectors (420a, 420b, 420c) in the sequence of feature vectors (420) and / or the order of the images (410a, 410b, 410b) in the sequence of images (410). The stacked feature vectors (430) can be classified according to the action by determining an action score (440). Each classification can have a related action, and the classification with the highest action score can be selected to determine the action performed in the sequence.
[0075] In various embodiments, the classification of the stacked feature vectors (430) can include inputting the stacked feature vectors (430) for action classification for classification by a second neural network into the second neural network. The second neural network can include one or more convolutional layers, such as a 1D convolutional layer. The second neural network can additionally or alternatively include one or more fully connected layers configured to output an action score (440).
[0076] Based on the action score (440), the action performed by a person in a sequence can be determined, for example, by an action module. For example, the action score (442a) can be higher than the action score (442b) and other different action scores. Thus, the action module can determine that the action corresponding to the action score (442a) was sequentially performed. In FIG. 4, the action score (442a) corresponds to "playing basketball".
[0077] In various embodiments, the action module can require that the highest action score be greater than a predetermined value (e.g., 0.5) in order to be considered a valid action. In such an embodiment, if the highest action score is less than the predetermined value, the action module can determine that an unknown action was performed in the sequence.
[0078] A system for action recognition can include a part of a first neural network configured to generate feature vectors (420a, 420b, 420c) based on a sequence, and a second neural network configured to determine an action in the sequence. The second neural network configured to encode time information of a sequence of images or feature vectors can directly use the feature vectors or human pose features and may not require additional data such as a depth component of an image / frame (e.g., an RGB - d image or a grayscale - d image). This provides an improved system for action recognition that is robust to context.
[0079] In various embodiments, single - frame detection and deep mid - level 3D pose features can be generated using a 3D pose detector module. Temporal integration can be accomplished by generating tube proposals, along with which implicit 3D pose - based features are concatenated and used to directly classify actions performed using a 1D temporal convolutional layer.
[0080] The method described herein generates less inherent pose - estimation noise compared to other 3D action - recognition methods used to estimate and classify a sequence of 3D human poses. Further, other methods for 3D action recognition can include a complex architecture for extracting time information from a sequence of images compared to the embodiments described herein that use 1D temporal convolution.
[0081] For example, an input video is at least partially processed by a pose detector module, such as LCR-Net++, to detect a person tube. Mid-level pose features can be extracted from the pose detector module that processes the input video. The mid-level pose features that can be generated based on the person tube can be stacked / concatenated along the time stream. Then, a single 1D temporal convolution can be applied to the stacked pose features to obtain an action score. The single 1D temporal convolution applied to the mid-level pose features can be executed in a 1D temporal convolution layer of a neural network. The 1D temporal convolution layer provides a simple yet not complex architecture of the neural network to extract temporal information.
[0082] FIG. 5 is a functional block diagram including an exemplary architecture (500) of a neural network for person pose determination. The architecture (500) includes a neural network (module) (502). The neural network (502) can be or include the first neural network described herein. The neural network (502) is configured to determine the pose for each person fully or partially shown in the image (510) input to the neural network (502).
[0083] The neural network (502) includes one or more convolutional layers (520) and can include a Region Proposal Network (RPN) module (530), a classification module (550), and a regression module (560). Further, the neural network (502) can include one or more fully connected (fc) layers (540). The fully connected layer (540a) is configured to generate an implicit representation that is used as an input to a second neural network (580) based on the output of the convolutional layer (520). The fully connected layers (540b and 540c) are configured to generate explicit representations based on the output of the fully connected layer (540a). The classification and regression modules (550, 560) can determine a classification and perform regression on the outputs of the fully connected layers (540b and 540c). The fully connected layer (540a) is used to generate an implicit representation and estimate a pose because it cannot be used to directly specify a pose as can be performed by the fully connected layers (540b and 540c).
[0084] The RPN module (530) is configured to extract candidate boxes around a person. In an embodiment, a pose proposal is determined by the RPN module (530) by matching / fitting the pose of the person within the candidate box to a pre-determined anchor pose and placing the anchor pose in the candidate box. The pose proposal is scored by the classification module (550) and refined / regressed by the regression module (560). The regression module (560) can include class-specific regressors that are trained independently for each anchor pose.
[0085] An anchor pose can be a key pose corresponding to a person's standing pose, sitting pose, and other common human poses. Each anchor pose includes a plurality of key points. The regression module (560) can take as input the same features used for classification in the classification module (550). The anchor poses can be jointly defined in 2D and 3D, and refinement / regression can be performed in a common 2D-3D pose space.
[0086] In various embodiments, the neural network (502) can be configured to detect a large number of people in a scene and an image, or to output a full-body pose even in the case of occlusion or cutting by an image boundary, or to be configured to perform all of these. In an example of cutting at an image boundary, the neural network (502) can cut a full-body pose at the same key points. The results can be generated in real time. The neural network (502) can include a Faster R-CNN similar to the architecture.
[0087] The neural network (502) can be configured to extract a person tube or track a person from a sequence of images such as a video. In this example, each image of the sequence of images or a subset of the images of the sequence of images is input into the neural network (502), and candidate boxes around the person for the subset of images are detected. Then, the candidate boxes of subsequent images are linked, for example, using the IoU (Intersection-over-Union) between the candidate boxes.
[0088] In various embodiments, the implicit representation of a human pose includes features or feature vectors that are used as input to the final layer for pose classification and to the final layer for joint 2D and 3D pose refinement / regression. Such features can be selected and / or extracted and used for the classification of human actions. By way of example, such features can include 2048 dimensions (or other suitable dimensions), be generated by a ResNet50 backbone or other suitable type of image classification algorithm, or can correspond to all of the above.
[0089] Features generated by a neural network for pose detection, similar to neural network (502), can encode (or include) information for both 2D and 3D poses. Thus, such poses can include pose ambiguity. The methods described herein can directly classify such features to determine human actions and can eliminate the tracking or temporal reasoning required by pure pose-based methods to clarify complex cases and situations. Features can be simply stacked over time by a human tube, and a temporal convolution of kernel size T can be applied on the resulting matrix, where T corresponds to the number of images in the sequence of images. This convolution can output an action score for the sequence of images.
[0090] FIG. 6 is a flowchart showing an exemplary method (600) for training at least a portion of a neural network for human action recognition and pose estimation. Optionally, at reference numeral 610, the neural network is trained to classify an image according to a human pose displayed in the image based on a training image and a label corresponding to the training image. The training image can display one or more humans or at least a portion of one or more humans, and the label can indicate one or more poses of one or more humans. The neural network can be trained to estimate 2D and / or 3D human poses of one or more humans or at least a portion of a human displayed in the image.
[0091] At reference numeral 620, a first set of one or more images and a first label corresponding to the first set are obtained. Each set of the first image set can include a sequence of images or frames such as a video sequence. One of the first labels can define one or more actions performed by a human in one of the one or more images of the first set.
[0092] At reference numeral 630, one or more sets of at least one feature vector are generated by inputting at least one image of each of the first set of one or more images into a first neural network. The set of at least one feature vector is generated by at least a portion of a first neural network such as one or more convolutional layers and / or an RPN module. In various embodiments, the first neural network is the neural network trained at reference numeral 610. Alternatively, the first neural network can be a neural network that is obtained together with the first set of one or more images and the first label and is initially trained to determine human poses.
[0093] Reference number 640 can be selective as shown by the dotted line. In this case, each of the first set of one or more images includes a sequence of images, each of at least one set of feature vectors includes a sequence of feature vectors that at least partially corresponds to the sequence of images, and a plurality of feature vectors of the sequence of feature vectors can be concatenated or stacked at reference number 640.
[0094] Referring to FIG. 5 at reference number 650, at least a portion of the second neural network for the classification of human actions is trained based on at least one feature vector corresponding to at least one of at least one image and at least one corresponding label of the first label. In various embodiments, the training of the second neural network (e.g., the exemplary second neural network (580) of FIG. 5) includes training a convolutional layer, such as a 1D temporal convolutional layer of the second neural network, with a plurality of concatenated feature vectors (e.g., the feature vectors of the fully connected layer (540a) of FIG. 5). In various embodiments, the second neural network (580) can include or be composed of a single convolutional layer that can be a 1D temporal convolutional layer.
[0095] After training, a second neural network is coupled to at least a portion of the first neural network illustrated in FIG. 5, and a neural network configured to predict actions in a video / sequence can be generated, and the corresponding portion can be configured to generate, for example, feature vectors or implicit pose features. At least a portion (575) of the first neural network (502) or the first neural network (502) (omitting the portion identified by reference number 570 used to train the first neural network (502)) can be configured to generate a feature vector for each person track / tube (e.g., the feature vectors 420a, 420b, 420c for one person track / tube shown in FIG. 4), and the feature vectors are stacked / concatenated (as with reference number 430 in FIG. 4) and classified at reference number 590 by a second neural network (580) (e.g., 440 in FIG. 4 showing classifications 442a and 442b).
[0096] In various embodiments, the second neural network is trained after the training of the first neural network. In this case, the weights of the first neural network can be fixed during the training of the second neural network, which means that such weights are not changed during training. The weights of the first neural network can be fixed during training, for example, due to graphics processing unit (GPU) memory constraints or to enable a larger temporal window to be considered. In various embodiments, at least a portion of the first neural network is trained together with the second neural network based on a first image set and its corresponding first labels.
[0097] During training, random clips of T consecutive frames are sampled, and cross-entropy loss can be used, where T is the number of frames. At test time, a fully convolutional architecture can be used for the first neural network, and class probabilities can be averaged by the softmax of scores for all clips of the video.
[0098] (Experimental results) Various embodiments, such as the example of FIG. 4, are compared respectively with two baselines that use a spatiotemporal graph convolutional network for explicit 3D or 2D pose sequences. To evaluate the results for the mimicked actions, experimental results for the standard action recognition dataset and the mimicked action "Mimetics" dataset are presented.
[0099] FIG. 7 shows an overview of the first baseline (700) based on explicit 3D pose information. Given an input video, a person tube is detected, and 2D / 3D poses are estimated using LCR-Net++ (only shown for one tube for readability in FIG. 7). Then, a 3D pose sequence is generated, and an action score is obtained using an action recognition method based on a spatiotemporal graph neural network. More precisely, 3D poses estimated by LCR-Net++ are generated for each candidate box, and a 3D human pose skeleton sequence for each tube is constructed. Then, a 3D action recognition method is used to process the 3D human pose skeleton sequence. In this idea, it includes creating a spatiotemporal graph from the pose sequence to which spatiotemporal convolution is applied. Hereinafter, the first baseline is referred to as STGCN3D.
[0100] The second criterion is based on 2D poses. The transformation of the previous pipeline replaces the 3D poses estimated by LCR-Net++ with 2D poses. On the other hand, since 3D poses provide more information than inherently ambiguous 2D poses, such a transformation may degrade performance. On the other hand, 2D poses extracted from images and videos are more accurate than 3D poses and are more susceptible to noise. Below, the second criterion is STGCN2D.
Table 1
[0101] Diverse action recognition datasets with various levels of measured materials are illustrated in Table 1, which summarizes the datasets in terms of the number of classes, number of videos, number of splits, and frame-level GT (ground-truth). For datasets with multiple splits, some results are reported only for the first split, e.g., denoted as JHMDB-1 for split 1 of JHMDB. The implementation of the present invention can be used to perform action recognition in actual videos, but the results can be verified using the NTU 3D action recognition dataset including 2D and 3D GT poses. The standard CS (cross-subject) split is used. Also, the experiments are conducted on the JHMDB and PennAction datasets that have GT2D poses but no 3D poses, because the datasets include in-the-wild videos. Additional datasets are HMDB51, UCF101, and Kinetics, which do not contain more information than the measured labels of each video. As a metric, the standard mean accuracy averaged over all classes, i.e., the proportion of correctly classified videos per class, is calculated.
[0102] For datasets with GT2D poses, the performance when using Ground-Truth Tubes (GT Tubes) obtained from the GT2D poses is compared with the performance when using the estimated tubes (LCR Tubes) made from the estimated 2D poses. In the latter case, if the spatio-temporal IoU of the GT tube is greater than 0.5, the tube is displayed as a positive number, and otherwise as a negative number. When there is no tube annotation, it is assumed that a video class label is assigned to all tubes. If no tubes are extracted, the video is ignored during training and considered misclassified in the test videos. The Kinetics dataset contains many videos where only the head is visible and many first-person perspective clips where only one hand or the main object being manipulated during the action is visible. Tubes cannot be obtained for 2% of the videos in PennAction, 8% in JHMDB, 17% in HMDB51, 27% in UCF101, and 34% in Kinetics.
[0103] Figures 8A - 8F show plots of the average precision of the exemplification described herein (hereinafter also referred to as the Stack Implicit Pose Network (SIP-Net)) for different numbers of video frames T for all datasets for different tubes (GT or LCR) and features (poses or actions extracted at low or high resolution). Generally, the larger the clip size T, the higher the classification accuracy. This is especially the case for datasets with long videos such as NTU and Kinetics. This is maintained for both when using GT tubes and LCR tubes. In various embodiments, the number of frames T can be maintained at T = 32 in the experiments described below.
[0104] In various embodiments, the temporal convolution for the LCR pose function was compared with the features extracted in a Faster R-CNN model with a ResNet50 backbone trained to classify actions. Such frame-level detectors were used in various embodiments. LCR-Net++ can be used to extract LCR pose features. LCR-Net++ is based on Faster R-CNN using ResNet50. Thus, only the learned weights of the network trained to classify actions are changed compared to LCR-Net++ features.
[0105] In this experiment, features of two videos with "baseball swing" and "bench press" actions were extracted using LCR-Net++ and the network trained to classify actions along the tube. For each sequence, the difference (distance) between certain features along the tube when using Faster R-CNN action or LCR pose features is shown through the feature correlation relationship. The results show that the implicit pose features (LCR pose features) can show more deformations inside the tube than the Faster R-CNN action features. In particular, when using action features instead of pose features, a decrease in accuracy (about 20% in JHMDB-1 and PennAction, about 5% in NTU for T = 32) is shown. The pose features are significantly improved in performance compared to the action features because the difference or distance between features in the tube increases. In particular, when training a frame-by-frame detector for an action, most of the features of a given tube are correlated with each other. Thus, it may be difficult to utilize the temporal information from them. Conversely, the LCR-Net++ pose features are significantly changed over time like the pose, and more advantages of temporal integration can be obtained.
[0106] Finally, compare the impact of image resolution on performance (during testing) using (a) features extracted at low resolution after resizing the input image such that the smallest side is 320 pixels (by the real-time model of LCR-Net++), and (b) features extracted at high resolution with the smallest image size set to 800 pixels (used for training all of the Faster R-CNN and LCR-Net++ models). In JHMDB-1, it can be seen that the higher the resolution, the lower the performance, which can be explained by the relatively small dataset size and low video quality. In the PennAction and NTU datasets, which include high-resolution videos, features extracted at higher resolutions can bring improvements of approximately 1% and 10% respectively. The small-resolution setting is maintained in some embodiments for speed, especially in the case of large datasets such as Kinetics.
Table 2
[0107] Table 2 shows the results of comparing the embodiments of the present invention with the first and second criteria using GT and LCR tubes in the JHMDB-1, PennAction, and NTU datasets. For the JHMDB-1 and PennAction datasets, this application outperforms the criteria based on explicit 2D-3D pose representations in all of the GT and LCR tubes, despite having a much simpler architecture compared to the criteria. The estimated 3D pose sequences are generally noisy and may lack temporal consistency. Also, the first criterion confirms that it contains more ambiguous information rather than being discriminatory by 2D poses, as it is much more performant than the corresponding 2D counterparts.
[0108] In the NTU dataset, the 3D pose criterion can achieve an accuracy of 75.4% when using GT tubes and predicted poses, and 81.5% when using GT 3D poses. Such a 6% difference in a limited environment can increase in the case of actual captured videos. The implementation performance of the present invention may be lower at 66.7% in GT tubes, but as shown in FIGS. 8A-8F, this can be because the features are extracted at a low resolution. At a higher resolution, 77.5% equivalent to STGCN3D can be obtained.
Table 3
[0109] Table 3 compares the classification accuracy for all datasets when using the LCR tube (see the first three rows). The present invention (SIP-Net) based on implicit pose features exhibits far superior performance to other methods using explicit 2D and 3D poses with a difference of more than 10% in HMDB51, UCF101, and Kinetics. This shows the robustness of this representation when compared with explicit body keypoint coordinates. Interestingly, in HMDB51, UCF101, and Kinetics, the 2D pose criterion exhibits slightly better performance than the 3D pose criterion, which can mean that there is noise in the 3D pose estimates.
[0110] Table 3 also shows the comparison between the present application and the pose-based method. Compared with PoTion, in the case of the present application, higher accuracy can be obtained with a difference of 3% in JHMDB, 0.5% in HMDB51, and 6% in Kinetics, and lower accuracy can be obtained in UCF101. In some videos of the UCF101 dataset, the people are too small to be detected and no tubes can be created. Higher accuracy can be obtained in the first split of JHMDB and HMDB51 compared to some pose models, and lower accuracy can be obtained not only in NTU but also in UCF101 for the reasons as described above due to low-resolution processing. In NTU and PennAction, some methods obtain higher accuracy because they also utilize appearance features. When the present application is combined with the standard RGB stream using the 3D ResNeXt-101 backbone, an average accuracy of 98.9% can be obtained in the PennAction dataset.
[0111] To evaluate the bias of the action recognition algorithm for scenes and objects and to evaluate the generalizability in the absence of such visual context, this application can be evaluated using mimetic, i.e., a dataset of mimicked actions. Mimetic includes short YouTube video clips containing mimicked human actions including interactions or manipulations with specific objects. Here, it includes sports activities such as playing tennis or juggling a soccer ball, daily activities such as drinking and personal hygiene (e.g., brushing teeth), or playing musical instruments including bass guitar, accordion or violin. Such classes are selected in the action labels of the Kinetics dataset and models trained in Kinetics can be evaluated. Mimetic is used only for testing purposes, which means that the neural network being tested has not been trained on this dataset. Mimetic includes 619 short videos for a subset of 45 human action classes, with 10 - 20 clips for each action. Such actions are performed by pantomime artists on stage or at a distance, but are also commonly performed during pantomime games in daily life or captured and shared for fun on social media. For example, the video shows a soccer player mimicking the "bowling" action to celebrate an indoor surfing or a goal. Clips for each class were obtained by searching for candidates with keywords such as imitation or mimicry where the desired action continues, or using query words such as virtual and intangible where a specific object category follows. The dataset was constructed so that a human observer could recognize the mimicked actions. Care was taken so that clips of the same class do not overlap and do not contain common material (e.g., the same person mimicking the same action in the same background). The videos have diverse resolutions and frame rates and were manually trimmed to lengths between 2 and 10 seconds based on the Kinetics dataset.
Table 4
[0112] Table 4 shows the results for the Mimeticus dataset. All methods are trained on 400 Kinetics classes and tested on the videos of the Mimeticus dataset. Table 4 shows the Top-1, Top-5, and Top-10 accuracies and mAP (mean average-precision). Since each video has a single label, the average precision is calculated as the inverse of the average of the GT label ranks across all videos of the class for each class. The two baselines and the implementation of the present invention were compared not only with the spatio-temporal 3D deep convolutional network trained on clips of 64 consecutive RGB frames, but also with other methods based on OpenPose. A code release was used along with the spatio-temporal 3D convolution.
[0113] The performance is relatively low for all methods, with a Top-1 accuracy of less than 12% and an mAP of less than 20%, indicating that it is difficult to recognize mimicked actions. In fact, all methods cannot be properly performed for some actions such as climbing stairs, reading a newspaper, driving a car, and bowling. In the case of the action of driving a car, there is a tendency for the person to be blocked by the car in the training video, resulting in incorrect pose estimation. In the case of bowling, the pantomime often faces the camera when mimicking the action, but the training video can be captured from the back or side. In most cases (for example, when playing air guitar or reading a journal), the pantomime tends to be exaggerated, making it difficult to classify correctly. Another difficulty for all methods is that some Kinetics actions are subdivided (for example, different classes corresponding to eating various types of food), and it is particularly difficult to distinguish them when mimicking. This implies a problem related to the training data.
[0114] Including the temporal convolution applied to the stacked implicit pose representation, the present application can be utilized to obtain the best overall performance, and it may be possible to reach an accuracy of 11.3% for Top-1 and 19.3% for mAP. Generally, when people are too small or too close, 3% of the mimetic videos have no LCR tube, which affects the performance of the present application based on LCR-Net++ and the two criteria. When there are multiple people in a scene, other failure cases occur. The tube can mismix multiple people or other people (e.g., onlookers) and obtain a higher score than those that mimic the action of interest.
[0115] Also, all pose-based methods, on average, outperform the 3D-ResNeXt baseline. For some classes such as archery, violin playing, bass guitar playing, trumpet playing, or accordion playing, 3D-ResNeXt can obtain 0%, while other methods can be reasonably well performed. Therefore, the action recognition method can learn a method of detecting the object being operated or the scene where the video is captured rather than the action performed.
[0116] 3D-ResNeXt can perform well on classes where objects (such as cigarettes) are too small or objects (such as baseballs, toothbrushes, hairbrushes) blocked by hands are not clearly visible in training videos. In such cases, 3D-ResNeXt can be well performed in mimicked actions as it focuses on the face and hands (for example, for brushing teeth, smoking, hair brushing) or the body (such as throwing a baseball). Also, it reaches good performance for actions captured in scenes related to mimicked actions for most mimetics videos. For example, 70% of the mimetics videos of dunk shooting basketball can be those that mimic (without a ball) at the basketball hoop. The 3D-ResNeXt method can successfully classify such videos (all Top-5), but may fail in other videos of classes that do not occur near the basketball backboard, which mainly indicates that it utilizes the context of the scene.
[0117] Finally, when comparing this application with other pose-based methods, better performance can be obtained for actions where a large object blocks one of the subject's arms (such as violin playing, accordion playing) compared to other OpenPose-based methods. This may be due to the fact that unlike OpenPose which only generates predictions for visible hands and feet, LCR-Net++ outputs the full body pose even when blocked. Also, this application can exceed the 3D criteria for classes such as trumpet playing where the actor is partially visible for most of the time. The 3D pose is noisy while this application is more robust.
[0118] While some specific embodiments have been described in detail, it will be apparent to those skilled in the art that various modifications, changes, and improvements of the embodiments can be made without departing from the intended scope of the embodiments, in light of the teachings described above, and within the content of the appended claims. Also, areas that are judged to be well-known to those skilled in the art have been omitted here in order not to unnecessarily obscure the embodiments described herein. Therefore, it must be understood that the embodiments are not limited by specific exemplary embodiments, but only by the scope of the appended claims.
[0119] While the embodiments have been described in the context of method steps, these also represent descriptions of corresponding components, modules, or features of the corresponding apparatus or system.
[0120] Some or all of the above-described method steps or functions can be executed (or used) by one or more processors, one or more microprocessors, one or more electronic circuits, or processing circuits, and thus can be executed by a computer.
[0121] The above-described embodiments can be executed in hardware or a combination of hardware and software. The above-described embodiments can be implemented, for example, using non-transitory storage media such as floppy disks, DVDs, Blu-Ray discs, CDs (compact discs), read-only memories (ROMs), PROMs, and EPROMs, EEPROMs, RAMs (Random Access Memories), FLASH memories, computer-readable storage media. Such computer-readable media can be any available media that can be accessed by a general-purpose or special-purpose computer system.
[0122] Generally, an embodiment can be implemented by program code or computer-executable commands that are operated to perform one of the methods when the computer program product is executed on a computer, or by a computer program product having the program code or computer-executable commands. The program code or computer-executable commands can be stored on a computer-readable storage medium.
[0123] In one embodiment, the storage medium (or data carrier, or computer-readable medium) includes a computer program or computer-executable commands stored on the storage medium for performing one of the methods described herein when executed by a processor. In additional embodiments, the apparatus includes one or more processors and the storage medium described above.
[0124] In additional embodiments, the apparatus includes means configured or adapted to perform one of the methods described herein, such as, for example, a processing circuit (e.g., a processor communicating with a memory).
[0125] Additional embodiments include a computer on which a computer program or commands for performing one of the methods described herein are installed.
[0126] The embodiments have been described in relation to human action recognition, but those skilled in the art will understand that the above embodiments can be used for individual action recognition, where the individual is used to include inanimate objects such as humans, animals, and robots. The individual can be a living object and / or a moving object that can belong to a particular class, species, or part of a collection.
[0127] The foregoing methods and examples can be implemented within an architecture such as that of FIG. 9, which includes a server (900) communicating via a network (904) (which can be wireless and / or wired), such as the Internet, for data exchange, and one or more client devices (902). The server (900) and the client devices (902) include a data processor (912) (or more simply a processor) (e.g., 912a, 912b, 912c, 912d, 912e) and a memory (913) (e.g., 913a, 913b, 913c, 913d, 913e), such as a hard disk. The client device (902) can be any type of computing device that communicates with the server (900), such as an autonomous vehicle (902b), a robot (902c), a computer (902d), or a mobile phone (902e).
[0128] More specifically, in an example, the training and classification methods according to the illustrations of FIGS. 1-3 and 6 can be performed at the server (900). In other examples, the training and classification methods according to the examples of FIGS. 1-3 and 6 can be performed at the client device (902). In still other examples, the training and classification methods can be performed in a distributed manner across different servers or multiple servers.
[0129] A new approach is disclosed that utilizes implicit pose representations that can be performed by mid-level features learned by a pose detector without using explicit body keypoint coordinates.
[0130] Depending on the operations performed in such an image or video, the function of classifying the image or video is related to multiple technical operations. For example, the present application can index a video based not only on metadata such as the title of the video, but also on the content of the video that may not be related to the content of the video or may be subject to misinterpretation. For example, in the context of indexing a large number of videos for a search application such as a search engine, search results can be provided as a response to a search request based not only on the metadata of the video, but also on the relevance of the search request to the actual content of the video. Further, the present application can provide advertisements related to the content of the video. For example, a video sharing website such as Snow can provide specific advertisements for video content. Alternatively or additionally, when viewing a video on a web page, other videos can be recommended based on the content of the video. Also, in the context of robotics and autonomous driving, the present application can be used to recognize the actions performed by nearby people. For example, a person can interact with a robot that can determine the actions and gestures performed by the person. Also, an autonomous vehicle can detect dangerous human actions including actions classified as dangerous, and the autonomous vehicle can adjust its speed or perform an emergency brake accordingly. Also, embodiments of the present invention can be applied to the field of video games. The user can play without a remote control.
[0131] The method described herein can outperform a neural network-based method that operates on 3D skeleton data or the coordinates of body keypoints in terms of the accuracy of predicting the results for actual 3D action recognition. Implicit pose representations such as deep pose-based features of a pose detector, such as LCR-Net++, can exhibit much better performance than deep features specially trained for action classification.
[0132] The foregoing description is essentially merely exemplary and is not restrictive of the application or use of the present disclosure. The broad teachings of the present disclosure can be implemented in a variety of forms. Accordingly, although the present disclosure includes specific examples, the scope of the present disclosure should not be limited to the specific examples since modifications within the scope of the drawings, the specification, and the claims are possible. It should be understood that one or more steps of the method can be performed in a different order (or simultaneously) without changing the principles of the present disclosure. Also, although each embodiment has been described as having specific features, any one or more of these features described in connection with any embodiment of the present disclosure can be implemented and / or combined with the features of other embodiments even if the combination is not explicitly described. In other words, the described embodiments are not mutually exclusive, and substituting one or more embodiments for other embodiments is within the scope of the present disclosure.
[0133] The disclosed embodiments describe a system and method for generating full-body poses for a person, but this application is also applicable to generating full-body poses through appropriate training and determining the behavior of other types of animals (e.g., dogs, cats, etc.).
[0134] The spatial and functional relationships between components (e.g., modules, circuit elements, semiconductor layers, etc.) are described using a variety of terms including "connected", "cooperated", "coupled", "adjacent", "next", "above", "below", "first". Unless explicitly referred to as "direct", when the relationship between a first and a second element is described in the above disclosure, the relationship can be a direct relationship with no other intermediate elements between the first and second elements, or can be an indirect relationship with one or more intermediate elements existing (spatially or functionally) between the first and second elements. As used herein, the expression "at least one of A, B, and C" should be construed to mean A or B or C using a non-exclusive logical OR, and should not be construed to mean "at least one of A, at least one of B, and at least one of C".
[0135] In the drawings, the direction of an arrow indicated by an arrow generally indicates the flow of information (e.g., data or instruction words) that is the subject of the illustration. For example, if elements A and B exchange various information and the information transmitted from element A to element B is relevant to the illustration, the arrow can point from element A to element B. A unidirectional arrow does not mean that there is no other information transmitted from element B to element A. Also, for the information transmitted from element A to element B, element B can transmit a request for information or an acknowledgment of receipt to element A.
[0136] In this application, terms such as "module" or "controller" can be replaced with the term "circuit", including the following definitions. The term "module" can include an ASIC (Application Specific Integrated Circuit), digital, analog or mixed analog / digital discrete circuit, digital, analog or mixed analog / digital integrated circuit, combinational logic circuit, FPGA (Field Programmable Gate Array), a processor circuit that executes code (shared, dedicated or grouped), a memory circuit that stores code executed by the processor circuit (shared, dedicated or grouped), other suitable hardware components that provide the described functionality, or combinations of some or all of the above, such as in a system-on-chip, or can be a part thereof.
[0137] A module can include one or more interface circuits. In some examples, the interface circuit can include a wired or wireless interface connected to a local area network (LAN), the Internet, a wide area network (WAN), or combinations thereof. The functionality of any module of the present disclosure can be distributed among a number of modules connected via the interface circuit. For example, load regulation can be possible by a plurality of modules. In other examples, a server (also referred to as remote or cloud) module can perform some functions instead of a client module.
[0138] The code used may include software, firmware, and / or microcode, and may refer to programs, routines, functions, classes, data structures, and / or objects. A shared processor circuit includes a single processor circuit that executes some or all of the code in a plurality of modules. A group processor circuit includes a processor circuit that, in combination with additional processor circuits, executes some or all of the code from one or more modules. The expression multiple processor circuit includes multiple processor circuits on discrete die, multiple processor circuits on a single die, multiple cores of a single processor circuit, multiple threads of a single processor circuit, or combinations thereof. The term shared memory circuit includes a single memory circuit that stores some or all of the code of a plurality of modules. The term group memory circuit includes a memory circuit that, in combination with additional memory, stores some or all of the code from one or more modules.
[0139] The term memory circuit is a subset of computer-readable media. The term computer-readable media as used herein does not include transient electrical or electromagnetic signals propagated through a medium (e.g., on a carrier wave), and thus the term computer-readable media can be regarded as tangible and non-transitory. Non-limiting examples of non-transitory and tangible computer-readable media include non-volatile memory circuits (e.g., flash memory circuits, erasable programmable read-only memory circuits, or mask read-only memory circuits), volatile memory circuits (e.g., static random access memory circuits or dynamic random access memory circuits), magnetic storage media (e.g., analog or digital magnetic tape or hard disk drive), and optical storage media (e.g., CD, DVD, or Blu-ray disk).
[0140] The apparatus and methods described in this application can be executed, partially or fully, by a special-purpose computer generated by configuring a general-purpose computer to perform one or more specific functions executed by a computer program. The functional blocks, flowchart components, and other elements described above function as software specifications that can be converted into a computer program by routine work of a skilled technician or programmer.
[0141] The computer program includes processor-executable instructions stored in at least one non-transitory and tangible computer-readable medium. Also, the computer program can include or depend on stored data. The computer program can include a basic input / output system (BIOS) that interacts with the hardware of the special-purpose computer, device drivers that interact with specific devices of the special-purpose computer, one or more operating systems, user applications, background services, background applications, and the like.
[0142] A computer program can include the following: (i) a parsed predicate such as HTML (hypertext markup language), XML (extensible markup language), or JSON (JavaScript Object Notation); (ii) assembly code; (iii) object code generated from source code by a compiler; (iv) source code executed by an interpreter; (v) source code for compilation and execution by a JIT (just-in-time) compiler, etc. For example, the source code can be created using the syntax of languages including C, C++, C#, ObjectiveC, Swift, Haskell, Go, SQL, R, Lisp, Java (registered trademark), Fortran, Perl, Pascal, Curl, OCaml, Javascript (registered trademark), HTML5 (hypertext markup language 5), Ada, ASP (Active Server Pages), PHP (PHP: hypertext preprocessor), Scala, Eiffel, Smalltalk, Erlang, Ruby, Flash (registered trademark), VisualBasic (registered trademark), Lua, MATLAB, SIMULINK, and Python (registered trademark).
Claims
1. A computer-implemented method for recognizing actions performed by an individual, comprising: obtaining, by one or more processors, an image including at least a portion of the individual; generating, by the one or more processors, an implicit representation of a pose of the individual in the image based on the image; determining, by the one or more processors, an action performed by the individual and captured in the image by classifying the implicit representation of the pose of the individual based on an action score; wherein the implicit representation does not include keypoints, does not include skeletal data of a 2D or 3D individual, and includes a sequence of feature vectors; wherein the feature vectors of the sequence of feature vectors are mid-level feature vectors learned by a pose detector; wherein the feature vectors of the sequence of feature vectors are stacked, and the action score is determined based on the stacked feature vectors; A computer-implemented method.
2. The computer-implemented method according to claim 1, further comprising deriving, from the implicit representation, keypoints having 3D coordinates corresponding to a pose of the skeleton of the individual.
3. The computer-implemented method according to claim 1, wherein the step of generating the implicit representation includes generating the implicit representation of the pose of the individual in the image using a first neural network.
4. The computer-implemented method according to claim 3, further comprising training the first neural network to determine at least one of 2D and 3D poses of the individual in the image.
5. The step of determining the action performed by the individual includes the step of determining the action by classifying the implicit representation of the pose using a second neural network, the computer-implemented method according to claim 1. **Claim 6** The computer-implemented method according to claim 5, further comprising the step of training the second neural network to classify a feature vector based on the action of the individual. **Claim 7** The computer-implemented method according to claim 1, wherein the individual includes a person, the pose includes a pose of a person, and the action includes an action performed by the person. **Claim 8** The computer-implemented method according to claim 1, wherein the image includes a sequence of images from a video. **Claim 9** The step of generating the implicit representation includes the step of generating each of the implicit representations based on the image, The computer-implemented method further includes the step of concatenating the implicit representations to generate a concatenated implicit representation, The step of determining the action includes the step of performing convolution on the concatenated implicit representation to generate a final implicit representation, and the step of determining the action performed by the individual by classifying the final implicit representation. The computer-implemented method according to claim 1. **Claim 10** The computer-implemented method according to claim 9, wherein the convolution corresponds to one-dimensional (1D) convolution. **Claim 11** The computer-implemented method according to claim 9, wherein the convolution corresponds to one-dimensional (1D) temporal convolution. **Claim 12** The method further includes the step of determining, by the one or more processors, a candidate box around the individual captured in the image. The step of generating the implicit expression includes the step of generating the implicit expression based on the candidate box. The computer-implemented method according to claim 1.
13. The computer-implemented method according to claim 12, further comprising the step of extracting a tube from the image based on the candidate box by the one or more processors.
14. The step of determining the candidate box includes the step of determining the candidate box using an RPN (regional proposal network) module, and the computer-implemented method according to claim 12.
15. A system, One or more processors, A memory including code that performs a function when executed by the one or more processors, The function includes Obtaining an image including at least a part of an individual, Based on the image, generating an implicit expression of the pose of the individual in the image, Determining an action performed by the individual and captured in the image by classifying the implicit expression of the pose of the individual based on an action score. The implicit expression does not include key points, does not include 2D (two-dimensional) or 3D (three-dimensional) skeleton data of an individual, and includes a sequence of feature vectors. The feature vectors of the sequence of feature vectors are mid-level feature vectors learned by a pose detector. The feature vectors of the sequence of feature vectors are stacked, and the action score is determined based on the stacked feature vectors. System.
Citation Information
Patent Citations
Behavior identification apparatus, behavior identification method and behavior identification program
JP2009134600A
Method and recording medium for contactless input interface with real-time hand pose recognition
KR1020150075648A
Active learning method for temporal action localization in untrimmed videos
US10726313B2