Skeleton-based action recognition method, processor and program

KR103022137B1Active Publication Date: 2026-09-29SUPERGATE CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
KR1020230130574
Authority / Receiving Office
KR · KR
Patent Type
Patents
Current Assignee / Owner
Filing Date
2023-09-27
Publication Date
2026-09-29
Estimated Expiration
2043-09-27

Smart Images

  • Figure 112023107797246-PAT00002_ABST
    Figure 112023107797246-PAT00002_ABST
Patent Text Reader

Abstract

The present invention may include a skeleton-based behavior recognition method for a processor comprising: a step of acquiring a behavior video capturing the behavior of an object; a step of generating first classification data reflecting the temporal characteristics of the object by tracking the behavior of the object in the behavior video; a step of generating second classification data reflecting the spatial characteristics of the object to be recognized by distinguishing between the object to be recognized and an object other than the object to be recognized in the behavior video; and a step of recognizing the behavior of the object based on the first classification data and the second classification data.
Need to check novelty before this filing date? Find Prior Art

Description

Technology Field

[0001] The present invention relates to a skeleton-based behavior recognition method, a processor, and a computer program. Background Technology

[0002] The technology to accurately recognize various human behaviors has been considered an important research topic in the field of computer vision. This enables applications in diverse areas such as security systems, games, and virtual character motion, and recent research is recognizing human behavior in real time using skeleton-based behavior recognition methods.

[0003] Skeleton-based behavior recognition is a method that represents a person in the form of a skeleton and classifies behaviors based on the movement of that skeleton.

[0004] However, while conventional methods are intuitive and effective, they ignore the spatial characteristics of actions, which leads to the problem of being unable to distinguish between similar skeleton movements.

[0005] Due to the limitations of such simple behavior recognition, there are restrictions in distinguishing and recognizing various complex behaviors (lying down, falling down).

[0006] Therefore, to overcome the limitations of conventional skeleton-based behavior recognition methods, behavior recognition technology that considers spatial and temporal characteristics is required. The problem to be solved

[0007] The present invention aims to provide a skeleton-based behavior recognition method, a processor, and a program.

[0008] Specifically, the purpose is to provide a skeleton-based behavior recognition method, processor, and program that consider spatial and temporal characteristics. means of solving the problem

[0009] A skeleton-based behavior recognition method for a processor according to the present invention for solving the above technical problem may include the steps of: acquiring a behavior video capturing the behavior of an object; tracking the behavior of an object in the behavior video to generate first classification data reflecting the temporal characteristics of the object; distinguishing between an object to be recognized and an object other than the object to be recognized in the behavior video to generate second classification data reflecting the spatial characteristics of the object to be recognized; and recognizing the behavior of the object based on the first classification data and the second classification data.

[0010] Additionally, the step of generating the second classification data may generate second classification data of the object using an algorithm, and the algorithm may be composed of a first partial algorithm that spatially classifies objects within the behavioral video and a second partial algorithm that individually recognizes objects within the behavioral video.

[0011] In addition, the second classification data may include predicted values ​​that classify all individual objects at the pixel level within the behavioral video.

[0012] In addition, the step of recognizing the behavior of the object can recognize the behavior of the object by considering the correlation between the object to be recognized and surrounding objects.

[0013] Additionally, the step of generating the first classification data may further include the step of generating instances of objects in the behavior video based on object information obtained by tracking objects in the behavior video, and the step of generating the first classification data by classifying the instances using a previously trained behavior classification model.

[0014] In addition, the step of generating the instance can recognize objects within the action video, generate skeleton data according to the pose of each object, and track objects in each frame of the action video based on the skeleton data for each object.

[0015] In addition, the above instance may include coordinate information regarding the temporal change of frame-by-frame skeleton data within the action video.

[0016] In addition, the first classification data may include predicted values ​​of the object's behavior according to temporal changes. Effects of the invention

[0017] The present invention can accurately recognize the behavior of an object by considering temporal and spatial characteristics. Brief explanation of the drawing

[0018] FIG. 1 is a conceptual diagram illustrating a skeleton-based behavior recognition method according to one embodiment of the present invention. FIG. 2 is a flowchart illustrating a skeleton-based behavior recognition method according to an embodiment of the present invention. FIG. 3 is an exemplary diagram showing a behavioral video according to one embodiment of the present invention. FIG. 4 is a flowchart illustrating a skeleton-based behavior recognition method according to one embodiment of the present invention in more detail. FIG. 5 is an exemplary diagram illustrating a method for generating skeleton data using a neural network model according to an embodiment of the present invention. FIG. 6 is a conceptual diagram of an algorithm according to one embodiment of the present invention. FIG. 7 is an example diagram showing a behavioral image recognized through an algorithm according to an embodiment of the present invention. FIG. 8 is a block diagram showing the configuration of a processor according to one embodiment of the present invention. Specific details for implementing the invention

[0019] The following description merely illustrates the principles of the invention. Therefore, those skilled in the art may invent various devices that embody the principles of the invention and are included within the concept and scope of the invention, even if they are not explicitly described or illustrated in this specification. Furthermore, all conditional terms and embodiments listed in this specification are, in principle, explicitly intended only for the purpose of enabling an understanding of the concept of the invention and should be understood as not being limited to the embodiments and conditions specifically listed elsewhere.

[0020] The aforementioned objectives, features, and advantages will become clearer through the following detailed description in conjunction with the attached drawings, and accordingly, a person skilled in the art to which the invention pertains will be able to easily implement the technical concept of the invention.

[0021] In addition, in describing the invention, if it is determined that a detailed description of known technology related to the invention may unnecessarily obscure the essence of the invention, such detailed description will be omitted. Hereinafter, preferred embodiments of the present invention will be described in detail with reference to the attached drawings.

[0022] FIG. 1 is a conceptual diagram illustrating a skeleton-based behavior recognition method according to one embodiment of the present invention.

[0023] Referring to FIG. 1, a camera (20) can capture the actions of at least a plurality of objects (10) and transmit the captured action video (or a plurality of action images) to a processor (100). At this time, at least a portion of the object may be captured in the action video (or action image) so that the action of the object can be recognized. At this time, the camera (20) may be a 3D depth camera, and in this case, the camera (20) may map depth data corresponding to the action video and transmit it to the processor (100).

[0024] Here, the object refers to an object that is the subject of behavior recognition and is capable of acting, such as a person or an animal, and the present invention describes the object based on a person.

[0025] And, the processor (100) that receives the behavior video can recognize the behavior of an object using the first classification data and spatial information within the behavior video. In this regard, the method of recognizing the behavior of the processor (100) is described in detail with reference to FIG. 2.

[0026] FIG. 2 is a flowchart illustrating a skeleton-based behavior recognition method according to an embodiment of the present invention.

[0027] Referring to FIG. 2, the processor (100) can acquire a behavior video or a plurality of behavior images that capture the behavior of a plurality of objects (S100). Here, the plurality of objects may refer to an object that is the target of behavior recognition and an object other than the target of behavior recognition (e.g., background).

[0028] For example, the processor (100) may receive a video of multiple objects being filmed as in (a) of FIG. 3, or receive multiple images of multiple actions (31, 33, 35) from a camera as in (b) of FIG. 3.

[0029] Meanwhile, the processor (100) may receive action videos or multiple action images from a database or an external server in addition to the camera (20).

[0030] Next, the processor (100) can track the behavior of an object in a behavior video to generate first classification data that reflects the temporal characteristics of the object (S200). At this time, the processor (100) can track the object by generating skeleton data of the object. This is explained with further reference to FIGS. 4 and FIG. 5.

[0031] FIG. 4 is a flowchart illustrating a skeleton-based behavior recognition method according to one embodiment of the present invention in more detail.

[0032] FIG. 5 is an exemplary diagram illustrating a method for generating skeleton data using a neural network model according to an embodiment of the present invention.

[0033] Referring to FIG. 4, the processor (100) can generate an instance of an object in a behavioral video based on object information obtained by tracking an object in a behavioral video (S210).

[0034] Specifically, the processor (100) can recognize objects in action video in chronological order for each frame using a neural network model and generate key points according to the pose of each object within the frame. Here, the neural network model can be trained to distinguish and recognize objects within the frames of the action video, extract object key points (or feature points) for each part of the object, and generate skeleton data consisting of key points (or feature points) and lines connecting the key points.

[0035] Here, the neural network model is based on a deep learning convolutional neural network (CNN) or a transformer-based neural network, and a library capable of extracting key points of various objects in real time from a photo can be used.

[0036] For example, the processor (100) inputs a behavioral video frame by frame in chronological order to a pre-trained neural network model, and the neural network model can extract key points corresponding to the poses of objects within each frame of the behavioral video. For example, the key points may include feature parts of the human body such as the nose, eyes, ears, shoulders, elbows, wrists, hips, knees, and ankles.

[0037] In addition, skeleton data corresponding to each object can be generated by setting some central key points among multiple key points, establishing a skeleton central axis centered on the set central key points, and connecting the key points.

[0038] For example, as shown in FIG. 5, the processor (100) can sequentially input action images (51) into a neural network model frame by frame to generate skeleton data (54) corresponding to an object (53) and generate an action image (52) that reflects the skeleton data (54). At this time, if there are multiple objects to be recognized, skeleton data can be generated for each object.

[0039] That is, the processor (100) can recognize objects in each frame of the action video and generate skeleton data corresponding to the pose of each object.

[0040] And, the processor (100) can track the skeleton data of each object in each frame of the action video by comparing the distance traveled by the skeleton data of each object in each frame of the action video, and can create an instance by combining the object information obtained through tracking in chronological order for each object in the action video. At this time, the processor (100) can track the skeleton data of each object by considering the distance between key points in the skeleton data and the position of the central axis, etc. Here, the object information may include information about the skeleton data corresponding to the object, the action video frame number containing the skeleton data, etc.

[0041] Additionally, the instance is data formed by combining frame-by-frame skeleton data within the action video in chronological order, and may be a set of coordinate information that can display the frame-by-frame skeleton data on the screen.

[0042] For example, an instance may contain frame numbers (frame_number) in chronological order of frames, coordinate information for skeleton data within the frame (e.g., human body keypoints (L-eye, R-eye, neck ...), and ID information for each object). When visualizing such coordinate information within an instance, it can be displayed in the form of a graph.

[0043] In other words, an instance is data representing the behavioral patterns of objects within a behavioral video, consisting of coordinate information regarding the temporal changes in the skeleton data of objects frame by frame within the video.

[0044] And, the processor (100) can generate first classification data by classifying instances using a previously learned behavior classification model (S220).

[0045] Specifically, the processor (100) can input instances of each object within a behavior video into a behavior classification model, either by object or by inputting the entire instance, to produce a behavior prediction value corresponding to the instance as first classification data. Here, the behavior prediction value is a probability value indicating which behavior of the object corresponds to the instance, and the first classification data may include behavior prediction values ​​that predict the behavior of the object in succession according to temporal change.

[0046] At this time, the behavior classification model may be a model trained to produce the behavior prediction value of an instance as a classification result using past instances and the past behavior classification results of those past instances as a single training dataset.

[0048] Again, referring to FIG. 2, the processor (100) can distinguish between objects to be recognized and objects not to be recognized in an action video and generate second classification data that reflects the spatial characteristics of the objects to be recognized (S300). Here, objects not to be recognized refer to all objects excluding the objects to be recognized in action, and may include, for example, backgrounds such as the sky, trees, floor, and buildings, or other objects.

[0049] Specifically, the processor (100) can generate second classification data of objects by using an algorithm that classifies objects, including the boundaries of the objects, and identifies the class of objects and other objects for each pixel in the image that are subject to action recognition. Here, the algorithm may consist of a first sub-algorithm that spatially classifies objects in an action video and a second sub-algorithm that individually recognizes objects in an action video.

[0050] Specifically, the first part algorithm can classify classes for every pixel in the action video based on spatial differences. In this case, the first part algorithm can segment regions of an uncountable class (stuff class) without segmenting overlapping parts. For example, the first part algorithm can be applied to the background within the action video.

[0051] In other words, the first part of the algorithm can apply an Auto Encoder method using U-NET for Semantic Segmentation.

[0052] Furthermore, the second part algorithm can individually separate and classify each object within the action video. In this case, the second part algorithm can identify independent objects and capture tight pixel boundaries to segment regions into countable classes (Thing class). For example, the second part algorithm can be applied to objects within the action video that are the subject of action recognition.

[0053] In other words, the second part of the algorithm is an instance segmentation method, with Mask R-CNN being a major model example; it uses detection techniques to acquire specific regions and predict a mask for each specific region.

[0054] Next, the algorithm can generate second classification data from the behavioral video by combining the region segmented and classified by the first partial algorithm and the region segmented and classified through the second partial algorithm.

[0055] In other words, the algorithm can be panoptic segmentation that integrates semantic segmentation and instance segmentation.

[0056] Meanwhile, regarding the algorithm, further explanation is provided with reference to FIGS. 6 and FIGS. 7.

[0057] FIG. 6 is a conceptual diagram of an algorithm according to one embodiment of the present invention.

[0058] FIG. 7 is an example diagram showing a behavioral image recognized through an algorithm according to an embodiment of the present invention.

[0059] Referring to FIGS. 6 and FIGS. 7, when an action video (61) is input, the algorithm (62) can spatially divide and classify all objects within the action video (61) through the first partial algorithm, such as the first classified action video (a in FIG. 7). At this time, the first classified action video (a in FIG. 7) can divide and classify backgrounds (walls, floors) with objects to be recognized, but may not classify overlapping objects (barbell racks and barbells, etc.).

[0060] Additionally, the algorithm (62) can individually divide and classify all objects within the behavior video (61), such as the second classified behavior video (b in FIG. 7) through the second partial algorithm.

[0061] In this case, the second classified action video (Fig. 7b) can classify independent objects (barbell, person exercising, barbell, etc.) by dividing them individually, but the background (wall, floor, etc.) may not be classified by dividing it.

[0062] And, the algorithm (62) can combine the first classified action image (Fig. 7a) and the second classified action image (Fig. 7b) from the first part algorithm and the second part algorithm to classify classes on a pixel basis for objects that are targets for action recognition within the action image and other objects, as in Fig. 7c, while simultaneously distinguishing individual objects, and can generate second classification data using this.

[0063] Here, the second classification data is data containing spatial characteristics within the action video, and may include classification prediction values ​​that classify all individual objects within the action video at the pixel level. Here, the classification prediction values ​​may be probability values ​​indicating what type of objects are within the action video.

[0064] Meanwhile, in addition to the method of generating second classification data containing spatial characteristics of the behavior video through segmentation as described above, the processor (100) may generate second classification data containing spatial characteristics of the behavior video using depth data corresponding to the behavior video. Here, the depth data may be obtained from the camera (20).

[0065] For example, the processor (100) can generate second classification data containing spatial characteristics of the action video by distinguishing between objects that are subject to recognition and objects that are not recognized within the action video, using depth data corresponding to the action video.

[0067] Again, referring to FIG. 2, the processor (100) can recognize the behavior of an object in a behavior video based on first classification data and second classification data (S400). At this time, the processor (100) can recognize the behavior of an object by considering the correlation between the object to be recognized and surrounding objects. At this time, surrounding objects may be the background in the behavior video.

[0068] For example, the processor (100) can determine that the object is not sitting but is exercising (squatting posture) by considering the correlation between the object to be recognized and surrounding objects based on the first classification data and the second classification data corresponding to the action video of FIG. 3.

[0069] Additionally, the processor (100) can distinguish between skydiving behavior and lying behavior.

[0070] That is, the processor (100) can perceive the same action differently depending on the surrounding objects.

[0071] Additionally, the processor (100) can determine an interaction object among the objects in the behavioral video according to a reference condition and recognize the behavior of the object to be recognized by considering the interaction between the object to be recognized and the interaction object. At this time, the reference condition is a criterion for determining the interaction object, and may be, for example, whether there is contact with the object to be recognized, within a predetermined distance, whether there is a linked action, etc.

[0073] Next, the configuration of the processor (100) is described.

[0074] FIG. 8 is a block diagram showing the configuration of a processor (100) according to one embodiment of the present invention.

[0075] Referring to FIG. 8, the processor (100) may include some or all of the communication unit (110), the first classification data generation unit (120), the second classification data generation unit (130), the control unit (140), and the storage unit (150).

[0076] The communication unit (110) can transmit and receive various data required by the processor (100).

[0077] Specifically, the communication unit (110) can receive action videos or action images from a camera or an external database.

[0078] The first classification data generation unit (120) can generate first classification data that reflects the temporal characteristics of an object by tracking the action of an object in an action video. At this time, the first classification data generation unit (120) may use a neural network or an action classification model.

[0079] The second classification data generation unit (130) can generate second classification data that reflects the spatial characteristics of the object by distinguishing between the object and the background other than the object in the action video. At this time, the second classification data generation unit (120) may use an algorithm.

[0080] The control unit (150) can control the overall operation of the processor (100).

[0081] Specifically, the control unit (150) can control the communication unit (110) to receive action videos or action images.

[0082] In addition, the control unit (150) can recognize the behavior of an object in a behavior video using the first classification data and the second classification data.

[0083] The storage unit (150) can store various data required by the processor (100).

[0084] Specifically, the storage unit (150) can store various neural networks and behavior classification models.

[0085] Additionally, the storage unit (150) may store a program recorded on a computer-readable recording medium on which program code for executing a skeleton-based behavior recognition method is recorded.

[0086] Meanwhile, according to the present invention, by considering temporal and spatial characteristics, the behavior of an object can be accurately recognized.

[0088] Furthermore, the various embodiments described herein may be implemented, for example, in a recording medium readable by a computer or similar device using software, hardware, or a combination thereof.

[0089] According to hardware implementation, the embodiments described herein may be implemented using at least one of ASICs (application specific integrated circuits), DSPs (digital signal processors), DSPDs (digital signal processing devices), PLDs (programmable logic devices), FPGAs (field programmable gate arrays), processors, controllers, microcontrollers, microprocessors, and other electrical units for performing functions. In some cases, the embodiments described herein may be implemented as the control module itself.

[0090] According to software implementation, embodiments such as the procedures and functions described herein may be implemented in separate software modules. Each of the software modules may perform one or more functions and operations described herein. Software code may be implemented as a software application written in a suitable programming language. The software code may be stored in a memory module and executed by a control module.

[0091] The above description is merely an illustrative explanation of the technical concept of the present invention, and those skilled in the art to which the present invention pertains will be able to make various modifications, changes, and substitutions within the scope of the essential characteristics of the present invention without departing from its nature.

[0092] Accordingly, the embodiments disclosed in this invention and the accompanying drawings are intended to illustrate, not limit, the technical concept of the invention, and the scope of the technical concept of the invention is not limited by such embodiments and accompanying drawings. The scope of protection of this invention shall be interpreted by the claims below, and all technical concepts within an equivalent scope shall be interpreted as being included within the scope of rights of this invention.

Claims

Claim 1 A skeleton-based behavior recognition method for a processor comprises: a step of acquiring a behavior video capturing the behavior of an object; a step of generating first classification data reflecting the temporal characteristics of the object by tracking the behavior of the object in the behavior video; a step of generating second classification data reflecting the spatial characteristics of the object to be recognized by distinguishing between the object to be recognized and objects other than the object to be recognized in the behavior video; and a step of recognizing the behavior of the object based on the first classification data and the second classification data. The step of recognizing the behavior of the object is characterized by recognizing the behavior of the object, but recognizing the same behavior of the object differently by considering the correlation between the object and the objects other than the object to be recognized, and determining the objects to be interacted according to criteria conditions including whether there is contact with the object to be recognized among the objects other than the object to be recognized, within a predetermined distance, and whether there is a linked operation. Claim 2 A behavior recognition method according to claim 1, wherein the step of generating the second classification data generates the second classification data of the object using an algorithm, and the algorithm is composed of a first partial algorithm that spatially classifies the object within the behavior video and a second partial algorithm that individually recognizes the object within the behavior video. Claim 3 A behavior recognition method according to claim 1, characterized in that the second classification data includes predicted values ​​that classify all individual objects in a behavior video at the pixel level. Claim 4 delete Claim 5 The behavior recognition method according to claim 1, further comprising the step of generating first classification data, the step of generating instances of objects within the behavior video based on object information obtained by tracking objects within the behavior video; and the step of generating first classification data by classifying the instances using a previously trained behavior classification model. Claim 6 In claim 5, the step of generating the instance is characterized by recognizing an object in the action video, generating skeleton data according to the pose of each object, and tracking the object for each frame in the action video based on the object-specific skeleton data. Claim 7 A behavior recognition method according to claim 5, characterized in that the instance includes coordinate information regarding temporal changes of frame-by-frame skeleton data within a behavior video. Claim 8 A behavior recognition method according to claim 1, characterized in that the first classification data includes predicted behavior values ​​of an object according to temporal changes. Claim 9 A program stored on a computer-readable recording medium comprising program code for executing a behavior recognition method according to any one of claims 1 to 3 and claims 5 to 8. Claim 10 A computer-readable recording medium storing a program that performs an action recognition method according to any one of claims 1 to 3 and claims 5 to 8.

Citation Information

Patent Citations

  • Apparatus and method for recognizing movement of object

    KR1020200056602A