Interaction recognition program for multiple objects using association between individual behavior feature and interaction feature

The multi-object interaction recognition program addresses the challenge of recognizing interactions between multiple objects by using a program that detects and tracks objects, updates groups based on behavior features, and predicts interaction characteristics, thereby enhancing recognition accuracy.

WO2025121952A1PCT designated stage expired Publication Date: 2025-06-12SAFEMOTION INC
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2024/019975
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-08
Filing Date
2024-12-06
Publication Date
2025-06-12

AI Technical Summary

Technical Problem

Existing technologies struggle to efficiently recognize and analyze interactions between multiple objects in CCTV footage, as they can only recognize individual actions and lack the capability to update object groups for interaction recognition.

Method used

A multi-object interaction recognition program that detects and tracks human objects, extracts skeleton information, generates relationship graphs, updates object groups based on individual behavior features, and predicts interaction characteristics using updated group features.

Benefits of technology

The program significantly improves the accuracy of interaction recognition by learning individual and interactive actions simultaneously, allowing for the updating of object groups and enhancing recognition accuracy over time.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2024019975_12062025_PF_FP_ABST
    Figure KR2024019975_12062025_PF_FP_ABST
Patent Text Reader

Abstract

The present invention relates to an interaction recognition program for multiple objects and, more specifically, to an interaction recognition program for multiple objects, wherein the program, which is stored in a computer-readable recording medium, can recognize each of individual actions and interactions of human objects in a video, and can increase the accuracy of interaction recognition with a structure for learning the individual actions together when recognizing the interactions while updating a group of the objects for the interaction recognition.
Need to check novelty before this filing date? Find Prior Art

Description

A multi-object interaction recognition program using the correlation between individual and interaction features.

[0001] The present invention relates to a multi-object interaction recognition program, and more specifically, to a multi-object interaction recognition program stored in a computer-readable recording medium, which can recognize individual actions and interaction actions of human objects in an image, and which can increase the accuracy of interaction recognition by recognizing a structure in which individual actions are learned together with the group of objects for interaction recognition while recognizing interaction actions.

[0002] Behavior recognition is a technology that recognizes user actions and target actions by continuously observing human actions and their environmental conditions.

[0003] Since the 1980s, behavior recognition has made significant progress as a personal assistance technology, and is particularly utilized in medicine, social sciences, and human-computer interaction. Due to its multifaceted applicability, it is sometimes referred to by various names, such as intent recognition and location-based services, depending on the application.

[0004] Meanwhile, CCTV installation is mandatory in multi-use facilities such as public transportation facilities, performance halls, exhibition halls, public educational facilities, and childcare facilities.

[0005] In places with such a large number of people, the amount of CCTV footage is not only enormous, but the number of people appearing in the footage is also significant, so it takes a tremendous amount of time and money for people to analyze the footage and analyze the behavior of the people in the footage.

[0006] Therefore, a method is needed to analyze CCTV footage more efficiently, and a method is needed to recognize and track the interactions of people appearing in the footage.

[0007] Korean Patent Publication No. 10-2022-0142673, 'LSTM-based action recognition method using human body joint coordinate system' discloses a technology for an LSTM-based action recognition method that can perform action recognition through a learned stacked LSTM by extracting feature vectors having spatiotemporal feature information from the human body joint coordinate system of a human object detected in an image.

[0008] This technology has the advantage of being able to show a high recognition rate for the target's actions even in various motion environments by generating feature vectors based on direction vectors and motion vectors that have spatiotemporal information about movement using a normalized human joint coordinate system, training these in a stacked LSTM, and then performing action recognition. However, it can only recognize individual actions of objects in an image, and there is no configuration for recognizing multiple objects interacting with each other.

[0009] In addition, Korean Patent Publication No. 10-2022-0078893, 'Device and method for recognizing human behavior in a video' discloses a technology that can detect abnormal behavior of a person in a CCTV video by analyzing the CCTV video in real time, thereby preventing safety accidents through detection of unauthorized intrusion, detection of abnormal behavior such as violence, etc., and can recognize interactive behaviors between multiple people by combining the continuous postures of each of multiple human objects detected for an image frame.

[0010] However, this technology can only recognize interactions between objects in a single image frame, and there is no configuration that can recognize interactions in situations where the relationships between individual objects change.

[0011] The present invention has been devised to solve the above-described problem, and the purpose of the present invention is to provide a multi-object interaction recognition program capable of recognizing both individual and interactive actions of objects within an image.

[0012] In addition, the purpose of the present invention is to provide a multi-object interaction recognition program that can greatly improve the accuracy of recognition by learning not only the interaction but also individual actions while updating the object group for recognizing the interaction of objects within an image.

[0013] In order to achieve the above object, the present invention provides a computer comprising: an object detection and tracking means for detecting a human object from an input image, assigning an identifier thereto, and tracking the object; a skeleton extraction means for extracting skeleton information by detecting joint parts of the detected object; a group generation means for generating a relationship graph including vertex information, which is position information of the detected object, and edge information, which is line information connecting vertex information of neighboring objects, and for setting objects connected by a single edge as a single group; a visual feature extraction means for extracting visual features using pixel values ​​of the detected object; an individual behavior recognition means for extracting individual behavior features, which are behavioral features of each object, using the skeleton information and the visual features, and predicting individual behavior using the extracted individual behavior features; a group update means for calculating a probability that objects belong to a predetermined group using the individual behavior features of objects included in the predetermined group, and deleting objects whose probability is below a predetermined probability threshold to update the group; And the present invention provides a multi-object interaction recognition program stored in a computer-readable recording medium, characterized in that it functions as an interaction recognition means that predicts interaction characteristics, which are behavioral characteristics between objects, by using individual behavioral characteristics of each object in the updated group.

[0014] In a preferred embodiment, the group generating means converts positional information of an object in a two-dimensional image into coordinates of actual positional information in a three-dimensional space to generate vertex information, and connects objects located within a preset actual distance to each other to generate edge information.

[0015] In a preferred embodiment, the group creation means sets objects connected by one edge as one group, and among the set groups, groups in which objects belong to the same group are duplicated to remove duplicates, leaving only one group and deleting the remaining groups.

[0016] In a preferred embodiment, the individual action recognition means extracts the individual action features using a backbone network that inputs skeletal information and visual features of an object, and inputs the extracted individual action features into a personal action classifier (personal_action_classifier) ​​to predict the individual action.

[0017] In a preferred embodiment, the visual feature extraction means extracts visual features using a feature extraction network such as ResNet or HRNet.

[0018] In a preferred embodiment, the group updating means calculates, for each preset group, the probability that objects belonging to each group belong to the group, inputs individual behavioral features of objects belonging to the group into a multi-layer perceptron (MLP) neural network to transform the features, creates the transformed individual behavioral features as one group feature through pooling, connects the created group feature to the individual behavioral feature of each object to create an individual behavioral feature to which the group feature is connected, and then inputs the individual behavioral feature connected to the group feature into the MLP neural network to transform the feature and calculates the probability that the transformed individual behavioral feature belongs to the group.

[0019] In a preferred embodiment, among the updated groups, groups with identical objects are deduplicated to leave only one group and the remaining groups are deleted.

[0020] In addition, the present invention further provides a method for recognizing interaction of multiple objects in an image, wherein the method comprises the steps of: detecting a human object in an input image, assigning an identifier thereto, and tracking it; detecting joints of the detected object to extract skeleton information; generating a relationship graph including vertex information, which is location information of the detected object, and edge information, which is line information connecting vertex information of neighboring objects, and setting objects connected by a single edge as a single group; extracting visual features using pixel values ​​of the detected object; extracting individual behavioral features, which are behavioral features of each object, using the skeleton information and the visual features, and predicting individual behaviors using the extracted individual behavioral features; calculating a probability that objects belong to a predetermined group using the individual behavioral features of the objects included in the predetermined group, and deleting objects whose values ​​are below a predetermined probability threshold, thereby updating the group; and predicting interactional features, which are behavioral features between the objects, using the individual behavioral features of each object in the updated group.

[0021] The present invention has the following excellent effects.

[0022] According to the multi-object interaction recognition program of the present invention, individual actions and mutual actions of objects within an image can be recognized together, and when recognizing mutual actions, the initially set group can be updated and recognized, so that the accuracy of recognition can be greatly improved.

[0023] In addition, since it is possible to recognize mutual and individual actions in conjunction, the accuracy of recognition can be improved as learning is repeated together.

[0024] FIG. 1 is a diagram showing the configuration of a multi-object interaction recognition program according to one embodiment of the present invention;

[0025] FIG. 2 is a drawing for explaining a group creation means of a multi-object interaction recognition program according to one embodiment of the present invention;

[0026] FIG. 3 is a drawing for explaining a group update means of a multi-object interaction recognition program according to one embodiment of the present invention;

[0027] FIG. 4 is a drawing for explaining a group probability calculator of a group update means of a multi-object interaction recognition program according to one embodiment of the present invention.

[0028] FIG. 5 is a drawing for explaining an interaction recognition means of a multi-object interaction recognition program according to one embodiment of the present invention.

[0029] [Explanation of symbols]

[0030] 100: Object detection and tracking means 200: Skeleton extraction means

[0031] 300: Group creation means 310: Relationship graph generator

[0032] 320: Group setting device 400: Visual feature extraction device

[0033] 500: Individual action recognition means 600: Group update means

[0034] 610, 610a, 610b: Group Probability Calculator 620: Group Updater

[0035] 700: Means of Recognizing Interaction

[0036] The terms used in the present invention are selected from the most widely used general terms as much as possible, but in certain cases, there are terms arbitrarily selected by the applicant. In such cases, the meaning of the terms should be understood by considering the meaning described or used in the detailed description of the invention, rather than the simple name of the term.

[0037] Hereinafter, the technical configuration of the present invention will be described in detail with reference to preferred embodiments illustrated in the attached drawings.

[0038] However, the present invention is not limited to the embodiments described herein and may be embodied in other forms. Like reference numbers designate like elements throughout the specification.

[0039] FIG. 1 is a diagram showing the configuration of a multi-object interaction recognition program according to one embodiment of the present invention.

[0040] Referring to FIG. 1, a multi-object interaction recognition program according to one embodiment of the present invention is a program capable of recognizing both individual and interactive actions of human objects in an input image, and can repeatedly learn individual and interactive actions by updating a group of objects for interaction recognition, thereby increasing the accuracy of recognition.

[0041] Here, individual behavior means when a child and a teacher hug each other, for the child it is 'receiving a hug' and for the teacher it is 'giving a hug', and the reciprocal behavior is 'hugging'.

[0042] Also, as another example, when a teacher feeds a child, the child's individual action can be 'receiving food', the teacher's individual action can be 'feeding', and the reciprocal action can be 'feeding'.

[0043] In this way, the present invention is a program that can recognize individual behavior of each object in an image and interactive behavior between objects.

[0044] In addition, the multi-object interaction recognition program of the present invention is stored in a computer and performs the interaction recognition method by causing the computer to function, and the computer is a broad device including not only a general personal computer but also an embedded system, a smart device, etc.

[0045] In addition, the above-mentioned interactive recognition program may be provided by being stored on a separate recording medium, and the recording medium may be one that is specially designed and configured for the present invention or one that is known and available to a person having ordinary knowledge in the computer software field.

[0046] For example, the recording medium may be a hardware device specifically configured to store and execute program instructions, either singly or in combination, such as a magnetic medium such as a hard disk, a floppy disk, and a magnetic tape, an optical recording medium such as a CD or a DVD, a magneto-optical recording medium capable of both magnetic and optical recording, a ROM, a RAM, a flash memory, etc.

[0047] In addition, the above-mentioned interaction recognition program may be a program composed of program commands, local data files, local data structures, etc. alone or in combination, and may be a program written in high-level language code that can be executed by a computer using an interpreter or the like, as well as machine language code created by a compiler.

[0048] Meanwhile, the above-mentioned interaction recognition program is not provided by being stored on a separate removable storage medium, but is stored on a server system and then transmitted to the computer via a communication network and installed to perform the interaction recognition method.

[0049] Below, the functions of the multi-object interaction recognition program of the present invention are described in detail.

[0050] The above multi-object interaction recognition program causes a computer to function as an object detection and tracking means (100), a skeleton extraction means (200), a group creation means (300), a visual feature extraction means (400), an individual action recognition means (500), a group update means (600), and an interaction recognition means (700).

[0051] The above object detection and tracking means (100) detects a human object in an input image (10) and tracks it by assigning an identifier.

[0052] In addition, the image (10) above means a frame image within a video, and the object is not necessarily limited to a person, and can be applied to moving objects such as animals.

[0053] In addition, object detection within the image (10) can utilize various known methods such as Yolo (you only look once), SSD (Single Shot Multibox Detector), and Faster R-CNN, and a process of tracking the object is performed after object detection. In addition, object tracking can utilize techniques such as Sort (Simple Online and Realtime Tracking) and ByteTracker.

[0054] In addition, in Fig. 1, it is assumed that four human objects (1, 2, 3, 4) exist in the image (10) and are detected and tracked as bounding boxes.

[0055] The above skeleton extraction means (200) cuts out the bounding box area of ​​the detected object and extracts skeleton information (S) of the object within the box area.

[0056] Here, skeleton information (S) is position information regarding the joint position (a) of the object.

[0057] The above group creation means (300) uses the skeleton information of the detected object to create a relationship graph including vertex information, which is the location information of the object, and edge information, which is the line information connecting the vertex information of the neighboring object, and sets objects connected by one edge as one group (G).

[0058] Referring to FIG. 2, a more detailed description is given. FIG. 2 shows the group creation means (300), which includes a relationship graph generator (310) and a group setting device (320).

[0059] The above relationship graph generator (310) calculates the vertex information and the edge information and generates a relationship graph which is a set of calculated information.

[0060] In addition, taking the image (10) as an example, there are four objects (1, 2, 3, 4) in the image (10), and the relationship graph generator (310) generates vertex information (v1, v2, v3, v4) and edge information (e) of each object. 12 ,e 13 ,e 23 ,e 24 ) to generate a relationship graph.

[0061] Here, vertex information is calculated using the position information included in the skeleton information of each object, and can be calculated as a point between the positions of the object's two feet, for example.

[0062] However, the above vertex information can be calculated as an average position of all joint positions of the object, etc.

[0063] In addition, it is preferable that the relationship graph generator (310) reduce positional errors by converting the positional information of an object in a two-dimensional image into coordinates of actual positional information in a three-dimensional space and generating the vertex information.

[0064] Here, the coordinate transformation can use a pixel distance measurement method, a 3D distance measurement method using multiple cameras, and a 3D distance measurement method using a single camera. In the present invention, a 3D distance measurement method using a single camera is used. This method is a method of creating a coordinate system through calibration between the floor surface of the real world and the camera, and calculating the position by projecting the position between the positions of the two feet of the detected object onto the coordinate system. In addition, the 3D distance measurement method using the single camera can be referred to Korean Patent No. 10-2555645, 'Method for Calibrating a Coordinate System Using a Laser Pointer'.

[0065] In addition, the edge information is information of a line connecting the vertex information of the objects, and in the example of the image (10), the first edge information (e) connecting the vertex information (v1) of the first object (1) and the vertex information (v2) of the second object (2)12 ), second edge information (e) connecting the vertex information (v1) of the first object (1) and the vertex information (v3) of the third object (3) 13 ), third edge information (e) connecting the vertex information (v2) of the second object (2) and the vertex information (v3) of the third object (3) 23 ), the fourth edge information (e) connecting the vertex information (v2) of the second object (2) and the vertex information (v4) of the fourth object (4) 24 ) is generated.

[0066] The above group setting device (320) sets objects connected to one edge among the objects (1, 2, 3, 4) of the image (10) as one group.

[0067] Here, being connected by one edge means that there is one edge information between objects, and taking the first object (1) as an example, the first object (1) and the second object (2) have the first edge information (e 12 ) are connected as one, and the third object (3) and the second edge information (e 13 ) are connected as one. That is, based on the first object (1), the first object (1), the second object (2), and the third object (3) are set as one group (G1).

[0068] However, the first object (1) has first edge information (e) with the fourth object (4) 12 ) and the fourth edge information (e 24 ) are connected by two edges, so they cannot be set as one group.

[0069] In the same way, based on the second object (2), the second object (2) is set as one group (G2) with the first object (1), the third object (3), and the fourth object (4), based on the third object (3), the third object (3) is set as one group (G3) with the first object (1) and the second object (2), and based on the fourth object (4), the fourth object (4) is set as one group (G4) with the second object (2).

[0070] In this way, four groups (G1, G2, G3, G4) are set based on each object.

[0071] Meanwhile, in the case of the first group (G1) and the third group (G3), the objects belonging to the group are identical, with elements of {1,2,3}. In this case, duplicate removal is performed by removing one of the groups, and the mutual actions of the same group are not predicted to be duplicated. In the drawings, the removal of the third group (G3) is exemplified.

[0072] The above visual feature extraction means (400) extracts visual features (F1, F2, F3, F4) using pixel values ​​within the detected objects (1, 2, 3, 4).

[0073] Additionally, the above visual features are extracted for each object, and can be extracted using a feature extraction network such as ResNet or HRNet.

[0074] The above individual action recognition means (500) extracts the skeleton information (S) extracted by the above skeleton extraction means (200). j ) and the visual features (F) extracted from the visual feature extraction means (400) j ) to obtain individual behavioral features (F), which are behavioral characteristics of each object. j * ) and extract the extracted individual behavioral features (F j * ) to use individual actions (P P ) is predicted.

[0075] In addition, the above individual behavioral characteristics (F j * ) is the skeleton information 'S' as shown in Equation 1 below. j ' and visual features 'F j ' can be extracted using a backbone network that takes as input.

[0076] [Formula 1]

[0077]

[0078] In addition, the above individual behavioral characteristics (F j * ) is an individual action (P) through an individual action classifier (personal_action_classifier) ​​as shown in Equation 2 below. P ) is predicted, and individual action classifiers can use pose recognition algorithms such as STGCN and PoseC3D.

[0079] [Formula 2]

[0080]

[0081] The above group update means (600) is used to update the individual behavioral characteristics (F) of objects included in the group (G) initially set in the above group creation means (300). j * ) is used to calculate the probability that objects belong to the group, and the group is updated by deleting objects whose calculated probability is below a preset probability threshold.

[0082] This function can improve the accuracy of interaction recognition by updating the group initially set as the distance between objects using the individual behavioral characteristics of the objects.

[0083] Referring to FIG. 3, a detailed description is given. FIG. 3 shows the group update means (600), which includes a group probability calculator (610) and a group updater (620).

[0084] The above group probability calculator (610) calculates the individual behavioral characteristics (F) of objects within each preset group (G1, G2, G3). j * ) is used to calculate the probability that objects belong to the group.

[0085] In addition, FIG. 3 illustrates that multiple group probability calculators (610, 610a, 610b) are provided to calculate the group probabilities of each of the first group (G1), the second group (G2), and the third group (G3), but one group probability calculator may also calculate the group probabilities.

[0086] Referring to FIG. 4, which is a drawing for explaining the group probability calculator (610), the group probability calculator (610) calculates individual behavioral characteristics (F1) of objects belonging to a group (G1). * ,F2 * ,F3 * ) is input into the MLP (Multi-layer Perceptron) neural network (611) to transform the features, pool the transformed individual behavioral features to create one group behavioral feature, transform the created group behavioral feature again through the MLP neural network (613), and use the concat function (614) to transform the transformed group behavioral feature into the individual behavioral feature (F1) of each object. * ,F2 * ,F3 * ) and the individual behavioral features combined with the group behavioral features are transformed through the MLP neural network (615) to obtain the probability (Conf) that each object belongs to the group. i j ) is calculated.

[0087] Additionally, group probability calculations are calculated for all initially set groups (G1, G2, G3).

[0088] Referring again to FIG. 3, the group updater (620) removes objects with low probability for each group, then generates a new group and outputs the updated group.

[0089] In the example of Fig. 3, the individual behavioral characteristics of the fourth object (4) belonging to the initial second group (G2) are shown to be removed because the probability of belonging to the group is calculated to be below the threshold value.

[0090] That is, the group updater (620) performs group update by maintaining the objects (1,2,3) of the initial first group (G1), removing the fourth object among the objects (1,2,3,4) of the second group (G2), and maintaining the objects (4,2) of the fourth group (G4).

[0091] Meanwhile, in this case, after the update, the first group (G1) * ) and the second group (G2) after renewal * ) objects are identical to {1,2,3}, but one overlapping update group is removed so that the interaction of the same group is not duplicated and predicted.

[0092] The above-mentioned interaction recognition means (700) uses the individual behavioral characteristics of each object in the updated group as in Equation 3 below to generate interaction characteristics (P), which are behavioral characteristics between objects. m ) is predicted.

[0093] [Formula 3]

[0094]

[0095] In addition, the above interaction characteristics are updated by group (G i * ) is predicted.

[0096] In addition, referring to FIG. 5, the interactive behavior recognition means (700) is configured to recognize individual behavioral characteristics (F1) of objects belonging to the updated group. * ,F2 * ,F3 * ) is input into the MPL neural network (710) to transform the features, pool the transformed features (720) to create a single feature, and operate the created feature through the MPL neural network (730) to predict the interactive behavior features.

[0097] That is, according to the multi-object interaction recognition program of the present invention, individual actions and interaction actions of objects within an image can be recognized together, and when recognizing interaction actions, the initially set group can be updated and recognized, so the accuracy of recognition can be greatly improved.

[0098] As described above, the present invention has been illustrated and described with reference to preferred embodiments, but is not limited to the above embodiments, and various changes and modifications may be made by a person skilled in the art to which the invention pertains within a scope that does not depart from the spirit of the present invention.

Claims

1. Computer, An object detection and tracking means for detecting a human object in an input image and tracking it by assigning an identifier; A skeleton extraction means for extracting skeleton information by detecting joint parts of a detected object; A group creation means for creating a relationship graph including vertex information, which is location information of a detected object, and edge information, which is line information connecting vertex information of neighboring objects, and setting objects connected by one edge as one group; A visual feature extraction means for extracting visual features using pixel values ​​of a detected object; An individual action recognition means for extracting individual action features, which are action features of each object, using the above skeleton information and the above visual features, and predicting individual actions using the extracted individual action features; A group update means for calculating the probability that objects belong to a group by using the individual behavioral characteristics of objects included in a preset group and updating the group by deleting objects below a preset probability threshold; and A multi-object interaction recognition program stored in a computer-readable recording medium, characterized in that it functions as an interaction recognition means that predicts interaction characteristics, which are behavioral characteristics between objects, by using individual behavioral characteristics of each object in an updated group; 2. In paragraph 1, A program for recognizing the interaction of multiple objects stored in a computer-readable recording medium, characterized in that the group generating means converts the positional information of objects in a two-dimensional image into coordinates of actual positional information in a three-dimensional space to generate the vertex information, and generates edge information by connecting objects located within a preset actual distance to each other.

3. In paragraph 2, A program for recognizing the interaction of multiple objects stored in a computer-readable recording medium, characterized in that the group creation means sets objects connected by one edge as one group, and among the set groups, groups in which objects belonging to the same group are duplicated to leave only one group and delete the remaining groups.

4. In paragraph 1, The above individual action recognition means extracts the individual action features by using a backbone network that inputs the object's skeleton information and visual features. A multi-object interaction recognition program stored in a computer-readable storage medium, characterized in that it predicts the individual action by inputting extracted individual action features into a personal action classifier (personal_action_classifier).

5. In paragraph 1, A program for recognizing multi-object interaction stored in a computer-readable recording medium, characterized in that the visual feature extraction means extracts visual features using a feature extraction network such as ResNet or HRNet.

6. In any one of paragraphs 1 to 5, A program for recognizing multi-object interactions stored in a computer-readable recording medium, characterized in that the group updating means calculates, for each preset group, the probability that objects belonging to each group belong to the group, inputs individual behavioral features of the objects belonging to the group into a multi-layer perceptron (MLP) neural network to transform the features, generates the transformed individual behavioral features as one group feature through pooling, connects the generated group features to the individual behavioral features of each object to generate individual behavioral features with linked group features, and then inputs the individual behavior features with linked group features into the MLP neural network to transform the features and calculates the probability that the transformed individual behavioral features belong to the group.

7. In paragraph 6, A multi-object interaction recognition program stored in a computer-readable storage medium, characterized in that among the updated groups, groups in which the objects belonging to the same group are duplicated are removed to leave only one group and the remaining groups are deleted.

8. The computer, A step of detecting a human object in an input image and assigning an identifier to track it; A step of extracting skeleton information by detecting joint parts of a detected object; A step of creating a relationship graph including vertex information, which is location information of a detected object, and edge information, which is line information connecting vertex information of neighboring objects, and setting objects connected by one edge as one group; Step of extracting visual features using pixel values ​​of detected objects A step of extracting individual action features, which are behavioral characteristics of each object, using the above skeleton information and the above visual features, and predicting individual actions using the extracted individual action features; A step of calculating the probability that objects belong to a group by using the individual behavioral characteristics of objects included in a preset group, and updating the group by deleting objects below a preset probability threshold; and A method for recognizing interaction between multiple objects in an image by performing a step of predicting interaction features, which are interaction features between objects, using individual behavior features of each object in an updated group;

Citation Information

Patent Citations

  • Apparatus and method for extracting person domain based on RGB-Depth image

    KR1020170028605A

  • Electronic device, method, and computer-readable storage medium for displaying visual object related to application in virtual space

    KR1020250015663A

  • High-rate composting method using compost composition of livestock sludge

    KR1020250076278A

  • Point detection systems and methods for object identification and targeting

    US20230252624A1

  • KR20230039468A