An interactive behavior recognition method
By establishing a new coordinate system with the interaction center point as the origin in interactive behavior recognition, calculating and identifying the coordinates of both parties in the new coordinate system, and using a neural network model to solve the perspective error problem of interactive behavior recognition from different perspectives, the accuracy of interactive behavior recognition is improved.
Patent Information
- Application Number
- CN202310998576.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-08-09
- Publication Date
- 2025-11-18
- Estimated Expiration
- 2043-08-09
AI Technical Summary
In existing technologies, depth camera-based interactive behavior recognition methods struggle to match the same action from different viewpoints, resulting in large viewpoint errors that affect the classification and recognition of interactive behavior features.
By identifying the first and second target objects of the interaction, a new coordinate system is established with the center point of the interaction as the origin. The coordinates of both parties in the new coordinate system are calculated and identified, and the interaction behavior is identified using a neural network model.
It reduces the difficulty of recognizing the same interactive behavior caused by perspective error, and improves the accuracy and recognition effect of interactive behavior feature classification.
Smart Images

Figure CN117011941B_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of image data processing technology, and in particular relates to an interactive behavior recognition method. Background Technology
[0002] Human behavior recognition technology, as a type of behavior recognition technology, can monitor and identify human behavior in order to predict or detect any unexpected or dangerous events in advance. Human behavior recognition technology is also a challenging task with broad application prospects, such as intelligent video surveillance, human-computer interaction, automatic identification alarm, public safety, etc. Human behavior recognition has become a research hotspot in related fields and has its potential economic value.
[0003] Current research on human behavior recognition focuses on using the coordinates of individuals obtained from the original coordinate system of a depth camera for behavior recognition. However, camera devices typically only capture images of individuals from the perspective of one or a few cameras. The recognition and classification of the same action varies greatly from different perspectives, making it difficult to achieve matching for the same action. This is not conducive to classifying interactive behavior features and thus affects the effectiveness of human interaction behavior recognition. Summary of the Invention
[0004] The purpose of this application is to provide an interactive behavior recognition method, which aims to improve the accuracy of recognizing the interactive behavior of two parties.
[0005] Firstly, this application provides an interactive behavior recognition method, including:
[0006] Determine the first and second target objects for interactive behavior recognition;
[0007] A new coordinate system is established with the interaction center point between the first target object and the second target object as the origin.
[0008] Calculate the new coordinates of the first target object in the new coordinate system, and the new coordinates of the second target object in the new coordinate system.
[0009] The specific interaction behaviors between the first target object and the second target object are identified based on the new coordinates of the first target object and the new coordinates of the second target object.
[0010] Secondly, this application provides an interactive behavior recognition system, comprising:
[0011] The determining unit is used to determine the first target object and the second target object for interactive behavior recognition.
[0012] A unit is established to create a new coordinate system with the interaction center point between the first target object and the second target object as the origin.
[0013] The calculation unit is used to calculate the new coordinates of the first target object in the new coordinate system, and the new coordinates of the second target object in the new coordinate system.
[0014] The identification unit is used to identify the specific interaction behavior between the first target object and the second target object based on the new coordinates of the first target object and the new coordinates of the second target object.
[0015] Thirdly, this application provides a computer device, comprising:
[0016] Processor, memory, bus, input / output interfaces, network interfaces;
[0017] The processor is connected to the memory, the input / output interface, and the network interface via the bus;
[0018] The memory stores a program;
[0019] When the processor executes the program stored in the memory, it implements the interactive behavior recognition method as described in any one of the first aspects above.
[0020] Fourthly, this application provides a computer-readable storage medium storing instructions that, when executed on a computer, cause the computer to perform the interactive behavior recognition method as described in any of the first aspects above.
[0021] Fifthly, this application provides a computer program product that, when executed on a computer, causes the computer to perform the interactive behavior recognition method as described in any of the first aspects above.
[0022] As can be seen from the above technical solutions, the embodiments of this application have the following advantages:
[0023] The interactive behavior recognition method in this application identifies a first target object and a second target object to be identified, thereby knowing the two parties to be identified. Then, a new coordinate system is established with the interaction center point of the first target object and the second target object as the origin. The new coordinates of the first target object and the second target object in the new coordinate system are calculated. The new coordinate system is used to describe the two parties to be interacting. Since the new coordinate system is at the same distance from the two parties to be interacting, it can help reduce the difficulty of action recognition of the same interactive behavior caused by perspective error, and facilitate the classification of interactive behavior features, thereby affecting the effect of behavior recognition. The specific interactive behavior of the first target object and the second target object can be identified quickly and accurately based on the new coordinates of the first target object and the second target object. Attached Figure Description
[0024] Figure 1 This is a schematic flowchart of an embodiment of the interactive behavior recognition method of this application;
[0025] Figure 2 This is a schematic flowchart of another embodiment of the interactive behavior recognition method of this application;
[0026] Figure 3 This is a schematic diagram of the structure of an embodiment of the interactive behavior recognition system of this application;
[0027] Figure 4 This is a schematic diagram of another embodiment of the interactive behavior recognition system of this application;
[0028] Figure 5 This is a schematic diagram of the structure of one embodiment of the computer device of this application;
[0029] Figure 6 This is a schematic diagram of an embodiment of "the same two-person interaction behavior" under the camera view of different depth cameras in this application;
[0030] Figure 7 This is a schematic diagram of an example of a human skeleton sequence using the default human model "NTU RGB+D 60" dataset in the experiments of this application;
[0031] Figure 8 This is a schematic diagram illustrating the effect of an embodiment of interactive behavior recognition using geometric features in this application.
[0032] Figure 9 This is a schematic diagram illustrating the effect of an embodiment of individual B "kicking" individual A in this application;
[0033] Figure 10 A schematic diagram illustrating the effect of an embodiment in which individual B "pushes down" individual A in this application;
[0034] Figure 11 This is a schematic diagram illustrating the effect of an embodiment of the translation from the original coordinate system O-XYZ to the new coordinate system I-XYZ in this application;
[0035] Figure 12 This is a schematic flowchart of an embodiment of the expression S' of the skeleton sequence coordinates of the individual weights, time weights, and spatial weights in the first target object and the second target object in this application;
[0036] Figure 13 This is a schematic diagram of the confusion matrix for 11 interactive behavior categories using the default human model “NTU RGB+D 60” dataset in the experiments of this application. Detailed Implementation
[0037] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0038] It's important to note that current research on human behavior recognition based on RGB video has found that the success rate is significantly affected by factors such as lighting, scene, and camera lens angle. This is because, under the influence of various factors such as differences in lighting conditions, background complexity, and diverse viewpoints, it is difficult to accurately describe the skeletal sequence information of the human body, leading to inaccurate classification of human behavior. However, with the development of technologies such as depth cameras and pose estimation, obtaining human skeletal sequence information from videos captured by depth cameras using pose estimation algorithms has become increasingly easier and more accurate. Specifically, by processing depth images captured frame by frame from the depth camera sensor, a human skeleton composed of multiple joints can be inferred using pose estimation algorithms. This allows the recognition of human behavior in videos to be transformed into a description of the human skeletal sequence information at different times. It should be noted that currently mature depth cameras include Kinect and RealSense.
[0039] In actual depth camera footage, most scenarios involving human behavior recognition involve interactions with others. Even when multiple people are present, an interaction by one person at a particular moment is often directed at another person. Therefore, human behavior recognition typically involves interactions between two people, also known as two-person interactions. Consequently, research on behavior recognition for two-person interactions has significant practical application value.
[0040] Currently, videos captured using depth cameras to observe two-person interactions can identify and obtain information related to the human skeleton sequence of the two individuals (e.g., joint coordinates and names). This information also includes more complex spatial relationships (represented by the pose and relative distance between the two individuals) and temporal relationships (represented by the pose and relative distance between the two individuals in different frames of the video). A key issue lies in the numerical differences in skeleton sequence information related to the same behavior from different viewpoints, such as... Figure 6 As shown, the scenes in viewpoint 1, viewpoint 2, and viewpoint 3 all depict the same two-person interaction behavior: "one person kicks the other person". It can be seen that the joint coordinate values of the human skeleton sequences of the two target persons obtained by the depth camera from different viewpoints are significantly different under the same original coordinate system O-XYZ, making it difficult for the network model to classify and identify them as the same interaction behavior.
[0041] Assuming that numerous interactions between two parties are recorded by a depth camera in a pre-defined original coordinate system O-XYZ, we can first identify the two parties involved in the interaction, referring to them as the first target object and the second target object. We then record video of the first and second target objects within the pre-defined original coordinate system O-XYZ of the depth camera, obtaining a video containing F frames. We then determine the three-dimensional coordinates of the first target object contained within these images. Three-dimensional coordinates of the second target object Specifically,
[0042]
[0043]
[0044] In the above formula, n, i, and F are all positive integers greater than 0; the three-dimensional coordinates of the first target object. In this context, "A" refers to the first target object, "n" refers to the joint numbered n in the default human model corresponding to the first target object, and "i" refers to the i-th frame of the video composed of F frames recorded by the depth camera. This allows for a comprehensive representation of the three-dimensional coordinates of the first target object. Similarly, the three-dimensional coordinates of the second target object In this context, "B" refers to the second target object, "n" refers to the joint numbered n in the default human model corresponding to the second target object, and "i" refers to the i-th frame of the video composed of F frames recorded by the depth camera. This allows for a comprehensive representation of the three-dimensional coordinates of the second target object.
[0045] It is worth noting that, to simplify the complexity of personnel identification and reduce the computational load for identifying personnel interaction behaviors, a simplified default human body model is typically used to replace the actual skeletal joints of the person as the basis for interaction behavior recognition in this embodiment. For the rules governing the labeling in the default human body model of this embodiment, please refer to [link to relevant documentation]. Figure 7 . Figure 7 The figure shown is a default human body model that can be used in this embodiment. This default human body model contains 25 joint labels, where: joint label 1 represents the hip center, joint label 2 represents the spine, joint label 3 represents the neck, joint label 4 represents the head, joint label 5 represents the left shoulder, joint label 6 represents the left elbow, joint label 7 represents the left wrist, joint label 8 represents the left hand, joint label 9 represents the right shoulder, joint label 10 represents the right elbow, joint label 11 represents the right wrist, joint label 12 represents the right hand, joint label 13 represents the left hip, joint label 14 represents the left knee, joint label 15 represents the left ankle, joint label 16 represents the left foot, and joint label 17 represents the right hip. Hip, joint number 18 represents the right knee, joint number 19 represents the right ankle, joint number 20 represents the right foot, joint number 21 represents the shoulder center, joint number 22 represents the tip of the left hand, joint number 23 represents the left thumb, joint number 24 represents the tip of the right hand, and joint number 25 represents the right thumb. It can be seen that... Figure 7 If each person is simplified to a default human body model with 25 joint labels as the whole, then the identification of the two parties performing interactive behavior can be regarded as the identification of the joint combination classification between two human body models. The coordinates of the positions corresponding to the 25 joint labels of each person can be calculated by the depth camera in the original coordinate system through the pose estimation algorithm.
[0046] Specifically, the depth camera can determine the skeleton sequence coordinates of the first target object in the original coordinate system (i.e., the set of coordinates of all joints of the human model corresponding to the first target object in the video image frames) as S.A The skeleton sequence coordinates of the second target object in the original coordinate system (i.e., the set of coordinates of all joints of the human model corresponding to the second target object in the video image frames) is S. B ;
[0047]
[0048]
[0049] S=(S A ,S B )
[0050] Where N is a positive integer greater than 0 (when using the default human model with only 25 joint labels, N is also a positive integer less than or equal to 25). This represents the three-dimensional coordinates of the nth joint in the i-th frame of the first target object. S represents the three-dimensional coordinates of the nth joint in the i-th frame of the second target object; S represents the expression for the skeleton sequence coordinate combination of the two parties in the interaction (the first target object and the second target object).
[0051] From the above Figure 6 It can be seen that the joint coordinate values of the human skeleton sequences of two target persons acquired by depth cameras from different viewpoints differ significantly under the same original coordinate system O-XYZ. That is, the coordinates S of the skeleton sequence of the first target object performing the same action under different viewpoints are different. A The skeleton sequence coordinates S of the second target object B Significant numerical differences make it difficult to match the same action numerically, hindering feature classification and thus affecting the effectiveness of behavior recognition. Currently, many researchers studying interactive behavior recognition still calculate interactive behavior features such as relative distance and relative position using skeleton coordinate sequences obtained from the original coordinate axes. This approach is not conducive to interactive behavior classification and recognition.
[0052] Please see Figure 1 An embodiment of the interactive behavior recognition method of this application includes:
[0053] 101. Determine the first target object and the second target object for interactive behavior recognition.
[0054] This step can identify the two parties interacting using a depth camera, such as the Kinect depth camera and RealSense depth camera mentioned above. This step has relatively mature existing technologies, so we will not go into too much detail here.
[0055] 102. Establish a new coordinate system with the interaction center point between the first target object and the second target object as the origin.
[0056] Specifically, for the interaction center point between the first target object and the second target object, this step can use the midpoint of the line connecting a certain joint point on the torso of the first target object and a corresponding joint point on the torso of the second target object as the interaction center point. For example, this joint point could be the hip center (joint number 1), the spine (joint number 2), or the shoulder center (joint number 21) in the default human model. Specifically, this step can use the hip center (joint number 1) to confirm the three-dimensional coordinates J of the interaction center point I between the first and second target objects in the original coordinate system. I ,as follows:
[0057]
[0058] in, This represents the 3D coordinates of the first joint in the first frame of the first target object, where the first joint is the individual center joint of the first target object (i.e., the center of the hip joint with joint number 1); similarly, This represents the 3D coordinates of the first joint in the first frame of the second target object, which is the individual center joint of the second target object (i.e., the center of the hip joint with joint number 1). Therefore, this step uses the 3D coordinates J of the interaction center point I in the original coordinate system. I Establishing a new coordinate system with the origin means that, in this embodiment, the coordinates of the skeleton sequences of the first and second target objects are translated from the original coordinate system O-XYZ to the new coordinate system I-XYZ for related calculations. Figure 11 As shown.
[0059] 103. Calculate the new coordinates of the first target object in the new coordinate system, and the new coordinates of the second target object in the new coordinate system.
[0060] Specifically, it is necessary to first determine the skeleton sequence coordinates of the first target object in the original coordinate system as S. A And the skeleton sequence coordinates of the second target object in the original coordinate system are S. B ,as follows:
[0061]
[0062]
[0063] Where N is a positive integer greater than 0 (when using the default human model with only 25 joint labels, N is also a positive integer less than or equal to 25). This represents the three-dimensional coordinates of the nth joint in the i-th frame of the first target object. This represents the three-dimensional coordinates of the nth joint in the i-th frame of the second target object.
[0064] Then calculate the skeleton sequence coordinates of the first target object in the new coordinate system. And calculate the skeleton sequence coordinates of the second target object in the new coordinate system.
[0065]
[0066]
[0067]
[0068] Among them, J I The three-dimensional coordinates of the interaction center point I in the original coordinate system in step 101 above.
[0069] 104. Identify the specific interaction behaviors between the first target object and the second target object based on the new coordinates of the first target object and the second target object.
[0070] Specifically, this step can directly identify the interaction behavior between the first target object and the second target object described by the new coordinate system in step 103 using a preset neural network model. This allows for the direct identification of the specific interaction behavior between the two objects, such as "kicking" or "pushing over." Since the new coordinate system maintains the same distance from both interacting objects, it helps reduce the difficulty of recognizing the same interaction behavior due to perspective errors, thereby improving the accuracy of the neural network model in identifying the interaction behavior.
[0071] The neural network model preset here is a trained neural network model capable of recognizing the interaction behavior types of two parties. For example, the neural network model could be one or more of VGG (visual geometry group), residual networks (ResNets), etc. The trained neural network model can be stored in the memory of the depth camera in the above embodiment so that it can be accessed by the neural network chip of the depth camera to quickly identify image frames from the video captured by the depth camera. The training of the neural network model can be achieved by legally authorized collection of videos of people in public places (such as shopping malls, airports, amusement parks, train stations, etc.) as training samples. If necessary, it can be combined with manual labeling and classification of the interaction behaviors of the two parties in different training samples (e.g., categorized as communication behavior, no interaction behavior, kicking / hitting behavior, pushing / pushing behavior, etc.). In this embodiment, the description of the skeleton sequence coordinates of the two parties in the new coordinate system can be used as the input features of the neural network model to facilitate training, resulting in a trained neural network model capable of recognizing the interaction behaviors of two parties.
[0072] Based on the above Figure 1 The description of the embodiments, in other embodiments, such as Figure 8 As shown, the identification of the current interaction behavior between two parties can be described using the geometric features present in the interaction behavior, including the skeletal sequence coordinates of both parties and the relative distances between different parts of both parties (e.g., relative distances between corresponding joints, relative distances between hand joints, relative distances between torso joints, relative distances between the closest joints, etc.). These geometric features describe the relative geometric features between individuals and the individual's geometric features. However, the importance of one individual's participation in the interaction behavior is not yet described, indicating that the two individuals have the same weight in terms of the importance of the interaction behavior. This obviously ignores the fact that in many interaction behaviors, the degree of participation of different individuals may differ significantly. Please refer to [link to relevant documentation]. Figure 9 as well as Figure 10 , Figure 9 The image shows an interaction where individual B on the right "kicks and punches" individual A on the left. Figure 10 The example demonstrates an interaction where individual B on the right "pushes down" individual A on the left. In these scenarios, individual B exhibits more significant action features, such as "kicking" and "pushing down," than individual A. This suggests that the descriptive features of individual B's interaction behavior should be far more important than those of individual A to be consistent with human cognition. This approach compensates for the shortcomings of interaction behavior recognition technology, which uses the same weight for each individual, as it is not conducive to the salience of interaction behavior features and thus hinders the improvement of interaction behavior recognition results.
[0073] Please see Figure 2 Another embodiment of the interactive behavior recognition method of this application includes:
[0074] 201. Determine the first target object and the second target object for interactive behavior recognition.
[0075] The execution of this step is the same as described above. Figure 1 Step 101 in the embodiment is similar, and the repeated parts will not be described again here.
[0076] 202. Establish a new coordinate system with the interaction center point between the first target object and the second target object as the origin.
[0077] The execution of this step is the same as described above. Figure 1 Step 102 in the embodiment is similar, and the repeated parts will not be described again here.
[0078] 203. Calculate the new coordinates of the first target object in the new coordinate system, and the new coordinates of the second target object in the new coordinate system.
[0079] The execution of this step is the same as described above. Figure 1 Step 103 in the embodiment is similar, and the repeated parts will not be described again here.
[0080] 204. Assign a first weight to the first target object and a second weight to the second target object. The sum of the first weight and the second weight is 1.
[0081] This embodiment addresses the technical problem that using equal weights for individuals in interactive behavior recognition technology is detrimental to the saliency description of interactive behavior features. It aims to achieve the classification of specific interactive behaviors as "certain specific interactive behaviors are actions performed by one person on another person," thereby improving the accuracy of interactive behavior recognition. For example... Figure 9 as well as Figure 10 Individual B exhibits more pronounced action characteristics of "kicking" and "pushing" than individual A. Therefore, it should be specifically identified as "individual B kicks individual A" and "individual B pushes individual A down", rather than simply as "the interaction between individual B and individual A is kicking" and "the interaction between individual B and individual A is pushing".
[0082] To achieve the above objective, this step requires assigning a first weight W to the first target object. A Assign a second weight W to the second target object. B The first weight W A With the second weight W B The sum is 1, as detailed below:
[0083] W A =L A / (LA +L B )
[0084] W B =L B / (L A +L B )
[0085] Wherein, the expression for the maximum position change of the first target object is L. A And the expression L for the maximum position change corresponding to the second target object. B ,as follows:
[0086]
[0087]
[0088] Here, max() represents taking the maximum value, and norm() represents the magnitude of the orientation quantity.
[0089] 205. Establish the first behavioral feature expression of the behavioral characteristics of the first target object. The first behavioral feature expression is equal to the new coordinates of the first target object multiplied by the first weight.
[0090] It is worth noting that, in addition to the individual importance of different individuals (the first target object or the second target object) in the interactive behavior, this embodiment also considers the temporal importance of the skeleton sequence coordinates of an individual in different image frames of the video, as well as the spatial importance of the skeleton sequence coordinates of an individual in the same image frame of the video. This embodiment provides a unified description of the individual importance in the interactive behavior, the temporal importance of different individual skeleton sequences in different image frames of the video, and the spatial importance of individual skeleton sequences in different image frames of the video.
[0091] For details, please refer to Figure 12 First, based on the skeleton sequence related data of different individuals A (assumed to be the first target object in this embodiment) and B (assumed to be the second target object in this embodiment) in the interaction, the skeleton sequence coordinates of the first target object in the new coordinate system are determined. And the skeleton sequence coordinates of the second target object in the new coordinate system It can form a three-dimensional spatial geometry of size C×F×N, where C represents the dimension of the skeleton sequence coordinates, i.e., C=3; as described in the previous embodiment, F represents the number of image frames in the video recorded by the depth camera, representing the dimension of time; N represents the number of joints of individual A and / or individual B in one frame of the video recorded by the depth camera, representing the dimension of space.
[0092] To reflect the feature weights of the skeleton sequence coordinates of different individuals in the time and space dimensions, this step involves resizing the skeleton sequence coordinates of the first target object in the new coordinate system. And the skeleton sequence coordinates of the second target object in the new coordinate system Pooling is performed in dimension N to obtain time-weighted features of size C×F×1; the skeleton sequence coordinates of the first target object in the new coordinate system are then calculated. And the skeleton sequence coordinates of the second target object in the new coordinate system Pooling is performed in dimension F to obtain spatial weight features of size C×1×N.
[0093] Establish the spatial weight feature expression of the first target object And / or, establish the time-weighted feature expression for the first target object. as follows:
[0094]
[0095]
[0096] in, This indicates that pooling is performed on the dimension F of the three-dimensional vector represented by the coordinates of the first target object skeleton sequence. This indicates that the dimension N of the three-dimensional vector represented by the coordinates of the first target object skeleton sequence is pooled.
[0097] The spatial weight feature expression for the first target object is established as follows: The time-weighted feature expression for the first target object is:
[0098] 206. Establish the second behavioral feature expression of the behavioral characteristics of the second target object. The second behavioral feature expression is equal to the new coordinates of the second target object multiplied by the second weight.
[0099] Similarly, based on step 205, the spatial weight feature expression of the second target object is established. And / or, establish the time-weighted feature expression for the second target object. as follows:
[0100]
[0101]
[0102] Similarly, The expression indicates that the dimension F of the three-dimensional vector represented by the coordinates of the second target object skeleton sequence is pooled, and pooling(;N) indicates that the dimension N of the three-dimensional vector represented by the coordinates of the second target object skeleton sequence is pooled.
[0103] The spatial weight feature expression for the second target object is established as follows: The time-weighted feature expression for the second target object is:
[0104] 207. Establish the interaction behavior expression between the first target object and the second target object. The interaction behavior expression is equal to the first behavior feature expression plus the second behavior feature expression.
[0105] Specifically, a spatial weighted feature expression U is established for the interaction behavior skeleton sequence between the first target object and the second target object. N And / or, establish the time-weighted feature expression U for the skeleton sequence of the interaction behavior between the first target object and the second target object. F ,as follows:
[0106]
[0107]
[0108] 208. By recognizing the interaction behavior expression through the trained network model, the specific interaction behavior between the first target object and the second target object is obtained. The trained network model can identify the specific interaction behavior of the target object with significant interaction behavior among the two target objects.
[0109] Specifically, the network model in this embodiment can be a multi-layer deep graph convolutional network model. Before recognizing the interaction behavior expression between the first target object and the second target object through the trained network model, the spatiotemporal weights U of the interaction behavior skeleton sequence between the first target object and the second target object can be calculated first:
[0110] U = U N ×U F
[0111] Then, the spatiotemporal weights are normalized:
[0112]
[0113] Where f1() represents passing through the fully connected layer f1, bn() represents passing through the normalized layer bn, and f2 represents passing through the fully connected layer f2; then, an expression S' is established to represent the skeleton sequence coordinates of the individual weights, temporal weights, and spatial weights in the first target object and the second target object, as follows:
[0114]
[0115]
[0116] The expression S' of the skeleton sequence is used as the input feature to train the network model, enabling it to extract preset features and classify preset interaction behaviors. This yields a trained network model that can then recognize interaction behavior expressions and determine the specific interaction behaviors between the first and second target objects. The trained network model can identify the specific interaction behaviors of target objects with salient interaction behaviors among the two target objects. The training process for this network model is a relatively mature technique and can be referenced in conjunction with the above. Figure 1 The relevant description of step 104 in the embodiment will not be repeated here.
[0117] It should be noted that the extraction and classification of interactive behavior features in the multi-layer deep graph convolutional network model takes into account the skeleton sequence-related data with individual attention weights as the input features of the multi-layer deep graph convolutional network model. The model extracts the depth features of the skeleton sequence-related data, passes them through a global pooling layer and a fully connected layer, and then through a SoftMax layer to obtain the recognition results of each multi-layer deep graph convolutional network model. The multi-layer deep graph convolutional network model contains nine temporal and spatial graph convolutional layers. Each graph convolutional layer includes spatial and temporal convolutions. Each temporal and spatial convolution is followed by a BatchNorm layer and a ReLU layer, with residual mechanisms applied in each layer.
[0118] To verify the recognition rate of the above network model, this embodiment conducts experiments on the most widely used skeleton dataset, "NTURGB+D 60," to test the interaction behaviors in the dataset. This dataset contains 10,347 interaction behavior skeleton sequences, encompassing 11 interaction behavior types (such as...). Figure 13 As shown in the confusion matrix), it is executed by 40 objects, each object's skeleton containing 25 joints (as shown in the confusion matrix). Figure 7 As shown in the image, three depth cameras at different positions and angles are used to capture interactive behavior: depth camera 2 faces the interaction directly, depth camera 1 captures from a 45-degree angle to the right, and depth camera 3 captures from a 45-degree angle to the left. This dataset provides two evaluation protocols: Cross-View (CV) and Cross-Subject (CS) cross-validation. For the CS protocol, the objects are divided into two equal parts, each containing 20 objects for training and testing respectively, with 7319 and 3028 samples for training and testing. For the CV protocol, depth camera 1 is used for testing, and depth cameras 2 and 3 are used for training, with 6889 and 3458 samples for training and testing respectively.
[0119] This embodiment utilizes a multi-layer deep graph convolutional network model to extract deep features from the skeleton sequence information of different interaction behaviors. The extracted deep features are then used for behavior classification to obtain the recognition rate. The multi-layer deep graph convolutional network model consists of 9 spatial-temporal graph convolutional layers. All experiments were implemented using the PyTorch framework on an NVIDIA GeForce P4000 GPU, and the experimental results are shown in Table 1 below.
[0120]
[0121] Table 1
[0122] Table 1 presents an experimental result of the interactive behavior recognition method in this embodiment. According to the experimental results, the skeleton sequence obtained based on the interaction center point coordinate transformation... Interactive behavior recognition was performed, achieving a recognition rate of 95.16% under the CV verification method, which is 2.75% higher than the 92.41% recognition rate achieved by the current technical solution based on the original skeleton sequence S. This demonstrates the effectiveness of the strategy based on interaction center point transformation proposed in this invention. Furthermore, after considering individual weights, temporal weights, and spatial weights, the recognition rate under the CV verification method reached 96.44%, which is 1.28% higher than the 95.16% recognition rate achieved by the skeleton sequence without considering individual, temporal, and spatial weights. This further illustrates the importance and effectiveness of considering individual, temporal, and spatial weights.
[0123] In the CV verification method, the confusion matrix for the 11 interaction behavior categories is as follows: Figure 12 As shown.
[0124] The following comparison of the interactive behavior recognition method of this embodiment with other interactive behavior recognition methods in the prior art demonstrates that the interactive behavior recognition method proposed in this embodiment has a higher recognition rate than other methods under different verification methods.
[0125] The comparison results are shown in Table 2 below:
[0126]
[0127] Table 2
[0128] The above embodiments describe the interactive behavior recognition method of this application. The interactive behavior recognition system of this application is described below. Please refer to [link / reference]. Figure 3 An embodiment of the interactive behavior recognition system of this application includes:
[0129] The determining unit 301 is used to determine the first target object and the second target object for interactive behavior recognition;
[0130] Establishment unit 302 is used to establish a new coordinate system with the interaction center point between the first target object and the second target object as the origin;
[0131] The calculation unit 303 is used to calculate the new coordinates of the first target object in the new coordinate system and the new coordinates of the second target object in the new coordinate system.
[0132] The identification unit 304 is used to identify the specific interaction behavior between the first target object and the second target object based on the new coordinates of the first target object and the new coordinates of the second target object.
[0133] The operations performed by the interactive behavior recognition system in this embodiment are the same as those described above. Figure 1 The operations performed in the embodiments are similar and will not be described again here.
[0134] Please see Figure 4 Another embodiment of the interactive behavior recognition system of this application includes:
[0135] The determining unit 401 is used to determine the first target object and the second target object for interactive behavior recognition;
[0136] Establishment unit 402 is used to establish a new coordinate system with the interaction center point between the first target object and the second target object as the origin;
[0137] The calculation unit 403 is used to calculate the new coordinates of the first target object in the new coordinate system and the new coordinates of the second target object in the new coordinate system.
[0138] The identification unit 404 is used to identify the specific interaction behavior between the first target object and the second target object based on the new coordinates of the first target object and the new coordinates of the second target object.
[0139] Optionally, when the establishing unit 402 establishes a new coordinate system with the interaction center point between the first target object and the second target object as the origin, it is specifically used for:
[0140] The first target object and the second target object are recorded in the original coordinate system preset by the camera device to obtain F-frame images;
[0141] The image is determined to contain three-dimensional coordinates of the first target object. The three-dimensional coordinates of the second target object
[0142]
[0143]
[0144] Wherein, n, i, and F are all positive integers greater than 0;
[0145] Confirm the three-dimensional coordinates J of the interaction center point I between the first target object and the second target object in the original coordinate system. I ;
[0146]
[0147] Among them, the This indicates the three-dimensional coordinates of the first joint in the first frame image of the first target object, where the first joint is the individual center joint of the first target object;
[0148] The This indicates the three-dimensional coordinates of the first joint in the first frame image of the second target object, where the first joint is the individual center joint of the second target object;
[0149] The three-dimensional coordinates J of the interaction center point I in the original coordinate system I Establish a new coordinate system with the origin as the origin.
[0150] Optionally, when the calculation unit 403 calculates the new coordinates of the first target object in the new coordinate system and the new coordinates of the second target object in the new coordinate system, it is specifically used for:
[0151] The skeleton sequence coordinates of the first target object in the original coordinate system are determined to be S. A The skeleton sequence coordinates of the second target object in the original coordinate system are S. B ;
[0152]
[0153]
[0154] Wherein, N is a positive integer greater than 0, and the This represents the three-dimensional coordinates of the nth joint in the i-th frame image of the first target object. This represents the three-dimensional coordinates of the nth joint in the i-th frame image of the second target object;
[0155] Calculate the skeleton sequence coordinates of the first target object in the new coordinate system as follows: And calculate the skeleton sequence coordinates of the second target object in the new coordinate system.
[0156]
[0157]
[0158] Wherein, J I Let I be the three-dimensional coordinates of the interaction center point I in the original coordinate system.
[0159] Optionally, when the identification unit 404 identifies the interaction behavior between the first target object and the second target object based on the new coordinates of the first target object and the new coordinates of the second target object, it is specifically used for:
[0160] Assign a first weight W to the first target object A Assign a second weight W to the second target object. B The first weight W A With the second weight W B The sum is 1;
[0161] Establish a first behavioral feature expression for the behavioral characteristics of the first target object, wherein the first behavioral feature expression is equal to the new coordinates of the first target object multiplied by the first weight;
[0162] Establish a second behavioral feature expression for the behavioral characteristics of the second target object, wherein the second behavioral feature expression is equal to the new coordinates of the second target object multiplied by the second weight;
[0163] Establish an interaction behavior expression between the first target object and the second target object, wherein the interaction behavior expression is equal to the first behavior feature expression plus the second behavior feature expression;
[0164] The trained network model identifies the interaction behavior expression to obtain the specific interaction behavior between the first target object and the second target object. The trained network model can identify the specific interaction behavior of the target object with significant interaction behavior among the two target objects.
[0165] Optionally, the identification unit 404 assigns a first weight W to the first target object. A Assign a second weight W to the second target object. B The first weight W A With the second weight W B The sum is 1, specifically including:
[0166] W A =L A / (L A +L B )
[0167] WB =L B / (L A +L B )
[0168] Wherein, the expression for the maximum position change of the first target object is L. A And the expression L for the maximum position change corresponding to the second target object. B ,as follows:
[0169]
[0170]
[0171] Wherein, max() represents taking the maximum value, and norm() represents the magnitude of the orientation quantity.
[0172] Optionally, the system further includes:
[0173] The establishment unit 402 is further configured to establish the spatial weight feature expression of the first target object.
[0174]
[0175] And / or,
[0176] The establishment unit 402 is further configured to establish the time weight feature expression of the first target object.
[0177]
[0178] Wherein, pooling(;F) represents pooling in the dimension F of the three-dimensional vector;
[0179] The pooling(;N) method represents pooling in dimension N of a three-dimensional vector.
[0180] Optionally, the system further includes:
[0181] The establishment unit 402 is further configured to establish the spatial weight feature expression of the second target object.
[0182]
[0183] And / or,
[0184] The establishment unit 402 is further configured to establish the time weight feature expression of the second target object.
[0185]
[0186] Wherein, pooling(;F) represents pooling in the dimension F of the three-dimensional vector;
[0187] The pooling(;N) method represents pooling in dimension N of a three-dimensional vector.
[0188] Optionally, the establishing unit 402 establishes an interaction behavior expression between the first target object and the second target object, wherein the interaction behavior expression is equal to the first behavior feature expression plus the second behavior feature expression, specifically including:
[0189] Establish the spatial weight feature expression U of the interaction behavior skeleton sequence between the first target object and the second target object. N ;
[0190]
[0191] And / or,
[0192] Establish the time-weighted feature expression U for the skeleton sequence of the interaction behavior between the first target object and the second target object. F ;
[0193]
[0194] Optionally, the system further includes:
[0195] The computing unit 403 is also used to calculate the spatiotemporal weight U of the interaction behavior skeleton sequence between the first target object and the second target object:
[0196] U = U N ×U F
[0197] Normalization unit 405 is used to normalize the spatiotemporal weights:
[0198]
[0199] Where f1() represents passing through the fully connected layer f1, bn() represents passing through the normalization layer bn, and f2 represents passing through the fully connected layer f2;
[0200] The establishment unit 402 is further configured to establish an expression S' that represents the skeleton sequence of individual weights, temporal weights, and spatial weights in the first target object and the second target object:
[0201]
[0202] Among them, the
[0203] Optionally, the system further includes:
[0204] Training unit 406 is used to train the network model by using the expression S' of the skeleton sequence as the input feature of the network model, so that the network model can extract preset features and classify preset interaction behaviors, thereby obtaining the trained network model.
[0205] The operations performed by the interactive behavior recognition system in this embodiment are the same as those described above. Figure 2 The operations performed in the embodiments are similar and will not be described again here.
[0206] The computer device in the embodiments of this application is described below. Please refer to [link / reference]. Figure 5 One embodiment of the computer device in this application includes:
[0207] The computer device 500 may include one or more central processing units (CPUs) 501 and memory 502, wherein the memory 502 stores one or more application programs or data. The memory 502 is volatile or persistent storage. The program stored in the memory 502 may include one or more modules, each module including a series of instruction operations on the computer device. Furthermore, the processor 501 may be configured to communicate with the memory 502 and execute the series of instruction operations stored in the memory 502 on the computer device 500. The computer device 500 may also include one or more wireless network interfaces 503, one or more input / output interfaces 504, and / or one or more operating systems, such as Windows Server, Mac OS, Unix, Linux, FreeBSD, etc. The processor 501 can execute the aforementioned... Figure 1 , Figure 2 The specific operations performed in the illustrated embodiment will not be described in detail here.
[0208] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods in the various embodiments of this application. The aforementioned storage medium includes: USB flash drive, mobile hard drive, read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk, and other media capable of storing program code.
[0209] The above description is merely a preferred embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of this application should be included within the protection scope of this application.
Claims
1. An interactive behavior recognition method, characterized in that, include: Determine the first and second target objects for interactive behavior recognition; A new coordinate system is established with the interaction center point between the first target object and the second target object as the origin. Calculate the new coordinates of the first target object in the new coordinate system, and the new coordinates of the second target object in the new coordinate system. Identify the specific interaction behaviors between the first target object and the second target object based on the new coordinates of the first target object and the second target object; Establishing a new coordinate system with the interaction center point between the first target object and the second target object as the origin includes: The first target object and the second target object are recorded in the original coordinate system preset by the depth camera to obtain F-frame images; The image is determined to contain three-dimensional coordinates of the first target object. The three-dimensional coordinates of the second target object Wherein, n, i, and F are all positive integers greater than 0; Confirm the three-dimensional coordinates J of the interaction center point I between the first target object and the second target object in the original coordinate system. I ; Among them, the This indicates the three-dimensional coordinates of the first joint in the first frame image of the first target object, where the first joint is the individual center joint of the first target object; The This indicates the three-dimensional coordinates of the first joint in the first frame image of the second target object, where the first joint is the individual center joint of the second target object; The three-dimensional coordinates J of the interaction center point I in the original coordinate system I Establish a new coordinate system with the origin as the origin; Calculating the new coordinates of the first target object in the new coordinate system and the new coordinates of the second target object in the new coordinate system includes: The skeleton sequence coordinates of the first target object in the original coordinate system are determined to be S. A The skeleton sequence coordinates of the second target object in the original coordinate system are S. B ; Wherein, N is a positive integer greater than 0, and the This represents the three-dimensional coordinates of the nth joint in the i-th frame image of the first target object. This represents the three-dimensional coordinates of the nth joint in the i-th frame image of the second target object; Calculate the skeleton sequence coordinates of the first target object in the new coordinate system as follows: And calculate the skeleton sequence coordinates of the second target object in the new coordinate system. Wherein, J I The three-dimensional coordinates of the interaction center point I in the original coordinate system; Identifying the interaction behavior between the first target object and the second target object based on their new coordinates includes: Assign a first weight W to the first target object A Assign a second weight W to the second target object. B The first weight W A With the second weight W B The sum is 1; Establish a first behavioral feature expression for the behavioral characteristics of the first target object, wherein the first behavioral feature expression is equal to the new coordinates of the first target object multiplied by the first weight; Establish a second behavioral feature expression for the behavioral characteristics of the second target object, wherein the second behavioral feature expression is equal to the new coordinates of the second target object multiplied by the second weight; Establish an interaction behavior expression between the first target object and the second target object, wherein the interaction behavior expression is equal to the first behavior feature expression plus the second behavior feature expression; The trained network model identifies the interaction behavior expression to obtain the specific interaction behavior between the first target object and the second target object. The trained network model can identify the specific interaction behavior of the target object with significant interaction behavior among the two target objects.
2. The interactive behavior recognition method according to claim 1, characterized in that, Assign a first weight W to the first target object A Assign a second weight W to the second target object. B The first weight W A With the second weight W B The sum is 1, specifically including: W A =L A / (L A +L B ) W B =L B / (L A +L B ) Wherein, the expression for the maximum position change of the first target object is L. A And the expression L for the maximum position change corresponding to the second target object. B ,as follows: Wherein, max() represents taking the maximum value, and norm() represents the magnitude of the orientation quantity.
3. The interactive behavior recognition method according to claim 2, characterized in that, Before establishing the first behavioral feature expression of the behavioral characteristics of the first target object, the method further includes: Establish the spatial weight feature expression of the first target object And / or, Establish the time weight feature expression of the first target object Wherein, pooling(;F) represents pooling in the dimension F of the three-dimensional vector; The pooling(;N) method represents pooling in dimension N of a three-dimensional vector.
4. The interactive behavior recognition method according to claim 3, characterized in that, Before establishing the second behavioral feature expression for the behavioral characteristics of the second target object, the method further includes: Establish the spatial weight feature expression of the second target object And / or, Establish the time weight feature expression of the second target object Wherein, pooling(;F) represents pooling in the dimension F of the three-dimensional vector; The pooling(;N) method represents pooling in dimension N of a three-dimensional vector.
5. The interactive behavior recognition method according to claim 4, characterized in that, Establish an interaction behavior expression between the first target object and the second target object, wherein the interaction behavior expression is equal to the first behavior feature expression plus the second behavior feature expression, specifically including: Establish the spatial weight feature expression U of the interaction behavior skeleton sequence between the first target object and the second target object. N ; And / or, Establish the time-weighted feature expression U for the skeleton sequence of the interaction behavior between the first target object and the second target object. F ; 6. The interactive behavior recognition method according to claim 5, characterized in that, Before recognizing the interaction behavior expression through a trained network model, the method further includes: Calculate the spatiotemporal weights U of the skeleton sequence of the interaction behaviors between the first target object and the second target object: U=U N ×U F The spatiotemporal weights are normalized: Where f1() represents passing through the fully connected layer f1, bn() represents passing through the normalization layer bn, and f2 represents passing through the fully connected layer f2; Establish an expression S' for the skeleton sequence coordinates that reflects the individual weights, temporal weights, and spatial weights of the first target object and the second target object: Among them, the 7. The interactive behavior recognition method according to claim 6, characterized in that, Before recognizing the interaction behavior expression through a trained network model, the method further includes: The expression S' of the skeleton sequence coordinates is used as the input feature of the network model to train the network model, so that the network model can extract preset features and classify preset interaction behaviors, thus obtaining the trained network model.
Citation Information
Patent Citations
Method for extracting spatio-temporal characteristics of interaction of two persons
CN116665308A