A semi-supervised skeleton point-based behavior recognition method and system

By introducing a multi-grained anchor comparison representation learning model in behavior recognition, extracting multi-grained features and optimizing the loss function, the problems of insufficient local motion information response and unclear positive/negative pairs in the prior art are solved, and more efficient behavior recognition performance is achieved.

CN114708656BActive Publication Date: 2025-06-24NANJING UNIV OF SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210317678.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-29
Publication Date
2025-06-24
Estimated Expiration
2042-03-29

AI Technical Summary

Technical Problem

In behavior recognition tasks based on semi-supervised skeleton points, how to obtain more discriminant information from labeled and unlabeled data is a challenging problem. The existing methods face limitations such as insufficient local motion information response, unclear positive/negative pairings, and neglect of granularity contrast.

Method used

A multi-grained anchor comparison representation learning model is proposed. Local, global and contextual features are extracted through graph convolutional neural networks and context graph convolutional neural networks, and the total loss function of comparison loss and identification loss is optimized through multi-grained anchor comparison loss to improve behavior recognition performance.

Benefits of technology

This method can more effectively capture soft positive/negative pairs of high confidence, avoid noise and outlier samples, and improve the accuracy and robustness of behavior recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114708656B_ABST
    Figure CN114708656B_ABST
Patent Text Reader

Abstract

The present invention relates to a semi-supervised skeleton point-based behavior recognition method and system, belonging to the field of human behavior recognition. The method includes obtaining human skeleton data; using a multi-granularity anchor contrast learning model to extract features from the human skeleton data to obtain local features, global features, and context features; the multi-granularity anchor contrast learning model includes a graph convolutional neural network and a context graph convolutional neural network; fusing the local features, the global features, and the context features to obtain a fused feature; and obtaining a human behavior recognition result according to the fused feature. The present invention improves the performance of behavior recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of human behavior recognition in the field of computer vision, and particularly to a behavior recognition method and system based on semi-supervised skeleton points. Background Art

[0002] Human behavior recognition is an important issue in the fields of computer vision and pattern recognition, and is currently developing rapidly due to its wide applications in video retrieval, video surveillance, virtual reality, human-computer interaction, etc. Compared with behavior recognition tasks based on RGB data and depth data, behavior recognition based on skeleton points has received increasing attention due to the advantages of data robustness and easy acquisition. Currently, most deep learning models for behavior recognition based on skeleton points are trained in a fully supervised manner to learn discriminative representations of skeletons. Although some remarkable performances have been achieved, a large amount of labeled data is usually required, and annotating skeleton sequence data is always time-consuming and laborious. Therefore, how to learn discriminative representations from unlabeled and labeled skeleton sequences, a method called behavior recognition based on semi-supervised skeleton points has become a work receiving much attention.

[0003] In the task of behavior recognition based on semi-supervised skeleton points, how to obtain more discriminative information from labeled and unlabeled data is a challenging problem. As the current mainstream method, contrastive learning can learn more enhanced data representations and can be used as a pre-task for behavior recognition. However, this method still faces three main limitations: 1) It usually learns global granularity features that cannot well reflect local motion information. 2) Its positive / negative pairs are usually predefined, and some of the positive / negative pairs are not clear. 3) It usually only measures the distance between positive / negative pairs within the same granularity, which ignores the contrast between positive and negative pairs of different granularities. Summary of the Invention

[0004] The purpose of the present invention is to provide a behavior recognition method and system based on semi-supervised skeleton points, and a new multi-granularity anchor contrast representation learning model is proposed to improve behavior recognition performance.

[0005] To achieve the above purpose, the present invention provides the following solutions:

[0006] A behavior recognition method based on semi-supervised skeleton points, comprising:

[0007] Obtaining human skeleton data; the human skeleton data includes a skeleton graph with learned connections and a skeleton graph with structural connections;

[0008] Feature extraction is performed on the human skeleton data using a multi-granularity anchor contrast learning model to obtain local features, global features, and context features; the multi-granularity anchor contrast learning model includes a graph convolutional neural network and a context graph convolutional neural network; the total loss function of the multi-granularity anchor contrast learning model includes a multi-granularity anchor contrast loss and an identification loss; the local features include a first local feature and a second local feature; the global features include a first global feature and a second global feature; the context features include a first context feature and a second context feature;

[0009] The local features, the global features, and the context features are fused to obtain a fused feature;

[0010] A human behavior recognition result is obtained based on the fused feature.

[0011] Optionally, the performing feature extraction on the human skeleton data using a multi-granularity anchor contrast learning model to obtain local features, global features, and context features specifically includes:

[0012] The first global feature and the first local feature are extracted from the learned-connected skeleton graph using the graph convolutional neural network;

[0013] The first context feature is extracted from the learned-connected skeleton graph using the context graph convolutional neural network;

[0014] The second global feature and the second local feature are extracted from the structurally-connected skeleton graph using the graph convolutional neural network;

[0015] The second context feature is extracted from the structurally-connected skeleton graph using the context graph convolutional neural network.

[0016] Optionally, the method for determining the total loss function specifically includes:

[0017] An anchor point adjacency matrix and a sample adjacency matrix are determined based on the fused feature;

[0018] A multi-granularity anchor contrast loss is determined based on the anchor point adjacency matrix and the sample adjacency matrix;

[0019] An identification loss is determined based on the fused feature;

[0020] A total loss function is determined based on the contrast loss and the identification loss.

[0021] Optionally, the determining the anchor point adjacency matrix and the sample adjacency matrix based on the fused feature specifically includes:

[0022] The fused feature is used as a node to construct an anchor graph;

[0023] Determine anchor points using a clustering algorithm based on the fusion features;

[0024] Determine an anchor point adjacency matrix and a sample adjacency matrix based on the fusion features, the anchor graph, and the anchor points.

[0025] Optionally, the fusing the local features, the global features, and the context features to obtain fusion features specifically includes:

[0026] Fuse the first global feature, the first local feature, the first context feature, the second global feature, the second local feature, and the second context feature to obtain fusion features.

[0027] A semi-supervised skeleton-based behavior recognition system, comprising:

[0028] A data acquisition module for acquiring human skeleton data; the human skeleton data includes a skeleton graph with learned connections and a skeleton graph with structural connections;

[0029] A feature extraction module for extracting local features, global features, and context features using a multi-granularity anchor contrast learning model based on the human skeleton data; the multi-granularity anchor contrast learning model includes a graph convolutional neural network and a context graph convolutional neural network; the total loss function of the multi-granularity anchor contrast learning model includes a multi-granularity anchor contrast loss and an identification loss; the local features include a first local feature and a second local feature; the global features include a first global feature and a second global feature; the context features include a first context feature and a second context feature;

[0030] A fusion module for fusing the local features, the global features, and the context features to obtain fusion features;

[0031] An identification module for obtaining a human behavior recognition result based on the fusion features.

[0032] Optionally, the feature extraction module specifically includes:

[0033] A first extraction unit for extracting a first global feature and a first local feature using the graph convolutional neural network based on the skeleton graph with learned connections;

[0034] A second extraction unit for extracting a first context feature using the context graph convolutional neural network based on the skeleton graph with learned connections;

[0035] A third extraction unit for extracting a second global feature and a second local feature using the graph convolutional neural network based on the skeleton graph with structural connections;

[0036] A fourth extraction unit, configured to extract second context features by using the context graph convolutional neural network according to the skeleton graph connected by the structure.

[0037] Optionally, the method for determining the total loss function specifically includes:

[0038] Determine an anchor adjacent matrix and a sample adjacent matrix according to the fusion features;

[0039] Determine a multi-granularity anchor contrast loss according to the anchor adjacent matrix and the sample adjacent matrix;

[0040] Determine an identification loss according to the fusion features;

[0041] Determine the total loss function according to the contrast loss and the identification loss.

[0042] Optionally, the determining the anchor adjacent matrix and the sample adjacent matrix according to the fusion features specifically includes:

[0043] Construct an anchor graph by using the fusion features as nodes;

[0044] Determine anchor points by using a clustering algorithm according to the fusion features;

[0045] Determine the anchor adjacent matrix and the sample adjacent matrix according to the fusion features, the anchor graph and the anchor points.

[0046] Optionally, the fusion module specifically includes:

[0047] A fusion unit, configured to fuse the first global feature, the first local feature, the first context feature, the second global feature, the second local feature and the second context feature to obtain fusion features.

[0048] According to the specific embodiments provided by the present invention, the present invention discloses the following technical effects:

[0049] Obtain human skeleton data in the present invention; the human skeleton data includes a skeleton graph with learned connections and a skeleton graph with structural connections; use a multi-granularity anchor contrast learning model to extract features based on the human skeleton data to obtain local features, global features, and context features; the multi-granularity anchor contrast learning model includes a graph convolutional neural network and a context graph convolutional neural network; the total loss function of the multi-granularity anchor contrast learning model includes a multi-granularity anchor contrast loss and an identification loss; the local features include a first local feature and a second local feature; the global features include a first global feature and a second global feature; the context features include a first context feature and a second context feature; fuse the local features, global features, and context features to obtain a fused feature; obtain a human behavior recognition result based on the fused feature. The multi-granularity anchor contrast loss in the multi-granularity anchor contrast learning model measures the consistency between high-confidence soft positive pairs based on the anchor graph, thereby improving the model recognition performance. Brief Description of the Drawings

[0050] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required to be used in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0051] Figure 1 It is a flowchart of the behavior recognition method based on semi-supervised skeleton points provided by the present invention;

[0052] Figure 2 It is a schematic diagram of the flow of the behavior recognition method based on semi-supervised skeleton points provided by the present invention;

[0053] Figure 3 It is a schematic diagram of the GCN block structure;

[0054] Figure 4 It is a schematic diagram of the Context GCN block structure;

[0055] Figure 5 It is a schematic diagram of the semi-supervised multi-granularity anchor contrast representation learning model provided by the present invention. Detailed Description of the Embodiments

[0056] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0057] The object of the present invention is to provide a semi-supervised skeleton point-based behavior recognition method and system to improve behavior recognition performance.

[0058] To make the above objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0059] As Figure 1 shown, a semi-supervised skeleton point-based behavior recognition method provided by the present invention includes:

[0060] Step 101: Obtain human skeleton data; the human skeleton data includes a skeleton graph with learned connections and a skeleton graph with structural connections. Among them, the connections between skeleton points in the skeleton graph with learned connections are learnable and not fixed connection structures.

[0061] Step 102: Use a multi-granularity anchor contrast learning model to extract local features, global features, and context features according to the human skeleton data; the multi-granularity anchor contrast learning model includes a graph convolutional neural network and a context graph convolutional neural network; the total loss function of the multi-granularity anchor contrast learning model includes a multi-granularity anchor contrast loss and an identification loss; the local features include a first local feature and a second local feature; the global features include a first global feature and a second global feature; the context features include a first context feature and a second context feature.

[0062] Step 102 specifically includes:

[0063] Extract the first global feature and the first local feature according to the skeleton graph with learned connections using the graph convolutional neural network; extract the first context feature according to the skeleton graph with learned connections using the context graph convolutional neural network; extract the second global feature and the second local feature according to the skeleton graph with structural connections using the graph convolutional neural network; extract the second context feature according to the skeleton graph with structural connections using the context graph convolutional neural network.

[0064] Step 103: Fuse the local features, the global features, and the context features to obtain a fused feature. Step 103 specifically includes: Fuse the first global feature, the first local feature, the first context feature, the second global feature, the second local feature, and the second context feature to obtain a fused feature.

[0065] Step 104: Obtain a human behavior recognition result according to the fused feature.

[0066] In practical applications, the method for determining the total loss function specifically includes: determining an anchor adjacent matrix and a sample adjacent matrix according to the fused features; determining a multi-granularity anchor contrast loss according to the anchor adjacent matrix and the sample adjacent matrix; determining an identification loss according to the fused features; and determining the total loss function according to the contrast loss and the identification loss.

[0067] In practical applications, determining the anchor adjacent matrix and the sample adjacent matrix according to the fused features specifically includes:

[0068] Taking the fused features as nodes to construct an anchor graph.

[0069] Determining anchor points according to the fused features by using a clustering algorithm.

[0070] Determining the anchor adjacent matrix and the sample adjacent matrix according to the fused features, the anchor graph, and the anchor points.

[0071] The multi-granularity anchor contrast representation learning model based on semi-supervised skeleton point behavior recognition of the present invention includes four processes: extracting multi-granularity features, establishing an anchor graph to calculate the sample adjacent matrix and the anchor adjacent matrix, calculating the multi-granularity anchor contrast loss, and obtaining the model objective function.

[0072] As Figure 2 and Figure 5 shown, extracting multi-granularity features includes the following steps:

[0073] Step 1: Obtaining a human skeleton data set The elements of which are a skeleton data, C represents the number of channels, T represents the total number of frames, Q represents the number of joint points of each person, P represents the number of people in each frame, S is the total number of skeleton data samples, and s is the subscript representing the s-th skeleton data. The human skeleton data is constructed into two graphs with learning connections and structural connections.

[0074] Step 2: On the skeleton graph with learning connections, inputting the data v s obtained in Step 1 into the graph convolutional network (GCNs) G1(·), followed by a global average pooling (GAP) to obtain the global feature f G = GAP(G1(v s )); inputting the skeleton data v s into the graph convolutional network (GCNs) G2(·) to obtain the local feature f L = G2(v s ); G1(·) and G2(·) both represent the graph convolutional network GCNs. Inputting the skeleton data v sInput into the context graph convolutional network (Context GCNs) G3(·) to obtain the context feature f C = GAP(G3(v s ))). G3(·) represents the context graph convolutional network Context GCNs. Among them, GCNs consists of 5 attached Figure 3 shown GCN blocks stacked and a fully connected layer. Attached Figure 3 The SGCN in f in and f out are the input and output of the SGCN, W k is the parameter of the network, K v is defined as the kernel size of the spatial dimension, is the adjacency matrix of the human skeleton graph, the diagonal matrix k represents the k-th spatial partition, i represents the i-th row of the matrix, and j represents the j-th column of the matrix; Attached Figure 3 The TGCN in Figure 4 shown: A common L×1 convolutional layer used to aggregate the context representations embedded in adjacent frames, where L represents the length of the time window. Context GCNs is similar to GCNs, the only difference being that Context GCN blocks also contain an attention module for capturing key joints as context joints, as shown in

[0075] Step 3: On the structurally connected skeleton graph, input the skeleton data v s obtained in Step 1 into Context GCNs G4(·) followed by GAP, GCNs G5(·), and GCNs G6(·) followed by GAP respectively to obtain the context feature h C = GAP(G4(v s ))), the local feature h L = G5(v s ), and the global feature h G = GAP(G6(v s ))). The structure is the same as that of the graph convolutional network and the context graph convolutional network mentioned in the previous step. G4(·) represents the context graph convolutional network Context GCNs, and G5(·), G6(·) represent the graph convolutional network GCNs.

[0076] Building the anchor graph to calculate the sample adjacency matrix and the anchor point adjacency matrix includes the following steps:

[0077] Step 4: Combine the multiple features {f s , f G , f L , fC , h C , h L , h G} are fused into a feature D s , that is:

[0078] D s = concat(f L , f C , f G , h L , h C , h G )

[0079] Therefore the fused features corresponding to all elements in can be defined as Set the elements in as nodes to build an anchor graph, and perform a clustering algorithm on m , that is: to represent the distribution of all samples.

[0080] Step 5: Calculate the anchor adjacency matrix Z and the sample adjacency matrix W based on the anchor graph obtained in Step 4. Z represents the relationship between samples and anchors, and its elements:

[0081]

[0082] where Z s,m is the distance between sample v s and anchor A m , K h (·) uses the Gaussian kernel function, and h is a hyperparameter, <s>is the index set of the nearest anchor points to the sample v s The set of indices of the nearest anchor points to the sample v. W represents the relationship between samples: where the diagonal matrix is defined as

[0083] The calculation of the multi-granularity anchor contrast loss includes the following steps:

[0084] Step 6: Calculate the inter-granularity and intra-granularity contrast losses based on the sample adjacency matrix W and the anchor adjacency matrix Z obtained in Step 5. Assume that the batch size during the training process is N. For skeleton data The corresponding multi-granularity features calculated can be defined as

[0085]

[0086]

[0087] where is the set of local features of the skeleton data extracted on the learned-connected skeleton graph, is the set of context features of the skeleton data extracted on the learned-connected skeleton graph, is the set of global features of the skeleton data extracted on the learned-connected skeleton graph; is the set of local features of the skeleton data extracted on the structurally-connected skeleton graph, is the set of context features of the skeleton data extracted on the structurally-connected skeleton graph, is the set of global features of the skeleton data extracted on the structurally-connected skeleton graph. n is a subscript representing the nth skeleton data feature. is the local feature of the nth skeleton data extracted on the learned-connected skeleton graph, is the context feature of the nth skeleton data extracted on the learned-connected skeleton graph, is the global feature of the nth skeleton data extracted on the learned-connected skeleton graph; is the local feature of the nth skeleton data extracted on the structurally-connected skeleton graph, is the context feature of the nth skeleton data extracted on the structurally-connected skeleton graph, is the global feature of the nth skeleton data extracted on the structurally-connected skeleton graph.

[0088] Define as the global-context feature set with 2N elements. Then the global-context contrast loss between the global feature and the context feature ​ Defined as:

[0089]

[0090] Where g i 、g j 、g k are elements in the set, representing the global or context features extracted; the set is the set of global features extracted from the learned connected skeleton graph and the set of context features is the union of, g i 、g j 、g k are elements of the set In the formula, all elements of the set are traversed (from 1 to 2N), so g i 、g j 、g k can take either global features or context features. Z' is obtained by setting all elements except the maximum value in each row of the anchor adjacency matrix Z to 0, that is, each sample retains only one nearest anchor point; In the MAC-Loss formula, the value of the hyperparameter τ is set to 0.07; H u and H v are projection matrices; is an indicator function, which is 1 when the condition in the square brackets is true and 0 otherwise; W i,j is an element of the sample adjacent matrix W, and N is the number of skeleton data samples. Calculate as shown in Table 1.

[0091] Table 1 Definition table of inter-granularity and intra-granularity contrast losses in MAC-loss

[0092]

[0093] Step 7: To integrate all inter-granularity and intra-granularity contrast losses, the definition of MAC-loss is as follows:

[0094]

[0095] The acquisition of the model objective function includes the following steps:

[0096] Step 8: Calculate the recognition loss according to D s obtained in Step 4 That is:

[0097]

[0098] Among them, y is the true label of the behavior.

[0099] Step 9: According to the and obtained in Step 7 and Step 8, define the objective function of MAC-learning. This model uses MAC-loss and Recognition loss to jointly train MAC-learning, and defines the objective function Ψ(θ) of MAC-learning as follows:

[0100]

[0101] Among them, θ is the parameter set of MAC-learning.

[0102] The model proposed by the present invention has better performance than other current representative methods on NTU RGB+D and NW-UCLA. The experimental results are shown in Tables 2 and 3 in the appendix. The excellent performance of this model is mainly because MAC-learing creatively captures high-confidence soft positive / negative pairs in contrastive learning, avoiding interference from fuzzy pairs of noise and outlier samples. MAC-Learing uses the multi-granularity anchor contrast loss (MAC-loss) including inter-granularity and intra-granularity contrast losses to measure the difference / consistency between soft negative / positive pairs among three granularities on the learnable and structurally connected skeletons.

[0103] Table 2 Comparison table of recognition accuracies (%) obtained by different methods on the NW-UCLA dataset with 5%, 10%, 30%, and 40% labeled data in the training set

[0104]

[0105] Table 3 Comparison table of recognition accuracies (%) obtained by different methods on the NTU RGB+D dataset (Cross-Subject (CS) and Cross-View (CV)) with 5%, 10%, 20%, and 40% labeled data in the training set

[0106]

[0107] The present invention also provides a semi-supervised skeleton point-based behavior recognition system, including:

[0108] A data acquisition module for acquiring human skeleton data; the human skeleton data includes a skeleton graph of learned connections and a skeleton graph of structural connections.

[0109] A feature extraction module, configured to use a multi-granularity anchor contrast learning model to extract features based on the human skeleton data, so as to obtain local features, global features, and context features; the multi-granularity anchor contrast learning model includes a graph convolutional neural network and a context graph convolutional neural network; the total loss function of the multi-granularity anchor contrast learning model includes a multi-granularity anchor contrast loss and an identification loss; the local features include a first local feature and a second local feature; the global features include a first global feature and a second global feature; the context features include a first context feature and a second context feature.

[0110] A fusion module, configured to fuse the local features, the global features, and the context features to obtain a fused feature.

[0111] An identification module, configured to obtain a human behavior recognition result according to the fused feature.

[0112] As an optional implementation manner, the feature extraction module specifically includes:

[0113] A first extraction unit, configured to use the graph convolutional neural network to extract a first global feature and a first local feature according to the skeleton graph of the learned connection.

[0114] A second extraction unit, configured to use the context graph convolutional neural network to extract a first context feature according to the skeleton graph of the learned connection.

[0115] A third extraction unit, configured to use the graph convolutional neural network to extract a second global feature and a second local feature according to the skeleton graph of the structural connection.

[0116] A fourth extraction unit, configured to use the context graph convolutional neural network to extract a second context feature according to the skeleton graph of the structural connection.

[0117] As an optional implementation manner, the method for determining the total loss function specifically includes:

[0118] Determine an anchor point adjacency matrix and a sample adjacency matrix according to the fused feature.

[0119] Determine a multi-granularity anchor contrast loss according to the anchor point adjacency matrix and the sample adjacency matrix.

[0120] Determine an identification loss according to the fused feature.

[0121] Determine a total loss function according to the contrast loss and the identification loss.

[0122] As an optional implementation manner, the determining of the anchor point adjacency matrix and the sample adjacency matrix according to the fused feature specifically includes:

[0123] Construct an anchor graph with the fusion feature as a node.

[0124] Determine anchor points according to the fusion feature using a clustering algorithm.

[0125] Determine an anchor point adjacency matrix and a sample adjacency matrix according to the fusion feature, the anchor graph, and the anchor points.

[0126] As an optional implementation manner, the fusion module specifically includes:

[0127] A fusion unit for fusing the first global feature, the first local feature, the first context feature, the second global feature, the second local feature, and the second context feature to obtain a fusion feature.

[0128] The present invention proposes a new multi-granularity anchor contrast representation learning model (MAC-learning), aiming to learn the potential semantic connections of human joints and then obtain multi-granularity action representations. To avoid the interference of noise and abnormal samples on fuzzy pairs, the present invention for the first time uses an anchor graph of anchor points and samples adjacent to capture high-confidence soft positive / negative pairs in contrastive representation learning, and designs a more reliable multi-granularity anchor contrast loss (MAC-loss), which measures the (in)consistency between high-confidence soft (negative) positive pairs based on the anchor graph, rather than the hard (negative) positive pairs in traditional contrastive losses, and achieves good performance.

[0129] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The same or similar parts among the various embodiments can be referred to each other. For the system disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the description in the method part.

[0130] Specific examples are used in this article to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation on the present invention.< / s>

Claims

1. A semi-supervised skeleton point-based behavior recognition method, characterized in that, Including: Obtain human body skeleton data; the human body skeleton data includes a skeleton graph of learning connections and a skeleton graph of structural connections; Extract features using a multi-granularity anchor contrast learning model based on the human body skeleton data to obtain local features, global features, and context features; The multi-granularity anchor contrast learning model includes a graph convolutional neural network and a context graph convolutional neural network; the total loss function of the multi-granularity anchor contrast learning model includes a multi-granularity anchor contrast loss and an identification loss; the local features include a first local feature and a second local feature; the global features include a first global feature and a second global feature; the context features include a first context feature and a second context feature; Fuse the local features, the global features, and the context features to obtain a fused feature; Determine an anchor point adjacency matrix and a sample adjacency matrix based on the fused feature; Determine the multi-granularity anchor contrast loss based on the anchor point adjacency matrix and the sample adjacency matrix; Determine the identification loss based on the fused feature; Determine the total loss function based on the contrast loss and the identification loss; Obtain the human behavior recognition result based on the fused feature.

2. The behavior recognition method based on semi-supervised skeleton points according to claim 1, wherein, The step of extracting features using a multi-granularity anchor contrast learning model based on the human body skeleton data to obtain local features, global features, and context features specifically includes: Extract the first global feature and the first local feature using the graph convolutional neural network based on the skeleton graph of learning connections; Extract the first context feature using the context graph convolutional neural network based on the skeleton graph of learning connections; Extract the second global feature and the second local feature using the graph convolutional neural network based on the skeleton graph of structural connections; Extract the second context feature using the context graph convolutional neural network based on the skeleton graph of structural connections.

3. The behavior recognition method based on semi-supervised skeleton points according to claim 1, wherein, The step of determining an anchor point adjacency matrix and a sample adjacency matrix based on the fused feature specifically includes: Construct an anchor graph with the fused feature as nodes; Determine anchor points using a clustering algorithm based on the fused feature; Determine the anchor point adjacency matrix and the sample adjacency matrix based on the fused feature, the anchor graph, and the anchor points.

4. The method for behavior recognition based on semi-supervised skeleton points according to claim 1, characterized in that, The step of fusing the local features, the global features, and the context features to obtain a fused feature specifically includes: Fuse the first global feature, the first local feature, the first context feature, the second global feature, the second local feature, and the second context feature to obtain a fused feature.

5. A semi-supervised skeleton point-based behavior recognition system, characterized in that, Including: A data acquisition module for obtaining human body skeleton data; The human body skeleton data includes a skeleton graph of learning connections and a skeleton graph of structural connections; A feature extraction module, which is used to extract features by using a multi-granularity anchor contrast learning model according to the human skeleton data, so as to obtain local features, global features and context features; the multi-granularity anchor contrast learning model includes a graph convolutional neural network and a context graph convolutional neural network; the total loss function of the multi-granularity anchor contrast learning model includes a multi-granularity anchor contrast loss and an identification loss; the local features include a first local feature and a second local feature; the global features include a first global feature and a second global feature; the context features include a first context feature and a second context feature; A fusion module, which is used to fuse the local features, the global features and the context features to obtain a fused feature; An identification module, which is used to obtain a human behavior recognition result according to the fused feature; Determine an anchor point adjacency matrix and a sample adjacency matrix according to the fused feature; Determine a multi-granularity anchor contrast loss according to the anchor point adjacency matrix and the sample adjacency matrix; Determine an identification loss according to the fused feature; Determine a total loss function according to the contrast loss and the identification loss.

6. The semi-supervised skeleton point-based behavior recognition system according to claim 5, wherein The feature extraction module specifically includes: A first extraction unit, which is used to extract a first global feature and a first local feature by using the graph convolutional neural network according to the skeleton graph of the learned connection; A second extraction unit, which is used to extract a first context feature by using the context graph convolutional neural network according to the skeleton graph of the learned connection; A third extraction unit, which is used to extract a second global feature and a second local feature by using the graph convolutional neural network according to the skeleton graph of the structural connection; A fourth extraction unit, which is used to extract a second context feature by using the context graph convolutional neural network according to the skeleton graph of the structural connection.

7. The semi-supervised skeleton point-based behavior recognition system according to claim 5, wherein The determination of the anchor point adjacency matrix and the sample adjacency matrix according to the fused feature specifically includes: Taking the fused feature as a node to construct an anchor graph; Determining anchor points according to the fused feature by using a clustering algorithm; Determining an anchor point adjacency matrix and a sample adjacency matrix according to the fused feature, the anchor graph and the anchor points.

8. The semi-supervised skeleton point-based behavior recognition system according to claim 5, wherein, The fusion module specifically includes: A fusion unit, which is used to fuse the first global feature, the first local feature, the first context feature, the second global feature, the second local feature and the second context feature to obtain a fused feature.

Citation Information

Patent Citations

  • Construction method of behavior recognition deep network model and behavior recognition method

    CN111985343A

  • Calligraphy Chinese character judgment method based on feature fusion

    CN112597876A