Multi-label group activity learning and classification method based on equidistant depth embedding

Through an isometric deep embedding network, the original low-dimensional label vector is converted into a high-dimensional isometric embedding vector is solved, and the inequalities and dependencies in multi-label group activity recognition is achieved, and more accurate group activity recognition is achieved.

CN120298738APending Publication Date: 2025-07-11SICHUAN UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202410036173.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-10
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

The existing multi-label group activity recognition technology is difficult to effectively deal with the problems of inequality between multi-label vectors and label dependence, resulting in inaccurate description of group activity in real scenarios.

Method used

The original low-dimensional label vector is converted into high-dimensional equidistant embedding vectors by using an isometric depth embedding network, and the group activity recognition model is trained by minimizing the cosine similarity loss, and the cosine distance between high-dimensional vectors is classified.

Benefits of technology

It effectively solves the problems of inequal distance and dependence between multi-label vectors, and improves the accuracy and consistency of group activity recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure BDA0004657809280000042
    Figure BDA0004657809280000042
  • Figure BDA0004657809280000043
    Figure BDA0004657809280000043
  • Figure BDA0004657809280000048
    Figure BDA0004657809280000048
Patent Text Reader

Abstract

The invention provides a multi-label group activity learning and classification method based on equidistant depth embedding, and mainly relates to a multi-label group activity identification problem. The method comprises the following steps: firstly, converting an original low-dimensional label vector into a high-dimensional equidistant embedding vector through an equidistant depth embedding network; then, taking a high-dimensional equidistant embedded vector as a new label, taking minimization of cosine similarity loss as a target, and training a group activity recognition model; and finally, calculating the cosine distance between the high-dimensional vector output by the group activity identification network and each high-dimensional vector label, and taking the high-dimensional label vector with the shortest cosine distance as a classification result. Due to the fact that the original low-dimensional label vectors and the high-dimensional label vectors have the one-to-one correspondence relation, after the output high-dimensional vectors are obtained, the final activity type can be obtained according to the output high-dimensional vectors. The problem of non-equidistant label vector and label dependency in the current group activity identification task is effectively solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the problem of multi-label group activity recognition in the field of deep learning, and particularly to a multi-label group activity learning and classification method based on isometric depth embedding. Background Art

[0002] Currently, multi-label group activity recognition plays a crucial role in various computer vision systems. Existing research mainly extracts rich spatio-temporal features, such as individual behaviors, interaction relationships, and scene information in the group, to assist activity understanding. With the in-depth research, researchers found that group activities in real scenarios usually cannot be accurately described by a single word or phrase, and the traditional single-label annotation method is no longer sufficient to accurately describe group activities involving multiple individuals. To solve this problem, multi-label learning is introduced into this task. Compared with single-label learning, multi-label learning is more suitable for learning complex attributes in images, videos, and other modalities, and can better understand group activities.

[0003] As an important research task in the field of computer vision, multi-label group activity recognition has received extensive attention from relevant researchers at home and abroad. However, the current research on group activity recognition mainly focuses on model inference and cannot well handle the problems of non-isometric distances between multi-label vectors and label dependencies. To solve these problems, this patent proposes a multi-label group activity learning and classification method based on isometric depth embedding. The original low-dimensional label vectors are mapped into high-dimensional isometric embedding vectors through an isometric depth embedding network, and the obtained high-dimensional isometric embedding vectors are used as new labels to train the group recognition model. Then, a high-dimensional vector with the same dimension as the new label vector is obtained through the network. Finally, the cosine distance between the high-dimensional vector output by the recognition network and the labels of each high-dimensional vector is calculated, and the high-dimensional label vector with the shortest cosine distance is taken as the classification result. The final activity type is identified through the corresponding relationship between the original label vector and the high-dimensional vector. Summary of the Invention

[0004] The object of the present invention is to provide a multi-label group activity learning and classification method based on isometric depth embedding. First, the original low-dimensional label vectors are converted into high-dimensional isometric embedding vectors through an isometric depth embedding network. Then, with the high-dimensional isometric embedding vectors as new labels and minimizing the cosine similarity loss as the goal, the group activity recognition model is trained. Finally, the cosine distance between the high-dimensional vector output by the group activity recognition network and the labels of each high-dimensional vector is calculated, and the high-dimensional label vector with the shortest cosine distance is taken as the classification result. Since there is a one-to-one correspondence between the original low-dimensional label vectors and the high-dimensional label vectors, after obtaining the output high-dimensional vectors, the final activity type can be obtained accordingly, effectively solving the problems of non-isometric label vectors and label dependencies in the current group activity recognition task.

[0005] For the convenience of explanation, the following concepts are introduced first:

[0006] Vector equidistance: If the distances between vectors are equal pairwise, they are said to be equidistant; it should be noted that there are various definitions of vector distance, such as Euclidean distance, cosine distance, Hamming distance, etc. In this patent, cosine distance is taken as an example.

[0007] Pre-trained model: The training of a neural network requires a large amount of data, time, and sufficient computing resources. To avoid repeated training of the network, the model parameters of a model with good performance trained by other researchers are transferred to a model for a specific task and fine-tuned to meet the requirements of this task.

[0008] Inflated 3D ConvNet (I3D): A deep learning network architecture for image and video classification tasks. The design of I3D is based on the successful experience of two-dimensional convolutional neural networks. By extending it to three-dimensional convolutional operations, it can better process image and video data.

[0009] Group Activity Recognition (GAR): By analyzing and understanding the individual behaviors in a group to identify the overall group activity, it is an important issue in video understanding and is widely used in real-world scenarios such as sports competition video analysis, surveillance video recognition, and social behavior understanding.

[0010] Multi-Label Learning (MLL): As a type of machine learning method, multi-label learning aims to handle problems with multiple related labels. In multi-label learning, each sample can be associated with multiple labels, and there may be some association or dependence relationship between these labels; compared with traditional single-label learning, multi-label learning can more accurately describe the complexity and diversity of data.

[0011] The present invention specifically adopts the following technical solutions:

[0012] A multi-label group activity learning and classification method based on equidistant deep embedding, characterized in that:

[0013] a) Extract high-dimensional feature vectors for group activity recognition from video sequences or images through a feature extraction network;

[0014] b) Embed the original low-dimensional label vectors into a higher-dimensional space through an equidistant deep embedding network, and at the same time adopt a self-supervised learning method to ensure that the embedded high-dimensional label vectors are equidistant from each other;

[0015] c) Use the high-dimensional feature vectors and high-dimensional label vectors to implement the learning and classification of the group activity recognition model;

[0016] The method mainly includes the following steps:

[0017] (1) Label vector conversion: Convert the original low-dimensional label vector into a high-dimensional isometric embedded label vector through an isometric depth embedding network; this network is obtained by minimizing the isometric regularization loss;

[0018] (2) Classification model training: Use the high-dimensional isometric vector as the new label, and train the group activity recognition model with the goal of minimizing the cosine similarity loss; the output of this network is a high-dimensional vector with the same dimension as the new label vector;

[0019] (3) Group activity classification: Calculate the cosine distance between the high-dimensional vector output by the group activity recognition network and the labels of each high-dimensional vector, and take the high-dimensional label vector with the shortest cosine distance as the classification result; since there is a one-to-one correspondence between the original low-dimensional label vector and the high-dimensional label vector, after obtaining the output high-dimensional vector, the final activity type can be obtained accordingly.

[0020] The beneficial effects of the present invention are:

[0021] (1) By constructing an isometric depth embedding network, the problems caused by the non-isometry and dependence between multi-label vectors are effectively solved.

[0022] (2) The multi-label learning and classification method proposed by the present invention is not limited to the group activity recognition task, and is also effective for other multi-label classification tasks. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] Figure 1 It is the overall algorithm block diagram of the model.

[0024] Figure 2 It is the structural block diagram of the isometric depth embedding network.

[0025] Figure 3 It is the structural block diagram of the distance-aware classification network. DETAILED DESCRIPTION OF THE INVENTION

[0026] The present invention will be further described in detail below in conjunction with the accompanying drawings and embodiments. It is necessary to point out that the following embodiments are only used to further illustrate the present invention and cannot be understood as limiting the protection scope of the present invention. Those skilled in the art make some non-essential improvements and adjustments to the present invention according to the above-mentioned invention content for specific implementation, which should still fall within the protection scope of the present invention.

[0027] Such as Figure 1As shown, the method of the present invention mainly involves using a backbone network (the backbone is an I3D network, followed by an ROI module, and the cropping size is 5×5) to extract high-dimensional feature vectors for group activity recognition from video sequences or images; at the same time, converting the original label vector set into a high-dimensional isometric embedded label vector set through an isometric depth embedding network; and using the converted high-dimensional isometric embedded label vector as a new label to train a group activity recognition model, that is Figure 1 the high-dimensional feature vector learning shown in; subsequently, calculating the cosine distance between the high-dimensional vector output by the group activity recognition network and each high-dimensional vector label through a distance-aware classification module, and taking the high-dimensional label vector with the shortest cosine distance as the classification result. Through the one-to-one correspondence between the original label vector and the high-dimensional label vector, the activity type prediction result can be obtained. The detailed process is as follows:

[0028] To solve the problem of non-isometry between multi-label vectors in the original label vector set, an isometric depth embedding network is adopted in the present invention to convert the original label vector set into a high-dimensional isometric embedded label vector set. The network structure is as Figure 2 shown. For the convenience of understanding and explanation, the original label vector set and the embedded label vector set are defined as follows:

[0029]

[0030]

[0031] Among them, Θ is the original label vector set, Ω is the embedded label vector set, n represents the number of label combinations, c0 to c n-1 respectively represent the 0th to the (n - 1)th original label vectors, k represents the number of categories of the original label vectors, d represents the dimension of the embedded label vectors, and t0 to t n-1 respectively represent the 0th to the (n - 1)th embedded label vectors.

[0032] During the process of using the embedding network to convert the original low-dimensional sparse label vectors into high-dimensional embedded label vectors, to ensure stability, it is also necessary to perform a normalization operation on the embedded label vectors to eliminate amplitude fluctuations. The normalization formula is as follows:

[0033]

[0034] Among them, f em (·) represents the embedding network, and ∈ is a positive value close to zero to ensure that the denominator is not zero.

[0035] In the network, to achieve isometric embedded label vectors, the present invention uses an isometric regularization loss function, and the formula is as follows:

[0036]

[0037] Among them, n represents the number of embedded label vectors, represents the number of pairwise combinations of embedded label vectors; t i and t j respectively represent the i-th and j-th embedded label vectors; topk(Ω, δ) is a penalty term, which calculates the cosine distance between elements in the high-dimensional embedded label vector set and sums the top δ maximum distance values; λ is a balance coefficient.

[0038] After obtaining the isometric high-dimensional embedded label vectors through the isometric depth embedding network, it is necessary to use the obtained vectors to effectively classify group activities. For this purpose, a learning and classification strategy based on high-dimensional vectors is proposed in the present invention, which mainly involves high-dimensional feature vector learning and distance-aware classification network.

[0039] The purpose of high-dimensional feature vector learning is to use the isometric high-dimensional embedded vectors as new labels and minimize the cosine similarity loss as the goal to train the group activity recognition model. The function methods involved are as follows:

[0040] Given that the high-dimensional feature vector extracted by the backbone network is and the isometric high-dimensional embedded label vector obtained by the corresponding isometric depth embedding network is then the formula of the cosine similarity loss function used is expressed as:

[0041]

[0042] Among them, B and d respectively represent the batch size and dimension of the input vector samples, y i ∈Y and respectively represent the i-th output high-dimensional vector and its corresponding embedded label vector.

[0043] After obtaining the trained and optimized model, this model can be used for group activity recognition. The model outputs a high-dimensional vector with the same dimension as the new label vector, and then calculates the cosine distance between the high-dimensional vector output by the group activity recognition network and each high-dimensional vector label through the distance-aware classification network shown in Figure 3 and takes the high-dimensional label vector with the shortest cosine distance as the classification result. The detailed process is as follows:

[0044] As shown in Figure 3 in the distance-aware network, the formula (5) is used to calculate the cosine similarity between the vector y * output by the recognition network and each vector in the embedded label vector set Ω:

[0045]

[0046] Then search for the index i of the maximum cosine similarity *, and find the finally corresponding original label vector c * ∈Θ.

[0047]

[0048] According to the one-to-one correspondence between the original label vector and the high-dimensional label vector, the corresponding group activity type can be obtained, and the recognition of multi-label group activities can be realized.

Claims

1. A multi-label group active learning and classification method based on isometric depth embedding, characterized in that: a) Extract high-dimensional feature vectors for group activity recognition from video sequences or images through a feature extraction network; b) Embed the original low-dimensional label vectors into a higher-dimensional space through an isometric depth embedding network, and at the same time adopt a self-supervised learning method to ensure that the embedded high-dimensional label vectors are equidistant from each other; c) Use the high-dimensional feature vectors and high-dimensional label vectors to realize the learning and classification of the group activity recognition model; This method mainly includes the following steps: (1) Label vector conversion: Convert the original low-dimensional label vectors into high-dimensional isometric embedded label vectors through an isometric depth embedding network; this network is obtained by minimizing the isometric regularization loss; (2) Classification model training: Use the high-dimensional isometric embedded vectors as new labels, and train the group activity recognition model with the goal of minimizing the cosine similarity loss; the output of this network is a high-dimensional vector with the same dimension as the new label vector; (3) Group activity classification: Calculate the cosine distance between the high-dimensional vector output by the group activity recognition network and each high-dimensional vector label, and take the high-dimensional label vector with the shortest cosine distance as the classification result; since there is a one-to-one correspondence between the original low-dimensional label vectors and the high-dimensional label vectors, after obtaining the output high-dimensional vector, the final activity type can be obtained accordingly.

2. The multi-label group active learning and classification method based on isometric depth embedding according to claim 1, characterized in that In step (1), the isometric regularization loss function is: Among them, n represents the number of embedded label vectors, represents the number of pairwise combinations of high-dimensional embedded label vectors; t i and t j respectively represent the i-th and j-th high-dimensional embedded label vectors; topk(Ω,δ) is a penalty term, which calculates the cosine distance between elements in the high-dimensional embedded label vector set and sums the top δ maximum distance values; λ is a balance coefficient.

3. The multi-label group active learning and classification method based on isometric depth embedding according to claim 1, characterized in that In step (2), the cosine similarity loss function is: where B and d denote the batch size and dimension of the input vector samples, respectively, and y i and denote the i-th output high-dimensional vector and its corresponding embedded label vector, respectively.

4. The multi-label group active learning and classification method based on isometric depth embedding according to claim 1, characterized in that In step (3), the distance-based classification method is specifically: first calculate the cosine similarity between the output vector y* in the recognition network and each vector in the embedded label vector set, then find the high-dimensional label vector corresponding to the maximum cosine similarity as the classification result, and finally obtain the prediction result according to the correspondence between the high-dimensional label vector and the original label vector; the cosine similarity calculation formula is: The index i for finding the maximum cosine similarity * is as follows: Among them, n represents the number of embedded label vectors, and y * is the vector output in the recognition network, and s i is the cosine similarity between y * and each vector in the embedded label vector set, and t i represents the i-th high-dimensional embedded label vector.

Citation Information

Cited By

  • Network device classification

    US12483575B2