Teacher action recognition method based on cross-modal attention

By constructing a video dataset of teachers' daily behavior and a cross-modal attention method, and combining text semantics and skeleton information, the problems of low recognition accuracy and poor stability in existing teacher action recognition methods are solved, and efficient and accurate teacher action recognition is achieved.

CN120913273APending Publication Date: 2025-11-07SHAANXI NORMAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511021788.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-24
Publication Date
2025-11-07

AI Technical Summary

Technical Problem

Existing teacher action recognition methods suffer from low recognition accuracy and poor stability, making it difficult to achieve objective and continuous monitoring of the entire teaching process. Furthermore, existing multimodal fusion methods face challenges in cross-modal semantic alignment and efficient fusion.

Method used

A cross-modal attention approach is adopted. By constructing a video dataset of teachers' daily behavior, combining text semantics and skeleton information, skeleton sequences are extracted using HR-Net and a 3D human pose estimation network, and action text sequences are generated through a large language model. An action recognition network is constructed and trained using cross-domain and cross-modal contrastive loss functions to achieve parallel operation of the skeleton encoder and text encoder, thereby improving recognition accuracy.

Benefits of technology

It significantly improved the accuracy of teacher action recognition and the computing speed of the recognition network, shortened the computing time, and achieved efficient fusion of cross-modal features and semantic expression.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120913273A_ABST
    Figure CN120913273A_ABST
Patent Text Reader

Abstract

The invention discloses a teacher action recognition method based on cross-modal attention. The method comprises the steps of data set construction, data preprocessing, generation of an enhanced skeleton sequence, generation of an action text sequence, construction of an action recognition network, training of the action recognition network and testing of the action recognition network. Due to the fact that the action recognition network formed by connecting the skeleton encoder and the text encoder in parallel is adopted, in the training step, the operation speed and precision and the recognition accuracy are improved, and the operation time is shortened; in the step of training the action recognition network, a cross-modal comparison loss function Lc is adopted, and cross-modal feature expression fused with semantic text prompts is achieved. Compared with seven existing action recognition methods, a comparison experiment is carried out on a daily behavior data set and an action recognition public data set NTU-RGB + D-60 of a teacher, the experiment result shows that the Top-1 accuracy of the method is improved, the method is remarkably superior to all comparison experiments, the beneficial effects of the method are verified, and the method can be used for teacher action recognition.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the technical field of computer, and particularly relates to a teacher action recognition method based on cross-modal attention. BACKGROUND

[0002] With the continuous development of intelligent education and education informatization, teacher behavior recognition, as a key link in the intelligent teaching system, has gradually attracted people's attention. The traditional teacher action recognition method mainly relies on manual observation and experience judgment, which has strong subjectivity and low efficiency, and it is difficult to realize objective and continuous monitoring of the whole teaching process. The existing methods mainly use single-view video acquisition or sensor input, which leads to weak behavior detail capture ability, is easily affected by occlusion and environmental changes, and the recognition accuracy and stability are difficult to guarantee. There are obvious deficiencies in the aspects of intelligence, refinement and scalability, and there is an urgent need for more efficient, robust and adaptable new teacher action recognition methods and technologies. In recent years, the skeleton-based action recognition method has become one of the research hotspots due to its strong robustness to interference such as light and background. However, the skeleton data in the current mainstream method only contains the geometric structure information of the joint nodes, and the semantic expression is insufficient, which is difficult to meet the demand of complex action fine-grained modeling. On the other hand, the feature space difference between different modalities is large, and the existing multi-modal fusion method often uses simple splicing or independent encoding module, which has problems such as cross-modal semantic alignment and multi-modal efficient fusion.

[0003] In the field of action recognition technology, one of the technical problems to be solved urgently is to provide a teacher action recognition method with high prediction accuracy and recognition rate. SUMMARY

[0004] The technical problem to be solved by the present application is to overcome the shortcomings of the prior art and provide a teacher action recognition method based on cross-modal attention, which is more suitable for text semantics and skeleton information and has accurate recognition results.

[0005] The technical solution adopted to solve the above technical problems is composed of the following steps:

[0006] (1) Data set construction

[0007] The behaviors that may occur in the daily life of teachers are sorted and collected to form a teacher daily behavior video data set, which includes drinking water f1, sitting down f2, standing up f3, clapping f4, reading f5, writing f6, making a phone call f7, pointing to something with fingers f8, looking at the wristwatch time f9, nodding or bowing f 10 , shaking head f 11 , crossing hands to show pause f 12 , sneezing or coughing f 13 , touching the neck to show discomfort f 14 , fanning with hands / paper or feeling hot f 15, and a total of 15 action labels; each action contains 948 videos, each video is 1-2 seconds, and the dataset is divided into training set and test set according to 8:2.

[0008] (2) Data preprocessing

[0009] The video in the teacher daily behavior video dataset is input into the HR-Net 2D human pose estimation network to extract a two-dimensional human skeleton, and then a 3D human pose estimation network is used to extract a three-dimensional human skeleton, so as to obtain an original skeleton sequence (N, 3, T, V, M) corresponding to each video, wherein N is the number of samples, 3 is the number of channels, T is the number of frames, V is the number of joints, and M is the number of human bodies.

[0010] (3) Generating enhanced skeleton sequence

[0011] Random occlusion and time sequence disturbance conventional data enhancement operations are applied to the original skeleton sequence, and structure-constrained time and space masks are applied to obtain an enhanced skeleton sequence.

[0012] (4) Generating action text sequence

[0013] The 15 action labels are input into the large language model GPT-3 to generate descriptions containing 15 action features, and the skeleton sequences of the same category in the original skeleton sequence are mapped to the same text description, so as to obtain a text sequence consistent with the number of original skeleton sequences.

[0014] (5) Constructing action recognition network

[0015] The action recognition network is composed of a skeleton encoder and a text encoder in parallel.

[0016] The skeleton encoder is composed of a query encoder and a key encoder in parallel.

[0017] (6) Training action recognition network

[0018] 1) Constructing loss function

[0019] The loss function L includes a cross-domain contrast loss function L d , a cross-modal contrast loss function L c , and the loss function L is constructed according to formula (1):

[0020] L=L d +λL c (1)

[0021] Wherein, λ is a weight, and λ∈(0,1].

[0022] 2) Pre-training action recognition network

[0023] The original skeleton sequence of the training set is generated into an enhanced skeleton sequence through step (3), and the action text sequence of step (4) is obtained, and the enhanced skeleton sequence and the action text sequence are input into the action recognition network for pre-training, and the pre-training parameters are: the total number of pre-training rounds is 450 rounds, the temperature coefficient is 0.2, and the momentum coefficient is 0.999; during the pre-training process, the training batch is 64, the initial learning rate is 0.01, and the learning rate is decayed to 0.001 at the 350th round.

[0024] 3) Training the action recognition network

[0025] The original skeleton sequence of the training set is input into the pre-trained action recognition network for training and fine-tuning; the training parameters are: the total number of training rounds is 80 rounds, the training batch is 1024, the initial learning rate is 2, the hardware GPU is NVIDIA RTXA6000 GPU, and the training is performed until the loss function L converges.

[0026] (7) Testing the action recognition network

[0027] The original skeleton sequence of the test set is input into the trained action recognition network, and f1, f2, …, f 15 The action recognition accuracy of the action category.

[0028] In step (2) of the present application, the query encoder is composed of feature extraction layer 1, dimension transformation layer 1, feature mapping layer 1 and Transformer encoding layer 1 connected in sequence.

[0029] The key encoder of the present application is composed of feature extraction layer 2, dimension transformation layer 2, feature mapping layer 2 and Transformer encoding layer 2 connected in sequence.

[0030] The feature extraction layer 1 of the present application is composed of a spatial graph convolution network and a temporal graph convolution network connected in sequence. The structure of the feature extraction layer 2 of the present application is the same as that of the feature extraction layer 1.

[0031] The text encoder of the present application is composed of an embedding layer, a Transformer encoding layer 3 and a linear projection layer connected in sequence.

[0032] In step (5) of the present application, the loss function of the action recognition network is constructed in 1) as follows: c The construction method of the cross-modal contrast loss function L

[0033] The cross-modal contrast loss function L c is constructed according to formula (2):

[0034]

[0035]

[0036]

[0037]

[0038]

[0039]

[0040] q′ t =MLP(q t )

[0041]

[0042] where MLP(·) denotes a three-layer perceptron, the first layer is a linear layer for dimension compression, the second layer is a ReLU nonlinear activation function, and the third layer is a linear layer; Softmax(·) is an activation function, d is a projection dimension, and τ c is a temperature hyperparameter for cross-modal contrast; I is a diagonal label vector; CrossEntropy(·) is a cross-entropy loss function; q t represents a text feature.

[0043] In formula (2), d is a projection dimension, and d is in a range of 2 0 ~ 2 12 ; τ c is a temperature hyperparameter for cross-modal contrast, and τ c ∈(0,1].

[0044] In formula (2), d is a projection dimension, and d is best in a range of 2 9 ; τ c is a temperature hyperparameter for cross-modal contrast, and v c is best in a range of 0.5.

[0045] Since the motion recognition network composed of the skeleton encoder and the text encoder in parallel is adopted in the application, the operation speed, the precision, and the recognition accuracy are improved, and the operation time is shortened in the training step; in the training motion recognition network step, the cross-modal contrast loss function L c is adopted, and the cross-modal feature expression of the fused semantic text prompt is realized. The comparative experiments of the application and seven existing motion recognition methods are carried out on the teacher daily behavior data set and the public data set NTU-RGB+D-60, and the experimental results show that the Top-1 accuracy of the application is improved, is significantly better than each comparative experiment, and verifies the beneficial effects of the application, and the application can be used for teacher motion recognition. BRIEF DESCRIPTION OF DRAWINGS

[0046] Figure 1 is a flowchart of embodiment 1 of the present application.

[0047] Figure 2 is a structural diagram of the action recognition network.

[0048] Figure 3 is Figure 2 is a structural diagram of the skeleton encoder.

[0049] Figure 4 is Figure 3 is a structural diagram of the feature extraction layer.

[0050] Figure 5 is Figure 2 is a structural diagram of the Chinese text encoder. DETAILED DESCRIPTION

[0051] The present application will be further described below in conjunction with the accompanying drawings and examples, but the present application is not limited to the following embodiments.

[0052] Embodiment 1

[0053] The teacher action recognition method based on cross-modal attention of the present embodiment is composed of the following steps (see Figure 1 ):

[0054] (1) Data set construction

[0055] The behaviors that may occur in the daily life of teachers are sorted and collected to form a teacher daily behavior video data set, which includes drinking water f1, sitting down f2, standing up f3, clapping f4, reading f5, writing f6, making a phone call f7, pointing to something f8, looking at the wristwatch time f9, nodding or bowing f 10 , shaking head f 11 , crossing hands to show pause f 12 , sneezing or coughing f 13 , touching the neck to show discomfort f 14 , fanning with hands / paper or feeling hot f 15 , a total of 15 action labels; each action contains 948 videos, each video is 1-2 seconds, and the data set is divided into a training set and a test set in a ratio of 8:2.

[0056] (2) Data preprocessing

[0057] The videos in the teacher daily behavior video data set are input into the HR-Net 2D human pose estimation network to extract the two-dimensional human skeleton, and then the 3D human pose estimation network is used to extract the three-dimensional human skeleton, so as to obtain the 3D human skeleton of each

[0058] The original skeleton sequence (N,3,T,V,M) corresponding to the video, where N is the number of samples, 3 is the number of channels, T is the number of frames, V is the number of joints, and M is the number of human bodies.

[0059] (3) Generate enhanced backbone sequence

[0060] The original skeleton sequence is subjected to random occlusion and temporal perturbation conventional data augmentation operations, while temporal and spatial masks with structural constraints are applied to obtain the augmented skeleton sequence.

[0061] (4) Generate action text sequence

[0062] The 15 action labels are input into the large language model GPT-3 to generate descriptions containing the features of the 15 action categories. The skeleton sequences of the same category in the original skeleton sequence are mapped to the same text descriptions to obtain a text sequence with the same number of skeleton sequences as the original.

[0063] (5) Constructing an action recognition network

[0064] Figure 2 A schematic diagram of the action recognition network in this embodiment is provided. Figure 2 In this embodiment, the action recognition network consists of a skeleton encoder and a text encoder connected in parallel.

[0065] Figure 3 Given Figure 2 A schematic diagram of the skeleton encoder. Figure 3 In this embodiment, the skeleton encoder is composed of a query encoder and a key encoder connected in parallel.

[0066] The query encoder in this embodiment is composed of a feature extraction layer 1, a dimension transformation layer 1, a feature mapping layer 1, and a Transformer encoding layer 1 connected in series.

[0067] The key encoder in this embodiment is composed of a feature extraction layer 2, a dimension transformation layer 2, a feature mapping layer 2, and a Transformer encoding layer 2 connected in series.

[0068] Figure 4 Given Figure 3 A schematic diagram of the structure of feature extraction layer 1. Figure 4 In this embodiment, feature extraction layer 1 is composed of a spatial graph convolutional network and a temporal graph convolutional network connected in series. The structure of feature extraction layer 2 in this embodiment is the same as that of feature extraction layer 1.

[0069] Figure 5 Given Figure 2 A schematic diagram of the structure of a Chinese text encoder. Figure 5 In this embodiment, the text encoder is composed of an embedding layer, a Transformer encoding layer 3, and a linear projection layer connected in series.

[0070] Since the motion recognition network constituted by the skeleton encoder and the text encoder in parallel is adopted in the application, the operation speed, precision and recognition accuracy are improved, and the operation time is shortened in the training step.

[0071] (6) Training the motion recognition network

[0072] 1) Constructing the loss function

[0073] The loss function L includes a cross-domain contrast loss function L d , a cross-modal contrast loss function L c , and the loss function L is constructed according to formula (1):

[0074] L = L d + λL c (1)

[0075] Wherein, λ is a weight, λ ∈ (0, 1], and λ of the embodiment is 0.5.

[0076] The cross-modal contrast loss function L c of the embodiment is constructed as follows:

[0077] The cross-modal contrast loss function L c is constructed according to formula (2):

[0078]

[0079]

[0080]

[0081]

[0082]

[0083]

[0084] q′ t = MLP(q t )

[0085]

[0086] Wherein, MLP(·) represents a three-layer perception machine, the first layer is a linear layer for dimension compression, the second layer is a ReLU nonlinear activation function, and the third layer is a linear layer; Softmax(·) is an activation function, d is a projection dimension, d takes a value range of 2 0 ~ 2 12 , and d of the embodiment takes a value of 2 9 ; τ cτ is the temperature hyperparameter for cross-modal comparison. c ∈(0,1], τ in this embodiment c The value is 0.5, I is the diagonal label vector, CrossEntropy(·) is the cross-entropy loss function, and q t Represents text features.

[0087] Because this invention employs a cross-modal contrast loss function L c This invention achieves cross-modal feature representation that integrates semantic text prompts. Comparative experiments were conducted on a teacher daily behavior dataset and the public dataset NTU-RGB+D-60 using this invention and seven existing action recognition methods. The experimental results show that the Top-1 accuracy of this invention is significantly improved compared to the comparative experiments, validating the beneficial effects of this invention and its applicability to teacher action recognition.

[0088] Construct the cross-domain contrast loss function L according to equation (3). d :

[0089]

[0090]

[0091]

[0092]

[0093]

[0094] h(q g ,k s ) = exp(q g ·k s / τ)

[0095] h(q g ,k t ) = exp(q g ·k t / τ)

[0096] h(q s ,k g ) = exp(q s ·k g / τ)

[0097] h(q g ,k g ) = exp(q g ·k g / τ)

[0098] q s =F s (z s),

[0099] q t =F t (z t ),

[0100] q g =F g ([z t ,z s ])

[0101] k s =F s (z s )

[0102] k t =F t (z t )

[0103] k g =F g ([z t ,z s ])

[0104] Among them, F s F t F g is the projection function; [,] is the stitching operation; τ is the temperature hyperparameter, τ∈(0,1], and in this embodiment, the value of τ is 0.5; M is the extracted feature queue, z s Represents the generated spatial features, z t Representing the time feature, q g and k g Represents skeletal features.

[0105] 2) Pre-trained action recognition network

[0106] The original skeleton sequence of the training set is used to generate an enhanced skeleton sequence in step (3), and the action text sequence of step (4) is obtained. The enhanced skeleton sequence and the action text sequence are input into the action recognition network for pre-training. The pre-training parameters are: a total of 450 pre-training rounds, a temperature coefficient of 0.2, and a momentum coefficient of 0.999. During the pre-training process, the training batch size is 64, the initial learning rate is 0.01, and the learning rate decays to 0.001 in the 350th round.

[0107] 3) Training the action recognition network

[0108] The original skeleton sequences of the training set were input into the pre-trained action recognition network for training and fine-tuning. The training parameters were: a total of 80 training rounds, a training batch size of 1024, an initial learning rate of 2, and an NVIDIA RTX A6000 GPU. Training continued until the loss function L converged.

[0109] (7) Test action recognition network

[0110] The original skeleton sequence of the test set is input into the trained action recognition network, and f1, f2, …, f 15 The action recognition accuracy of the human action category.

[0111] The teacher action recognition method based on cross-modal attention is completed.

[0112] Embodiment 2

[0113] The teacher action recognition method based on cross-modal attention of this embodiment consists of the following steps:

[0114] (1) Data set construction

[0115] This step is the same as embodiment 1.

[0116] (2) Data preprocessing

[0117] This step is the same as embodiment 1.

[0118] (3) Generate enhanced skeleton sequence

[0119] This step is the same as embodiment 1.

[0120] (4) Generate action text sequence

[0121] This step is the same as embodiment 1.

[0122] (5) Construct action recognition network

[0123] This step is the same as embodiment 1.

[0124] (6) Train action recognition network

[0125] 1) Construct loss function L

[0126] The loss function L includes the cross-domain contrast loss function L d , the cross-modal contrast loss function L c , and the loss function L is constructed according to formula (1).

[0127] The expression of formula (1) is the same as embodiment 1.

[0128] In formula (1), λ is the weight, λ ∈ (0, 1], and λ of this embodiment is 0.1.

[0129] The cross-modal contrast loss function L c of this embodiment is constructed as follows:

[0130] The cross-modal contrast loss function L c is constructed according to formula (2):

[0131] The expression of equation (2) is the same as that in Example 1.

[0132] In equation (2), d is the projection dimension, and the value of d ranges from 2. 0 ~2 12 In this embodiment, d is taken as 2. 0 ;τ c τ is the temperature hyperparameter for cross-modal comparison. c ∈(0,1], τ in this embodiment c The value is 0.1. The other parameters and variables in this step, as well as their value ranges, are the same as in Example 1. The other steps in this step are the same as in Example 1.

[0133] Construct the cross-domain contrast loss function L according to equation (3). d :

[0134] The expression of equation (3) is the same as that in Example 1.

[0135] In equation (3), τ is the temperature hyperparameter, τ∈(0,1], and in this embodiment, the value of τ is 0.1. The parameters, variables, and value ranges in this step are the same as in embodiment 1. The other steps in this step are the same as in embodiment 1.

[0136] The other steps are the same as in Example 1. This completes the teacher action recognition method based on cross-modal attention.

[0137] Example 3

[0138] The teacher action recognition method based on cross-modal attention in this embodiment consists of the following steps:

[0139] (1) Dataset Construction

[0140] The steps are the same as in Example 1.

[0141] (2) Data preprocessing

[0142] The steps are the same as in Example 1.

[0143] (3) Generate enhanced backbone sequence

[0144] The steps are the same as in Example 1.

[0145] (4) Generate action text sequence

[0146] The steps are the same as in Example 1.

[0147] (5) Constructing an action recognition network

[0148] The steps are the same as in Example 1.

[0149] (6) Training the action recognition network

[0150] 1) constructing a loss function L

[0151] The loss function L includes a cross-domain contrast loss function L d , a cross-modal contrast loss function L c , and the loss function L is constructed according to formula (1):

[0152] The expression of formula (1) is the same as that of embodiment 1.

[0153] In formula (1), λ is a weight, λ ∈ (0, 1], and λ of the embodiment is 1.

[0154] The cross-modal contrast loss function L c of the embodiment is constructed as follows:

[0155] The cross-modal contrast loss function L c is constructed according to formula (2):

[0156] The expression of formula (2) is the same as that of embodiment 1.

[0157] In formula (2), d is a projection dimension, d takes a value in a range of 2 0 ~ 2 12 , and d of the embodiment takes a value of 2 12 ; τ c is a temperature hyperparameter for cross-modal contrast, τ c ∈ (0, 1], and τ c of the embodiment takes a value of 1. The other parameters and variables and value ranges in this step are the same as those of embodiment 1. The other steps of this step are the same as those of embodiment 1.

[0158] The cross-domain contrast loss function L d is constructed according to formula (3):

[0159] The expression of formula (3) is the same as that of embodiment 1.

[0160] In formula (3), τ is a temperature hyperparameter, τ ∈ (0, 1], and τ of the embodiment takes a value of 1. The other parameters and variables and value ranges in this step are the same as those of embodiment 1. The other steps of this step are the same as those of embodiment 1.

[0161] The other steps are the same as those of embodiment 1. The cross-modal attention-based teacher action recognition method is completed.

[0162] To verify the beneficial effects of the application, a comparative experiment is performed by using the cross-modal attention-based teacher action recognition method of embodiment 1 of the application.

[0163] Experiment 1

[0164] The teacher action recognition method based on cross-modal attention of the embodiment 1 of the present application (referred to as the present application method) and "Sun S, Liu D, Dong J, et al. Unified multi-modal unsupervised representation learning for skeleton-based action understanding [C]. Proceedings of the ACM International Conference on Multimedia. 2023: 2973-2984." (referred to as comparative experiment 1), "Wu C, Wu X J, Kittler J, et al. Scd-net: Spatiotemporal clues disentanglement network for self-supervised skeleton-based action recognition [C]. Proceedings of the AAAI Conference on Artificial Intelligence. 2024, 38(6): 5949-5957." (referred to as comparative experiment 2) were compared on the teacher daily behavior dataset, and the teacher action recognition results were evaluated by Top-1 accuracy. The experimental results are shown in Table 1.

[0165] Table 1 Experimental results of the present application method and comparative experiments 1 and 2

[0166] Experimental group Top-1 accuracy (%) Comparative experiment 1 88.6 Comparative experiment 2 90.7 The method of the present application 91.1

[0167] As can be seen from Table 1, compared with comparative experiments 1 and 2, the Top-1 accuracy of the present application is improved. The Top-1 accuracy of the present application is improved by 2.5% compared with comparative experiment 1, and the Top-1 accuracy is improved by 0.4% compared with comparative experiment 2. The above results show that the present application is superior to each of the comparative experiments in the Top-1 accuracy of teacher action recognition, verifying the beneficial effects of the proposed method, which can be used for teacher action recognition.

[0168] Experiment 2

[0169] The teacher action recognition method based on cross-modal attention of the embodiment 1 of the present application (referred to as the method of the present application) is used, and "Guo T, Liu H, Chen Z, et al. Contrastive learning from extremely augmented skeleton sequences for self-supervised action recognition [C]. Proceedings of the AAAI Conference on Artificial Intelligence. 2022, 36(1): 762-770." (referred to as comparative experiment 3), "Sun S, Liu D, Dong J, et al. Unified multi-modal unsupervised representation learning for skeleton-based action understanding [C]. Proceedings of the ACM International Conference on Multimedia. 2023: 2973-2984." (referred to as comparative experiment 4), "Zhang J, Lin L, Liu J. Prompted contrast with masked motion modeling: Towards versatile 3d action representation learning [C]. Proceedings of the ACM International Conference on Multimedia. 2023: 7175-7183." (referred to as comparative experiment 5), "Lin L, Zhang J, Liu J. Mutual information driven equivariant contrastive learning for 3d action representation learning [J]. IEEE Transactions on Image Processing, 2024." (referred to as comparative experiment 6), "Weng W, Wang H, Wang J, et al. Usdrl: Unified skeleton-based dense representation learning with multi-grained feature decorrelation [C].Proceedings of the AAAI Conference on Artificial Intelligence.2025, 39(8):8332-8340.” (referred to as Comparative Experiment 7) carried out comparative experiments on the public data set NTU-RGB+D-60, and the action recognition results were evaluated by Top-1 accuracy, and the experimental results are shown in Table 2.

[0170] Table 2 Experimental results of the method of the present application and Comparative Experiments 3-7

[0171] Experimental group Top-1 accuracy (%) Comparative experiment 3 74.3 Comparative experiment 4 82.3 Comparative experiment 5 83.9 Comparative experiment 6 83.9 Comparative experiment 7 85.2 The method of the present application 86.8

[0172] As can be seen from Table 2, compared with Comparative Experiments 3-7, the Top-1 accuracy of the present application is improved. The Top-1 accuracy of the present application is improved by 12.5% compared with Comparative Experiment 3, by 4.5% compared with Comparative Experiment 4, by 2.9% compared with Comparative Experiment 5, by 2.9% compared with Comparative Experiment 6, and by 1.6% compared with Comparative Experiment 7. The above results show that compared with the comparative experiments, the present application is significantly superior to each comparative experiment in the Top-1 accuracy of action recognition, further verifying the beneficial effects of the proposed method, which can be used for teacher action recognition.

Claims

1. A teacher action recognition method based on cross-modal attention, characterized in that Comprise the following steps: (1) Data set construction The behaviors that may occur in the daily life of teachers are sorted and collected to form a teacher daily behavior video dataset, which includes drinking water f1, sitting down f2, standing up f3, clapping f4, reading f5, writing f6, making a phone call f7, pointing to something with fingers f8, looking at a wristwatch f9, nodding or bowing f 10 , shaking head f 11 , crossing hands to show stop f 12 , sneezing or coughing f 13 , touching neck to show discomfort f 14 , fanning with hands / paper or feeling hot f 15 , a total of 15 action tags; each action contains 948 videos, each video is 1-2 seconds, and the dataset is divided into a training set and a test set in a ratio of 8:2; (2) Data preprocessing Input the video in the teacher daily behavior video data set into the HR-Net 2D human body posture estimation network to extract the two-dimensional human body skeleton, and then use the 3D human body posture estimation network to extract the three-dimensional human body skeleton, to obtain the original skeleton sequence (N, 3, T, V, M) corresponding to each video, wherein N is the sample number, 3 is the channel number, T is the frame number, V is the number of joints, and M is the number of human bodies; (3) Generating enhanced skeleton sequence Random occlusion and time sequence disturbance conventional data enhancement operation are applied to the original skeleton sequence, and structure constraint time and space mask is applied at the same time, to obtain the enhanced skeleton sequence; (4) Generating action text sequence Input the 15 types of action labels into the large language model GPT-3 to generate descriptions containing 15 types of action features, map the skeleton sequences of the same category in the original skeleton sequence to the same text description, and obtain a text sequence consistent with the number of original skeleton sequences; (5) Constructing action recognition network The action recognition network is composed of a skeleton encoder and a text encoder in parallel; The skeleton encoder is composed of a query encoder and a key encoder in parallel; (6) Training the action recognition network 1) Constructing loss function The loss function L includes a cross-domain contrast loss function L d , a cross-modal contrast loss function L c The loss function L is constructed according to formula (1): L = L d + λL c (1) Wherein, λ is a weight, λ∈(0, 1]; 2) Pre-training the action recognition network The original skeleton sequence of the training set is generated into an enhanced skeleton sequence through step (3), and the action text sequence of step (4) is obtained, the enhanced skeleton sequence and the action text sequence are input into the action recognition network for pre-training, the pre-training parameters are: the total number of pre-training rounds is 450, the temperature coefficient is 0.2, and the momentum coefficient is 0.999; during the pre-training process, the training batch is 64, the initial learning rate is 0.01, and the learning rate is decayed to 0.001 at the 350th round; 3) Training the action recognition network The original skeleton sequence of the training set is input into the pre-trained action recognition network for training and fine-tuning; the training parameters are: the total number of training rounds is 80, the training batch is 1024, the initial learning rate is 2, the hardware GPU is NVIDIA RTXA6000 GPU, and the training is stopped when the loss function L converges; (7) Testing the action recognition network The original skeleton sequence of the test set is input into the trained action recognition network, and f1, f2, …, f 15 Action recognition accuracy of action categories. 2.The teacher action recognition method based on cross-modal attention according to claim 1, characterized in that: In step (2) of constructing the action recognition network, the query encoder is composed of a feature extraction layer 1, a dimension transformation layer 1, a feature mapping layer 1, and a Transformer encoding layer 1 connected in sequence; The key encoder is composed of a feature extraction layer 2, a dimension transformation layer 2, a feature mapping layer 2, and a Transformer encoding layer 2 connected in sequence. 3.The teacher action recognition method based on cross-modal attention according to claim 2, characterized in that: The feature extraction layer 1 is composed of a spatial graph convolution network and a temporal graph convolution network connected in series; the structure of the feature extraction layer 2 is the same as that of the feature extraction layer 1. 4.The teacher action recognition method based on cross-modal attention according to claim 1, characterized in that: The text encoder is composed of an embedding layer, a Transformer encoding layer 3, and a linear projection layer connected in sequence. 5.The teacher action recognition method based on cross-modal attention according to claim 1, characterized in that In the step (5) of constructing the loss function 1) for training the action recognition network, the cross-modal contrast loss function L c is constructed as follows: The cross-modal contrast loss function L is constructed according to formula (2) c : q' t = MLP(q t ) where MLP(·) represents a three-layer perceptron, the first layer is a linear layer for dimension compression, the second layer is a ReLU nonlinear activation function, and the third layer is a linear layer; Softmax(·) is an activation function, d is the projection dimension, and τ c is the temperature hyperparameter for cross-modal contrast; I is a diagonal label vector; CrossEntropy(·) is a cross-entropy loss function; q t represents the text feature. 6.The teacher action recognition method based on cross-modal attention according to claim 5, characterized in that In equation (2), d is the projection dimension, and the value of d ranges from 2. 0 ~2 12 The τ mentioned above c τ is the temperature hyperparameter for cross-modal comparison. c ∈(0,1).

7. The cross-modal attention based teacher action recognition method according to claim 5 or 6, characterized in that In equation (2), d is the projection dimension, and d takes a value of 2. 9 The τ mentioned above c τ is the temperature hyperparameter for cross-modal comparison. c The value is 0.5.