Scalable thinking chain guided few-sample continuous teaching behavior identification method

Through thinking chain guidance large language models (LLMs) for semantic expansion and triple relationship extraction, combined with the freezing of the backbone network of the visual language model, the teaching behavior recognition model is optimized, and the problem of insufficient semantic capture of verbs is solved, and high-accuracy, small-sample continuous teaching behavior recognition is achieved.

CN120564264APending Publication Date: 2025-08-29UNIV OF ELECTRONICS SCI & TECH OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510673483.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-23
Publication Date
2025-08-29

AI Technical Summary

Technical Problem

In the existing method of continuous teaching behavior recognition of small samples, the visual language model has weak ability to match behavior labels with its corresponding pictures, and lacks ability to capture verbs and their semantic details, resulting in insufficient accuracy in the recognition of new teaching behaviors.

Method used

Through thinking chain guidance large language models (LLMs) to mine semantic information in behavioral labels, perform semantic expansion and triple relationship extraction, combine the freezing of the backbone network of pre-trained visual language models, use prompt learning optimization model, and design multi-level cross-modal matching strategies for classification.

Benefits of technology

It improves the accuracy of the model's identification of teaching behaviors, solves the problem of catastrophic forgetting in continuous learning in small samples, and improves the ability to identify new teaching behaviors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120564264A_ABST
    Figure CN120564264A_ABST
Patent Text Reader

Abstract

The invention discloses a few-sample continuous teaching behavior recognition method guided by a scalable thinking chain, and relates to the field of image processing. Semantic knowledge of different levels in a behavior tag is mined by guiding a large language model (LLMs) through a thinking chain, and the semantic knowledge is condensed into triple knowledge of a (main, called and guest) structure, so that the problem that an existing pre-training visual language model is relatively weak in verb understanding capability is solved, and accurate understanding and recognition of behaviors are realized. Compared with a common few-sample continuous learning method, the method freezes the backbone network of the pre-trained visual language model, the model is trained only through prompt learning, and compared with traditional backbone network characterization adaptation tuning, the method has few training parameters, and the calculation complexity is greatly reduced. According to the method, few-sample continuous teaching behavior recognition tasks are carried out on a classroom scene data set, and compared with other advanced methods, the optimal recognition result is obtained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image processing, and in particular to a method for recognizing continuous teaching behaviors using a small number of samples and guided by a scalable thinking chain in the field of teaching behaviors. Background Art

[0002] Teaching behavior recognition has widespread application in smart classrooms. It can analyze student and teacher status in real time, while also intelligently monitoring and evaluating teachers' teaching behaviors, helping them adjust their teaching strategies in a timely manner. However, in real-world teaching scenarios, new teaching behaviors constantly emerge, and due to the cost of annotation, the number of labeled samples for these behaviors is quite limited. Under these circumstances, ensuring the model's ability to continuously learn from a small number of samples, accurately identifying new teaching behaviors while avoiding catastrophic forgetting, becomes a key challenge in this task.

[0003] Most existing few-shot continuous learning methods are based on pre-trained visual language models (such as CLIP), which adapt visual and textual representations by fine-tuning the network. However, these studies often ignore the fact that behavior labels themselves contain rich semantic information. In addition, existing pre-trained visual language models have a low understanding of verbs. This is because verbs are more abstract than nouns, and their semantics often depend on context and dynamic interactions. Existing models are insufficient in capturing such dynamic associations and temporal relationships. To this end, we propose a scalable thought chain-guided few-shot continuous teaching behavior recognition method. By guiding large language models (LLMs) through thought chains, we can mine the verb semantic knowledge in behavior labels, help the model better understand and recognize behavioral actions, and perform satisfactorily on the teaching behavior recognition task. Summary of the Invention

[0004] Existing methods for identifying actions in continuous instruction using few-shot learning based on pre-trained visual language models suffer from a problem: current visual language models are poorly able to match action labels with their corresponding images, and are insufficiently able to capture verbs and their semantic details. This is because a single verb label struggles to represent complex action semantics, let alone the subject and object information associated with the action. To address this issue, we propose a scalable thought chain-guided approach for identifying actions in continuous instruction using few-shot learning. This approach uses thought chains to guide large language models (LLMs) to semantically expand and condense action labels, making them more specific and easier for the model to understand.

[0005] The technical solution of the present invention is as follows:

[0006] A method for identifying continuous teaching behaviors with a small number of samples guided by a scalable thought chain includes the following steps:

[0007] Step 1: The classroom scene dataset is an image dataset built based on real classroom monitoring scenes, including different classroom teaching behaviors for teachers and students;

[0008] Step 2: Perform task division for few-shot continuous learning. Divide the preprocessed dataset into multiple tasks, including a base task and multiple incremental tasks. The base task has more categories and more samples than the incremental tasks. In the incremental tasks, the training dataset size, number of categories, and number of samples for each task are the same. Furthermore, the category labels of the samples in the base task and each incremental task do not overlap.

[0009] Step 3: Use the three-part question-answering method of the thought chain to guide the large language model GPT-4 to mine the semantic information in the action labels and obtain three levels of semantically expanded sentences;

[0010] Step 4: After obtaining the semantically expanded sentences of different levels of the action labels, the large language model GPT-4 is used to extract triple relationships to obtain triple knowledge of the (subject, predicate, object) structure. This knowledge is condensed and redundant information is removed to make it more structured and more intuitively display the noun entities related to the action labels and the relationships between them.

[0011] Step 5: Use the existing open-source pre-trained visual language model CLIP. The visual language model CLIP includes a visual encoder, a text encoder, and a cue learning module. The visual language model CLIP can extract semantically consistent image features and text features. Freeze the backbone networks of the visual encoder and text encoder of the pre-trained visual language model CLIP, and optimize the models using only cue learning.

[0012] Step 6: Input the image into the visual encoder to obtain different layers of visual features;

[0013] Step 7: Input the triplets of different levels obtained in step 4 into the text encoder to obtain text features of different levels. At the same time, embed the behavior label into a text prompt template and input it into the text encoder to obtain an overall text feature.

[0014] Step 8: Design a multi-level cross-modal matching strategy to calculate the similarity between the different layers of visual features obtained in step 6 and the different layers of text features. Then, adopt a hierarchical weighting strategy to set different weights based on the contribution of each layer of features in the action recognition task, and finally obtain the total similarity used for classification. The total similarity is used to calculate the probability prediction distribution used for classification, and the cross-entropy loss function is obtained.

[0015] Step 9: Use the cross-entropy loss function as the overall loss function of the visual language model CLIP, perform backpropagation, and optimize network training. After completing a task training, use the obtained visual language model CLIP to perform recognition on the test set to obtain its recognition accuracy. The test set should contain all categories that have been seen. Repeat this process until the last task is completed.

[0016] Furthermore, the specific method of step 3 is:

[0017] The first is the basic semantic level, which requires the large language model GPT-4 to convert the action label into a basic sentence by adding necessary nouns, determining the subject and location of the action;

[0018] This is followed by semantic extension of the entity interaction level, which introduces other visual entities related to actions and emphasizes the interactivity of the behavior;

[0019] Finally, the semantic expansion at the scene perception level allows the large language model GPT-4 to generate more detailed scene descriptions, capture the dynamic changes and reactions of the subjects in the scene, and display the details of the behavior in a more fine-grained manner.

[0020] Furthermore, the CLIP structure of the visual language model in step 5 includes:

[0021] Specifically, the visual language model includes the following structural components and connections: visual encoder, text encoder, and prompt learning module;

[0022] The visual encoder adopts a visual Transformer (ViT) structure, which is used to encode the input image into a high-dimensional feature representation and output a visual feature vector;

[0023] The text encoder uses a Transformer structure to encode the input text into the corresponding semantic feature representation and outputs a text feature vector;

[0024] The prompt learning module adds a learnable prompt vector before the image input and text input, and inputs it into the visual and text encoders together with the image and text to obtain the final visual feature vector and text feature vector. The prompt learning module parameters are the only trainable part, and the rest of the backbone network remains frozen;

[0025] The connection relationship of the entire model structure is as follows: the input image and the learnable prompt vector are used together through the visual encoder to extract visual features, and the input text and the learnable prompt vector are used together through the text encoder to extract text features; the classification result is obtained by calculating and matching the similarity of visual features and text features.

[0026] Furthermore, the total similarity sim in step 8 is calculated as follows:

[0027]

[0028] Among them, cosine_sim represents the calculation of cosine similarity, and <·> represents the vector f v and f t The inner product of ||·|| represents the modulus of the vector, α is a hyperparameter used to balance the influence of different feature similarities, f v Represents the visual features of the last layer, f t It represents the overall text feature obtained in step 7, and the subscript i represents the number of the corresponding feature;

[0029] Calculate the probability prediction distribution y' used for classification, the formula is as follows:

[0030] y'=softmax(sim)

[0031] Calculate the cross entropy loss function L, the formula is as follows:

[0032]

[0033] Where N is the total number of samples, C is the total number of categories, and y i,c is the one-hot encoded true label of the i-th image belonging to the c-th category, y' i,c is the predicted probability of the i-th image.

[0034] The present invention proposes a scalable thought chain-guided few-shot continuous teaching behavior recognition method. The thought chain guides large language models (LLMs) to mine different levels of semantic knowledge in behavior labels and condense them into triple knowledge of (subject, predicate, object) structure, which solves the problem that the existing pre-trained visual language model has weak understanding of verbs and realizes accurate understanding and recognition of behaviors. Compared with the common few-shot continuous learning method, our method freezes the backbone network of the pre-trained visual language model and only trains the model through prompt learning. Compared with the traditional backbone network representation adaptation and tuning, our method has very few training parameters, which greatly reduces the computational complexity. The present invention performs a few-shot continuous teaching behavior recognition task on a classroom scene dataset. Compared with other advanced methods, the present invention achieves the best recognition results. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] Figure 1 This is an overall flow chart of the method for identifying continuous teaching behaviors with a small number of samples guided by a scalable thinking chain in an embodiment of the present invention.

[0036] Figure 2 This is a schematic diagram of the overall structure of a method for identifying continuous teaching behaviors using a small number of samples guided by a scalable thinking chain in an embodiment of the present invention.

[0037] Figure 3 This is a curve chart comparing the recognition effects in a small number of samples of continuous teaching behavior recognition tasks in an embodiment of the present invention. DETAILED DESCRIPTION

[0038] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.

[0039] In order to address the problems of weak matching ability between behavior labels and their corresponding images and insufficient ability to capture verbs and their semantic details when performing small-sample continuous teaching behavior recognition using existing pre-trained visual language models, the present invention proposes a small-sample continuous teaching behavior recognition method guided by a scalable thinking chain.

[0040] like Figure 1 As shown, the embodiment of the present invention includes the following steps:

[0041] Step 1: Preprocess the image data in the classroom scene dataset. The classroom scene dataset is an image dataset built based on real classroom monitoring scenes, including different classroom teaching behaviors for teachers and students. The preprocessing used includes: transforming the image size to 224×224, randomly flipping the image, and calculating the pixel values ​​of the three RGB channels according to the formula Perform normalization processing, where x is the original input data, μ is the mean of all training samples, σ is the standard deviation of all training samples, and x' is the data obtained after normalization;

[0042] Step 2: Based on the setting of few-shot continuous learning, divide the dataset obtained in step 1 into training sets and test sets for different tasks. The division method is as follows:

[0043] The preprocessed dataset is divided into training and test sets for T = 5 tasks, including one base task and four incremental tasks. The base task has 20 categories and a large number of samples in the training set. In the incremental tasks, the training datasets for each task are identical, with 3 categories and 5 samples. Furthermore, the category labels of the samples in the base task and each incremental task do not overlap. During the training phase of task t, the model only has access to the training set for task t and sees the categories in task t. During the testing phase of task t, the model must be tested on all previously seen categories. This means that the test set must contain test samples from all categories in the current task and all previous tasks.

[0044] Step 3: If Figure 2As shown in the first part on the left side of the lower half, the three-part question-answering of the thinking chain guides the large language model GPT-4 to mine the semantic information in the behavior labels, obtain rich contextual association information at different levels, and conduct a semantic association process.

[0045] First, at the basic semantic level, the large language model GPT-4 is asked to convert the action label into a basic sentence by adding the necessary nouns, identifying the subject and location of the action. The prompt is as follows: Expand the verb '[class]' into a full sentence by adding the subject and object, specifying the person performing the action and their immediate surroundings in a classroom.

[0046] This is followed by semantic expansion at the entity interaction level, which further emphasizes the interactivity of the action by introducing other visual entities related to the action. The prompt is as follows: Enhance the sentence by including objects or people directly interacting with the action.

[0047] Finally, semantic extensions at the scene-aware level allow large language models (LLMs) to generate more detailed scene descriptions, capturing the dynamic changes and reactions of the subject in the scene and displaying more fine-grained details of the behavior. The prompt is as follows: Describe how the person is positioned or interacting with their surroundings while performing the action, including small gestures or reactions that reflect engagement with the classroom setting.

[0048] Step 4: If Figure 2 As shown in the second part on the left of the lower half, after obtaining sentences of different levels after semantic expansion of the behavior label, the large language model GPT-4 is used to extract triple relationships, obtaining triple knowledge of the (subject, predicate, object) structure, such as (students, hold on, bottle). Through the knowledge condensation process, redundant information is removed, making it more structured and more intuitively displaying the noun entities related to the behavior label and the relationships between them.

[0049] Step 5: Use the current mainstream open-source pre-trained visual language model CLIP, which utilizes large-scale image-text data for comparative learning. By maximizing the similarity between images and matching text and minimizing the similarity with mismatched text, the visual encoder and text encoder are trained to extract semantically consistent image features and text features, respectively. Freeze the backbone networks of the visual encoder and text encoder of the pre-trained visual language model and optimize the models using only cue learning. Specifically, the visual language model includes the following structural components and connection methods:

[0050] ① Visual encoder: uses the visual Transformer (ViT) structure to encode the input image into a high-dimensional feature representation and output a visual feature vector;

[0051] ② Text encoder: uses the Transformer structure to encode the input text into the corresponding semantic feature representation, and outputs the text feature vector;

[0052] ③ Hint learning module: A learnable hint vector is added before the image input and text input, and is input into the visual and text encoders together with the image and text to obtain the final visual feature vector and text feature vector. The hint learning module parameters are the only trainable part, and the rest of the backbone network remains frozen;

[0053] The connection relationship of the entire model structure is as follows: the input image and the learnable prompt vector are used together through the visual encoder to extract visual features, and the input text and the learnable prompt vector are used together through the text encoder to extract text features; the classification result is obtained by calculating and matching the similarity of visual features and text features.

[0054] Step 6: Figure 2 As shown in the upper part, the image is input into the visual encoder to obtain different layers of visual features f v ,in Represent the visual features of the 6th, 9th, and 11th layers of the visual encoder, respectively, v Represents the visual features of the last layer.

[0055] Step 7: Input the obtained triples at different levels into the text encoder to obtain text features at different levels At the same time, the behavior label is embedded in a text prompt template and input into the text encoder to obtain an overall text feature f t .

[0056] Step 8: Get the prediction results of the model;

[0057] like Figure 2As shown in the right half, we designed a multi-level cross-modal matching strategy to calculate the similarity between visual features at different levels and text features at different levels. We then adopted a hierarchical weighting strategy, setting different weights based on the contribution of each layer of features in the action recognition task, to obtain the final total similarity used for classification. The calculation is as follows:

[0058]

[0059] Where cosine_sim represents the calculation of cosine similarity, <·> represents the vector f v and f t The inner product of ||·|| represents the modulus of the vector, and α is a hyperparameter used to balance the influence of different feature similarities, α = 0.1.

[0060] Calculate the probability prediction distribution for classification, the formula is as follows:

[0061] y'=softmax(sim)

[0062] Calculate the cross entropy loss function, the formula is as follows:

[0063]

[0064] Where N is the total number of samples, C is the total number of categories, and y i,c is the one-hot encoded true label of the i-th image belonging to the c-th category, y' i,c is the predicted probability of the i-th image.

[0065] Step 9: Use the cross entropy loss function as the overall loss function of the model to optimize network training. After the training of a task is completed, use the obtained model to classify the test set and evaluate its classification accuracy. Repeat this process until the last task is completed, and obtain the average recognition accuracy of all tasks ACC avg And the average recognition accuracy ACC of the last task last .

[0066] The present invention conducts a few-sample continuous teaching behavior recognition experiment on the classroom scene dataset. The experimental results are shown in Table 1 below. The experimental results on each task are as follows: Figure 3 The experimental results show that the proposed method of continuous teaching behavior recognition guided by scalable thinking chain has an average recognition accuracy of ACC in all tasks compared with the existing advanced methods. avg And the average recognition accuracy ACC of the last task lastWe achieved the highest results in all tasks and the best recognition performance in each task stage, which fully demonstrates that our method is very helpful in improving the model's ability to understand and recognize behaviors, and can effectively solve the problem of catastrophic forgetting in few-shot continuous learning tasks.

[0067] Table 1

[0068] Model 0 1 2 3 4 <![CDATA[ACC avg ]]> <![CDATA[ACC last ]]> CLIP-Zero Sampling 28.45 26.82 26.13 21.63 20.34 24.67 20.34 CLIP-fine-tuning 77.79 73.16 70.73 68.79 62.42 70.58 62.42 L2P 67.68 64.25 62.74 62.29 62.15 63.82 62.15 Dualprompt 71.31 67.45 65.96 65.49 65.28 67.09 65.28 Privilege 68.79 65.29 62.33 55.87 55.20 61.50 55.20 ZSCL 74.00 69.21 65.83 64.70 61.60 67.07 61.60 CPE-CLIP 75.61 71.27 69.38 69.14 69.32 70.94 69.32 Method of the present invention 76.79 74.09 71.84 71.23 70.68 72.92 70.68

Claims

1. A method for identifying continuous teaching behaviors with a small number of samples guided by a scalable thought chain, comprising the following steps: Step 1: The classroom scene dataset is an image dataset built based on real classroom monitoring scenes, including different classroom teaching behaviors for teachers and students; Step 2: Perform task division for few-shot continuous learning, dividing the preprocessed dataset into multiple tasks, including a base class task and multiple incremental tasks; The number of categories and samples of the base task are more than those of the incremental task; In the incremental tasks, the training datasets for each task are of the same size, with the same number of categories and samples. At the same time, the category labels of the samples in the base task and each incremental task do not overlap. Step 3: Use the three-part question-answering method of the thought chain to guide the large language model GPT-4 to mine the semantic information in the action labels and obtain three levels of semantically expanded sentences; Step 4: After obtaining the semantically expanded sentences of different levels of the action labels, the large language model GPT-4 is used to extract triple relationships to obtain triple knowledge of the (subject, predicate, object) structure. This knowledge is condensed and redundant information is removed to make it more structured and more intuitively display the noun entities related to the action labels and the relationships between them. Step 5: Use the existing open-source pre-trained visual language model CLIP. The visual language model CLIP includes a visual encoder, a text encoder, and a prompt learning module. The visual language model CLIP can extract semantically consistent image features and text features. Freeze the backbone networks of the visual encoder and text encoder of the pre-trained visual language model CLIP, and use only prompt learning to optimize the model; Step 6: Input the image into the visual encoder to obtain different layers of visual features; Step 7: Input the triplets of different levels obtained in step 4 into the text encoder to obtain text features of different levels. At the same time, embed the behavior label into a text prompt template and input it into the text encoder to obtain an overall text feature. Step 8: Design a multi-level cross-modal matching strategy to calculate the similarity between the different layers of visual features obtained in step 6 and the different layers of text features. Then, adopt a hierarchical weighting strategy to set different weights based on the contribution of each layer of features in the action recognition task, and finally obtain the total similarity used for classification. The total similarity is used to calculate the probability prediction distribution used for classification, and the cross-entropy loss function is obtained. Step 9: Use the cross-entropy loss function as the overall loss function of the visual language model CLIP, perform backpropagation, and optimize network training. After completing a task training, use the obtained visual language model CLIP to perform recognition on the test set to obtain its recognition accuracy. The test set should contain all categories that have been seen. Repeat this process until the last task is completed.

2. The method for identifying continuous teaching behaviors using a small number of samples guided by a scalable thought chain as claimed in claim 1, characterized in that: The specific method of step 3 is: The first is the basic semantic level, which requires the large language model GPT-4 to convert the action label into a basic sentence by adding necessary nouns, determining the subject and location of the action; This is followed by semantic extension of the entity interaction level, which introduces other visual entities related to actions and emphasizes the interactivity of the behavior; Finally, the semantic expansion at the scene perception level allows the large language model GPT-4 to generate more detailed scene descriptions, capture the dynamic changes and reactions of the subjects in the scene, and display the details of the behavior in a more fine-grained manner.

3. The method for identifying continuous teaching behaviors using a small number of samples guided by a scalable thought chain as claimed in claim 1, characterized in that: The CLIP structure of the visual language model in step 5 shown includes: Specifically, the visual language model includes the following structural components and connections: visual encoder, text encoder, and prompt learning module; The visual encoder adopts a visual Transformer (ViT) structure, which is used to encode the input image into a high-dimensional feature representation and output a visual feature vector; The text encoder uses a Transformer structure to encode the input text into the corresponding semantic feature representation and outputs a text feature vector; The prompt learning module adds a learnable prompt vector before the image input and text input, and inputs it into the visual and text encoders together with the image and text to obtain the final visual feature vector and text feature vector. The prompt learning module parameters are the only trainable part, and the rest of the backbone network remains frozen; The connection relationship of the entire model structure is as follows: the input image and the learnable prompt vector are used together through the visual encoder to extract visual features, and the input text and the learnable prompt vector are used together through the text encoder to extract text features; the classification result is obtained by calculating and matching the similarity of visual features and text features.

4. The method for identifying continuous teaching behaviors using a small number of samples guided by a scalable thought chain as claimed in claim 1, characterized in that: The total similarity sim in step 8 is calculated as follows: Among them, cosine_sim represents the calculation of cosine similarity, and <·> represents the vector f v and f t The inner product of ||·|| represents the modulus of the vector, α is a hyperparameter used to balance the influence of different feature similarities, f v Represents the visual features of the last layer, f t It represents the overall text feature obtained in step 7, and the subscript i represents the number of the corresponding feature; Calculate the probability prediction distribution y' used for classification, the formula is as follows: y'=softmax(sim) Calculate the cross entropy loss function L, the formula is as follows: Where N is the total number of samples, C is the total number of categories, and y i,c is the one-hot encoded true label of the i-th image belonging to the c-th category, y' i,c is the predicted probability of the i-th image.