Method and device for identifying classroom behavior, computer device and storage medium

By inputting classroom images into multiple pre-trained artificial intelligence models and combining imaging distortion features for optimization processing, the problem of insufficient accuracy in AI prediction of classroom behavior is solved, and higher prediction accuracy and fairness are achieved.

CN119541041BActive Publication Date: 2025-06-20JIUJIANG DIGITAL IND DEV CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411452653.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-17
Publication Date
2025-06-20
Estimated Expiration
2044-10-17

AI Technical Summary

Technical Problem

The accuracy of AI prediction of classroom behavior in the prior art needs to be improved, and unfair behavior predictions may occur due to deviations in input data or algorithm design.

Method used

By obtaining classroom audio and video data collected by multiple audio and video acquisition devices during the teaching process, the classroom images are input to multiple pre-trained artificial intelligence models for behavior prediction, and combining the audio and video acquisition devices to optimize the prediction results for each character to improve the prediction accuracy.

Benefits of technology

Combining the classroom behavior prediction results and imaging distortion characteristics of a variety of pre-trained artificial intelligence models, the prediction results are optimized, which improves the accuracy of classroom behavior prediction and avoids prediction deviations caused by imaging distortion.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119541041B_ABST
    Figure CN119541041B_ABST
Patent Text Reader

Abstract

Embodiments of the present invention disclose a method and device for identifying classroom behaviors, a computer device, and a storage medium. The method includes: obtaining classroom audio-visual data collected by multiple audio-visual acquisition devices during the teaching process, where the classroom audio-visual data at least includes each classroom image collected by each audio-visual acquisition device; respectively inputting the classroom images into multiple pre-trained artificial intelligence models for behavior prediction to obtain multiple first prediction results for each person in the classroom images, and the prediction results are used to reflect the types of classroom behaviors of the persons; determining the imaging distortion characteristics of each audio-visual acquisition device for each person; and using the imaging distortion characteristics and the multiple first prediction results to perform optimization processing on the prediction results to obtain the target prediction results for each person. Through the above method, not only the respective model advantages of multiple pre-trained artificial intelligence models are integrated, but also the influence of imaging distortion on the prediction results is considered, so that the accuracy of the obtained target prediction results is higher.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular, to a method and device for identifying classroom behaviors, a computer device, and a storage medium. Background Art

[0002] Artificial Intelligence (AI) is abbreviated as AI in English. It is a new technical science that studies, develops theories, methods, technologies, and application systems for simulating, extending, and expanding human intelligence. Artificial intelligence is a very broad science, including robots, speech recognition, image recognition, natural language processing, expert systems, machine learning, computer vision, etc.

[0003] The application of AI in the field of education is becoming increasingly common, especially in classroom behavior management. The combination of AI and the classroom can bring many benefits, such as real-time feedback: the AI system can monitor the behaviors and engagement of students in real time and provide instant feedback to teachers, which enables teachers to adjust teaching strategies in a timely manner to better attract students' attention and improve classroom management efficiency.

[0004] Although the application of AI in predicting classroom behaviors has great potential, there are also some drawbacks and challenges. Classroom teaching behaviors are very complex, including various factors such as speech, body movements, and facial expressions. The AI system needs to be able to accurately identify and classify these behaviors, which is a challenge technically; moreover, the AI system may generate unfair behavior predictions due to biases in input data or algorithm design, which may lead to unfair treatment of certain students. Therefore, the accuracy of current AI in predicting classroom behaviors still needs to be improved. Summary of the Invention

[0005] The main purpose of the present invention is to provide a method and device for identifying classroom behaviors, a computer device, and a storage medium, which can solve the problem that the accuracy of current AI in predicting classroom behaviors needs to be improved.

[0006] To achieve the above object, in a first aspect of the present invention, a method for identifying classroom behaviors is provided, and the method includes:

[0007] Obtain classroom audio-visual data collected by multiple audio-visual collection devices during the teaching process, where the classroom audio-visual data includes at least each classroom image collected by each audio-visual collection device;

[0008] Input the classroom images into a variety of pre-trained artificial intelligence models for behavior prediction respectively to obtain multiple first prediction results for each person in the classroom images, and the prediction results are used to reflect the classroom behavior types of the persons.

[0009] Determine the imaging distortion characteristics of the audio-visual acquisition device for each person;

[0010] Use the imaging distortion characteristics and the multiple first prediction results to perform prediction result optimization processing to obtain the target prediction result for each person.

[0011] In a feasible implementation manner, the audio-visual acquisition device at least includes a camera. Then, determining the imaging distortion characteristics of the audio-visual acquisition device for each person includes:

[0012] Use the first position coordinates of each person in the world coordinate system and the second position coordinates of the camera in the world coordinate system to determine the respective target azimuth angles between each person and each of the cameras. The imaging distortion characteristics include the target azimuth angles.

[0013] In a feasible implementation manner, using the imaging distortion characteristics and the multiple first prediction results to perform prediction result optimization processing to obtain the target prediction result for each person includes:

[0014] Determine the target distortion type of each camera;

[0015] Use a preset database of distortion type and deformation coefficient list and the target distortion type to determine the target deformation coefficient list corresponding to the target distortion type. The deformation coefficient list includes the correspondence between the deformation coefficient and the azimuth angle under the distortion type;

[0016] Search in the target deformation coefficient list based on the target azimuth angle to obtain the target deformation coefficient corresponding to the azimuth angle that is the same as the target azimuth angle;

[0017] Use the target deformation coefficient as a weight to perform prediction result optimization processing on the multiple first prediction results of each person to obtain the target prediction result for each person.

[0018] In a feasible implementation manner, the first prediction result includes the first prediction label of the classroom behavior type and the credibility of the first prediction label. Then, using the target deformation coefficient as a weight to perform prediction result optimization processing on the multiple first prediction results of each person to obtain the target prediction result for each person includes:

[0019] Use the target azimuth angle as the clustering condition to respectively cluster the multiple first prediction results of each person to obtain the target clustering set corresponding to each target azimuth angle of each person. The target clustering set includes the respective second prediction results corresponding to the target azimuth angle;

[0020] Using the credibility as a weight, perform weighted decision-making processing on the second prediction labels included in the second prediction results of each of the target clustering sets to obtain third prediction labels after weighted decision-making;

[0021] Using the target deformation coefficient as a weight, perform weighted decision-making processing on the third prediction labels of each person to obtain the final prediction label of each person, and the target prediction result includes the final prediction label.

[0022] In a feasible implementation manner, the artificial intelligence model at least includes a Transformer model based on a self-attention mechanism, a three-dimensional convolutional neural network, a recurrent neural network, and a convolutional neural network. Then, inputting the classroom image into multiple pre-trained artificial intelligence models for behavior prediction to obtain multiple first prediction results of each person in the classroom image includes:

[0023] Input the classroom image into the trained Transformer model based on the self-attention mechanism, three-dimensional convolutional neural network, recurrent neural network, and convolutional neural network for behavior prediction to obtain multiple first prediction results of each person in the classroom image.

[0024] In a feasible implementation manner, the method further includes:

[0025] Obtain training sample data, where the training sample data at least includes a number of sample classroom images and the true labels of the classroom behavior types corresponding to the sample classroom images;

[0026] Input the sample classroom image into the initial convolutional neural network for training to obtain the fourth prediction label of each person;

[0027] Using the fourth prediction label, the true label, and a preset loss function, determine the loss value of the initial convolutional neural network;

[0028] If the loss value is lower than a preset loss threshold, determine that the initial convolutional neural network converges to obtain a trained convolutional neural network.

[0029] In a feasible implementation manner, the loss function is the following mathematical expression:

[0030]

[0031] In the formula, N is the total number of samples, M is the total number of classroom behavior types, and IOU i represents the ratio of the intersection and union of the predicted box and the true box of sample i, and D i is the distance between the center points of the predicted box and the true box of sample i, and L iis the diagonal length of the minimum bounding rectangle of the predicted box and the ground truth box for sample i, v i is the similarity of the width-to-height ratio between the predicted box and the ground truth box for sample i, and α is the influencing factor of v i y ic is an indicator variable used to indicate the truthfulness of the classroom behavior type c of the fourth predicted label of sample i. If the classroom behavior type c is the same as the classroom behavior type of the ground truth label of sample i, then y ic takes the value of 1, otherwise 0, p ic represents the predicted probability that the observed sample i belongs to the classroom behavior type c.

[0032] To achieve the above object, the second aspect of the present invention provides a recognition device for classroom behavior, and the device includes:

[0033] Data acquisition module: used to obtain classroom audio-visual data collected by multiple audio-visual acquisition devices during the teaching process, and the classroom audio-visual data at least includes each classroom image collected by each audio-visual acquisition device;

[0034] Behavior prediction module: used to input the classroom images into a variety of pre-trained artificial intelligence models for behavior prediction respectively, and obtain multiple first prediction results of each person in the classroom images, and the prediction results are used to reflect the classroom behavior type of the person;

[0035] Distortion determination module: used to determine the imaging distortion characteristics of the audio-visual acquisition device for each person;

[0036] Result optimization module: used to perform prediction result optimization processing by using the imaging distortion characteristics and the multiple first prediction results to obtain the target prediction result of each person.

[0037] To achieve the above object, the third aspect of the present invention provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the processor is caused to execute the steps as shown in the first aspect and any feasible implementation manner.

[0038] To achieve the above object, the fourth aspect of the present invention provides a computer device including a memory and a processor, the memory stores a computer program, and when the computer program is executed by the processor, the processor is caused to execute the steps as shown in the first aspect and any feasible implementation manner.

[0039] Adopting the embodiments of the present invention has the following beneficial effects:

[0040] The present invention provides a method for recognizing classroom behaviors, the method comprising: obtaining classroom audio-visual data collected by a plurality of audio-visual acquisition devices during the teaching process, the classroom audio-visual data at least including each classroom image collected by each audio-visual acquisition device; respectively inputting the classroom images into a plurality of pre-trained artificial intelligence models for behavior prediction to obtain a plurality of first prediction results for each person in the classroom images, the prediction results being used to reflect the types of classroom behaviors of the persons; determining the imaging distortion characteristics of each audio-visual acquisition device for each person; and performing prediction result optimization processing by using the imaging distortion characteristics and the plurality of first prediction results to obtain the target prediction result for each person. By the above method, the prediction results of classroom behaviors of a plurality of pre-trained artificial intelligence models and the imaging distortion characteristics can be comprehensively used to optimize the prediction results, so as to obtain the optimized target prediction result for each person. This not only combines the respective model advantages of the plurality of pre-trained artificial intelligence models, but also takes into account the influence of imaging distortion on the prediction results, and thus the accuracy of the obtained target prediction result is higher. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention, and those of ordinary skill in the art can obtain other drawings based on these drawings without creative efforts.

[0042] Wherein:

[0043] Figure 1 is a flowchart of a method for recognizing classroom behaviors in an embodiment of the present invention;

[0044] Figure 2 is another flowchart of a method for recognizing classroom behaviors in an embodiment of the present invention;

[0045] Figure 3 is a structural block diagram of a device for recognizing classroom behaviors in an embodiment of the present invention;

[0046] Figure 4 is a structural block diagram of a computer device in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0047] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art without creative efforts based on the embodiments of the present invention belong to the scope of protection of the present invention.

[0048] Please refer to Figure 1, Figure 1 is a flowchart of a method for identifying classroom behaviors in an embodiment of the present invention. This method can be applied to a terminal or a server. The terminal can specifically be a desktop terminal or a mobile terminal, and the mobile terminal can specifically be at least one of a mobile phone, a tablet computer, a laptop computer, etc. The server can be implemented by an independent server or a server cluster composed of multiple servers. In this embodiment, an example of applying it to a terminal is given, such as Figure 1 The method shown includes the following steps:

[0049] 101. Obtain classroom audio-visual data collected by multiple audio-visual collection devices during the teaching process. The classroom audio-visual data at least includes each classroom image collected by each audio-visual collection device;

[0050] It should be noted that in order to realize the intelligent monitoring of students' classroom behaviors during the teaching process, it is necessary to pre-arrange audio-visual collection devices in the classroom. The audio-visual collection devices can be integrated image collection devices and audio collection devices, or can be split image collection devices and audio collection devices. In this application, an example of integrated image collection devices and audio collection devices is given. Among them, the integrated image collection devices and audio collection devices can be monitors. The number of monitors in this application can be multiple, at least two. Each monitor is scattered and arranged at different positions in the classroom to achieve all-round and multi-angle monitoring of students. Among them, the audio-visual collection devices are used to collect classroom audio-visual data during the teaching process, including the images and audio of the classroom during the teaching process. The classroom images can be obtained by dividing the images frame by frame, and the classroom behaviors can be identified based on the classroom images. It can be understood that each camera can capture classroom images. Therefore, in this application, multi-angle classroom images can be collected at the same time, and correspondingly, the multi-faceted behavior performances of each student can be obtained at the same time, so as to perform subsequent classroom behavior recognition with higher accuracy.

[0051] 102. Input the classroom images into multiple pre-trained artificial intelligence models respectively for behavior prediction, and obtain multiple first prediction results of each person in the classroom images. The prediction results are used to reflect the types of classroom behaviors of the person;

[0052] Further, after obtaining the above classroom image, input the classroom image into a pre-trained artificial intelligence model for behavior prediction to obtain multiple first prediction results for each person in the classroom image. The prediction results are used to reflect the classroom behavior types of the person. Classroom behavior refers to the performance of students in the classroom, including aspects such as their learning attitude, participation, compliance with classroom rules, and respect for others. Classroom behavior types include, but are not limited to, behaviors such as listening, being distracted, standing, and turning the head. The prediction results include the prediction labels of the classroom behavior types and the credibility of the prediction labels. The prediction labels can be the above-mentioned behavior types such as listening, being distracted, standing, and turning the head, which are not limited herein.

[0053] It should be noted that the terms "first", "second", etc. in the specification of this application are used to distinguish different objects, rather than to describe a specific order.

[0054] Since different artificial intelligence models have different advantages and disadvantages, this application uses multiple artificial intelligence models for prediction, which can utilize the advantages of each artificial intelligence model. In a feasible implementation manner, the artificial intelligence model at least includes a Transformer model based on the self-attention mechanism, a three-dimensional convolutional neural network, a recurrent neural network, a convolutional neural network, and so on.

[0055] Among them, the Transformer model based on the self-attention mechanism, the three-dimensional convolutional neural network (3D CNN), the recurrent neural network (RNN), and the convolutional neural network (CNN) are all important models in the field of deep learning. 1) The Transformer model based on the self-attention mechanism is a deep learning model based on the self-attention mechanism. The core of the Transformer model is the multi-head self-attention mechanism, which allows the model to consider the information of the entire sequence when processing each element, thereby capturing long-range dependencies within the sequence. This mechanism enables the model to process sequence data in parallel, improving computational efficiency. 2) The three-dimensional convolutional neural network (3D CNN) is an extension of the convolutional neural network. It adds a depth dimension on the basis of two-dimensional convolution, enabling it to process three-dimensional data such as videos or three-dimensional medical images. 3D CNN extracts features by sliding the convolutional kernel in three dimensions, which enables it to capture patterns in space and time and is very suitable for video classification and three-dimensional image segmentation tasks. 3) The recurrent neural network (RNN) is a neural network suitable for processing sequence data. It can remember previous information and use it for the calculation of the current output. The hidden layer nodes of the RNN not only have connections from the input layer but also connections from the previous time step, which enables it to maintain context information in the sequence. 4) The convolutional neural network (CNN) is a deep learning model commonly used in image processing. It extracts local features of images through convolutional layers and reduces the spatial dimensions of features through pooling layers. CNN can identify features such as edges and textures in images through learned convolutional kernels and combine features at multiple levels to finally achieve the recognition and classification of image content.

[0056] Furthermore, the classroom images are respectively input into multiple pre-trained artificial intelligence models for behavior prediction, and multiple first prediction results for each person in the classroom image are obtained. Specifically, the classroom images are respectively input into the trained Transformer model based on the self-attention mechanism, the three-dimensional convolutional neural network, the recurrent neural network, and the convolutional neural network for behavior prediction, and multiple first prediction results for each person in the classroom image are obtained.

[0057] It can be understood that the cameras are distributed everywhere, so there are multiple shooting angles for the cameras. Therefore, each person has multiple classroom images taken from multiple angles, and each classroom image corresponds to multiple first prediction results of multiple models. That is, there are S cameras and M artificial intelligence models. For person 1, there are S * M first prediction results. For example, the cameras are set in the east, south, west, and north directions of the classroom, that is, S = 4, and there are the above four artificial intelligence models, that is, M = 4. In this way, there are 16 first prediction results for person 1.

[0058] Among them, the above models are all pre-trained using pre-collected training sample data, and the training sample data at least includes a number of sample classroom images and the true labels of the classroom behavior types corresponding to the sample classroom images.

[0059] Taking the training of a convolutional neural network as an example, in order to improve the prediction accuracy of the convolutional neural network, the convolutional neural network in this application at least includes an object detection network and a depthwise separable convolutional neural network. Among them, the output of the object detection network is the input of the depthwise separable convolutional neural network. The sample classroom image is input into the object detection network, and the object detection network is used to identify the human object in the sample classroom image. Therefore, the output of the object detection network is the human region in the sample classroom image. The depthwise separable convolutional neural network is used to predict the classroom behavior of the human object in the sample classroom image. The human region includes the face and the body. In this way, the depthwise separable convolutional neural network can perform behavior recognition through face features and body features. Face features include but are not limited to the physical structure of the face, facial expressions, face shapes, and the features of various parts of the face. Body features include but are not limited to human pose features, etc. The convolutional neural network trained in this way can input the classroom image into the trained object detection network for object detection, identify the human object in the classroom image. Further, input the human object in the classroom image into the trained depthwise separable convolutional neural network for behavior prediction to obtain the prediction result of the classroom behavior of the human in the classroom image.

[0060] Among them, the network compositions of the object detection network and the depthwise separable convolutional neural network can refer to the network compositions of the existing object detection network and the depthwise separable convolutional neural network, which will not be elaborated. Further, the training of the convolutional neural network including the object detection network and the depthwise separable convolutional neural network can refer to the following processes A01 to A04:

[0061] A01. Obtain training sample data, where the training sample data at least includes a number of sample classroom images and the true labels of the classroom behavior types corresponding to the sample classroom images;

[0062] It can be understood that if the convolutional neural network wants to learn the corresponding classroom behavior, corresponding sample data needs to be used as the sample data to train the convolutional neural network. Among them, the training sample data at least includes a number of sample classroom images and the true labels of the classroom behavior types corresponding to the sample classroom images.

[0063] A02. Input the sample classroom image into the initial convolutional neural network for training to obtain the fourth prediction label for each person;

[0064] A03. Use the fourth prediction label, the true label, and a preset loss function to determine the loss value of the initial convolutional neural network;

[0065] A04. If the loss value is lower than the preset loss threshold, it is determined that the initial convolutional neural network converges, and a trained convolutional neural network is obtained.

[0066] Furthermore, the sample classroom images are input into the initial convolutional neural network for training to obtain the fourth prediction labels for each person. The initial convolutional neural network is the original convolutional neural network that has not yet learned the classroom behavior characteristics. Then, by comparing the prediction results with the actual situations, the initial convolutional neural network gradually learns the classroom behavior characteristics and obtains a trained convolutional neural network that can recognize classroom behaviors. Among them, using the fourth prediction label, the true label, and the preset loss function, the loss value of the initial convolutional neural network is determined. The loss value is an index in machine learning used to measure the difference between the model's predicted value and the true value. The loss function is a mathematical function that takes the predicted value and the true value of the model as inputs and outputs a non-negative real number to represent the loss value. The smaller the loss value, the closer the predicted result of the model is to the true value, and the better the performance of the model.

[0067] In a feasible implementation, the loss function is the following mathematical expression:

[0068]

[0069] In the formula, is the loss calculation part of the object detection network, is the loss calculation part of the depthwise separable convolutional neural network, where N is the total number of samples, M is the total number of classroom behavior types, IOU i represents the ratio of the intersection and union of the predicted box and the true box of sample i, D i is the distance between the center points of the predicted box and the true box of sample i, L i is the diagonal length of the minimum bounding rectangle of the predicted box and the true box of sample i, v i is the aspect ratio similarity between the predicted box and the true box of sample i, α is the influencing factor of v i y ic is an indicator variable used to indicate the authenticity of the classroom behavior type c of the fourth prediction label of sample i. If the classroom behavior type c is the same as the classroom behavior type of the true label of sample i, then y ic takes the value of 1, otherwise 0, p icRepresents the predicted probability that the observed sample i belongs to the classroom behavior type c. The predicted probability is also called the confidence level. Among them, sample i is the sample classroom image, and the ground truth box is the smallest bounding rectangle that encloses the person. Among them, the target detection network outputs the predicted box of the person. The sample classroom image is pre-annotated with the ground truth box of the person and the true classroom behavior label for calculating the corresponding loss value during training.

[0070] It should be noted that the loss value of the convolutional neural network is obtained by combining the loss values of two networks. Among them, the output of the target detection network is used as the input of the depthwise separable convolutional neural network. The target detection network determines the accuracy of person recognition and indirectly affects the accuracy of classroom behavior prediction of the depthwise separable convolutional neural network. If the target detection network does not converge, then the depthwise separable convolutional neural network will surely be difficult to converge; and when the target detection network converges, the depthwise separable convolutional neural network may not converge simultaneously. Therefore, a random number is taken between the loss values of the two networks through the RAND() function. As long as any one network does not converge, the randomly obtained loss value is non-convergent. And if the target detection network does not converge, the depthwise separable convolutional neural network will surely be difficult to converge, and the randomly obtained loss value is still non-convergent; if the target detection network converges and the depthwise separable convolutional neural network does not converge, then the randomly obtained loss value is still non-convergent. In this way, only when both networks converge will a convergent random loss value be obtained between the two convergent loss values, that is, the obtained must be a convergent random number. Therefore, the final loss value of the convolutional neural network can be calculated through the above loss function to evaluate whether the initial convolutional neural network converges.

[0071] Furthermore, if the loss value loss is lower than the preset loss threshold, it is determined that the initial convolutional neural network converges, and the trained convolutional neural network is obtained; on the contrary, if the loss value loss is greater than the preset loss threshold, it is determined that it does not converge. After further adjusting and optimizing the model parameters of the initial convolutional neural network, return to step A02 to continue training the initial convolutional neural network until it converges.

[0072] 103. Determine the imaging distortion characteristics of the audio-visual acquisition device for each person;

[0073] 104. Use the imaging distortion characteristics and the multiple first prediction results to perform prediction result optimization processing to obtain the target prediction result for each person.

[0074] Further, the audio-visual acquisition device is fixedly installed. For each person, the same audio-visual acquisition device will produce different imaging distortions, which affect the accuracy of the expressions and postures of the people in the classroom images, and further affect the first prediction result. Therefore, in order to further improve the accuracy of the prediction result, the present application determines the imaging distortion characteristics of the audio-visual acquisition device for each person; uses the imaging distortion characteristics and the multiple first prediction results to perform prediction result optimization processing to obtain the target prediction result for each person. The imaging distortion characteristics are used to reflect the degree of imaging distortion. The optimization processing of the prediction result includes but is not limited to weighted decision-making, which will not be elaborated here.

[0075] Among them, the audio data can also be input into the AI model of natural language processing for classroom behavior prediction to obtain a new prediction result of classroom behavior, and further decision-making is based on the target prediction result and the new prediction result.

[0076] The present invention provides a method for identifying classroom behaviors. The method includes: obtaining classroom audio-visual data collected by multiple audio-visual acquisition devices during the teaching process. The classroom audio-visual data at least includes each classroom image collected by each audio-visual acquisition device; respectively inputting the classroom images into a variety of pre-trained artificial intelligence models for behavior prediction to obtain multiple first prediction results for each person in the classroom images. The prediction results are used to reflect the classroom behavior types of the people; determining the imaging distortion characteristics of the audio-visual acquisition device for each person; using the imaging distortion characteristics and the multiple first prediction results to perform prediction result optimization processing to obtain the target prediction result for each person. Through the above method, the classroom behavior prediction results of a variety of pre-trained artificial intelligence models and the imaging distortion characteristics can be comprehensively used to optimize the prediction results, and the optimized target prediction results for each person can be obtained. This not only combines the respective model advantages of a variety of pre-trained artificial intelligence models, but also takes into account the influence of imaging distortion on the prediction results, so that the accuracy of the obtained target prediction results is higher.

[0077] Please refer to Figure 2 , Figure 2 which is another flowchart of a method for identifying classroom behaviors in an embodiment of the present invention. As Figure 2 shown, the method includes the following steps:

[0078] 201. Obtain classroom audio-visual data collected by multiple audio-visual acquisition devices during the teaching process. The classroom audio-visual data at least includes each classroom image collected by each audio-visual acquisition device;

[0079] 202. Respectively input the classroom images into a variety of pre-trained artificial intelligence models for behavior prediction to obtain multiple first prediction results for each person in the classroom images. The prediction results are used to reflect the classroom behavior types of the people;

[0080] 203. Determine the imaging distortion characteristics of the audio-visual acquisition device for each person;

[0081] It should be noted that steps 201 to 203 are similar to Figure 1 the content of steps 101 to 103 shown. To avoid repetition, it will not be elaborated here. Specifically, reference can be made to Figure 1 the content of steps 101 to 103 shown.

[0082] It can be understood that the audio-visual acquisition device includes at least a camera. Then, step 203 can specifically be to determine each target azimuth angle between each person and each of the cameras by using the first position coordinates of each person in the world coordinate system and the second position coordinates of the camera in the world coordinate system. The imaging distortion characteristics include the target azimuth angle.

[0083] In this embodiment, electronic tags are pre-set on the desks of each person. The position coordinates of the desks in the world coordinate system are written in the electronic tags. By identifying the desk electronic tags of each person, the acquisition of the first position coordinates of the person is realized. The desk position coordinates of the person are used as the person position coordinates. Then, by using the second position coordinates of the camera in the world coordinate system, each target azimuth angle between each person and the camera is determined. Since the orientation of the camera is fixed, the azimuth of the person relative to the camera can be known through the two coordinates.

[0084] 204. Determine the target distortion type of each of the cameras;

[0085] Since different camera types will cause different types of distortion, the distortion types at least include barrel distortion and pincushion distortion. The target distortion type of the camera, whether it is barrel distortion or pincushion distortion, can be known through the camera type.

[0086] 205. Use the database of the preset distortion type and deformation coefficient list and the target distortion type to determine the target deformation coefficient list corresponding to the target distortion type. The deformation coefficient list includes the corresponding relationship between the deformation coefficient and the azimuth angle under the distortion type;

[0087] Furthermore, determine the impact on imaging according to the distortion type. Specifically, use the database of the preset distortion type and the deformation coefficient list and the target distortion type to determine the target deformation coefficient list corresponding to the target distortion type. The database stores multiple distortion types and the deformation coefficient list of each distortion type. For example, the deformation coefficient list 1 of barrel distortion and the deformation coefficient list 2 of pincushion distortion are pre-stored in the database. If the target distortion type is pincushion distortion, then the target deformation coefficient list is the deformation coefficient list 2 of pincushion distortion. Among them, the deformation coefficient list includes the corresponding relationship between the azimuth angle and the deformation coefficient under the distortion type, that is, the deformation coefficient list stores the corresponding relationship between the azimuth angle of the person relative to the camera and the deformation coefficient under this distortion type. This corresponding relationship can be obtained through a controlled variable experiment on the azimuth angle and deformation situation between the person and each camera. In this way, the deformation situation under each azimuth angle of each distortion type of each camera can be obtained, and a unique deformation coefficient is assigned to each situation according to the deformation situation to characterize the deformation situation.

[0088] Among them, the target deformation coefficient is used to reflect the influence degree of the target azimuth angle on the deformation of the person's imaging under the target distortion type. The value of the deformation coefficient is between 0 and 1. The higher the deformation coefficient, the greater the influence degree of the azimuth angle on the person's imaging under this distortion type, and the more deformed the image; the lower the deformation coefficient, the smaller the influence degree of the azimuth angle on the person's imaging under this distortion type, and the less deformed the image.

[0089] 206. Search in the target deformation coefficient list based on the target azimuth angle to obtain the target deformation coefficient corresponding to the azimuth angle that is the same as the target azimuth angle;

[0090] 207. Use the target deformation coefficient as a weight to perform prediction result optimization processing on multiple first prediction results of each person to obtain the target prediction result of each person.

[0091] After obtaining the target azimuth angle of the person, the corresponding target deformation coefficient can be searched in the target deformation coefficient list based on the target azimuth angle. Specifically, search in the target deformation coefficient list based on the target azimuth angle to obtain the target deformation coefficient corresponding to the azimuth angle that is the same as the target azimuth angle. Furthermore, use the target deformation coefficient as a weight to perform prediction result optimization processing on multiple first prediction results of each person to obtain the target prediction result of each person.

[0092] In a feasible implementation manner, the first prediction result includes the first prediction label of the classroom behavior type and the credibility of the first prediction label. Then, the step of using the target deformation coefficient as a weight to perform prediction result optimization processing on multiple first prediction results of each person to obtain the target prediction result of each person includes steps B01 to B03:

[0093] B01. Using the target azimuth angle as the clustering condition, cluster the multiple first prediction results of each person to obtain a target clustering set corresponding to each target azimuth angle of each person. The target clustering set includes the respective second prediction results corresponding to the target azimuth angle.

[0094] First, cluster the first prediction results. Using the target azimuth angle as the clustering condition, cluster the multiple first prediction results of each person to obtain a target clustering set corresponding to each target azimuth angle of each person. The target clustering set includes the respective second prediction results corresponding to the target azimuth angle. Exemplarily, cameras are set at 4 positions in the classroom, i.e., S = 4, and there are the above four artificial intelligence models, i.e., M = 4. Thus, there are 16 first prediction results for person 1. Among them, there are target azimuth angle 1, target azimuth angle 2, target azimuth angle 3, and target azimuth angle 4 between person 1 and the four cameras respectively. Among them, target azimuth angle 1 is the azimuth angle between person 1 and camera 1, target azimuth angle 2 is the azimuth angle between person 1 and camera 2, target azimuth angle 3 is the azimuth angle between person 1 and camera 3, and target azimuth angle 4 is the azimuth angle between person 1 and camera 4. Cluster the first prediction results of person 1 with the target azimuth angle as the clustering condition to obtain target clustering sets for the four target azimuth angles respectively, including target clustering set 1 for target azimuth angle 1, target clustering set 2 for target azimuth angle 2, target clustering set 3 for target azimuth angle 3, and target clustering set 4 for target azimuth angle 4. Among them, target clustering set 1 includes the respective second prediction results output by multiple artificial intelligence models based on the classroom images captured by camera 1 for classroom behavior prediction of person 1; target clustering set 2 includes the respective second prediction results output by multiple artificial intelligence models based on the classroom images captured by camera 2 for classroom behavior prediction of person 1; target clustering set 3 includes the respective second prediction results output by multiple artificial intelligence models based on the classroom images captured by camera 3 for classroom behavior prediction of person 1; target clustering set 4 includes the respective second prediction results output by multiple artificial intelligence models based on the classroom images captured by camera 4 for classroom behavior prediction of person 1. By analogy, if there are i persons, the number of target clustering sets is i * 4.

[0095] B02. Using the credibility as the weight, perform weighted decision-making processing on the second prediction labels included in the second prediction results of each target clustering set to obtain the third prediction label after weighted decision-making.

[0096] After obtaining the target clustering sets classified by the target azimuth angle for each person, a weighted decision-making process can be carried out based on the target clustering sets. Specifically, using the credibility as the weight, weighted decision-making is performed on the second prediction labels included in the second prediction results of each of the target clustering sets to obtain the third prediction labels after weighted decision-making. It can be understood that the prediction results of the artificial intelligence model include prediction labels and prediction probabilities, and the prediction probability is the credibility of the prediction label. The weighted decision-making method can be to select the second prediction label with the highest credibility as the third prediction label after weighted decision-making; if there are multiple second prediction labels with the highest credibility, it can be determined whether the second prediction labels refer to the same classroom behavior type. If they refer to the same classroom behavior type, randomly select one of the highest second prediction labels as the third prediction label after weighted decision-making. If they do not refer to the same classroom behavior type, select the category of second prediction labels with the largest number among several second prediction labels with the highest credibility as the third prediction label after weighted decision-making.

[0097] Continuing with the example where each person has four target clustering sets as above, at this time, there are four third prediction labels after weighted decision-making. Specifically, one second prediction label can be selected from the target clustering set 1 as the third prediction label, one second prediction label can be selected from the target clustering set 2 as the third prediction label, one second prediction label can be selected from the target clustering set 3 as the third prediction label, and one second prediction label can be selected from the target clustering set 4 as the third prediction label.

[0098] B03. Using the target deformation coefficient as the weight, perform weighted decision-making on the third prediction labels of each person to obtain the final prediction label of each person. The target prediction result includes the final prediction label.

[0099] Furthermore, through weighted decision-making again using the target deformation coefficient, select the final prediction label from multiple third prediction labels to obtain the target prediction result. Among them, the weighted decision-making method can be to select the third prediction label with the lowest target deformation coefficient as the final prediction label after weighted decision-making; if there are multiple third prediction labels with the lowest target deformation coefficient, it can be determined whether the third prediction labels refer to the same classroom behavior type. If they refer to the same classroom behavior type, randomly select one of the lowest third prediction labels as the final prediction label after weighted decision-making. If they do not refer to the same classroom behavior type, select the category of third prediction labels with the largest number among several third prediction labels with the lowest credibility as the final prediction label after weighted decision-making.

[0100] The present invention provides a method for identifying classroom behaviors. The method includes: obtaining classroom audio-visual data collected by multiple audio-visual acquisition devices during the teaching process, where the classroom audio-visual data at least includes various classroom images collected by each audio-visual acquisition device; inputting the classroom images into multiple pre-trained artificial intelligence models respectively for behavior prediction to obtain multiple first prediction results for each person in the classroom images, and the prediction results are used to reflect the classroom behavior types of the persons; determining the imaging distortion characteristics of each person by the audio-visual acquisition device, where the imaging distortion characteristics include the target azimuth angle; determining the target distortion type of each camera; using a database of a preset list of distortion types and deformation coefficients and the target distortion type to determine a target deformation coefficient list corresponding to the target distortion type, and the deformation coefficient list includes the corresponding relationship between the deformation coefficient and the azimuth angle under the distortion type; searching in the target deformation coefficient list based on the target azimuth angle to obtain the target deformation coefficient corresponding to the azimuth angle identical to the target azimuth angle; using the target deformation coefficient as a weight to perform prediction result optimization processing on the multiple first prediction results of each person to obtain the target prediction result of each person. Through the above method, the prediction results can be optimized by integrating the classroom behavior prediction results of multiple pre-trained artificial intelligence models and the imaging distortion characteristics, and the optimized target prediction results of each person can be obtained. This not only integrates the respective model advantages of multiple pre-trained artificial intelligence models, but also takes into account the imaging distortion of the person due to the azimuth angle between the person and the camera, and measures the influence of the imaging distortion on the prediction results. Therefore, the accuracy of the obtained target prediction results is higher.

[0101] Please refer to Figure 3 , Figure 3 which is a structural block diagram of an apparatus for identifying classroom behaviors in an embodiment of the present invention. As Figure 3 shown, the apparatus includes:

[0102] A data acquisition module 301: configured to obtain classroom audio-visual data collected by multiple audio-visual acquisition devices during the teaching process, where the classroom audio-visual data at least includes various classroom images collected by each audio-visual acquisition device;

[0103] A behavior prediction module 302: configured to input the classroom images into multiple pre-trained artificial intelligence models respectively for behavior prediction to obtain multiple first prediction results for each person in the classroom images, and the prediction results are used to reflect the classroom behavior types of the persons;

[0104] A distortion determination module 303: configured to determine the imaging distortion characteristics of each person by the audio-visual acquisition device;

[0105] A result optimization module 304: configured to perform prediction result optimization processing by using the imaging distortion characteristics and the multiple first prediction results to obtain the target prediction result of each person.

[0106] It should be noted that, as Figure 3 shown, the functions of the various modules in the device are similar to Figure 1 the content of the various steps in the method shown. To avoid repetition, no further elaboration will be provided here. Specifically, reference can be made to Figure 1 the content of the various steps in the method shown.

[0107] The present invention provides a device for recognizing classroom behaviors. The device includes: a data acquisition module: used to obtain classroom audio-visual data collected by multiple audio-visual acquisition devices during the teaching process. The classroom audio-visual data at least includes each classroom image collected by each audio-visual acquisition device; a behavior prediction module: used to input the classroom images into multiple pre-trained artificial intelligence models respectively for behavior prediction to obtain multiple first prediction results for each person in the classroom image. The prediction results are used to reflect the classroom behavior types of the persons; a distortion determination module: used to determine the imaging distortion characteristics of the audio-visual acquisition device for each person; a result optimization module: used to perform prediction result optimization processing by using the imaging distortion characteristics and the multiple first prediction results to obtain the target prediction result for each person. Through the above method, the prediction results of classroom behaviors of multiple pre-trained artificial intelligence models and the imaging distortion characteristics can be comprehensively used to optimize the prediction results, and the optimized target prediction results for each person can be obtained. This not only combines the respective model advantages of multiple pre-trained artificial intelligence models, but also takes into account the influence of imaging distortion on the prediction results, so that the accuracy of the obtained target prediction results is higher.

[0108] In a feasible implementation manner, the above device further includes:

[0109] a sample acquisition module: used to obtain training sample data. The training sample data at least includes a number of sample classroom images and the true labels of the classroom behavior types corresponding to the sample classroom images;

[0110] a model training module: used to input the sample classroom images into an initial convolutional neural network for training to obtain a fourth prediction label for each person;

[0111] a loss determination module: used to determine the loss value of the initial convolutional neural network by using the fourth prediction label, the true label, and a preset loss function;

[0112] a convergence judgment module: used to determine that the initial convolutional neural network converges and obtain a trained convolutional neural network if the loss value is lower than a preset loss threshold.

[0113] The present invention also provides a recognition device for classroom behaviors. The device includes: a sample acquisition module for obtaining training sample data, where the training sample data at least includes a number of sample classroom images and the true labels of the classroom behavior types corresponding to the sample classroom images; a model training module for inputting the sample classroom images into an initial convolutional neural network for training to obtain the fourth prediction labels for each person; a loss determination module for determining the loss value of the initial convolutional neural network by using the fourth prediction labels, the true labels, and a preset loss function; and a convergence judgment module for determining that the initial convolutional neural network converges and obtaining a trained convolutional neural network if the loss value is lower than a preset loss threshold. By training the initial convolutional neural network in the above manner and judging the convergence of the training based on the loss value of the loss function, not only can a convolutional neural network that can predict classroom behaviors be obtained, but also the convergence judgment accuracy is higher, improving the training accuracy.

[0114] Figure 4 FIG. shows the internal structure diagram of a computer device in an embodiment. The computer device may specifically be a terminal or a server. As Figure 4 shown, the computer device includes a processor, a memory, and a network interface connected through a system bus. Among them, the memory includes a non-volatile storage medium and an internal memory. The non-volatile storage medium of the computer device stores an operating system and may also store a computer program. When the computer program is executed by the processor, the processor can implement the above method. The internal memory may also store a computer program. When the computer program is executed by the processor, the processor can execute the above method. Those skilled in the art can understand that Figure 4 the structure shown in FIG. is only a block diagram of some structures related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.

[0115] In an embodiment, a computer device is proposed, including a memory and a processor. The memory stores a computer program. When the computer program is executed by the processor, the processor is caused to execute the steps of the method as Figure 1 or Figure 2 shown.

[0116] In an embodiment, a computer-readable storage medium is proposed, storing a computer program. When the computer program is executed by the processor, the processor is caused to execute the steps of the method as Figure 1 or Figure 2 shown.

[0117] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The program can be stored in a non-volatile computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memories can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memories can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.

[0118] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.

[0119] The above-described embodiments only represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application should be subject to the appended claims.

Claims

1. A method for identifying classroom behavior, characterized in that: The method comprises: Acquire classroom audio and video data collected by multiple audio and video collection devices during the teaching process, wherein the classroom audio and video data at least includes each classroom image collected by each audio and video collection device; Inputting the classroom image into a plurality of pre-trained artificial intelligence models for behavior prediction, respectively, to obtain a plurality of first prediction results for each character in the classroom image, wherein the prediction results are used to reflect the classroom behavior type of the character; Determining imaging distortion characteristics of each person by the video and audio acquisition device; Utilizing the imaging distortion feature and the plurality of first prediction results to perform prediction result optimization processing to obtain a target prediction result for each person; Wherein, the video and audio acquisition device at least includes a camera, and the step of determining the imaging distortion characteristics of each person by the video and audio acquisition device includes: Determine each target azimuth between each person and each camera using the first position coordinates of each person in the world coordinate system and the second position coordinates of the camera in the world coordinate system, wherein the imaging distortion feature includes the target azimuth; The step of optimizing the prediction results by using the imaging distortion features and the plurality of the first prediction results to obtain the target prediction result for each person includes: Determining a target distortion type for each of the cameras; Determine a target deformation coefficient list corresponding to the target deformation type by using a preset database of deformation types and deformation coefficient lists and the target deformation type, wherein the deformation coefficient list includes a corresponding relationship between the deformation coefficient and the azimuth under the distortion type; Searching in the target deformation coefficient list based on the target azimuth angle, obtaining a target deformation coefficient corresponding to the azimuth angle that is the same as the target azimuth angle; Using the target deformation coefficient as a weight to perform prediction result optimization processing on the plurality of the first prediction results for each character, to obtain a target prediction result for each character; Wherein, the first prediction result includes the first prediction label of the classroom behavior type and the credibility of the first prediction label, then the target deformation coefficient is used as a weight to perform prediction result optimization processing on multiple first prediction results of each character to obtain the target prediction result of each character, including: Taking the target azimuth as a clustering condition, clustering the plurality of the first prediction results for each character respectively, to obtain a target cluster set corresponding to each target azimuth of each character, wherein the target cluster set includes each second prediction result corresponding to the target azimuth; Using the credibility as a weight, respectively perform weighted decision processing on the second prediction labels included in the second prediction results of each target cluster set to obtain a third prediction label after weighted decision; The target deformation coefficient is used as a weight to perform weighted decision processing on each third prediction label of each character to obtain a final prediction label of each character, and the target prediction result includes the final prediction label.

2. The method according to claim 1, characterized in that: The artificial intelligence model at least includes a Transformer model based on a self-attention mechanism, a three-dimensional convolutional neural network, a recurrent neural network, and a convolutional neural network. The classroom image is respectively input into a plurality of pre-trained artificial intelligence models for behavior prediction to obtain a plurality of first prediction results for each character in the classroom image, including: The classroom image is respectively input into the trained Transformer model based on the self-attention mechanism, the three-dimensional convolutional neural network, the recurrent neural network and the convolutional neural network for behavior prediction to obtain multiple first prediction results for each character in the classroom image.

3. The method according to claim 2, characterized in that: The method further comprises: Acquire training sample data, wherein the training sample data includes at least a number of sample classroom images and true labels of classroom behavior types corresponding to the sample classroom images; Input the sample classroom images into the initial convolutional neural network for training to obtain the fourth predicted label for each character; Determine a loss value of the initial convolutional neural network using the fourth predicted label, the true label, and a preset loss function; If the loss value is lower than a preset loss threshold, it is determined that the initial convolutional neural network has converged, and a trained convolutional neural network is obtained.

4. The method according to claim 3, characterized in that: The loss function is expressed as follows: In the formula, N is the total number of samples, M is the total number of classroom behavior types, and IOU i Represents the ratio of the intersection and union of the predicted box and the true box of sample i, D i is the distance between the predicted box and the center point of the real box of sample i, L i is the diagonal length of the minimum enclosing rectangle of the predicted box and the real box of sample i, v i is the similarity between the aspect ratio of the predicted box and the real box of sample i, and α is v i The impact factor, y ic is an indicator variable used to indicate the authenticity of the classroom behavior type c of the fourth predicted label of sample i. If the classroom behavior type c is the same as the classroom behavior type of the true label of sample i, then y ic The value is 1, otherwise it is 0. ic Represents the predicted probability that observation sample i belongs to classroom behavior type c.

5. A device for identifying classroom behavior, characterized in that: The device comprises: Data acquisition module: used to acquire classroom audio and video data collected by multiple audio and video acquisition devices during the teaching process, wherein the classroom audio and video data at least includes each classroom image collected by each audio and video acquisition device; Behavior prediction module: used for inputting the classroom image into a plurality of pre-trained artificial intelligence models for behavior prediction, and obtaining a plurality of first prediction results for each character in the classroom image, wherein the prediction results are used to reflect the classroom behavior type of the character; Distortion determination module: used to determine the imaging distortion characteristics of each person by the video and audio acquisition device; A result optimization module: used for optimizing the prediction results by using the imaging distortion features and the first prediction results to obtain a target prediction result for each person; Wherein, the video and audio acquisition device at least includes a camera, and the distortion determination module is specifically used to: determine each target azimuth between each person and each camera using the first position coordinates of each person in the world coordinate system and the second position coordinates of the camera in the world coordinate system, wherein the imaging distortion feature includes the target azimuth; Wherein, the result optimization module is specifically used to: determine the target distortion type of each of the cameras; determine the target distortion coefficient list corresponding to the target distortion type by using a preset database of distortion types and distortion coefficient lists and the target distortion type, wherein the distortion coefficient list includes the corresponding relationship between the distortion coefficient and the azimuth under the distortion type; search the target distortion coefficient list based on the target azimuth to obtain the target distortion coefficient corresponding to the azimuth that is the same as the target azimuth; use the target distortion coefficient as a weight to perform prediction result optimization processing on the multiple first prediction results of each character to obtain the target prediction result of each character; Among them, the first prediction result includes the first prediction label of the classroom behavior type and the credibility of the first prediction label, and the target deformation coefficient is used as a weight to perform prediction result optimization processing on multiple first prediction results of each character to obtain the target prediction result of each character, including: taking the target azimuth as the clustering condition, clustering the multiple first prediction results of each character respectively, and obtaining a target clustering set corresponding to each target azimuth of each character, and the target clustering set includes each second prediction result corresponding to the target azimuth; using the credibility as the weight, performing weighted decision processing on the second prediction labels included in the second prediction results of each target clustering set, and obtaining a third prediction label after weighted decision; using the target deformation coefficient as the weight, performing weighted decision processing on each third prediction label of each character, and obtaining a final prediction label for each character, and the target prediction result includes the final prediction label.

6. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the processor is caused to perform the steps of the method according to any one of claims 1 to 4.

7. A computer device comprising a memory and a processor, characterized in that: The memory stores a computer program, and when the computer program is executed by the processor, the processor is caused to perform the steps of the method according to any one of claims 1 to 4.

Citation Information

Patent Citations

  • Student classroom behavior detection method, system and terminal based on deep learning

    CN114359606A

  • Method for analyzing classroom interaction behaviors based on audio and video

    CN114998968A