Student learning state supervision method and device based on artificial intelligence, terminal equipment and storage medium

Through the combination of video and audio data acquisition and intelligent analysis, the accuracy of students' learning status monitoring in the training course is solved, and comprehensive monitoring and analysis of students' learning status is achieved, and learning efficiency and quality are improved.

CN119918016AActive Publication Date: 2025-05-02XUEERWEI INTELLIGENT VEHICLE (SHENZHEN) CO LTD

Patent Information

Application Number
CN202510414490.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-03
Publication Date
2025-05-02
Estimated Expiration
2045-04-03

AI Technical Summary

Technical Problem

The prior art cannot comprehensively and accurately capture the learning status of students in the training course, resulting in difficult reflection of learning efficiency and effects.

Method used

Through video and audio data acquisition, combined with keyframe recognition, abnormal behavior recognition, speech segmentation and text level sorting, behavior description and text analysis are integrated to generate students' learning status information.

Benefits of technology

It realizes timely monitoring and analysis of students' learning status during practical training, provides support for learning plans and teaching methods, and improves learning efficiency and quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119918016A_ABST
    Figure CN119918016A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a student learning state supervision method and device based on artificial intelligence, terminal equipment and a storage medium, and belongs to the technical field of artificial intelligence. The method comprises the following steps: performing key frame identification according to initial video data to obtain a target behavior sequence; obtaining a target abnormal behavior according to the target behavior sequence; obtaining a target association frame according to the target abnormal behavior, and obtaining a behavior description text according to the target association frame; performing voice segmentation on the initial audio data to obtain target audio data, and performing voice recognition on the target audio data to obtain target text data; performing text level identification on the target text data to obtain target text levels, and sorting the target text levels according to time to obtain a text level sequence; determining first state information according to the text level sequence; performing difference analysis according to the behavior description text and the target text data to obtain second state information; and fusing the first state information and the second state information to determine target state information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to an artificial intelligence-based student learning status supervision method, device, terminal equipment and storage medium. Background Art

[0002] With the continuous development of artificial intelligence, its influence has increasingly penetrated into the field of teaching. When artificial intelligence is introduced into the teaching scene, with the help of advanced artificial intelligence technology, students' learning status can be supervised from multiple dimensions, including students' class videos, homework, etc. In this way, we can have a more comprehensive understanding of students' mastery of knowledge and learning progress. However, in actual application, the existing technology has exposed obvious deficiencies in supervising students with practical training courses. Due to the limitations of current supervision methods, it is difficult to achieve precise supervision, and it is impossible to fully and accurately capture the rich and diverse learning status of students in practical training courses. As a result, there are large errors in the supervision of students' status in practical training courses, and it is difficult to truly reflect the actual learning effect of students in the practical training link. Summary of the invention

[0003] The main purpose of the embodiments of the present invention is to provide a student learning status supervision method, device, terminal device and storage medium based on artificial intelligence, aiming to solve the problem in the related technology that the learning status of students in practical training courses cannot be fully and accurately captured, thereby affecting the students' learning efficiency.

[0004] In a first aspect, an embodiment of the present invention provides a method for supervising a student's learning status based on artificial intelligence, comprising: Collecting initial video data corresponding to a target student at a training station using a video collector and collecting initial audio data corresponding to the target student at the training station using an audio collector; Performing key frame recognition according to the initial video data to obtain a target behavior sequence corresponding to the target student; Perform abnormal behavior recognition according to the target behavior sequence to obtain the target abnormal behavior corresponding to the target student; Obtaining a target-related frame from the target behavior sequence according to the target abnormal behavior, and performing a behavior description according to the target-related frame to obtain a behavior description text corresponding to the target student; Performing speech segmentation on the initial audio data to obtain target audio data corresponding to the target student, and performing speech recognition on the target audio data to obtain target text data; Performing text grade recognition on the target text data to obtain a corresponding target text grade, and sorting the target text grades according to time to obtain a text grade sequence corresponding to the target student; Determining first status information corresponding to the target student according to the text level sequence; Performing difference analysis on the behavior description text and the target text data to obtain second status information corresponding to the target student; The first state information and the second state information are integrated to determine the target state information corresponding to the target student.

[0005] In a second aspect, an embodiment of the present invention provides a student learning status monitoring device based on artificial intelligence, comprising: A data acquisition module, used to acquire initial video data corresponding to a target student at a training station using a video collector and to acquire initial audio data corresponding to the target student at the training station using an audio collector; A video processing module, used for performing key frame recognition according to the initial video data to obtain a target behavior sequence corresponding to the target student; An abnormality identification module, used for performing abnormal behavior identification according to the target behavior sequence to obtain the target abnormal behavior corresponding to the target student; A behavior analysis module, used for obtaining a target-related frame from the target behavior sequence according to the target abnormal behavior, and performing a behavior description according to the target-related frame to obtain a behavior description text corresponding to the target student; An audio processing module, used for performing speech segmentation on the initial audio data to obtain target audio data corresponding to the target student, and performing speech recognition on the target audio data to obtain target text data; A data sorting module, used to perform text grade recognition on the target text data to obtain the corresponding target text grade, and sort the target text grade according to time to obtain the text grade sequence corresponding to the target student; A state determination module, used for determining first state information corresponding to the target student according to the text level sequence; A difference analysis module, used for performing a difference analysis on the behavior description text and the target text data to obtain the second state information corresponding to the target student; A result determination module is used to fuse the first state information and the second state information to determine the target state information corresponding to the target student.

[0006] In the third aspect, an embodiment of the present invention further provides a terminal device, comprising a processor, a memory, a computer program stored in the memory and executable by the processor, and a data bus for realizing connection and communication between the processor and the memory, wherein when the computer program is executed by the processor, the steps of any one of the methods for supervising student learning status based on artificial intelligence provided in the specification of the present invention are implemented.

[0007] In a fourth aspect, an embodiment of the present invention further provides a storage medium for computer-readable storage, characterized in that the storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the steps of any one of the artificial intelligence-based student learning status supervision methods provided in the specification of the present invention.

[0008] The embodiment of the present invention provides a student learning status supervision method, device, terminal device and storage medium based on artificial intelligence. The method includes: collecting initial video data corresponding to a target student at a training station according to a video collector and collecting initial audio data corresponding to the target student at the training station according to an audio collector, and then performing key frame recognition according to the initial video data to obtain a target behavior sequence corresponding to the target student, and then performing abnormal behavior recognition according to the target behavior sequence to obtain a target abnormal behavior corresponding to the target student, so as to timely discover abnormal situations such as improper operations and illegal behaviors of the target student, and then obtaining a target-related frame from the target behavior sequence according to the target abnormal behavior, and performing a behavior description according to the target-related frame to obtain a behavior description text corresponding to the target student, so as to provide support for subsequent questions asked by the target student to the teacher to determine whether they match, and then performing voice segmentation on the initial audio data to obtain the target student's corresponding Target audio data, and perform speech recognition on the target audio data to obtain target text data, perform text level recognition on the target text data to obtain the corresponding target text level, and sort the target text level according to time to obtain the text level sequence corresponding to the target student, so as to determine the first state information corresponding to the target student according to the text level sequence; and perform difference analysis on the behavior description text and the target text data to obtain the second state information corresponding to the target student, and finally fuse the first state information and the second state information to determine the target state information corresponding to the target student, so as to timely discover the learning state of the target student in the practical training process, and then provide good support for the subsequent formulation of learning plans or teaching methods for the target students, thereby improving the learning efficiency and learning quality of the target students. This method also solves the problem that the related technology cannot comprehensively and accurately capture the learning state of students in practical training courses, thereby affecting students' learning efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0010] Figure 1A flowchart of a method for supervising student learning status based on artificial intelligence provided by an embodiment of the present invention; Figure 2 A schematic diagram of the module structure of a student learning status monitoring device based on artificial intelligence provided by an embodiment of the present invention; Figure 3 A schematic block diagram of the structure of a terminal device provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0011] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0012] The flowcharts shown in the accompanying drawings are only examples and do not necessarily include all the contents and operations / steps, nor must they be executed in the order described. For example, some operations / steps may also be decomposed, combined or partially merged, so the actual execution order may change according to actual conditions.

[0013] It should be understood that the terms used in this specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. As used in the specification of the present invention and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an" and "the" are intended to include plural forms.

[0014] The embodiment of the present invention provides a method, device, terminal device and storage medium for supervising the learning status of students based on artificial intelligence. The method for supervising the learning status of students based on artificial intelligence can be applied to a terminal device, which can be an electronic device such as a tablet computer, a laptop computer, a desktop computer, a personal digital assistant and a wearable device. The terminal device can be a server or a server cluster.

[0015] Some embodiments of the present invention are described in detail below in conjunction with the accompanying drawings. In the absence of conflict, the following embodiments and features in the embodiments can be combined with each other.

[0016] Please refer to Figure 1 , Figure 1 A flowchart of a method for supervising student learning status based on artificial intelligence is provided in an embodiment of the present invention.

[0017] like Figure 1 As shown, the artificial intelligence-based student learning status supervision method includes steps S101 to S109.

[0018] Step S101: collecting initial video data corresponding to a target student at a training station using a video collector and collecting initial audio data corresponding to the target student at the training station using an audio collector.

[0019] For example, in order to obtain the target status information corresponding to the target students in a timely manner, a video collector and an audio collector are installed at the training stations corresponding to the target students, wherein appropriate video acquisition equipment is selected according to the training environment and needs. For example, in a well-lit and open station, an ordinary high-definition network camera can meet the needs; in an environment with dim light or extremely high image quality requirements, a low-light, high-resolution camera is more suitable. And select appropriate audio equipment according to the noise level and pickup range of the training station. If the station environment is quiet, an ordinary directional microphone is sufficient; if the environment is noisy, a microphone with noise reduction function needs to be selected.

[0020] Exemplarily, each target student is assigned a unique identification, such as a badge, bracelet, etc., which may contain the target student's name, student number, and other information. The video collector and audio collector may combine image recognition or other technologies to accurately identify the target student, thereby associating the identity information of each target student with the corresponding video collector and audio collector to ensure that the collected data can accurately correspond to each target student, and then obtain the initial video data and initial audio data corresponding to the target student at the training station according to the corresponding identity identifier of the target student.

[0021] Step S102: performing key frame recognition based on the initial video data to obtain a target behavior sequence corresponding to the target student.

[0022] For example, firstly, unnecessary parts of the initial video data are cropped, such as blank edges, backgrounds unrelated to the target student, etc. Then, the target detection algorithms such as YOLO (You Only LookOnce) and Faster R-CNN are used on the cropped initial video data to detect the position of the target student in the video frame, such as the human body outline and bounding box. Then, after the target student is detected, the target tracking algorithms such as KCF (Kernelized Correlation Filters) and CSRT (Discriminative Correlation Filter with Channel and Spatial Reliability) are used to continuously track the target student. Ensure that the position of the target student can be accurately found in each frame of the video, and provide an accurate target area for subsequent key frame recognition.

[0023] Exemplarily, target features such as color features (such as color histogram), texture features (such as gray level co-occurrence matrix), shape features (such as contour moment), etc. are extracted from the target area. These features can describe the visual content of the video frame, so that the feature similarity between adjacent video frames is calculated using methods such as Euclidean distance and cosine similarity. When the similarity is lower than a certain set threshold, it is considered that a large change has occurred between the two frames and may contain key information, and the frame is marked as a candidate key frame.

[0024] For example, the identified candidate key frames are screened to remove some repeated or meaningless key frames. For example, if the student's movements in several consecutive frames are only slightly jittery and there is no substantial behavioral change, these frames can be merged or deleted, so that the screened key frames are sorted according to the time sequence of the video to ensure that the order of the key frames is consistent with the order in which the student's behavior occurs, so as to obtain the target behavior sequence corresponding to the target student.

[0025] Step S103: performing abnormal behavior identification according to the target behavior sequence to obtain the target abnormal behavior corresponding to the target student.

[0026] For example, based on the operation process, such as the correct operation sequence and action requirements of each training step, the behavior that does not conform to the process is defined as abnormal. For example, in assembly training, if the student skips a key part installation step, it is an abnormal behavior, or the abnormal behavior rules are determined based on the equipment usage specifications. For example, if the student starts the equipment without necessary equipment preheating, or performs illegal adjustment operations while the equipment is running, etc., the defined abnormal behavior rules are organized to form a structured rule base.

[0027] For example, key features that can reflect the characteristics of the behavior are extracted from the target behavior sequence, including the duration of the behavior, the frequency of the action, the amplitude of the action, the order in which the behavior occurs, etc., so as to combine with the abnormal behavior rule base to determine the features related to each rule, and quantify the extracted behavior features so that they can be expressed in numerical or other comparable forms. For example, the duration of the behavior is converted into a specific time value, and the amplitude of the action is converted into an angle or length value, so as to normalize the quantized features, eliminate the dimensional differences between different features, and ensure that the features are comparable in subsequent matching and analysis.

[0028] Exemplarily, abnormal behavior rules and quantified behavior features are used to perform abnormality judgment according to a machine learning classification algorithm, so as to determine which parts of the target behavior sequence belong to abnormal behaviors, thereby obtaining the target abnormal behavior corresponding to the target student.

[0029] In some embodiments, the abnormal behavior identification according to the target behavior sequence to obtain the target abnormal behavior corresponding to the target student includes: performing target identification on each target image in the target behavior sequence to obtain the target area corresponding to the target image; identifying the joint information corresponding to the target student in the target area to obtain the target joint information corresponding to the target image; obtaining the relevant joint information corresponding to each sub-joint from the target joint information corresponding to each target image in the target behavior sequence; performing curve fitting according to the relevant joint information to obtain the initial fitting curve corresponding to the sub-joint; performing change point identification on the initial fitting curve to obtain the joint change time corresponding to the sub-joint; fusing the relevant joint information according to the joint change times corresponding to multiple sub-joints to obtain the target behavior type corresponding to the target student; and performing abnormal behavior classification according to the target behavior type to obtain the target abnormal behavior corresponding to the target student.

[0030] For example, target detection algorithms such as Faster R-CNN and YOLO are used to identify each target image in the target behavior sequence to obtain the target area corresponding to the target student in each target image. Human body posture estimation algorithms such as OpenPose and AlphaPose are then used to perform joint recognition on the content corresponding to the target area in the target image to obtain the position information of each joint of the target student in the target area, including the coordinates of the joints, etc. This information is the target joint information corresponding to the target image.

[0031] Exemplarily, all identified joint information is classified and different sub-joints are marked, such as the left shoulder joint, the right knee joint, etc. According to the category of the sub-joint, the joint information of the sub-joint corresponding to each target image in the target behavior sequence is associated and sorted to form the relevant joint information corresponding to each sub-joint, and then the relevant joint information of the sub-joint (such as the change of the coordinates of the joint over time) is fitted by polynomial fitting, spline curve fitting, etc., to find a curve that can better approximate these data points, so as to obtain the initial fitting curve corresponding to each sub-joint.

[0032] Exemplarily, a statistical-based method or a machine learning-based method is used to find points in the initial fitting curve where the curve changes significantly, so that the time corresponding to the identified change point is used as the joint change time of the sub-joint.

[0033] Exemplarily, different target behavior types and their corresponding joint change patterns and joint information features are predefined. For example, the change patterns and information features of each sub-joint (such as shoulder joint, elbow joint, etc.) corresponding to the "raising hand" behavior are defined. Then, the corresponding joint position at the time is obtained according to the joint change time, and then the corresponding time clusters under the multiple joint change times are obtained by clustering according to the multiple joint change times, so as to obtain the first joint position corresponding to each sub-change time in each sub-cluster under the time cluster, and obtain the second joint position of other joints in the human body at the sub-change time, and then connect according to the order of human joints according to the first joint position and the second joint position, so as to obtain the joint pattern corresponding to the sub-cluster, and then match and fuse it with the predefined behavior pattern. By analyzing the change order, amplitude and other information of each sub-joint, the target behavior type corresponding to the target student is determined.

[0034] For example, according to the requirements of the training scenario and safety regulations, abnormal behavior classification rules corresponding to different target behavior types are formulated. For example, in a specific training operation, "raising your hand too quickly" may be defined as an abnormal behavior, and then the identified target behavior type is compared with the abnormal behavior classification rules to determine whether the behavior is an abnormal behavior. If it is an abnormal behavior, it is marked as a target abnormal behavior and the relevant information is recorded.

[0035] Specifically, by performing target recognition and joint information recognition on the target image, the target student can be accurately located and its detailed joint motion information can be obtained. Curve fitting and change point recognition can extract key change information from the time series data of joint motion, which helps to discover some difficult-to-detect abnormal behaviors, such as small movement deviations or irregular behavior start and end times. Then, from the perspective of multiple sub-joints, and by integrating relevant joint information, the behavior pattern of the target student can be fully described. Different sub-joints play different roles in behavior. Comprehensively considering the information of multiple sub-joints can provide a more comprehensive understanding of the student's behavior and avoid incomplete behavior analysis caused by focusing only on some joints.

[0036] In some embodiments, the method of fusing the relevant joint information according to the joint change times corresponding to multiple sub-joints to obtain the target behavior type corresponding to the target student includes: performing an intersection operation on the joint change times to obtain all joint times corresponding to all the sub-joints; traversing each sub-time in the total joint times, and obtaining the associated joint information that has changed at the sub-time from the relevant joint information according to the joint change time; obtaining the adjacent joint information corresponding to the previous time in the sub-time from the target joint information; determining the current joint information corresponding to the sub-time according to the adjacent joint information and the associated joint information; and performing behavior type classification according to the adjacent joint information and the current joint information to obtain the target behavior type corresponding to the target student at the sub-time.

[0037] Exemplarily, the joint change time data corresponding to multiple sub-joints are collected and sorted, and then the common time points of the joint change times of all sub-joints are found using the set intersection operation method. These common time points constitute the total joint time corresponding to all sub-joints. This process is like finding overlapping time scales in multiple time axes.

[0038] Exemplarily, each sub-time in all joint times is accessed in turn. For each sub-time, it is compared with the joint change time of each sub-joint. If the joint change time of a sub-joint includes this sub-time, it means that the sub-joint has changed at this sub-time point. The information of these changed sub-joints is extracted from the relevant joint information, and this information is the associated joint information that has changed under this sub-time.

[0039] For example, the target joint information is sorted in chronological order to ensure that the previous time point corresponding to each sub-time can be accurately found. For the currently traversed sub-time, its corresponding previous time point is found, and then the joint information of all sub-joints corresponding to the previous time point is extracted from the target joint information, and this information is the adjacent joint information.

[0040] Exemplarily, the current joint information corresponding to the sub-time is determined by combining the adjacent joint information and the associated joint information. For the sub-joints that have changed, their joint states are updated according to the associated joint information; for the sub-joints that have not changed, their states in the adjacent joint information are maintained. In this way, the complete current joint information at the sub-time is obtained.

[0041] Exemplarily, different behavior types and their corresponding joint motion patterns are predefined. For example, behavior types such as "screwing", "welding", and "hammering" are defined, as well as the characteristic performance of these behavior types on adjacent joint information and current joint information, such as the range of change of joint angles, the moving direction of joint positions, etc., and then the adjacent joint information and current joint information are compared and matched with the predefined behavior patterns. By analyzing the change characteristics of joint information, the behavior type of the target student at that sub-time is judged, thereby determining the corresponding target behavior type.

[0042] Specifically, by performing intersection operations and comprehensive analysis on the joint change times of multiple sub-joints, the changing moments of the target student's behavior can be captured more comprehensively and accurately. This makes it more precise to determine the joint information at each sub-time, thus providing a more reliable basis for the classification of behavior types. In addition, the behavior type classification is combined with the adjacent joint information and the current joint information, taking into account the dynamic change process of the behavior. Compared with classification based only on the joint information at a single moment, this method can better reflect the continuity and change trend of the behavior, reduce the possibility of misjudgment, and improve the accuracy of classification.

[0043] Step S104: obtaining a target-related frame from the target behavior sequence according to the target abnormal behavior, and performing a behavior description according to the target-related frame to obtain a behavior description text corresponding to the target student.

[0044] Exemplarily, a key frame determined as a target abnormal behavior in the target behavior sequence is determined as a target associated frame corresponding to the target abnormal behavior.

[0045] Exemplarily, key features related to the behavior are extracted from the target-associated frame. For example, the posture of the character (standing, bending over, raising hands, etc.), the amplitude of the movement (large swing, small movement), the relative position relationship of various parts of the body (the position of the hand and the head, the distance between the feet, etc.) and environmental information (whether the device is being operated, whether there are other people around, etc.). Therefore, the actions in the target-associated frame are classified according to the extracted features to determine the action type corresponding to the target-associated frame, and then combined with the frames arranged in chronological order, the sequence and logical relationship between the actions are analyzed. Determine whether the action is independent or part of a series of coherent actions. For example, "picking up a hammer" may be followed by "hitting an object", which constitutes a behavior sequence with a logical order.

[0046] Exemplarily, the identified actions and the training types corresponding to the training positions are integrated so as to combine individual actions into a behavior description text with complete meaning, such as "the student at the training station first picks up a hammer and then hits it on a screw."

[0047] Exemplarily, the behavior description text is used to characterize the corresponding content when the abnormal behavior of the target student determined according to the target association frame is described using text, thereby providing support for whether the description of the questions asked by the target student to the teacher is accurate when the teaching objectives cannot be achieved due to the abnormal behavior, and further judging the learning status of the target student.

[0048] In some embodiments, the behavior description according to the target association frame is performed to obtain the behavior description text corresponding to the target student, including: obtaining a training image and an image description text corresponding to the training image, and performing potential feature extraction on the training image according to the image feature extraction layer of the information association model to obtain a first image feature corresponding to the training image; performing text feature extraction on the image description text according to the text feature extraction layer of the information association model to obtain a first text feature corresponding to the image description text; performing feature mapping on the first image feature and the first text feature according to the feature association layer of the information association model to obtain an association triplet corresponding to the training image; and performing coding according to the text generation model. The code layer performs feature extraction on the target associated frame to obtain the corresponding second image feature; the data determination layer of the text generation model uses the target associated frame and the training image to obtain the target triplet corresponding to the target associated frame from the associated triplet; the self-attention layer of the text generation model performs feature extraction on the target triplet to obtain the corresponding second text feature; the structural attention layer of the text generation model uses the second image feature and the second text feature to obtain the target relationship between each entity in the target key frame; the long short-term memory network layer of the text generation model generates the behavior description text corresponding to the target student according to the target relationship and the target triplet.

[0049] For example, a large number of training images and their corresponding image description texts are collected. These training images should contain rich and diverse scenes and behaviors, and the image description texts should accurately and in detail describe the content in the images.

[0050] Exemplarily, the training image is processed using the image feature extraction layer of the information association model. This layer analyzes the visual information of the image, such as color, texture, shape, and object contour, and converts this information into potential feature representations, thereby obtaining the first image features corresponding to the training image. The image description text is analyzed by the text feature extraction layer of the information association model. This layer considers the vocabulary, grammar, semantics, and other information in the text and converts the text into the first text feature in the form of a vector.

[0051] Exemplarily, the feature association layer of the information association model maps the first image feature and the first text feature. It searches for the correspondence between the image features and the text features, determines which image features are associated with which text features, and then generates association triples corresponding to the training image based on the result of the feature mapping. The association triples usually contain image features, text features, and the association relationship between them, which helps to accurately establish the connection between the image and the text in the subsequent steps.

[0052] Exemplarily, the encoding layer of the text generation model performs feature extraction on the target-related frame. This layer extracts the visual features of the target-related frame like processing the training image to obtain the corresponding second image features.

[0053] Exemplarily, the data determination layer of the text generation model uses the target associated frame and the training image to search and match in the associated triples. It will find the associated triple that is most relevant to the target associated frame and determine it as the target triple corresponding to the target associated frame. The self-attention layer of the text generation model processes the target triple, extracts the text-related features therein, and obtains the corresponding second text features. The self-attention mechanism allows the model to focus on the important relationships between different parts in the target triple. The structural attention layer of the text generation model uses the second image features and the second text features to analyze the target relationship between each entity in the target key frame. This includes spatial relationships, action relationships, semantic relationships, etc. between entities.

[0054] Exemplarily, the long short-term memory network layer of the text generation model generates the behavior description text corresponding to the target student based on the target relationship and the target triple. The long short-term memory network can process sequence information and combine the various features and relationships obtained in the previous steps to generate smooth and accurate natural language text to describe the behavior of the target student in the target-related frame.

[0055] Specifically, by extracting features from images and text and associating them, the model can make comprehensive use of visual information and semantic information. This allows the generated behavior description text to more accurately reflect the actual behavior in the target-associated frame, reducing description bias caused by inaccurate single feature analysis. The structural attention layer's analysis of the relationship between entities can capture the details and logic of the behavior. In addition, through the process of determining the target triples, the model can find situations similar to the target-associated frame from the training data and draw on the existing associated information to generate descriptions. This approach enhances the model's adaptability to different inputs and improves generalization capabilities.

[0056] In some embodiments, the data determination layer according to the text generation model uses the target associated frame and the training image to obtain the target triplet corresponding to the target associated frame from the associated triplet, including: performing keyword extraction on the image description text according to the keyword extraction network of the data determination layer to obtain initial keywords; performing frequency statistics on the initial keywords according to the image description text according to the data statistics network of the data determination layer to obtain frequency information corresponding to the initial keywords; using the frequency information to filter the initial keywords according to the data screening network of the data determination layer to obtain image keywords corresponding to the image description text; using the image keywords to label the target associated frame according to the label determination layer to obtain the first associated keywords and the first associated keywords corresponding to the target associated frame. a first position information corresponding to an associated keyword; according to the label determination layer of the data determination layer, the training image is labeled using the image keyword to obtain a second associated keyword corresponding to the training image and a second position information corresponding to the second associated keyword; according to the data merging network of the data determination layer, the first associated keyword and the second associated keyword are subjected to intersection processing to obtain a target associated keyword; according to the similarity calculation network of the data determination layer, the image association degree between the target associated frame and the training image is determined according to the target associated keyword, the first position information and the second position information; according to the data determination network of the data determination layer, the target triplet corresponding to the target associated frame is obtained from the association triplet using the data image association degree; wherein the image association degree is obtained according to the following formula: ; in, represents the image association degree between the target associated frame and the i-th training image, num represents the number of words corresponding to the target associated keyword, Indicates the number of words corresponding to the first associated keyword or the second associated keyword, represents the data corresponding to the jth target-related keyword in the i-th training image in the second position information, Represents data corresponding to the j-th target-associated keyword in the target-associated frame in the first position information.

[0057] For example, the keyword extraction network of the data determination layer processes the image description text to obtain the initial keywords, and then the data statistics network performs frequency statistics on the initial keywords based on the image description text. It traverses the entire text and calculates the number of times each initial keyword appears, thereby obtaining the frequency information corresponding to each initial keyword.

[0058] For example, the data screening network uses frequency information to screen initial keywords. Usually a frequency threshold is set, and only initial keywords with a frequency higher than the threshold will be retained. These retained keywords become the image keywords corresponding to the image description text. For example, if the threshold is set to 2, then "student" with a frequency of 3 will be retained, while other initial keywords that only appear once may be screened out.

[0059] Exemplarily, the label determination layer uses image keywords to label the target associated frame. It searches for content related to the image keywords in the target associated frame, marks the location of these contents, and obtains the first associated keyword corresponding to the target associated frame and the first position information corresponding to the first associated keyword. For example, if the image keywords are "hammer" and "nails", the label determination layer will find the locations of the hammer and nails in the target associated frame and record them. Similarly, the label determination layer uses image keywords to label the training image to obtain the second associated keyword corresponding to the training image and the second position information corresponding to the second associated keyword.

[0060] Exemplarily, the data merging network performs intersection processing on the first associated keyword and the second associated keyword. It will find keywords that appear in both the target associated frame and the training image. These keywords are the target associated keywords. For example, if the first associated keywords of the target associated frame are "hammer" and "nail", and the second associated keywords of the training image are "hammer" and "chair", then the target associated keyword is "hammer".

[0061] Exemplarily, the similarity calculation network calculates the image association between the target-associated frame and the training image using a given formula based on the target-associated keywords, the first position information, and the second position information. The formula comprehensively considers factors such as the number of target-associated keywords and position information to obtain a quantitative association value, and then obtains the image association according to the following formula: ; in, represents the image association degree between the target associated frame and the i-th training image, num represents the number of words corresponding to the target associated keywords, Indicates the number of words corresponding to the first associated keyword or the second associated keyword. represents the data corresponding to the jth target-related keyword in the i-th training image in the second position information, Represents the data corresponding to the j-th target-associated keyword in the target-associated frame in the first position information.

[0062] Exemplarily, the data determination network uses the calculated image association degree to obtain the target triplet corresponding to the target associated frame from the association triplet. Usually, the association triplet corresponding to the training image with the highest association degree with the target associated frame image is selected as the target triplet.

[0063] Specifically, through the steps of keyword extraction, frequency statistics and screening, the key information in the image description text can be accurately found. The labeling process further corresponds these keywords to the specific positions in the target associated frame and the training image, making the subsequent matching process more accurate. In addition, the image association degree is calculated using the above formula to quantify the similarity between the target associated frame and the training image. This enables an objective standard to be used when selecting the target triple, reduces the influence of subjective factors, and improves the accuracy of matching. Furthermore, using keywords for matching instead of directly comparing all the features of the image enables the model to find similar semantic information in different images. Even if there are differences in the visual features of the image, as long as it contains the same key semantic information, it can be considered to be related. This enhances the model's adaptability to different images and improves its generalization ability, thereby fully utilizing text information to assist image matching by extracting and utilizing keywords in the image description text. This enables the model to mine useful information from a large amount of training data, improves the efficiency of data utilization, and further enhances generalization ability.

[0064] Step S105: performing speech segmentation on the initial audio data to obtain target audio data corresponding to the target student, and performing speech recognition on the target audio data to obtain target text data.

[0065] For example, speaker recognition technology is used to distinguish different speakers in the initial audio data. First, a speaker recognition model is trained, which can learn the speech features of different speakers. Then the initial audio is input into the model, and the model determines the speaker identity of each speech segment based on the speech features, thereby segmenting the speech segment of the target student, and then obtaining the target audio data corresponding to the target student.

[0066] Exemplarily, an open source speech recognition model such as DeepSpeech, Wav2Vec, etc. is used to collect corresponding speech data and text content corresponding to the speech data, and the speech recognition model is subjected to speech training to obtain a corresponding target speech recognition model, and then speech recognition is performed on the target audio data according to the target speech recognition model to obtain target text data corresponding to the target audio data.

[0067] Step S106: Perform text grade recognition on the target text data to obtain the corresponding target text grade, and sort the target text grades according to time to obtain a text grade sequence corresponding to the target student.

[0068] For example, the purpose of text level recognition is to identify the difficulty level of the questions asked by the target students to the teacher during the communication process between the target students and the teacher, or to identify the difficulty level of the practical training content involved in the communication process between the target students and the teacher.

[0069] Exemplarily, a practical training course is determined, and the knowledge points involved in the practical training course at different levels of difficulty and the knowledge keywords corresponding to the knowledge points are determined based on the practical training course, so as to perform keyword recognition on the target text data to obtain text keywords, and then calculate the similarity between the text keywords and the knowledge keywords to obtain the target text level corresponding to the target students.

[0070] Exemplarily, the time corresponding to the target text data in the initial audio data is obtained, and then the target text levels are sorted in chronological order to obtain a text level sequence corresponding to the target students.

[0071] In some embodiments, the text level recognition of the target text data to obtain the corresponding target text level includes: determining the practical training course corresponding to the target student, and obtaining the course-related texts corresponding to the practical training course at different levels; calculating the similarity between the target text data and the course-related text to obtain the target similarity between the target text data and the course-related text; and determining the target text level corresponding to the target text data according to the target similarity and the level information corresponding to the course-related text.

[0072] For example, the training schedule is obtained through the target student's course selection record to determine the practical training course that the target student participates in, and then the course-related texts corresponding to the course at different levels are collected. The course-related texts can come from course materials, teaching syllabi, and teachers' lecture notes. For example, for a programming practical training course, the course-related texts at the elementary level may be an introduction to basic programming grammar, the intermediate level may be the implementation ideas of simple projects, and the advanced level may be the analysis of complex algorithms.

[0073] Exemplarily, a text similarity calculation method such as a semantic-based method (such as using a pre-trained language model) is used to calculate the similarity between the target text data and the collected course-related texts one by one to obtain the target similarity between the target text data and the course-related texts.

[0074] Exemplarily, each course-related text is associated with its corresponding level information, and this level information can be elementary, intermediate, or advanced. For example, the introduction to basic programming syntax corresponds to the elementary level, and the complex algorithm analysis corresponds to the advanced level. Then, based on the calculated target similarity and the level information corresponding to the course-related text, the target text level corresponding to the target text data is determined. For example, the level corresponding to the course-related text with the highest similarity to the target text data is selected as the target text level. For example, if the target text data has the highest similarity to the course-related text of the intermediate level, then the target text level corresponding to the target text data is intermediate.

[0075] Step S107: Determine the first status information corresponding to the target student according to the text level sequence.

[0076] Exemplarily, the text level sequence is an ordered set of levels corresponding to a series of texts related to the target students. These levels can be divided according to the difficulty, knowledge mastery, etc., such as elementary, intermediate, and advanced. The first state information is information describing the state of the target students in a specific situation, which may include learning status (such as learning progress, learning stagnation, learning difficulties), ability level status (such as ability improvement, ability stability, ability decline), knowledge mastery status (such as comprehensive knowledge mastery, partial knowledge mastery, poor knowledge mastery), etc.

[0077] Exemplarily, observe whether the level in the text level sequence shows a trend of gradually increasing. If so, it means that the student is making continuous progress and may perform well in terms of knowledge mastery and ability improvement. For example, the student's text level gradually increases from elementary to intermediate, and then to advanced. When the level shows a trend of gradually decreasing, it indicates that the student's first state information may be getting worse, and may have encountered problems such as learning difficulties or knowledge forgetting. For example, a student who was originally at an intermediate level subsequently had his text level reduced to elementary. If the level does not fluctuate much within a certain range, and there is no obvious upward or downward trend, it means that the student's first state information is relatively stable, and the knowledge mastery and ability level may be in a relatively stable stage.

[0078] For example, if the change range between adjacent levels in the text level sequence is large, it means that the student's state is unstable and may be affected by external factors (such as changes in the learning environment, emergencies) or internal factors (such as fluctuations in learning attitudes, improper learning methods). Small fluctuations may be a normal error range or a small adjustment in the learning process, which has a relatively small impact on the overall state of the student, corresponding to positive first state information such as learning progress, ability improvement, and better knowledge mastery. For example, if the text level sequence shows a clear upward trend, it can be judged that the student is in a state of rapid progress in learning, his ability is constantly improving, and his knowledge mastery is becoming more and more comprehensive. Corresponding to negative first state information such as learning difficulties, decreased ability, and worse knowledge mastery. For example, when the text level continues to decline, it can be considered that the student has encountered obstacles in the learning process, his ability level has declined, and his knowledge mastery is not ideal. Corresponding to state information such as stable learning, stable ability, and stable knowledge mastery. If the level fluctuation is not large, it means that the student's learning state is relatively stable, and there is no obvious change in ability and knowledge mastery.

[0079] In addition, large fluctuations in an upward trend may indicate that the student has made a major breakthrough in the learning process, but there may also be a component of luck or a weak foundation. The corresponding first state information may be "great progress but unstable." Large fluctuations in a downward trend indicate that the student's learning state has deteriorated sharply, and they may face greater learning pressure or serious problems with their learning methods. The corresponding first learning state is a sharp decline. In an upward, downward or stable trend, small fluctuations generally do not change the overall state judgment, but can reflect a certain stability in the first state information. For example, in a stable trend, small fluctuations can mean "the learning state is stable with small adjustments."

[0080] Step S108: Perform difference analysis on the behavior description text and the target text data to obtain second status information corresponding to the target student.

[0081] For example, the behavior description text is the description of the abnormal behavior of the target student in the practical training course, and the target text data is the questions asked by the target student and the teacher during the communication process. Then, when the difference between the abnormal behavior description and the target text data is smaller, it means that the target student can accurately grasp the key points in the process of asking questions, and the questions raised by the target student or the content of the communication with the teacher are closely related to the root cause of the abnormal behavior in the practical training process, indicating that the student has a clearer understanding of the key points of the practical training course and can keenly perceive his own problems. Based on this, it can be reasonably inferred that the second state information corresponding to the target student is that the learning direction is correct.

[0082] For example, on the contrary, when the difference between the abnormal behavior description and the target text data is greater, it indicates that the target student may be in a confused state. The target student does not know where the problem lies. The questions raised are not strongly related to the actual abnormal behavior, reflecting that they lack insight and thinking about key issues in the learning process. In this case, the second state information corresponding to the target student is the wrong learning direction.

[0083] In some embodiments, the obtaining of the second status information corresponding to the target student by performing a difference analysis on the behavior description text and the target text data includes: performing keyword recognition on the behavior description text to obtain a first relationship between a first keyword and the first keyword; performing keyword recognition on the target text data to obtain a second relationship between a second keyword and the second keyword; calculating a first similarity between the first keyword and the second keyword, and obtaining a second similarity between the behavior description text and the target text data by combining the first relationship and the second relationship under the first similarity; performing a difference analysis on the behavior description text and the target text data according to the second similarity to obtain the second status information corresponding to the target student.

[0084] Exemplarily, a keyword extraction method such as a part-of-speech-based method is used to obtain the first keyword corresponding to the behavior description text, and the logical relationship between the first keywords is analyzed, such as causal relationship (such as "operational error leads to result deviation", "operational error" and "result deviation" have a causal relationship), parallel relationship (such as "illegal operation and non-compliance with procedures", "illegal operation" and "non-compliance with procedures" are parallel relationships), progressive relationship, etc., and these relationships are recorded as the first relationship corresponding to the first keyword.

[0085] Exemplarily, the second keyword is extracted from the target text data using the same or similar method as that used to extract the first keyword, and the logical relationship between the second keywords is analyzed in the same way as the first relationship, and recorded as the second relationship corresponding to the second keyword.

[0086] Exemplarily, the cosine similarity is used to calculate the first similarity between the first keyword and the second keyword, and on the basis of the first similarity, the first relationship and the second relationship are considered. If the logical relationship between the first keyword and the second keyword is consistent, for example, there is a causal relationship and the causal direction is the same, then the similarity can be appropriately increased; if the relationship is inconsistent, such as one is a parallel relationship and the other is a causal relationship, then the similarity is appropriately reduced. Comprehensively considering the consistency of the first similarity and the relationship, the second similarity between the behavior description text and the target text data is obtained.

[0087] Exemplarily, different similarity threshold ranges are pre-set, for example, a high similarity threshold range (such as 0.8 - 1), a medium similarity threshold range (such as 0.5 - 0.8) and a low similarity threshold range (such as 0 - 0.5). The calculated second similarity is compared with the set threshold range. If the second similarity is in the high similarity threshold range, it means that the difference between the behavior description text and the target text data is small; if it is in the low similarity threshold range, it means that the difference is large; if it is in the medium similarity threshold range, it means that there is a certain difference but it is not very large.

[0088] For example, the second state information corresponding to the target student is determined based on the results of the difference analysis. When the difference is small, it means that the target student has a good grasp of the key points when asking questions, and the corresponding second state information is that the learning direction is correct; when the difference is large, it means that the target student may not be clear about his or her own problems at present, and the corresponding second state information is that the learning direction is wrong; for the medium similarity situation, the difference can be further analyzed to determine that the student may have some deviations in direction or inaccurate understanding, and give a corresponding appropriate description of the second state information.

[0089] Step S109: integrating the first state information and the second state information to determine the target state information corresponding to the target student.

[0090] Exemplarily, when the first state information and the second state information agree in describing the learning state of the target student, this indicates that the same conclusion has been drawn from different evaluation dimensions, and the first state information and the second state information can be merged. For example, the first state information indicates that "the target student has made significant progress in learning and has a solid grasp of knowledge", and the second state information indicates that "the target student has a correct learning direction", then the merged target state information can be expressed as "the target student has made significant progress in learning, has a solid grasp of knowledge, and has a correct learning direction". Such a merger makes the information more concise and clear, while completely retaining the core content conveyed by the two state information, thereby obtaining the target state information that can fully reflect the learning state of the target student. However, when the first state information and the second state information describe the learning state of the target student inconsistently, the situation becomes more complicated. This inconsistency may be caused by different emphases of different evaluation criteria, errors in data sources, or the multifaceted nature of the student's learning state itself. At this point, simply merging the two state information may confuse the user about the learning state of the target student. In order to clearly prompt this situation, it is necessary to add a preset sentence "there is an abnormality in the current learning state judgment of the target student" after merging the first state information and the second state information. For example, the first status information shows that "the target student has difficulty learning and weak knowledge", while the second status information shows that "the target student has the right direction of learning and has made significant progress". After merging and adding a preset sentence, the target status information can be written as "the target student has difficulty learning and weak knowledge, but has the right direction of learning and has made significant progress. The target student's current learning status is judged to be abnormal." By adding this preset sentence, relevant personnel can be reminded to analyze and judge the actual learning status of the target student more carefully when referring to the target status information, and further explore the reasons for this inconsistency so as to take more targeted measures to help students improve their learning outcomes.

[0091] See also Figure 2 , Figure 2An artificial intelligence-based student learning status monitoring device 200 is provided in an embodiment of the present application. The artificial intelligence-based student learning status monitoring device 200 includes a data acquisition module 201, a video processing module 202, an abnormality recognition module 203, a behavior analysis module 204, an audio processing module 205, a data sorting module 206, a state determination module 207, a difference analysis module 208, and a result determination module 209, wherein the data acquisition module 201 is used to collect initial video data corresponding to a target student at a training station according to a video collector and initial audio data corresponding to the target student at the training station according to an audio collector; the video processing module 202 is used to perform key frame recognition according to the initial video data to obtain a target behavior sequence corresponding to the target student; the abnormality recognition module 203 is used to perform abnormal behavior recognition according to the target behavior sequence to obtain a target abnormal behavior corresponding to the target student; the behavior analysis module 204 is used to obtain a target abnormal behavior corresponding to the target student according to the target abnormal behavior. A target associated frame is obtained from the target behavior sequence, and a behavior description is performed based on the target associated frame to obtain a behavior description text corresponding to the target student; an audio processing module 205 is used to perform speech segmentation on the initial audio data to obtain the target audio data corresponding to the target student, and perform speech recognition on the target audio data to obtain target text data; a data sorting module 206 is used to perform text level recognition on the target text data to obtain the corresponding target text level, and sort the target text level according to time to obtain a text level sequence corresponding to the target student; a state determination module 207 is used to determine the first state information corresponding to the target student according to the text level sequence; a difference analysis module 208 is used to perform difference analysis on the behavior description text and the target text data to obtain the second state information corresponding to the target student; a result determination module 209 is used to fuse the first state information and the second state information to determine the target state information corresponding to the target student.

[0092] In some embodiments, the student learning status monitoring device 200 based on artificial intelligence can be applied to a terminal device.

[0093] It should be noted that technical personnel in the relevant field can clearly understand that, for the convenience and conciseness of description, the specific working process of the student learning status supervision device 200 based on artificial intelligence described above can refer to the corresponding process in the aforementioned embodiment of the student learning status supervision method based on artificial intelligence, and will not be repeated here.

[0094] See also Figure 3 , Figure 3 A schematic block diagram of the structure of a terminal device provided in an embodiment of the present invention.

[0095] like Figure 3As shown, the terminal device 300 includes a processor 301 and a memory 302 , and the processor 301 and the memory 302 are connected via a bus 303 , such as an I2C (Inter-integrated Circuit) bus.

[0096] Specifically, the processor 301 is used to provide computing and control capabilities to support the operation of the entire terminal device. The processor 301 can be a central processing unit (CPU), and the processor 301 can also be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc.

[0097] Specifically, the memory 302 may be a Flash chip, a read-only memory (ROM) disk, an optical disk, a USB flash drive, or a mobile hard disk.

[0098] Those skilled in the art will understand that Figure 3 The structure shown in the figure is only a block diagram of a partial structure related to the embodiment of the present invention, and does not constitute a limitation on the terminal device to which the embodiment of the present invention is applied. The specific server may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.

[0099] The processor is used to run a computer program stored in the memory, and implement any one of the artificial intelligence-based student learning status supervision methods provided in the embodiments of the present invention when executing the computer program.

[0100] In one embodiment, the processor is used to run a computer program stored in the memory, and implements the following steps when executing the computer program: Collecting initial video data corresponding to a target student at a training station using a video collector and collecting initial audio data corresponding to the target student at the training station using an audio collector; Performing key frame recognition according to the initial video data to obtain a target behavior sequence corresponding to the target student; Perform abnormal behavior recognition according to the target behavior sequence to obtain the target abnormal behavior corresponding to the target student; Obtaining a target-related frame from the target behavior sequence according to the target abnormal behavior, and performing a behavior description according to the target-related frame to obtain a behavior description text corresponding to the target student; Performing speech segmentation on the initial audio data to obtain target audio data corresponding to the target student, and performing speech recognition on the target audio data to obtain target text data; Performing text grade recognition on the target text data to obtain a corresponding target text grade, and sorting the target text grades according to time to obtain a text grade sequence corresponding to the target student; Determining first status information corresponding to the target student according to the text level sequence; Performing difference analysis on the behavior description text and the target text data to obtain second status information corresponding to the target student; The first state information and the second state information are integrated to determine the target state information corresponding to the target student.

[0101] It should be noted that technical personnel in the relevant field can clearly understand that, for the convenience and conciseness of description, the specific working process of the terminal device described above can refer to the corresponding process in the aforementioned artificial intelligence-based student learning status supervision method embodiment, and will not be repeated here.

[0102] An embodiment of the present invention also provides a storage medium for computer-readable storage, wherein the storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the steps of any artificial intelligence-based student learning status supervision method provided in the description of the embodiment of the present invention.

[0103] The storage medium may be an internal storage unit of the terminal device described in the foregoing embodiment, such as a hard disk or memory of the terminal device. The storage medium may also be an external storage device of the terminal device, such as a plug-in hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (SecureDigital, SD) card, a flash card (Flash Card), etc., equipped on the terminal device.

[0104] It will be appreciated by those skilled in the art that all or some of the steps, systems, and functional modules / units in the methods disclosed above may be implemented as software, firmware, hardware, and appropriate combinations thereof. In a hardware embodiment, the division between the functional modules / units mentioned in the above description does not necessarily correspond to the division of physical components; for example, a physical component may have multiple functions, or a function or step may be performed by several physical components in cooperation. Some or all physical components may be implemented as software executed by a processor, such as a central processing unit, a digital signal processor, or a microprocessor, or implemented as hardware, or implemented as an integrated circuit, such as an application-specific integrated circuit. Such software may be distributed on a computer-readable medium, which may include a computer storage medium (or non-transitory medium) and a communication medium (or transient medium). As is known to those skilled in the art, the term computer storage medium includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information (such as computer-readable instructions, data structures, program modules, or other data). Computer storage media include, but are not limited to, RAM, ROM, EEPROM, flash memory or other memory technology, CD-ROM, digital versatile disks (DVD) or other optical disk storage, magnetic cassettes, magnetic tapes, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired information and can be accessed by a computer. In addition, it is well known to those skilled in the art that communication media typically contain computer-readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transport mechanism, and may include any information delivery media.

[0105] It should be understood that the term "and / or" used in the present specification and the appended claims refers to any combination of one or more of the associated listed items and all possible combinations, including these combinations. It should be noted that, in this article, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or system including a series of elements includes not only those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such process, method, article or system. In the absence of further restrictions, an element defined by the sentence "including a..." does not exclude the presence of other identical elements in the process, method, article or system including the element.

[0106] The serial numbers of the embodiments of the present invention are only for description and do not represent the advantages and disadvantages of the embodiments. The above description is only a specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any technician familiar with the technical field can easily think of various equivalent modifications or replacements within the technical scope disclosed by the present invention, and these modifications or replacements should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention shall be based on the protection scope of the claims.

Claims

1. A student learning status supervision method based on artificial intelligence, characterized in that: The method comprises: Collecting initial video data corresponding to a target student at a training station using a video collector and collecting initial audio data corresponding to the target student at the training station using an audio collector; Performing key frame recognition according to the initial video data to obtain a target behavior sequence corresponding to the target student; Perform abnormal behavior recognition according to the target behavior sequence to obtain the target abnormal behavior corresponding to the target student; Obtaining a target-related frame from the target behavior sequence according to the target abnormal behavior, and performing a behavior description according to the target-related frame to obtain a behavior description text corresponding to the target student; Performing speech segmentation on the initial audio data to obtain target audio data corresponding to the target student, and performing speech recognition on the target audio data to obtain target text data; Performing text grade recognition on the target text data to obtain a corresponding target text grade, and sorting the target text grades according to time to obtain a text grade sequence corresponding to the target student; Determining first status information corresponding to the target student according to the text level sequence; Performing difference analysis on the behavior description text and the target text data to obtain second status information corresponding to the target student; The first state information and the second state information are integrated to determine the target state information corresponding to the target student.

2. The method according to claim 1, characterized in that The step of performing abnormal behavior identification according to the target behavior sequence to obtain the target abnormal behavior corresponding to the target student includes: Performing target recognition on each target image in the target behavior sequence to obtain a target area corresponding to the target image; Identify the joint information corresponding to the target student in the target area to obtain the target joint information corresponding to the target image; Obtain relevant joint information corresponding to each sub-joint from the target joint information corresponding to each target image in the target behavior sequence; Performing curve fitting according to the relevant joint information to obtain an initial fitting curve corresponding to the sub-joint; Identifying the change points of the initial fitting curve to obtain the joint change time corresponding to the sub-joint; According to the joint change times corresponding to the plurality of sub-joints, the relevant joint information is integrated to obtain the target behavior type corresponding to the target student; Abnormal behaviors are classified according to the target behavior types to obtain the target abnormal behaviors corresponding to the target students.

3. The method according to claim 2, characterized in that The step of fusing the relevant joint information according to the joint change times corresponding to the plurality of sub-joints to obtain the target behavior type corresponding to the target student includes: Performing an intersection operation on the joint change time to obtain all joint times corresponding to all the sub-joints; Traversing each sub-time in the total joint time, and obtaining the associated joint information that has changed in the sub-time from the related joint information according to the joint change time; Obtaining adjacent joint information corresponding to the previous time in the sub-time from the target joint information; Determine the current joint information corresponding to the sub-time according to the adjacent joint information and the associated joint information; Behavior type classification is performed according to the adjacent joint information and the current joint information to obtain the target behavior type corresponding to the target student at the sub-time.

4. The method according to claim 1, characterized in that: The step of performing behavior description according to the target associated frame to obtain a behavior description text corresponding to the target student includes: Obtaining a training image and an image description text corresponding to the training image, and performing potential feature extraction on the training image according to an image feature extraction layer of an information association model to obtain a first image feature corresponding to the training image; Performing text feature extraction on the image description text according to the text feature extraction layer of the information association model to obtain a first text feature corresponding to the image description text; Performing feature mapping on the first image feature and the first text feature according to the feature association layer of the information association model to obtain an association triplet corresponding to the training image; Performing feature extraction on the target associated frame according to the encoding layer of the text generation model to obtain corresponding second image features; Obtaining a target triplet corresponding to the target associated frame from the associated triplet using the target associated frame and the training image according to the data determination layer of the text generation model; Performing feature extraction on the target triplet according to the self-attention layer of the text generation model to obtain a corresponding second text feature; Obtaining a target relationship between each entity in the target key frame using the second image feature and the second text feature according to the structural attention layer of the text generation model; The behavior description text corresponding to the target student is generated according to the target relationship and the target triplet according to the long short-term memory network layer of the text generation model.

5. The method according to claim 4, characterized in that The data determination layer according to the text generation model obtains the target triplet corresponding to the target associated frame from the associated triplet by using the target associated frame and the training image, including: Extracting keywords from the image description text according to the keyword extraction network of the data determination layer to obtain initial keywords; The data statistics network of the data determination layer performs frequency statistics on the initial keywords according to the image description text to obtain frequency information corresponding to the initial keywords; The data screening network of the data determination layer uses the frequency information to screen the initial keywords to obtain image keywords corresponding to the image description text; The label determination layer of the data determination layer uses the image keyword to label the target associated frame to obtain a first associated keyword corresponding to the target associated frame and first position information corresponding to the first associated keyword; The label determination layer of the data determination layer uses the image keyword to label the training image to obtain a second associated keyword corresponding to the training image and second position information corresponding to the second associated keyword; Performing intersection processing on the first associated keyword and the second associated keyword according to the data merging network of the data determination layer to obtain a target associated keyword; Determine the image association degree between the target-associated frame and the training image according to the target-associated keyword, the first position information and the second position information based on the similarity calculation network of the data determination layer; Obtaining the target triplet corresponding to the target associated frame from the associated triplet using the data image association degree according to the data determination network of the data determination layer; The image correlation degree is obtained according to the following formula: ; in, represents the image association degree between the target associated frame and the i-th training image, num represents the number of words corresponding to the target associated keyword, Indicates the number of words corresponding to the first associated keyword or the second associated keyword, represents the data corresponding to the jth target-related keyword in the i-th training image in the second position information, Represents data corresponding to the j-th target-associated keyword in the target-associated frame in the first position information.

6. The method according to claim 1, characterized in that The step of performing text level recognition on the target text data to obtain a corresponding target text level includes: Determine the practical training courses corresponding to the target students, and obtain the course-related texts corresponding to the practical training courses at different levels; Calculating the similarity between the target text data and the course-related text to obtain a target similarity between the target text data and the course-related text; The target text grade corresponding to the target text data is determined according to the target similarity and the grade information corresponding to the course-related text.

7. The method according to claim 1, characterized in that The step of performing difference analysis on the behavior description text and the target text data to obtain the second status information corresponding to the target student includes: Performing keyword recognition on the behavior description text to obtain a first keyword and a first relationship corresponding to the first keyword; Performing keyword recognition on the target text data to obtain a second keyword and a second relationship corresponding to the second keyword; Calculating a first similarity corresponding to the first keyword and the second keyword, and obtaining a second similarity between the behavior description text and the target text data by combining the first relationship and the second relationship under the first similarity; The behavior description text and the target text data are subjected to difference analysis according to the second similarity to obtain the second status information corresponding to the target student.

8. A student learning status monitoring device based on artificial intelligence, characterized in that: include: A data acquisition module, used to acquire initial video data corresponding to a target student at a training station using a video collector and to acquire initial audio data corresponding to the target student at the training station using an audio collector; A video processing module, used for performing key frame recognition according to the initial video data to obtain a target behavior sequence corresponding to the target student; An abnormality identification module, used for performing abnormal behavior identification according to the target behavior sequence to obtain the target abnormal behavior corresponding to the target student; A behavior analysis module, used for obtaining a target-related frame from the target behavior sequence according to the target abnormal behavior, and performing a behavior description according to the target-related frame to obtain a behavior description text corresponding to the target student; An audio processing module, used for performing speech segmentation on the initial audio data to obtain target audio data corresponding to the target student, and performing speech recognition on the target audio data to obtain target text data; A data sorting module, used to perform text grade recognition on the target text data to obtain the corresponding target text grade, and sort the target text grade according to time to obtain the text grade sequence corresponding to the target student; A state determination module, used for determining first state information corresponding to the target student according to the text level sequence; A difference analysis module, used for performing a difference analysis on the behavior description text and the target text data to obtain the second state information corresponding to the target student; A result determination module is used to fuse the first state information and the second state information to determine the target state information corresponding to the target student.

9. A terminal device, characterized in that: The terminal device includes a processor and a memory; The memory is used to store computer programs; The processor is used to execute the computer program and implement the student learning status supervision method based on artificial intelligence as described in any one of claims 1 to 7 when executing the computer program.

10. A computer storage medium for computer storage, characterized in that: The computer storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the steps of the student learning status supervision method based on artificial intelligence as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Real-time classroom student state analysis and indication reminding system and method based on behavior and voice intelligent recognition

    CN110991381A

  • Learner learning state acquisition method based on multi-modal emotion feature fusion

    CN116244474A

  • Video theme-based knowledge graph generation method and device, equipment and medium

    CN117633241A

  • Video monitoring method and system based on multi-scene recognition and voice interaction

    CN117749995A

  • Application defect detection method, device and equipment and readable storage medium

    CN119377083A

Cited By

  • College laboratory patrol system and method based on edge calculation

    CN121392725A