Artificial Intelligence-Based Student Learning Status Supervision Method, Device, Terminal Device, and Storage Medium
Through video and audio data acquisition, students' learning status is solved, and the problem of the inability to fully capture students' learning status in the training courses in the existing technology is solved, and timely discovery and accurate evaluation of students' learning status is achieved, and learning efficiency and quality are improved.
Patent Information
- Application Number
- CN202510414490.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-03
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-04-03
AI Technical Summary
The existing technology cannot comprehensively and accurately capture the learning status of students in the training course, resulting in inefficiency in learning.
Through video and audio data acquisition, keyframes and abnormal behaviors are identified, combined with speech recognition and text level analysis, and a variety of state information is integrated to determine students' learning status.
It realizes timely discovery and accurate assessment of students' learning status during practical training, and improves learning efficiency and learning quality.
Smart Images

Figure CN119918016B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of artificial intelligence, and particularly to a method, device, terminal device and storage medium for supervising the learning state of students based on artificial intelligence. Background Art
[0002] With the continuous development of artificial intelligence, its influence has increasingly penetrated into the teaching field. When artificial intelligence is introduced into the teaching scenario, relying on advanced artificial intelligence technology, it can supervise the learning state of students from multiple dimensions, including the class videos and after-class assignments of students. In this way, it is possible to comprehensively understand the students' mastery of knowledge and learning progress. However, in actual applications, the existing technology has obvious deficiencies when supervising students with practical training courses. Due to the limitations of current supervision means, it is difficult to achieve precise supervision and comprehensively and accurately capture the rich and diverse learning states of students in practical training courses. As a result, there are large errors in the supervision of the states of students in practical training courses, and it is difficult to truly reflect the actual learning effects of students in the practical training session. Summary of the Invention
[0003] The main purpose of the embodiments of the present invention is to provide a method, device, terminal device and storage medium for supervising the learning state of students based on artificial intelligence, aiming to solve the problem in the related technology that the learning states of students in practical training courses cannot be comprehensively and accurately captured, thus affecting the learning efficiency of students.
[0004] In a first aspect, the embodiments of the present invention provide a method for supervising the learning state of students based on artificial intelligence, including:
[0005] Collecting initial video data corresponding to a target student at a practical training station according to a video collector and collecting initial audio data corresponding to the target student at the practical training station according to an audio collector;
[0006] Performing key frame recognition on the initial video data to obtain a target behavior sequence corresponding to the target student;
[0007] Performing abnormal behavior recognition on the target behavior sequence to obtain a target abnormal behavior corresponding to the target student;
[0008] Obtaining target associated frames from the target behavior sequence according to the target abnormal behavior, and performing behavior description according to the target associated frames to obtain a behavior description text corresponding to the target student;
[0009] Performing voice segmentation on the initial audio data to obtain target audio data corresponding to the target student, and performing speech recognition on the target audio data to obtain target text data;
[0010] Perform text level recognition on the target text data to obtain the corresponding target text level, and sort the target text level according to time to obtain the text level sequence corresponding to the target student;
[0011] Determine the first status information corresponding to the target student according to the text level sequence;
[0012] Perform difference analysis based on the behavior description text and the target text data to obtain the second status information corresponding to the target student;
[0013] Fuse the first status information and the second status information to determine the target status information corresponding to the target student.
[0014] In a second aspect, an embodiment of the present invention provides an artificial intelligence-based student learning status supervision device, including:
[0015] A data acquisition module, configured to acquire initial video data corresponding to a target student at a training station according to a video collector and acquire initial audio data corresponding to the target student at the training station according to an audio collector;
[0016] A video processing module, configured to perform key frame recognition on the initial video data to obtain a target behavior sequence corresponding to the target student;
[0017] An anomaly recognition module, configured to perform anomaly behavior recognition on the target behavior sequence to obtain a target anomaly behavior corresponding to the target student;
[0018] A behavior analysis module, configured to obtain target associated frames from the target behavior sequence according to the target anomaly behavior, and perform behavior description based on the target associated frames to obtain a behavior description text corresponding to the target student;
[0019] An audio processing module, configured to perform speech segmentation on the initial audio data to obtain target audio data corresponding to the target student, and perform speech recognition on the target audio data to obtain target text data;
[0020] A data sorting module, configured to perform text level recognition on the target text data to obtain the corresponding target text level, and sort the target text level according to time to obtain the text level sequence corresponding to the target student;
[0021] A status determination module, configured to determine the first status information corresponding to the target student according to the text level sequence;
[0022] A difference analysis module, configured to perform difference analysis based on the behavior description text and the target text data to obtain the second status information corresponding to the target student;
[0023] A result determination module, configured to fuse the first status information and the second status information to determine the target status information corresponding to the target student.
[0024] In a third aspect, an embodiment of the present invention further provides a terminal device, which includes a processor, a memory, a computer program stored on the memory and executable by the processor, and a data bus for implementing connection communication between the processor and the memory. When the computer program is executed by the processor, the steps of any one of the artificial intelligence-based student learning status supervision methods provided in the specification of the present invention are implemented.
[0025] In a fourth aspect, an embodiment of the present invention further provides a storage medium for computer-readable storage, characterized in that the storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the steps of any one of the artificial intelligence-based student learning status supervision methods provided in the specification of the present invention.
[0026] An embodiment of the present invention provides a method, device, terminal device, and storage medium for supervising the learning state of students based on artificial intelligence. The method includes: acquiring initial video data corresponding to a target student at a training station according to a video collector and initial audio data corresponding to the target student at the training station according to an audio collector. Then, performing key frame recognition on the initial video data to obtain a target behavior sequence corresponding to the target student, and performing abnormal behavior recognition on the target behavior sequence to obtain a target abnormal behavior corresponding to the target student. Thus, abnormal situations such as improper operations and violations of the target student can be discovered in a timely manner. Then, obtaining target associated frames from the target behavior sequence according to the target abnormal behavior, and performing behavior description based on the target associated frames to obtain a behavior description text corresponding to the target student, thereby providing support for subsequent matching of questions asked by the target student to the teacher. Next, performing voice segmentation on the initial audio data to obtain target audio data corresponding to the target student, performing voice recognition on the target audio data to obtain target text data, performing text level recognition on the target text data to obtain a corresponding target text level, and sorting the target text levels according to time to obtain a text level sequence corresponding to the target student, thereby determining first state information corresponding to the target student according to the text level sequence; and performing difference analysis on the behavior description text and the target text data to obtain second state information corresponding to the target student. Finally, fusing the first state information and the second state information to determine target state information corresponding to the target student, so that the learning state of the target student during the training process can be discovered in a timely manner, and thus provide good support for formulating a learning plan or setting a teaching method for the target student in the future to improve the learning efficiency and learning quality of the target student. This method also solves the problem in the related art that the learning state of students in training courses cannot be comprehensively and accurately captured, thus affecting the learning efficiency of students. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for the description of the embodiments will be briefly introduced below. Obviously, the drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0028] Figure 1 It is a flowchart showing a method for supervising the learning state of students based on artificial intelligence provided by an embodiment of the present invention;
[0029] Figure 2 It is a schematic block diagram of the module structure of a device for supervising the learning state of students based on artificial intelligence provided by an embodiment of the present invention;
[0030] Figure 3 It is a schematic block diagram of the structure of a terminal device provided by an embodiment of the present invention. Detailed implementation manners
[0031] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0032] The flowchart shown in the accompanying drawings is only an example illustration, and does not necessarily include all the contents and operations / steps, nor does it necessarily execute in the described order. For example, some operations / steps can also be decomposed, combined, or partially merged, so the actual execution order may be changed according to the actual situation.
[0033] It should be understood that the terms used in the specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. As used in the specification of the present invention and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an", and "the" are intended to include the plural forms.
[0034] The embodiments of the present invention provide a method, device, terminal device, and storage medium for supervising the learning state of students based on artificial intelligence. Among them, the method for supervising the learning state of students based on artificial intelligence can be applied to a terminal device, and the terminal device can be an electronic device such as a tablet computer, a notebook computer, a desktop computer, a personal digital assistant, and a wearable device. The terminal device can be a server or a server cluster.
[0035] Next, some embodiments of the present invention will be described in detail in conjunction with the accompanying drawings. Without conflict, the following embodiments and the features in the embodiments can be combined with each other.
[0036] Please refer to Figure 1 , Figure 1 , which is a schematic flowchart of a method for supervising the learning state of students based on artificial intelligence provided by an embodiment of the present invention.
[0037] As Figure 1 shown, the method for supervising the learning state of students based on artificial intelligence includes steps S101 to S109.
[0038] Step S101: Acquire the initial video data corresponding to the target student at the training station according to the video collector and acquire the initial audio data corresponding to the target student at the training station according to the audio collector.
[0039] Exemplarily, to obtain the target status information of the target student in a timely manner, a video collector and an audio collector are installed on the training workstation corresponding to the target student. Among them, a suitable video collection device is selected according to the training environment and requirements. For example, in a workstation with sufficient light and an open space, an ordinary high-definition network camera can meet the requirements; while in an environment with relatively dim light or extremely high requirements for image quality, a low-light, high-resolution camera is more suitable. And a suitable audio device is selected according to the noise level and sound pickup range of the training workstation. If the workstation environment is quiet, an ordinary directional microphone is sufficient; if the environment is noisy, a microphone with noise reduction function needs to be selected.
[0040] Exemplarily, a unique identity identifier, such as a name tag, a bracelet, etc., is assigned to each target student. Information such as the name and student number of the target student can be included on these identifiers. The video collector and the audio collector can combine image recognition or other technologies to accurately identify the target student, so as to associate the identity information of each target student with the corresponding video collector and audio collector, ensure that the collected data can accurately correspond to each target student, and then obtain the corresponding initial video data and initial audio data of the target student on the training workstation according to the identity identifier of the target student.
[0041] Step S102: Perform key frame recognition according to the initial video data to obtain the target behavior sequence corresponding to the target student.
[0042] Exemplarily, first, unnecessary parts in the initial video data are cropped, such as blank edges, backgrounds irrelevant to the target student, etc. Then, for the cropped initial video data, object detection algorithms such as YOLO (You Only Look Once), Faster R - CNN, etc. are used to detect the position of the target student in the video frame, such as the human body contour and bounding box. Then, after the target student is detected, object tracking algorithms such as KCF (Kernelized Correlation Filters), CSRT (Discriminative Correlation Filter with Channel and Spatial Reliability), etc. are used to continuously track the target student. Ensure that the position of the target student can be accurately found in each frame of the video, providing an accurate target area for subsequent key frame recognition.
[0043] Exemplarily, target features such as color features (e.g., color histograms), texture features (e.g., gray-level co-occurrence matrices), and shape features (e.g., contour moments) are extracted from the target region. These features can describe the visual content of video frames, and thus methods such as Euclidean distance and cosine similarity are used to calculate the feature similarity between adjacent video frames. Furthermore, when the similarity is lower than a certain set threshold, it is considered that a significant change has occurred between these two frames, which may contain key information, and this frame is marked as a candidate key frame.
[0044] Exemplarily, the identified candidate key frames are screened to remove some duplicate or meaningless key frames. For example, if the actions of a student in several consecutive frames only slightly jitter and there is no substantial behavioral change, these frames can be merged or deleted. Then, the screened key frames are sorted in the chronological order of the video to ensure that the order of the key frames is consistent with the order in which the student's behaviors occur, thereby obtaining the target behavior sequence corresponding to the target student.
[0045] Step S103, perform abnormal behavior recognition based on the target behavior sequence to obtain the target abnormal behavior corresponding to the target student.
[0046] Exemplarily, based on the operation process such as the correct operation sequence and action requirements of each training step, behaviors that do not conform to this process are defined as abnormal. For example, in an assembly training, if a student skips a certain key part installation step, it belongs to abnormal behavior, or abnormal behavior rules are determined based on equipment usage specifications. For example, if a student starts the equipment without necessary equipment preheating or makes illegal adjustment operations during equipment operation, etc. Then, the defined abnormal behavior rules are sorted out to form a structured rule library.
[0047] Exemplarily, key features that can reflect the behavior characteristics are extracted from the target behavior sequence, including the duration of the behavior, the frequency of actions, the amplitude of actions, the order in which behaviors occur, etc. Then, in combination with the abnormal behavior rule library, features related to each rule are determined, and the extracted behavior features are quantified so that they can be represented in numerical or other comparable forms. For example, the duration of the behavior is converted into a specific time value, and the amplitude of the action is converted into an angle or length value. Then, the quantified features are normalized to eliminate the dimensional differences between different features and ensure the comparability of the features in subsequent matching and analysis.
[0048] Exemplarily, based on the machine learning classification algorithm, abnormal judgment is performed using the abnormal behavior rules and the quantified behavior features, so as to judge which parts of the target behavior sequence belong to abnormal behaviors, thereby obtaining the target abnormal behavior corresponding to the target student.
[0049] In some embodiments, obtaining the target abnormal behavior corresponding to the target student by identifying abnormal behaviors according to the target behavior sequence includes: performing target recognition on each target image in the target behavior sequence to obtain the target area corresponding to the target image; identifying the joint information corresponding to the target student in the target area to obtain the target joint information corresponding to the target image; obtaining the relevant joint information corresponding to each sub-joint from the target joint information corresponding to each target image in the target behavior sequence; performing curve fitting according to the relevant joint information to obtain the initial fitting curve corresponding to the sub-joint; identifying the change points of the initial fitting curve to obtain the joint change time corresponding to the sub-joint; fusing the relevant joint information according to the joint change times corresponding to multiple sub-joints to obtain the target behavior type corresponding to the target student; and classifying the abnormal behaviors according to the target behavior type to obtain the target abnormal behavior corresponding to the target student.
[0050] Exemplarily, a target detection algorithm such as Faster R-CNN, YOLO, etc. is used to perform target recognition on each target image in the target behavior sequence to obtain the target area corresponding to the target student in each target image. Then, a human pose estimation algorithm such as OpenPose, AlphaPose, etc. is used to perform joint recognition on the content corresponding to the target area in the target image to obtain the position information of each joint of the target student in the target area, including the coordinates of the joints, etc. These information are the target joint information corresponding to the target image.
[0051] Exemplarily, all the identified joint information is classified to label different sub-joints, such as the left shoulder joint, the right knee joint, etc. According to the category of the sub-joints, the joint information of the sub-joint corresponding to each target image in the target behavior sequence is associated and sorted to form the relevant joint information corresponding to each sub-joint. Then, methods such as polynomial fitting and spline curve fitting are used to find a curve that can better approximate these data points for the relevant joint information of the sub-joint (such as the change of the joint coordinates over time), so as to obtain the initial fitting curve corresponding to each sub-joint.
[0052] Exemplarily, a statistical-based method or a machine learning-based method is used to find the points where the curve changes significantly in the initial fitting curve, and the time corresponding to the identified change points is used as the joint change time of the sub-joint.
[0053] Exemplarily, different target behavior types and their corresponding joint change patterns and joint information features are predefined. For example, the change patterns and information features of each sub-joint (such as the shoulder joint, elbow joint, etc.) corresponding to the "raising hand" behavior are defined. Then, according to the joint change time, the corresponding joint positions at that time are obtained. Furthermore, multiple joint change times are clustered to obtain time clusters corresponding to multiple joint change times. Thus, for each sub-change time in each sub-cluster under the time cluster, the first joint position is obtained, and the second joint position of other joints in the human body at this sub-change time is obtained. Then, according to the first joint position and the second joint position, they are connected in the order of human joints to obtain the joint pattern corresponding to this sub-cluster. Furthermore, it is matched and fused with the predefined behavior pattern. By analyzing information such as the change order and amplitude of each sub-joint, the target behavior type corresponding to the target student is judged.
[0054] Exemplarily, according to the requirements and safety specifications of the training scenario, abnormal behavior classification rules corresponding to different target behavior types are formulated. For example, in a specific training operation, "raising hand too fast" may be defined as an abnormal behavior. Then, the identified target behavior type is compared with the abnormal behavior classification rules to judge whether this behavior belongs to an abnormal behavior. If it belongs to an abnormal behavior, it is marked as the target abnormal behavior, and relevant information is recorded.
[0055] Specifically, through target recognition and joint information recognition of the target image, the target student can be accurately located and their detailed joint movement information can be obtained. Curve fitting and change point recognition can extract key change information from the time series data of joint movement, which helps to discover some imperceptible abnormal behaviors, such as minor movement deviations or irregular behavior start and end times. Furthermore, by analyzing from the perspective of multiple sub-joints and fusing relevant joint information, the behavior pattern of the target student can be comprehensively described. Different sub-joints play different roles in the behavior. Considering the information of multiple sub-joints comprehensively can help to understand the student's behavior actions more comprehensively and avoid incomplete behavior analysis caused by only focusing on some joints.
[0056] In some embodiments, fusing the relevant joint information according to the joint change times corresponding to the multiple sub-joints to obtain the target behavior type corresponding to the target student includes: performing an intersection operation on the joint change times to obtain all joint times corresponding to all the sub-joints; traversing each sub-time in all the joint times, and obtaining the associated joint information that changes at the sub-time from the relevant joint information according to the joint change times; obtaining the adjacent joint information corresponding to the previous time in the sub-time from the target joint information; determining the current joint information corresponding to the sub-time according to the adjacent joint information and the associated joint information; and performing behavior type classification according to the adjacent joint information and the current joint information to obtain the target behavior type corresponding to the target student at the sub-time.
[0057] Exemplarily, collect and organize the joint change time data corresponding to multiple sub-joints, and then use the intersection operation method of sets to find the common time points of the joint change times of all sub-joints. These common time points constitute all the joint times corresponding to all the sub-joints. This process is like finding the overlapping time scales in multiple time axes.
[0058] Exemplarily, sequentially access each sub-time in all the joint times. For each sub-time, compare it with the joint change times of each sub-joint. If the joint change time of a certain sub-joint contains this sub-time, it means that the sub-joint has changed at this sub-time point. Extract the information of these changed sub-joints from the relevant joint information, and this information is the associated joint information that changes at the sub-time.
[0059] Exemplarily, sort the target joint information in chronological order to ensure that the previous time point corresponding to each sub-time can be accurately found. For the currently traversed sub-time, find its corresponding previous time point, and then extract the joint information of all sub-joints corresponding to that previous time point from the target joint information, and this information is the adjacent joint information.
[0060] Exemplarily, combine the adjacent joint information and the associated joint information to determine the current joint information corresponding to the sub-time. For the changed sub-joints, update their joint states according to the associated joint information; for the unchanged sub-joints, keep their states in the adjacent joint information. In this way, the complete current joint information at the sub-time is obtained.
[0061] Exemplarily, different types of behaviors and their corresponding joint movement patterns are predefined. For example, behavior types such as "screwing a screw", "welding", "hammering", etc. are defined, as well as the characteristic manifestations of these behavior types on adjacent joint information and current joint information, such as the change range of joint angles, the moving direction of joint positions, etc. Then, the adjacent joint information and the current joint information are compared and matched with the predefined behavior patterns. By analyzing the change characteristics of the joint information, the behavior type of the target student at this sub-time is judged, so as to determine the corresponding target behavior type.
[0062] Specifically, by performing intersection operations and comprehensive analysis on the joint change times of multiple sub-joints, the change moments of the target student's behavior can be captured more comprehensively and accurately. This makes the determination of joint information at each sub-time more accurate, thus providing a more reliable basis for behavior type classification. In addition, classifying behavior types by combining adjacent joint information and current joint information takes into account the dynamic change process of the behavior. Compared with classifying based only on the joint information at a single moment, this method can better reflect the continuity and change trend of the behavior, reduce the possibility of misjudgment, and improve the accuracy of classification.
[0063] Step S104: Obtain target associated frames from the target behavior sequence according to the target abnormal behavior, and obtain the behavior description text corresponding to the target student according to the target associated frames.
[0064] Exemplarily, the key frames determined as the target abnormal behavior in the target behavior sequence are determined as the target associated frames corresponding to the target abnormal behavior.
[0065] Exemplarily, key features related to the behavior are extracted from the target associated frames. For example, the posture of the person (standing, bending, raising the hand, etc.), the amplitude of the movement (large swing, small movement), the relative position relationship of each part of the body (the position of the hand and the head, the distance between the feet, etc.), and the environmental information (whether operating a device, whether there are other people around, etc.). Then, according to the extracted features, the actions in the target associated frames are classified to determine the action type corresponding to the target associated frames, and then, combined with the frames arranged in chronological order, the sequence and logical relationship between the actions are analyzed. Judge whether the action is independent or part of a series of coherent actions. For example, after "picking up a hammer", it may be followed by "striking an object", which constitutes a behavior sequence with a logical order.
[0066] Exemplarily, the recognized actions and the training types corresponding to the training positions are integrated to combine individual actions into a behavior description text with complete meaning, such as "The student is at the training station, first picks up a hammer, and then strikes it on a screw".
[0067] Exemplarily, the behavior description text is used to characterize the content corresponding to the abnormal behavior of the target student determined according to the target association frame when using text description, so as to provide support for judging whether the question description of the target student to the teacher is accurate when the teaching purpose cannot be achieved due to the abnormal behavior, and further judging the learning state of the target student.
[0068] In some embodiments, obtaining the behavior description text corresponding to the target student by performing behavior description according to the target association frame includes: obtaining a training image and the image description text corresponding to the training image, and performing latent feature extraction on the training image according to the image feature extraction layer of the information association model to obtain the first image feature corresponding to the training image; performing text feature extraction on the image description text according to the text feature extraction layer of the information association model to obtain the first text feature corresponding to the image description text; performing feature mapping on the first image feature and the first text feature according to the feature association layer of the information association model to obtain the associated triple corresponding to the training image; performing feature extraction on the target association frame according to the encoding layer of the text generation model to obtain the corresponding second image feature; obtaining the target triple corresponding to the target association frame from the associated triple according to the data determination layer of the text generation model by using the target association frame and the training image; performing feature extraction on the target triple according to the self-attention layer of the text generation model to obtain the corresponding second text feature; obtaining the target relationship between each entity in the target key frame by using the second image feature and the second text feature according to the structure attention layer of the text generation model; generating the behavior description text corresponding to the target student according to the target relationship and the target triple by using the long short-term memory network layer of the text generation model.
[0069] Exemplarily, a large number of training images and the corresponding image description texts are collected. These training images should contain rich and diverse scenes and behaviors, and the image description texts should accurately and detailedly describe the content in the images.
[0070] Exemplarily, the image feature extraction layer of the information association model is used to process the training image. This layer analyzes visual information such as the color, texture, shape, and object contour of the image, and converts this information into a latent feature representation, thereby obtaining the first image feature corresponding to the training image. The image description text is analyzed by the text feature extraction layer of the information association model. This layer considers information such as vocabulary, grammar, and semantics in the text, and converts the text into the first text feature in vector form.
[0071] Exemplarily, the feature association layer of the information association model maps the first image feature and the first text feature. It searches for the corresponding relationship between the image feature and the text feature, determines which image features are associated with which text features, and then generates an associated triple corresponding to the training image based on the result of the feature mapping. The associated triple usually includes the image feature, the text feature, and the association relationship between them, which helps to accurately establish the connection between the image and the text in the subsequent steps.
[0072] Exemplarily, the encoding layer of the text generation model extracts features from the target associated frame. This layer extracts the visual features of the target associated frame like it processes the training image, and obtains the corresponding second image feature.
[0073] Exemplarily, the data determination layer of the text generation model searches and matches in the associated triple using the target associated frame and the training image. It finds the associated triple most relevant to the target associated frame and determines it as the target triple corresponding to the target associated frame. The self-attention layer of the text generation model processes the target triple, extracts the text-related features therein, and obtains the corresponding second text feature. The self-attention mechanism allows the model to focus on the important relationships between different parts of the target triple. The structure attention layer of the text generation model analyzes the target relationships between each entity in the target key frame using the second image feature and the second text feature. This includes spatial relationships, action relationships, semantic relationships, etc. between entities.
[0074] Exemplarily, the long short-term memory network layer of the text generation model generates the behavior description text corresponding to the target student according to the target relationship and the target triple. The long short-term memory network can process sequential information, combine various features and relationships obtained in the previous steps, and generate smooth and accurate natural language text to describe the behavior of the target student in the target associated frame.
[0075] Specifically, by extracting the features of the image and the text and associating them, the model can comprehensively utilize visual information and semantic information. This makes the generated behavior description text more accurately reflect the actual behavior in the target associated frame, reducing the description deviation caused by inaccurate single-feature analysis. The analysis of the relationships between entities by the structure attention layer can capture the details and logic in the behavior. In addition, through the process of determining the target triple, the model can find similar situations to the target associated frame in the training data and draw on the existing association information to generate the description. This way enhances the adaptability of the model to different inputs and improves the generalization ability.
[0076] In some embodiments, the data determination layer according to the text generation model obtains the target triple corresponding to the target associated frame from the associated triples by using the target associated frame and the training image, including: extracting initial keywords from the image description text by using the keyword extraction network of the data determination layer; obtaining frequency information corresponding to the initial keywords by performing frequency statistics on the initial keywords according to the image description text by using the data statistics network of the data determination layer; screening the initial keywords by using the frequency information by using the data screening network of the data determination layer to obtain image keywords corresponding to the image description text; performing tagging processing on the target associated frame by using the image keywords by using the tag determination layer of the data determination layer to obtain a first associated keyword corresponding to the target associated frame and first position information corresponding to the first associated keyword; performing tagging processing on the training image by using the image keywords by using the tag determination layer of the data determination layer to obtain a second associated keyword corresponding to the training image and second position information corresponding to the second associated keyword; performing an intersection process on the first associated keyword and the second associated keyword by using the data merging network of the data determination layer to obtain a target associated keyword; determining an image association degree between the target associated frame and the training image according to the target associated keyword, the first position information, and the second position information by using the similarity calculation network of the data determination layer; obtaining the target triple corresponding to the target associated frame from the associated triples by using the data determination network of the data determination layer according to the data image association degree; wherein, the image association degree is obtained according to the following formula:
[0077] ;
[0078] wherein, represents the image association degree between the target associated frame and the i-th training image, num represents the number of words corresponding to the target associated keyword, represents the number of words corresponding to the first associated keyword or the second associated keyword, represents the data corresponding to the j-th target associated keyword in the i-th training image in the second position information, represents the data corresponding to the j-th target associated keyword in the target associated frame in the first position information.
[0079] Exemplarily, the keyword extraction network of the data determination layer processes the image description text to obtain initial keywords, and then the data statistics network performs frequency statistics on the initial keywords according to the image description text. It traverses the entire text and calculates the number of times each initial keyword appears, thereby obtaining the frequency information corresponding to each initial keyword.
[0080] Exemplarily, the data screening network uses frequency information to screen the initial keywords. Usually, a frequency threshold is set, and only the initial keywords with a frequency higher than this threshold will be retained, and these retained keywords become the image keywords corresponding to the image description text. For example, if the threshold is set to 2, then the keyword "student" with a frequency of 3 will be retained, while other initial keywords that only appear once may be screened out.
[0081] Exemplarily, the label determination layer uses the image keywords to label the target associated frame. It will search for content related to the image keywords in the target associated frame and mark the positions where these contents are located, so as to obtain the first associated keywords corresponding to the target associated frame and the first position information corresponding to the first associated keywords. For example, if the image keywords are "hammer" and "nail", the label determination layer will find the positions of the hammer and the nail in the target associated frame and record them. Similarly, the label determination layer uses the image keywords to label the training image, obtaining the second associated keywords corresponding to the training image and the second position information corresponding to the second associated keywords.
[0082] Exemplarily, the data merging network performs an intersection process on the first associated keywords and the second associated keywords. It will find the keywords that appear in both the target associated frame and the training image, and these keywords are the target associated keywords. For example, if the first associated keywords of the target associated frame are "hammer" and "nail", and the second associated keywords of the training image are "hammer" and "chair", then the target associated keyword is "hammer".
[0083] Exemplarily, the similarity calculation network calculates the image correlation degree between the target associated frame and the training image according to the target associated keywords, the first position information, and the second position information, using a given formula. This formula comprehensively considers factors such as the number of target associated keywords and the position information, and obtains a quantified correlation degree value. Then, the image correlation degree is obtained according to the following formula:
[0084] ;
[0085] where represents the image correlation degree between the target associated frame and the i-th training image, num represents the number of words corresponding to the target associated keywords, represents the number of words corresponding to the first associated keywords or the second associated keywords, represents the data corresponding to the j-th target associated keyword in the i-th training image in the second position information, represents the data corresponding to the j-th target associated keyword in the target associated frame in the first position information.
[0086] Exemplarily, the data determination network obtains the target triple corresponding to the target associated frame from the associated triples by using the calculated image correlation degree. Usually, the associated triple corresponding to the training image with the highest image correlation degree with the target associated frame is selected as the target triple.
[0087] Specifically, through steps such as keyword extraction, frequency statistics, and screening, the key information in the image description text can be accurately found. The tagging process further corresponds these keywords to the specific positions in the target associated frame and the training image, making the subsequent matching process more accurate. In addition, the image correlation degree is calculated using the above formula to quantify the similarity between the target associated frame and the training image. This enables an objective criterion to be available when selecting the target triple, reduces the influence of subjective factors, and improves the accuracy of matching. Furthermore, using keywords for matching instead of directly comparing all features of the images enables the model to find similar semantic information in different images. Even if the visual features of the images are different, as long as they contain the same key semantic information, they can be considered relevant. This enhances the adaptability of the model to different images and improves the generalization ability. Thus, by extracting and utilizing the keywords in the image description text, the text information is fully utilized to assist image matching. This enables the model to mine useful information from a large amount of training data, improves the utilization efficiency of the data, and further enhances the generalization ability.
[0088] Step S105: Perform speech segmentation on the initial audio data to obtain the target audio data corresponding to the target student, and perform speech recognition on the target audio data to obtain the target text data.
[0089] Exemplarily, speaker recognition technology is used to distinguish different speakers in the initial audio data. First, a speaker recognition model is trained, which can learn the speech features of different speakers. Then the initial audio is input into the model, and the model will judge the speaker identity of each speech segment according to the speech features, so as to segment out the speech segments of the target student, and further obtain the target audio data corresponding to the target student.
[0090] Exemplarily, open-source speech recognition models such as DeepSpeech, Wav2Vec, etc. are used to collect the corresponding speech data and the text content corresponding to the speech data to perform speech training on the speech recognition model, so as to obtain the corresponding target speech recognition model, and then perform speech recognition on the target audio data according to the target speech recognition model to obtain the target text data corresponding to the target audio data.
[0091] Step S106: Perform text level recognition on the target text data to obtain the corresponding target text level, and sort the target text level according to time to obtain the text level sequence corresponding to the target student.
[0092] Exemplarily, the purpose of text level recognition is to identify the difficulty level of the questions asked by the target student to the teacher during the communication with the teacher, or to identify the difficulty level of the training content involved in the communication between the target student and the teacher.
[0093] Exemplarily, determine the training course, and based on the training course, determine the knowledge points involved at different difficulty levels and the knowledge keywords corresponding to the knowledge points, so as to perform keyword recognition on the target text data to obtain text keywords, and then calculate the similarity between the text keywords and the knowledge keywords to obtain the target text level corresponding to the target student.
[0094] Exemplarily, obtain the time corresponding to the target text data in the initial audio data, and then sort the target text level in chronological order to obtain the text level sequence corresponding to the target student.
[0095] In some embodiments, the performing text level recognition on the target text data to obtain the corresponding target text level includes: determining the training course corresponding to the target student, and obtaining the course-related texts corresponding to the training course at different levels; calculating the similarity between the target text data and the course-related texts to obtain the target similarity between the target text data and the course-related texts; and determining the target text level corresponding to the target text data according to the target similarity and the level information corresponding to the course-related texts.
[0096] Exemplarily, obtain the training arrangement table through the course selection record of the target student to determine the training course participated by the target student, and then collect the course-related texts corresponding to the course at different levels. The course-related texts can be sourced from course textbooks, teaching syllabi, and teachers' lecture notes. For example, for a programming training course, the course-related texts at the primary level may be introductions to basic programming syntax, at the intermediate level may be the implementation ideas of simple projects, and at the advanced level may be the analysis of complex algorithms, etc.
[0097] Exemplarily, use text similarity calculation methods such as semantic-based methods (such as using pre-trained language models) to calculate the similarity between the target text data and the collected course-related texts one by one to obtain the target similarity between the target text data and the course-related texts.
[0098] Exemplarily, for each piece of course-related text, its corresponding level information is associated. This level information can be beginner, intermediate, or advanced. For example, the introduction of basic programming syntax corresponds to the beginner level, and the analysis of complex algorithms corresponds to the advanced level. Then, based on the calculated target similarity and the level information corresponding to the course-related text, the target text level corresponding to the target text data is determined. For example, the level corresponding to the course-related text with the highest similarity to the target text data is selected as the target text level. For example, if the target text data has the highest similarity to the course-related text at the intermediate level, then the target text level corresponding to the target text data is intermediate.
[0099] Step S107: Determine the first status information corresponding to the target student according to the text level sequence.
[0100] Exemplarily, the text level sequence is an ordered set composed of the levels corresponding to a series of texts related to the target student. These levels can be divided according to criteria such as difficulty and knowledge mastery level, for example, beginner, intermediate, advanced. The first status information is information describing the status of the target student in a specific situation, which may include learning status (such as learning progress, learning stagnation, learning difficulties), ability level status (such as ability improvement, ability stability, ability decline), knowledge mastery status (such as comprehensive knowledge mastery, partial mastery, poor mastery), etc.
[0101] Exemplarily, observe whether the levels in the text level sequence show a gradually increasing trend. If so, it indicates that the student is making continuous progress and may perform well in terms of knowledge mastery, ability improvement, etc. For example, the student's text level gradually increases from beginner to intermediate and then to advanced. When the levels show a gradually decreasing trend, it indicates that the first status information of the student may be deteriorating, and the student may encounter learning difficulties or knowledge forgetting problems. For example, a student who was originally at the intermediate level has a subsequent text level drop to the beginner level. If the levels fluctuate little within a certain range and there is no obvious upward or downward trend, it means that the first status information of the student is relatively stable, and the knowledge mastery and ability level may be in a relatively stable stage.
[0102] Exemplarily, if the change range between adjacent levels in the text level sequence is large, it indicates that the student's state is unstable and may be affected by external factors (such as changes in the learning environment, unexpected events) or internal factors (such as fluctuations in learning attitude, inappropriate learning methods). Small fluctuations may be within the normal error range or small adjustments during the learning process, having relatively little impact on the overall state of the student, corresponding to positive first state information such as continuous learning progress, improving ability, and better knowledge mastery. For example, if the text level sequence shows an obvious upward trend, it can be judged that the student is in a learning state of rapid progress, with their ability continuously improving and their knowledge mastery becoming more comprehensive. Corresponding to negative first state information such as learning difficulties, declining ability, and deteriorating knowledge mastery. For instance, when the text level continuously drops, it can be considered that the student has encountered obstacles in the learning process, with their ability level declining and their knowledge mastery not being satisfactory. Corresponding to state information such as stable learning, stable ability, and stable knowledge mastery. If the level fluctuations are not significant, it indicates that the student's learning state is relatively stable, and there are no obvious changes in their ability and knowledge mastery level.
[0103] In addition, for large fluctuations in the upward trend, it may indicate that the student has made a major breakthrough in the learning process, but there may also be elements of luck or a weak previous foundation. The corresponding first state information may be "great progress but unstable". For large fluctuations in the downward trend, it shows that the student's learning state has deteriorated sharply, and they may be facing greater learning pressure or have serious problems with their learning methods. The corresponding first learning state is a sharp decline. In the upward, downward, or stable trend, small fluctuations generally do not change the overall state judgment, but can reflect a certain degree of stability in the first state information. For example, in the stable trend, small fluctuations can indicate "stable learning state with minor adjustments".
[0104] Step S108: Perform a difference analysis based on the behavior description text and the target text data to obtain the second state information corresponding to the target student.
[0105] Exemplarily, the behavior description text is an abnormal behavior description of the target student in a training course, and the target text data is the questions asked by the target student during the communication process with the teacher. Then, when the difference between the abnormal behavior description and the target text data is smaller, it means that during the process of asking questions, the target student can accurately grasp the key points, and the questions raised by the target student or the content of the communication with the teacher closely revolves around the root cause of their abnormal behavior during the training process, indicating that the student has a relatively clear understanding of the key points of the training course and can keenly perceive their own problems. Based on this, it can be reasonably inferred that the second state information corresponding to the target student is that the learning direction is correct.
[0106] Exemplarily, on the contrary, when the difference between the abnormal behavior description and the target text data is greater, it indicates that the target student may currently be in a confused state. The target student is not clear about where their own problems lie, and the questions raised have a weak correlation with the actual abnormal behavior that occurred, reflecting their lack of insight and thinking about key issues during the learning process. In this case, the second status information corresponding to the target student is that the learning direction is incorrect.
[0107] In some embodiments, obtaining the second status information corresponding to the target student by performing a difference analysis based on the behavior description text and the target text data includes: identifying first keywords and a first relationship corresponding to the first keywords from the behavior description text; identifying second keywords and a second relationship corresponding to the second keywords from the target text data; calculating a first similarity between the first keywords and the second keywords, and obtaining a second similarity between the behavior description text and the target text data by combining the first relationship and the second relationship under the first similarity; and performing a difference analysis on the behavior description text and the target text data according to the second similarity to obtain the second status information corresponding to the target student.
[0108] Exemplarily, use a keyword extraction method such as a part-of-speech-based method to obtain the first keywords corresponding to the behavior description text, and analyze the logical relationships between the first keywords, such as causal relationships (e.g., "operational errors lead to result deviations", there is a causal relationship between "operational errors" and "result deviations"), parallel relationships (e.g., "illegal operations and non-compliance with procedures", "illegal operations" and "non-compliance with procedures" are in a parallel relationship), progressive relationships, etc., and record these relationships as the first relationship corresponding to the first keywords.
[0109] Exemplarily, apply the same or similar method as extracting the first keywords to the target text data to extract the second keywords from the target text data, and analyze the logical relationships between the second keywords in the same way as analyzing the first relationship, and record it as the second relationship corresponding to the second keywords.
[0110] Exemplarily, use cosine similarity to calculate the first similarity between the first keywords and the second keywords, and on the basis of obtaining the first similarity, consider the first relationship and the second relationship. If the logical relationships of the first keywords and the second keywords are consistent, for example, both have causal relationships and the causal directions are the same, then the similarity can be appropriately increased; if the relationships are inconsistent, such as one is a parallel relationship and the other is a causal relationship, then the similarity is appropriately decreased. Considering the first similarity and the consistency of the relationships comprehensively, obtain the second similarity between the behavior description text and the target text data.
[0111] Exemplarily, different similarity threshold ranges are preset. For example, a high similarity threshold range (such as 0.8 - 1), a medium similarity threshold range (such as 0.5 - 0.8), and a low similarity threshold range (such as 0 - 0.5) are set. The calculated second similarity is compared with the set threshold range. If the second similarity is within the high similarity threshold range, it indicates that the difference between the behavior description text and the target text data is small; if it is within the low similarity threshold range, it means the difference is large; being within the medium similarity threshold range indicates that there is a certain difference but not very large.
[0112] Exemplarily, the second status information corresponding to the target student is determined according to the result of the difference analysis. When the difference is small, it indicates that the target student has a good grasp of the key points when asking questions, and the corresponding second status information is that the learning direction is correct; when the difference is large, it means that the target student may not currently be clear about where their problems lie, and the corresponding second status information is that the learning direction is wrong; for the medium similarity situation, the location of the difference can be further analyzed to determine whether the student may have partial direction deviation or inaccurate understanding, etc., and an appropriate description of the second status information is given.
[0113] Step S109: Combine the first status information and the second status information to determine the target status information corresponding to the target student.
[0114] Exemplarily, when the first status information and the second status information are consistent in describing the learning status of the target student, this indicates that the same conclusion is drawn from different evaluation dimensions, and then a merging operation can be performed on the first status information and the second status information. For example, the first status information indicates that "the target student has made significant progress in learning and has a solid grasp of knowledge", and the second status information states that "the target student has the right learning direction". Then the merged target status information can be expressed as "the target student has made significant progress in learning, has a solid grasp of knowledge, and has the right learning direction". Such merging makes the information more concise and clear, while completely retaining the core content conveyed by the two status information, thus obtaining the target status information that can comprehensively reflect the learning status of the target student. However, when the first status information and the second status information describe the learning status of the target student inconsistently, the situation becomes more complex. This inconsistency may be caused by different emphases of different evaluation criteria, errors in data sources, or the multi-faceted nature of the student's learning status itself. At this time, simply merging the two status information may cause confusion for the user about the learning status of the target student. To clearly prompt this situation, it is necessary to add a preset statement "There is an abnormality in the current judgment of the learning status of the target student" after merging the first status information and the second status information. For example, the first status information shows that "the target student has learning difficulties and a weak grasp of knowledge", while the second status information indicates that "the target student has the right learning direction and has made obvious progress". After merging and adding the preset statement, the target status information can be written as "The target student has learning difficulties and a weak grasp of knowledge, but has the right learning direction and has made obvious progress. There is an abnormality in the current judgment of the learning status of the target student". By adding this preset statement, it can remind relevant personnel to analyze and judge the actual learning status of the target student more carefully when referring to this target status information, and further explore the reasons for this inconsistency in depth, so as to take more targeted measures to help the student improve the learning effect.
[0115] Please refer to Figure 2 , Figure 2A student learning status supervision device 200 provided by an embodiment of the present application. The student learning status supervision device 200 based on artificial intelligence includes a data acquisition module 201, a video processing module 202, an anomaly recognition module 203, a behavior analysis module 204, an audio processing module 205, a data sorting module 206, a status determination module 207, a difference analysis module 208, and a result determination module 209. Among them, the data acquisition module 201 is used to acquire initial video data corresponding to a target student at a training station according to a video collector and acquire initial audio data corresponding to the target student at the training station according to an audio collector; the video processing module 202 is used to perform key frame recognition on the initial video data to obtain a target behavior sequence corresponding to the target student; the anomaly recognition module 203 is used to perform anomaly behavior recognition on the target behavior sequence to obtain a target anomaly behavior corresponding to the target student; the behavior analysis module 204 is used to obtain target associated frames from the target behavior sequence according to the target anomaly behavior and perform behavior description according to the target associated frames to obtain a behavior description text corresponding to the target student; the audio processing module 205 is used to perform voice segmentation on the initial audio data to obtain target audio data corresponding to the target student and perform speech recognition on the target audio data to obtain target text data; the data sorting module 206 is used to perform text level recognition on the target text data to obtain a corresponding target text level and sort the target text levels according to time to obtain a text level sequence corresponding to the target student; the status determination module 207 is used to determine first status information corresponding to the target student according to the text level sequence; the difference analysis module 208 is used to perform difference analysis on the behavior description text and the target text data to obtain second status information corresponding to the target student; the result determination module 209 is used to fuse the first status information and the second status information to determine target status information corresponding to the target student.
[0116] In some embodiments, the student learning status supervision device 200 based on artificial intelligence can be applied to a terminal device.
[0117] It should be noted that those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working process of the student learning status supervision device 200 based on artificial intelligence described above can refer to the corresponding process in the embodiment of the student learning status supervision method based on artificial intelligence described above, and will not be repeated here.
[0118] Please refer to Figure 3 , Figure 3 A schematic block diagram of the structure of a terminal device provided by an embodiment of the present invention.
[0119] As Figure 3As shown in the figure, the terminal device 300 includes a processor 301 and a memory 302. The processor 301 and the memory 302 are connected through a bus 303, which is, for example, an I2C (Inter - integrated Circuit) bus.
[0120] Specifically, the processor 301 is used to provide computing and control capabilities to support the operation of the entire terminal device. The processor 301 can be a central processing unit (CPU). The processor 301 can also be other general - purpose processors, digital signal processors (DSPs), application - specific integrated circuits (ASICs), field - programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general - purpose processor can be a microprocessor or the processor can also be any conventional processor, etc.
[0121] Specifically, the memory 302 can be a Flash chip, read - only memory (ROM), magnetic disk, optical disc, USB flash drive, or mobile hard disk, etc.
[0122] Those skilled in the art can understand that Figure 3 the structure shown in the figure is only a block diagram of some parts of the structure related to the solution of the embodiment of the present invention, and does not constitute a limitation on the terminal device to which the solution of the embodiment of the present invention is applied. The specific server may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0123] The processor is used to run a computer program stored in the memory and, when executing the computer program, implement any one of the artificial - intelligence - based student learning status supervision methods provided by the embodiments of the present invention.
[0124] In one embodiment, the processor is used to run a computer program stored in the memory and, when executing the computer program, implement the following steps:
[0125] Acquire the initial video data corresponding to the target student at the training station according to the video collector and the initial audio data corresponding to the target student at the training station according to the audio collector;
[0126] Perform key - frame recognition on the initial video data to obtain the target behavior sequence corresponding to the target student;
[0127] Anomaly behavior recognition is performed according to the target behavior sequence to obtain the target anomaly behavior corresponding to the target student;
[0128] Target associated frames are obtained from the target behavior sequence according to the target anomaly behavior, and a behavior description text corresponding to the target student is obtained by performing a behavior description based on the target associated frames;
[0129] The initial audio data is segmented into speech to obtain the target audio data corresponding to the target student, and the target audio data is subjected to speech recognition to obtain target text data;
[0130] The target text data is subjected to text level recognition to obtain a corresponding target text level, and the target text levels are sorted according to time to obtain a text level sequence corresponding to the target student;
[0131] A first state information corresponding to the target student is determined according to the text level sequence;
[0132] Difference analysis is performed according to the behavior description text and the target text data to obtain a second state information corresponding to the target student;
[0133] The first state information and the second state information are fused to determine a target state information corresponding to the target student.
[0134] It should be noted that those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working process of the above-described terminal device can refer to the corresponding process in the embodiment of the method for supervising the learning state of students based on artificial intelligence described above, and will not be elaborated here.
[0135] The embodiment of the present invention also provides a storage medium for computer-readable storage. The storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the steps of any one of the methods for supervising the learning state of students based on artificial intelligence provided in the specification of the embodiment of the present invention.
[0136] Among them, the storage medium may be an internal storage unit of the terminal device described in the foregoing embodiment, such as the hard disk or memory of the terminal device. The storage medium may also be an external storage device of the terminal device, such as a plug-in hard disk equipped on the terminal device, a smart media card (SMC), a secure digital (SD) card, a flash card, etc.
[0137] Those of ordinary skill in the art can understand that all or some of the steps in the methods disclosed above, and the functional modules / units in systems and devices, can be implemented as software, firmware, hardware, and their appropriate combinations. In the hardware embodiment, the division between the functional modules / units mentioned above does not necessarily correspond to the division of physical components; for example, a physical component can have multiple functions, or a function or step can be executed by several physical components in cooperation. Some or all physical components can be implemented as software executed by a processor, such as a central processing unit, a digital signal processor, or a microprocessor, or can be implemented as hardware, or can be implemented as an integrated circuit, such as an application-specific integrated circuit. Such software can be distributed on a computer-readable medium, which can include a computer storage medium (or non-transitory medium) and a communication medium (or transitory medium). As is well known to those of ordinary skill in the art, the term computer storage medium includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information, such as computer-readable instructions, data structures, program modules, or other data. The computer storage medium includes, but is not limited to, RAM, ROM, EEPROM, flash memory, or other memory technologies, CD-ROM, digital versatile disk (DVD), or other optical disk storage, magnetic cartridges, tapes, magnetic disk storage, or other magnetic storage devices, or any other medium that can be used to store the desired information and can be accessed by a computer. In addition, as is well known to those of ordinary skill in the art, a communication medium typically contains computer-readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transmission mechanism, and can include any information delivery medium.
[0138] It should be understood that the term "and / or" used in the specification and appended claims of the present invention refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations. It should be noted that, in this document, the terms "include", "comprise", or any other variants thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or system including a series of elements not only includes those elements but also includes other elements not explicitly listed, or further includes elements inherent to such a process, method, article, or system. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article, or system including that element.
[0139] The serial numbers of the embodiments of the present invention above are only for description and do not represent the superiority or inferiority of the embodiments. As mentioned above, the above are only specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims.
Claims
1. A method for supervising students' learning status based on artificial intelligence, characterized in that The method includes: Collecting initial video data corresponding to a target student at a training station by a video collector and collecting initial audio data corresponding to the target student at the training station by an audio collector; Performing key frame recognition on the initial video data to obtain a target behavior sequence corresponding to the target student; Performing abnormal behavior recognition on the target behavior sequence to obtain a target abnormal behavior corresponding to the target student; Obtaining target associated frames from the target behavior sequence according to the target abnormal behavior, and performing behavior description according to the target associated frames to obtain a behavior description text corresponding to the target student; Performing speech segmentation on the initial audio data to obtain target audio data corresponding to the target student, and performing speech recognition on the target audio data to obtain target text data; Performing text level recognition on the target text data to obtain a corresponding target text level, and sorting the target text levels according to time to obtain a text level sequence corresponding to the target student; Determining first state information corresponding to the target student according to the text level sequence; Performing difference analysis according to the behavior description text and the target text data to obtain second state information corresponding to the target student; Fusing the first state information and the second state information to determine target state information corresponding to the target student; Among them, the performing behavior description according to the target associated frames to obtain a behavior description text corresponding to the target student includes: Obtaining a training image and an image description text corresponding to the training image, and performing latent feature extraction on the training image by an image feature extraction layer of an information association model to obtain a first image feature corresponding to the training image; Performing text feature extraction on the image description text by a text feature extraction layer of the information association model to obtain a first text feature corresponding to the image description text; Performing feature mapping on the first image feature and the first text feature by a feature association layer of the information association model to obtain an associated triple corresponding to the training image; Performing feature extraction on the target associated frames by an encoding layer of a text generation model to obtain corresponding second image features; Obtaining a target triple corresponding to the target associated frames from the associated triple according to the target associated frames and the training image by a data determination layer of the text generation model; Performing feature extraction on the target triple by a self-attention layer of the text generation model to obtain corresponding second text features; Obtaining a target relationship between each entity in the target key frame by a structure attention layer of the text generation model using the second image feature and the second text feature; Generating the behavior description text corresponding to the target student by a long short-term memory network layer of the text generation model according to the target relationship and the target triple.
2. The method according to claim 1, characterized in that, The performing abnormal behavior recognition on the target behavior sequence to obtain a target abnormal behavior corresponding to the target student includes: Performing target recognition on each target image in the target behavior sequence to obtain a target area corresponding to the target image; Identify the joint information corresponding to the target student in the target area to obtain the target joint information corresponding to the target image; Obtain the relevant joint information corresponding to each sub-joint from the target joint information corresponding to each target image in the target behavior sequence; Perform curve fitting based on the relevant joint information to obtain the initial fitting curve corresponding to the sub-joint; Identify the change points of the initial fitting curve to obtain the joint change time corresponding to the sub-joint; Fuse the relevant joint information according to the joint change times corresponding to multiple sub-joints to obtain the target behavior type corresponding to the target student; Classify abnormal behaviors according to the target behavior type to obtain the target abnormal behavior corresponding to the target student.
3. The method according to claim 2, wherein The step of fusing the relevant joint information according to the joint change times corresponding to multiple sub-joints to obtain the target behavior type corresponding to the target student includes: Perform an intersection operation on the joint change times to obtain all joint times corresponding to all sub-joints; Traverse each sub-time in all the joint times, and obtain the associated joint information that changes at the sub-time from the relevant joint information according to the joint change time; Obtain the adjacent joint information corresponding to the previous time in the sub-time from the target joint information; Determine the current joint information corresponding to the sub-time according to the adjacent joint information and the associated joint information; Perform behavior type classification according to the adjacent joint information and the current joint information to obtain the target behavior type corresponding to the target student at the sub-time.
4. The method according to claim 1, wherein The step that the data determination layer of the text generation model uses the target association frame and the training image to obtain the target triple corresponding to the target association frame from the association triples includes: Extract initial keywords from the image description text according to the keyword extraction network of the data determination layer; Obtain the frequency information corresponding to the initial keywords by performing frequency statistics on the initial keywords according to the data statistics network of the data determination layer based on the image description text; Filter the initial keywords using the frequency information according to the data filtering network of the data determination layer to obtain the image keywords corresponding to the image description text; Perform tagging on the target association frame using the image keywords according to the tag determination layer of the data determination layer to obtain the first associated keyword corresponding to the target association frame and the first position information corresponding to the first associated keyword; Perform tagging on the training image using the image keywords according to the tag determination layer of the data determination layer to obtain the second associated keyword corresponding to the training image and the second position information corresponding to the second associated keyword; Perform an intersection process on the first associated keyword and the second associated keyword according to the data merging network of the data determination layer to obtain the target associated keyword; The similarity calculation network for determining layers based on the data determines the image correlation degree between the target correlation frame and the training image according to the target correlation keyword, the first position information, and the second position information; The data determination network for determining layers based on the data obtains the target triple corresponding to the target correlation frame from the correlation triples by using the data image correlation degree; wherein, the image correlation degree is obtained according to the following formula: ; Among them, represents the image correlation degree between the target correlation frame and the i-th training image, num represents the number of words corresponding to the target correlation keyword, represents the number of words corresponding to the first correlation keyword or the second correlation keyword, represents the data corresponding to the j-th target correlation keyword in the i-th training image in the second position information, represents the data corresponding to the j-th target correlation keyword in the target correlation frame in the first position information.
5. The method according to claim 1, wherein The text level recognition of the target text data to obtain the corresponding target text level includes: Determine the training courses corresponding to the target student, and obtain the course-related texts corresponding to the training courses at different levels; Calculate the similarity between the target text data and the course-related texts to obtain the target similarity between the target text data and the course-related texts; Determine the target text level corresponding to the target text data according to the target similarity and the level information corresponding to the course-related texts.
6. The method according to claim 1, wherein The difference analysis based on the behavior description text and the target text data to obtain the second state information corresponding to the target student includes: Perform keyword recognition on the behavior description text to obtain the first keyword and the first relationship corresponding to the first keyword; Perform keyword recognition on the target text data to obtain the second keyword and the second relationship corresponding to the second keyword; Calculate the first similarity corresponding to the first keyword and the second keyword, and combine the first relationship and the second relationship under the first similarity to obtain the second similarity between the behavior description text and the target text data; Perform difference analysis on the behavior description text and the target text data according to the second similarity to obtain the second state information corresponding to the target student.
7. An artificial intelligence-based student learning status supervision device, characterized in that, including: A data acquisition module, configured to collect initial video data corresponding to a target student at a training station according to a video collector and collect initial audio data corresponding to the target student at the training station according to an audio collector; A video processing module, configured to perform key frame recognition on the initial video data to obtain a target behavior sequence corresponding to the target student; An anomaly recognition module, configured to perform anomaly behavior recognition on the target behavior sequence to obtain a target anomaly behavior corresponding to the target student; A behavior analysis module, configured to obtain target associated frames from the target behavior sequence according to the target abnormal behavior, and perform behavior description based on the target associated frames to obtain a behavior description text corresponding to the target student; wherein, the performing behavior description based on the target associated frames to obtain a behavior description text corresponding to the target student includes: obtaining a training image and an image description text corresponding to the training image, and performing latent feature extraction on the training image according to an image feature extraction layer of an information association model to obtain a first image feature corresponding to the training image; performing text feature extraction on the image description text according to a text feature extraction layer of the information association model to obtain a first text feature corresponding to the image description text; performing feature mapping on the first image feature and the first text feature according to a feature association layer of the information association model to obtain an associated triple corresponding to the training image; performing feature extraction on the target associated frames according to an encoding layer of a text generation model to obtain corresponding second image features; obtaining a target triple corresponding to the target associated frames from the associated triple according to the target associated frames and the training image by using a data determination layer of the text generation model; performing feature extraction on the target triple according to a self-attention layer of the text generation model to obtain corresponding second text features; obtaining a target relationship between each entity in the target key frames by using the second image features and the second text features according to a structural attention layer of the text generation model; generating the behavior description text corresponding to the target student according to the target relationship and the target triple by using a long short-term memory network layer of the text generation model; An audio processing module, configured to perform speech segmentation on the initial audio data to obtain target audio data corresponding to the target student, and perform speech recognition on the target audio data to obtain target text data; A data sorting module, configured to perform text level recognition on the target text data to obtain a corresponding target text level, and sort the target text levels according to time to obtain a text level sequence corresponding to the target student; A status determination module, configured to determine first status information corresponding to the target student according to the text level sequence; A difference analysis module, configured to perform difference analysis on the behavior description text and the target text data to obtain second status information corresponding to the target student; A result determination module, configured to fuse the first status information and the second status information to determine target status information corresponding to the target student.
8. A terminal device, characterized in that, The terminal device includes a processor and a memory; The memory is used for storing a computer program; The processor is configured to execute the computer program and, when executing the computer program, implement the artificial intelligence-based method for supervising a student's learning status according to any one of claims 1 to 6.
9. A computer storage medium for computer storage, characterized in that, The computer storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the steps of the artificial intelligence-based student learning status supervision method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Video monitoring method and system based on multi-scene recognition and voice interaction
CN117749995A
Knowledge framework automatic generation method and device, computer equipment and storage medium
CN119475079A
Multi-dimensional learning condition statistical analysis method based on deep learning
CN119671499A