Learning rejection behavior evaluation method and device, electronic equipment and readable storage medium

By analyzing multimodal features through multidimensional expert networks, the problem of large assessment errors in different course contexts was solved, and more accurate assessment of school refusal behavior was achieved.

CN122090345APending Publication Date: 2026-05-26BEIJING NORMAL UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING NORMAL UNIVERSITY
Filing Date
2026-02-03
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

Existing technologies fail to adequately consider the differences in behavioral characteristics across different course contexts when assessing student rejection rates, resulting in significant errors in assessment results and reduced accuracy.

Method used

By acquiring multiple course videos and utilizing a multi-dimensional expert network to analyze multimodal features, including specific courses, shared courses, and grade-level expert networks, and taking into account the influence of multiple factors, a learning status score is obtained.

Benefits of technology

This improves the accuracy and interpretability of the assessment results, avoids the one-sidedness of single-dimensional analysis, and can more accurately reflect students' learning status.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122090345A_ABST
    Figure CN122090345A_ABST
Patent Text Reader

Abstract

The invention discloses a learning rejection behavior evaluation method and device, electronic equipment and a readable storage medium. The method comprises the steps of obtaining at least two course videos corresponding to a to-be-evaluated object in a preset evaluation interval, and inputting the at least two course videos into a target evaluation model; for any course video in the at least two course videos, acquiring a multi-modal feature sequence corresponding to the to-be-evaluated object; performing state analysis on the multi-modal feature codes corresponding to the multi-modal feature sequence through expert networks of at least two dimensions in the target evaluation model to obtain learning state features of at least two dimensions; and based on the learning state features corresponding to the at least two dimensions corresponding to the at least two course videos, obtaining a learning state score corresponding to the to-be-evaluated object output by the target evaluation model. Therefore, the learning state score is obtained based on the multi-course video and the multi-dimensional learning state features, so that the evaluation result is more practical, and the evaluation accuracy and interpretability are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of computer science, and specifically relates to a method, apparatus, electronic device, and readable storage medium for assessing school refusal behavior. Background Technology

[0002] With the development of existing technologies, the commonly used technical approach in the field of student rejection assessment is to use a single model to process all course data uniformly according to a predetermined assessment logic to determine the degree of student rejection. In other words, all course data is placed within the same assessment framework, and student behavioral characteristics are analyzed according to preset rules to obtain the rejection assessment result.

[0003] However, in practical applications, the psychological meaning of the same behavioral trait can vary significantly across different course contexts. For example, in a math class, a student frowning might often indicate that they are thinking deeply; while in less stressful classes, the same frowning behavior might suggest boredom. Failing to fully consider these differences in course context and treating behavioral traits uniformly across different contexts when processing course data leads to significant discrepancies between the obtained learning status assessments and the actual situation, thus reducing the accuracy of the assessments. Summary of the Invention

[0004] To overcome the problems existing in related technologies, this application provides a method, apparatus, electronic device, and readable storage medium for assessing school refusal behavior.

[0005] In a first aspect, embodiments of this application provide a method for assessing school refusal behavior, the method comprising: Obtain at least two course videos corresponding to the object to be evaluated within a preset evaluation interval, and input the at least two course videos into the target evaluation model; For any one of the at least two course videos, obtain the multimodal feature sequence corresponding to the object to be evaluated; The multimodal feature encoding corresponding to the multimodal feature sequence is analyzed by using an expert network with at least two dimensions in the target evaluation model to obtain learning state features with at least two dimensions. Based on the learning status features corresponding to the at least two dimensions of the at least two course videos, the learning status score of the object to be evaluated, output by the target evaluation model, is obtained.

[0006] Secondly, embodiments of this application provide an assessment device for school refusal behavior, the device comprising: The first acquisition module is used to acquire at least two course videos corresponding to the object to be evaluated within a preset evaluation interval, and input the at least two course videos into the target evaluation model. The second acquisition module is used to acquire the multimodal feature sequence corresponding to the object to be evaluated for any one of the at least two course videos; The first analysis module is used to perform state analysis on the multimodal feature encoding corresponding to the multimodal feature sequence through an expert network of at least two dimensions in the target evaluation model, so as to obtain learning state features of at least two dimensions. The third acquisition module is used to acquire the learning status score of the object to be evaluated, output by the target evaluation model, based on the learning status features corresponding to the at least two dimensions of the at least two course videos.

[0007] Thirdly, embodiments of this application provide an electronic device including a processor, a memory, and a program or instructions stored in the memory and executable on the processor, wherein the program or instructions, when executed by the processor, implement the steps of the method described in the first aspect.

[0008] Fourthly, embodiments of this application provide a readable storage medium on which a program or instructions are stored, which, when executed by a processor, implement the steps of the method described in the first aspect.

[0009] Fifthly, embodiments of this application provide a chip, the chip including a processor and a communication interface, the communication interface being coupled to the processor, the processor being used to run programs or instructions to implement the method as described in the first aspect.

[0010] In this embodiment, at least two course videos corresponding to the object to be evaluated within a preset evaluation interval are obtained and input into a target evaluation model. For any one of the at least two course videos, a multimodal feature sequence corresponding to the course video is obtained. A state analysis is performed on the multimodal feature encoding corresponding to the multimodal feature sequence using an expert network with at least two dimensions in the target evaluation model to obtain learning state features with at least two dimensions. Based on the learning state features corresponding to at least two dimensions of the at least two course videos, the learning state score corresponding to the object to be evaluated, output by the target evaluation model, is obtained. Thus, by using an expert network with at least two dimensions to process the multimodal feature encoding in parallel, the parallel routing mechanism avoids the one-sidedness of single-dimensional analysis, comprehensively considers the influence of multiple factors on the learning state of the object to be evaluated, and reduces evaluation errors caused by limitations of a single perspective. Furthermore, obtaining the learning state score based on multiple course videos and multi-dimensional learning state features makes the evaluation results more realistic, improving evaluation accuracy and interpretability. Attached Figure Description

[0011] Figure 1 This is a flowchart illustrating the steps of an assessment method for school refusal behavior provided in an embodiment of this application; Figure 2 This is a schematic diagram of the architecture of a target evaluation model provided in an embodiment of this application; Figure 3 This is a schematic diagram of the architecture of a single-course hybrid expert model provided in an embodiment of this application; Figure 4 This is a structural diagram of an assessment device for school refusal behavior provided in an embodiment of this application; Figure 5 This is a structural diagram of an electronic device provided in an embodiment of this application; Figure 6 This is a schematic diagram of the hardware structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0012] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0013] The terms "first," "second," etc., used in the specification and claims of this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such use of data can be interchanged where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein. Furthermore, in the specification and claims, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship.

[0014] The method for assessing school refusal behavior provided in this application will be described in detail below with reference to the accompanying drawings, through specific embodiments and application scenarios.

[0015] Figure 1 This is a flowchart illustrating the steps of an assessment method for school refusal behavior provided in an embodiment of this application, as follows: Figure 1 As shown, the method may include: Step 101: Obtain at least two course videos corresponding to the object to be evaluated within the preset evaluation interval, and input the at least two course videos into the target evaluation model.

[0016] In this embodiment, at least two course videos corresponding to the subject to be evaluated within a preset evaluation interval are acquired using camera equipment installed in the teaching scenario. The subject to be evaluated can be an individual in an educational setting whose learning status needs to be assessed and analyzed, such as a student in a school or a trainee in a training setting. The course videos corresponding to the subject to be evaluated can be video data recording the subject's performance during the course. The at least two course videos are input into a target evaluation model. The target evaluation model can be a pre-built and trained model used to output a learning status score for the subject to be evaluated based on the course videos.

[0017] Step 102: For any one of the at least two course videos, obtain the multimodal feature sequence corresponding to the object to be evaluated.

[0018] In this embodiment, for any course video, the feature extraction module in the target evaluation model performs multimodal feature extraction on the course video to obtain feature information corresponding to at least two modalities. The feature information corresponding to the at least two modalities is then fused to obtain a multimodal feature sequence corresponding to the object to be evaluated. The at least two modalities may include facial features, posture features, speech features, emotion features, and attention features, etc. The multimodal feature sequence is a collection of various types of features extracted from the course video corresponding to the object to be evaluated. For example, a pre-trained ResNet50 network in the feature extraction module can be used to extract facial features; a large language model (such as Qwen) in the feature extraction module can be used to recognize student postures, converting posture key points into natural language descriptions (such as "looking down while writing" or "looking at the blackboard"); and then a pre-trained BERT model in the feature extraction module can be used to extract posture features. The speech information of the object to be evaluated is recognized by the speech recognition model in the feature extraction module, and speech features are extracted. The emotion features of the object to be evaluated are extracted by the pre-trained emotion recognition model in the feature extraction module.

[0019] Step 103: Perform state analysis on the multimodal feature encoding corresponding to the multimodal feature sequence through an expert network of at least two dimensions in the target evaluation model to obtain learning state features of at least two dimensions.

[0020] In this embodiment, after obtaining the multimodal feature sequence, it can be encoded using a feature encoder to obtain the multimodal feature code. The multimodal feature code is then input into an expert network of at least two dimensions in the target evaluation model for state analysis, yielding learning state features corresponding to at least two dimensions. The target evaluation model can be a hierarchical routing-based hybrid expert model, including a single-course layer and a cross-course layer. The single-course layer can be composed of a single-course hybrid expert model, which can contain an expert network of at least two dimensions. This expert network can include a specific course expert network, a shared course expert network, and a grade-level expert network. The expert network can be a pre-trained neural network structure, such as a fully connected network, a convolutional network, or a small Transformer, capable of analyzing and interpreting the input multimodal feature code from a specific dimension to extract representative learning state features under that specific dimension.

[0021] Furthermore, the expert network, which has at least two dimensions, can be further divided into a time-segment expert network (distinguishing between morning, afternoon, and evening) and a classroom environment expert network (distinguishing between laboratory and regular classroom). By inputting multimodal encoding into the time-segment expert network and the teacher environment expert network, which are time-matched to the course video, learning state features in the time-segment dimension and learning state features in the classroom environment dimension can be obtained.

[0022] Course-specific expert networks can be constructed for each subject, such as a mathematics expert network or a language arts expert network. Considering that different subjects have unique knowledge systems, teaching methods, and learning requirements, students' behavioral performance and psychological states also differ across subject courses. Course-specific expert networks can help understand the specific behavioral patterns and state characteristics of the assessed individuals within these subjects, thereby analyzing their learning status in those subjects.

[0023] Shared course expert networks can be expert networks built according to different categories, such as science and humanities expert networks. Different categories of courses differ in terms of thinking styles, learning content, and learning methods. Therefore, shared course expert networks can capture the common and different characteristics at the subject classification level and derive the learning status characteristics of the assessment object that include the characteristics of the subject classification.

[0024] Grade-level expert networks can be constructed for each grade level. Students in different grades differ in terms of knowledge reserves, cognitive abilities, and learning goals. By using grade-level expert networks, the learning status of the assessed individuals can be determined by considering the characteristics of each grade level.

[0025] By performing state analysis on multimodal feature encoding using an expert network with at least two dimensions, learning state features corresponding to at least two dimensions can be obtained. These learning state features, corresponding to at least two dimensions, can reflect the learning state of the subject being evaluated in the course from different perspectives.

[0026] Step 104: Based on the learning status features of the at least two dimensions corresponding to the at least two course videos, obtain the learning status score of the object to be evaluated output by the target evaluation model.

[0027] In this embodiment, a learning status score corresponding to the subject to be evaluated, output by the target evaluation model, can be obtained based on at least two dimensions of learning status features for each of at least two course videos. The learning status score reflects the subject's learning performance and status level within a preset evaluation interval, characterizing the degree of the subject's refusal to learn. For example, the learning status score can be used to characterize the degree of refusal to learn, represented by a continuous value from 1 to 5, such as 1 = no refusal, 5 = very refusal.

[0028] The learning status score can be obtained by further concatenating the learning status features of at least two dimensions corresponding to at least two course videos to obtain a cross-course feature sequence. By using the cross-course encoder in the target evaluation model and combining the time dependency between at least two course videos, the cross-course feature sequence is analyzed to predict the learning status score of the object to be evaluated.

[0029] In one possible implementation, the learning status features of at least two dimensions corresponding to at least two course videos can be averaged or weighted averaged, and the target evaluation model can predict the learning status score of the object to be evaluated based on the average features obtained by the weighted average calculation.

[0030] In summary, in this embodiment, at least two course videos corresponding to the object to be evaluated within a preset evaluation interval are obtained and input into the target evaluation model. For any one of the at least two course videos, a multimodal feature sequence corresponding to the object to be evaluated is obtained. A state analysis is performed on the multimodal feature encoding corresponding to the multimodal feature sequence using an expert network with at least two dimensions in the target evaluation model to obtain learning state features with at least two dimensions. Based on the learning state features corresponding to at least two dimensions of the at least two course videos, the learning state score corresponding to the object to be evaluated output by the target evaluation model is obtained. Thus, by using an expert network with at least two dimensions to process the multimodal feature encoding in parallel, the parallel routing mechanism avoids the one-sidedness of single-dimensional analysis, comprehensively considers the influence of multiple factors on the learning state of the object to be evaluated, and reduces evaluation errors caused by the limitations of a single perspective. Furthermore, obtaining the learning state score based on multiple course videos and multi-dimensional learning state features makes the evaluation results more realistic, improving the accuracy and interpretability of the evaluation.

[0031] Optionally, step 104 may include the following steps: Step 201: Based on the learning state features of at least two dimensions corresponding to each of the at least two course videos, determine the course state features corresponding to each course video.

[0032] In this embodiment, the course status features corresponding to each course video are determined based on at least two dimensions of learning status features corresponding to each course video in at least two course videos. For example, for any course video, the at least two dimensions of learning status features corresponding to that course video can be concatenated, and then the concatenated features can be input into a fully connected layer to obtain the course status features corresponding to that course video output by the fully connected layer. These course status features can be used to reflect the comprehensive learning status of the subject being evaluated in the course corresponding to that course video.

[0033] Step 202: Concatenate the course status features corresponding to the at least two course videos to obtain a cross-course feature sequence.

[0034] In this embodiment, different courses often influence each other during the actual learning process. For example, a student's frustration in math class may affect their motivation and performance in subsequent language arts classes. By concatenating the course status features corresponding to at least two course videos, a cross-course feature sequence is obtained to integrate learning status information from different courses and capture the dependencies between them. For instance, the course status features of each course video can be connected in a certain order, such as the course time sequence, to form a cross-course feature sequence containing more information. This cross-course feature sequence contains comprehensive learning status information of the student being evaluated across multiple courses within a preset evaluation period, providing a more comprehensive reflection of the student's learning situation.

[0035] Step 203: Based on the cross-course feature sequence, obtain the learning status score corresponding to the object to be evaluated.

[0036] In this embodiment, a learning status score for the subject to be evaluated is obtained based on a cross-course feature sequence. Specifically, this cross-course feature sequence can be input into a cross-course encoder (Transformer) in the cross-course layer. The cross-course encoder can automatically learn cross-course correlation information between different courses through a self-attention mechanism, such as the influence of performance in math class on emotions in language arts class, thereby gaining an understanding of the changes in the learning status of the subject to be evaluated in different courses. The output of the cross-course encoder is passed through a fully connected regression layer. The fully connected regression layer can further integrate and transform the features extracted by the cross-course encoder, mapping them to a suitable scoring range, and finally obtaining the learning status score for the subject to be evaluated.

[0037] In this embodiment, by determining at least two dimensions of learning status features corresponding to each course video and further obtaining course status features, the learning status of a single course can be characterized from multiple dimensions. By concatenating multiple course status features into a cross-course feature sequence and modeling the temporal relationship between courses at the cross-course layer, the development trajectory of a student's learning status over one or more days can be captured, comprehensively reflecting changes in learning status, avoiding isolated assessments, and making the obtained learning status scores more realistic and accurate.

[0038] Optionally, embodiments of this application may further include the following steps: Step 301: Obtain the break-time video corresponding to the object to be evaluated within the preset evaluation interval, and obtain the break-time status features corresponding to each break-time video.

[0039] In this embodiment, video footage of the subject being evaluated during recess is captured using a camera device within a preset evaluation period. The video footage records various behaviors of the subject during recess, such as interacting with classmates and engaging in activities outside the classroom. Recess can be considered a type of course, and a specific course expert network, namely a recess expert network, can be established for it. When processing the recess video, features can be extracted first through the recess expert network to obtain the first recess feature, and then features can be extracted again through the grade-level expert network corresponding to the subject being evaluated to obtain the second recess feature. The first and second recess features are then concatenated to obtain the recess state features.

[0040] Accordingly, step 202 may include the following steps: Step 302: Concatenate the course status features corresponding to the at least two course videos and the break-time status features corresponding to each of the break-time videos in chronological order to obtain a cross-course feature sequence.

[0041] In this embodiment, the course status features corresponding to at least two course videos and the break-time status features corresponding to each break-time video are concatenated according to their actual chronological order within a preset evaluation interval to obtain a cross-course feature sequence. This cross-course feature sequence can be used to characterize the continuous state changes of the subject being evaluated throughout the entire evaluation interval, from classroom to break-time and back to classroom. For example, in the subject's morning schedule, there is a first class, followed by a break, then a second class, and finally another break. Following this chronological order, the course status features of the first class, the break-time status features of the first break, the course status features of the second class, and the break-time status features of the second break are sequentially concatenated to form a cross-course feature sequence that reflects the complete state changes of the subject being evaluated that morning.

[0042] In this embodiment of the application, by acquiring the video of the break between classes and the characteristics of the break between classes, and splicing the characteristics of the course status and the characteristics of the break between classes in chronological order, it is helpful to analyze the changing pattern of the status of the object to be evaluated over time, as well as the correlation between different states, thereby improving the accuracy of the performance evaluation of the object to be evaluated in course-related scenarios.

[0043] Optionally, step 102 may include the following steps: Step 401: Extract images from the course video according to a preset time interval to obtain a set of human images corresponding to the object to be evaluated; the set of human images contains human images corresponding to multiple acquisition times.

[0044] In this embodiment, images are extracted from the course video according to a preset time interval to obtain a set containing human body images corresponding to multiple acquisition times. The preset time interval can be flexibly set according to actual needs, such as every 5 minutes, 10 minutes, etc.

[0045] Step 402: Extract features from each human image in the human image set to obtain the facial features corresponding to each human image.

[0046] In this embodiment, for each human image in the human image set, a specific feature extraction algorithm is used to extract features, converting each human image into a set of representative feature vectors to obtain the facial features corresponding to each human image. These facial features can be used to distinguish the facial features of different students, providing a basis for subsequent student identification and classroom performance analysis. For example, a pre-trained deep learning model, such as the ResNet50 network, can be used for feature extraction.

[0047] Step 403: Generate pose features corresponding to each of the human body images based on the pose description information corresponding to each of the human body images.

[0048] In this embodiment, based on each human body image, a large language model is used to identify the student's posture and obtain posture description information corresponding to the human body image. The large language model can be the Qwen model, which analyzes the posture in the human body image and converts the posture key points into natural language descriptions. For example, when a student is writing on paper with their head down, the large language model can describe it as "writing with head down"; when a student looks up at the blackboard, it can be described as "looking at the blackboard".

[0049] The pre-trained BERT model extracts posture text features based on posture description information. For example, it transforms natural language descriptions of posture into numerical features that computers can understand and process, thereby generating posture features corresponding to each human image. For instance, for two different posture descriptions, "looking down to write" and "looking at the blackboard," the BERT model will produce different posture feature vectors after processing, thus accurately reflecting the differences in students' postures.

[0050] In one possible implementation, the pose keypoint coordinates of the human body image can also be directly used as features, as pose features.

[0051] Step 404: For any acquisition time, the facial features and pose features corresponding to the human body image at the acquisition time are stitched together to obtain the multimodal features corresponding to the acquisition time.

[0052] In this embodiment of the application, for each acquisition time in the human body image set, the facial features extracted from the human body image corresponding to that acquisition time and the generated pose features are spliced ​​together to obtain the multimodal features corresponding to that acquisition time.

[0053] Step 405: Concatenate the multimodal features corresponding to each acquisition time in chronological order to obtain the multimodal feature sequence corresponding to the object to be evaluated.

[0054] In this embodiment, the multimodal features corresponding to each acquisition moment are concatenated according to the actual chronological order of their occurrence in the course video to obtain the multimodal feature sequence corresponding to the subject to be evaluated. The chronological order reflects the changes in the subject's state during the class, and concatenating them in chronological order preserves the continuity of these state changes. This allows the multimodal feature sequence to characterize the multimodal feature information of the subject at different times throughout the course, reflecting the student's state changes in the classroom.

[0055] In this embodiment, by extracting images at preset intervals, the student's classroom state can be captured evenly, avoiding information omissions and redundancy. By stitching together facial and pose features at any given acquisition moment, comprehensive multimodal features can be formed, obtaining the instantaneous state at the acquisition moment. Then, by stitching together the features from each moment in chronological order, the process of changes in the student's classroom state can be fully presented, improving the richness of the multimodal feature sequence.

[0056] Optionally, the expert network with at least two dimensions includes a course-specific expert network, a shared course expert network, and a grade-level expert network; step 103 may include the following steps: Step 501: Based on the feature encoder in the target evaluation model, encode the multimodal feature sequence to obtain the multimodal feature code.

[0057] In this embodiment of the application, the multimodal feature sequence is encoded based on the feature encoder in the target evaluation model to obtain the multimodal feature code.

[0058] Step 502: Input the multimodal feature encoding into a target-specific course expert network that matches the course video, and obtain the first learning state feature output by the target-specific course expert network.

[0059] In this embodiment, multimodal feature encoding is input into a target-specific course expert network (PSB) that matches the course video. The PSB can be one or more PSBs associated with the course video. Each PSB performs feature extraction and analysis on the input multimodal feature encoding, extracting learning state features related to the subject. For example, the corresponding PSB can be directly selected based on the course label, or one or more PSBs can be selected by calculating the matching degree between the multimodal feature encoding and each PSB.

[0060] Because different subjects have different learning focuses and characteristics, the feature content output by each target-specific course expert network will also differ. When there are at least two target-specific course expert networks, the feature vectors output by each target-specific course expert network are weighted and summed according to their matching weights to obtain the target-specific course expert output, i.e., the first learning state feature. The matching weights can be pre-assigned to the target-specific course expert networks or determined based on the matching degree between the target-specific course expert networks and the course videos.

[0061] The selection weights of the routing mechanism (i.e. which experts are activated and the magnitude of their weights) can be directly used as an explanation for the model's decisions. For example, when evaluating course videos of math lessons, the target assessment model mainly relies on math experts and shared expert 1, which helps relevant educators understand the basis for the model's judgment.

[0062] Step 503: Input the multimodal feature encoding into the target shared course expert network that matches the course video, and obtain the second learning state feature output by the target shared course expert network.

[0063] In this embodiment, the shared course expert network is designed for general, interdisciplinary knowledge and skills within the course, and can extract learning state features related to these general contents from the course videos. Therefore, by calculating the matching degree between the course videos and each shared course expert network, a target shared course expert network matching the course video is determined. Multimodal feature encoding is input into the target shared course expert network to obtain the second learning state features output by the target shared course expert network. These second learning state features reflect the learning state regarding general knowledge and skills in the course videos, complementing the first learning state features. This approach considers not only subject-specificity but also the generality of the course, making the assessment of the student's refusal to learn more comprehensive. For example, regardless of whether the course is mathematics or language arts, it may involve general features such as students' logical thinking ability and attention span, which the shared course expert network can capture.

[0064] Step 504: Input the multimodal feature encoding into the target grade expert network that matches the object to be evaluated, and obtain the third learning feature output by the target grade expert network.

[0065] In this embodiment, a target grade expert network is selected based on the grade label of the object to be evaluated. After the multimodal feature encoding is input into the target grade expert network, the target grade expert network processes the input multimodal feature encoding in combination with the characteristics of the grade and outputs the third learned feature.

[0066] In this embodiment, by acquiring the first learning characteristic, the second learning characteristic, and the third learning characteristic, the learning characteristics and needs of students at different stages of growth are considered. At the same time, the course specificity, cross-course universality, and grade commonality are also considered. The learning status of the subject to be evaluated is comprehensively characterized from three dimensions: subject, course universality, and grade, thereby improving the accuracy and authenticity of the assessment of school refusal behavior.

[0067] Optionally, embodiments of this application may further include the following steps: Step 601: Obtain the query vector of the specific course network corresponding to each subject.

[0068] In this embodiment, firstly, a specific course network is constructed for each subject. This specific course network can be trained on a large amount of data and used to extract feature information closely related to the subject. During the construction process, specific design and optimization are performed for factors such as teaching content, teaching methods, and learning focuses for different subjects. After the specific course network is constructed, a query vector can be extracted for each subject's specific course network. The query vector can be a feature representation of the subject's specific course network, containing key features and patterns of the subject's network when processing data. For example, for mathematics, the query vector of the specific course network may contain feature information related to mathematical concepts and operational rules; for language arts, the query vector may contain feature information related to language expression and literary understanding.

[0069] Step 602: For any subject-specific course network, calculate the matching degree between the query vector of the subject-specific course network and the multimodal feature code corresponding to the course video.

[0070] Step 603: According to the order of matching degree from high to low, determine the specific course network corresponding to the preset number of subjects with the highest matching degree as the target specific course expert network that matches the course video.

[0071] In this embodiment, for any subject-specific course network, the matching degree between the query vector of that subject-specific course network and the multimodal feature encoding corresponding to the course video is calculated to obtain the matching degree between each subject-specific course network and the course video. The matching degree can be calculated using similarity calculation methods such as cosine similarity and Euclidean distance. Then, according to the order of matching degree from high to low, a preset number of subject-specific course networks with the highest matching degree are selected and determined as the target subject-specific course expert networks matching the course video. For example, if the preset number is 2, then the two subject-specific course networks with the highest matching degree are selected as the target subject-specific course expert networks. For example, assuming the course video is a video corresponding to a mathematics course, the two subject-specific course networks with the highest matching degree could be the mathematics course network and the physics course network.

[0072] In this embodiment, filtering is performed from specific course networks of different disciplines based on matching degree. This allows for the selection of the subject network most suitable for the course video, ensuring that subsequent feature extraction aligns with the video content and improving the accuracy and relevance of the assessment of school refusal behavior. Furthermore, a routing mechanism can dynamically select the expert with the highest matching degree based on the input data, enabling the model to adapt to different courses and avoiding the "one-size-fits-all" problem of a uniform model.

[0073] Alternatively, the target evaluation model can be trained in the following way: Step 701: Input the sample dataset into the model to be trained and evaluated; the sample dataset contains multiple sets of sample data, and any sample data includes at least two sampled course videos corresponding to the sampled object within a preset sampling time and a self-score of learning status.

[0074] In this embodiment, a sample dataset is input into the evaluation model to be trained. The sample dataset contains multiple sets of sample data. Each set of sample data covers at least two sampled course videos collected from the same sampled object within a preset sampling period, along with the sampled object's self-rating of their learning status corresponding to these course videos. The evaluation model to be trained includes a specific course expert network to be trained for different courses, a shared course expert network to be trained for different course types, and a grade-level course network to be trained for different grade levels.

[0075] Step 702: For any sampled course video, based on the subject-specific expert network to be trained for the sampled course video, perform feature extraction on the sample data to obtain the first sample features.

[0076] In this embodiment, for any sampled course video, the corresponding expert network for the specific course to be trained is invoked based on the subject to which the sampled course video belongs. After the sample data is input into the network, the network analyzes and processes various information in the video, extracting subject-specific features, namely the first sample features.

[0077] Step 703: Based on the subject-specific shared course expert network corresponding to the sampled course videos, extract features from the sample data to obtain second sample features.

[0078] In this embodiment, for the sampled course videos, the expert network for the shared course to be trained corresponding to the subject is invoked. After the sample data is input into the network, the network extracts features related to general learning abilities, thinking methods, etc., which are the second sample features.

[0079] Step 704: Based on the course network of the grade to be trained corresponding to the sampled object, perform feature extraction on the sample data to obtain the third sample features.

[0080] In this embodiment, feature extraction is performed on the sample data based on the course network of the grade level to be trained corresponding to the sampled object. After the sample data is input into the network, the network will extract features related to the learning status of students in that grade, i.e., the third sample features, in combination with the characteristics of the grade.

[0081] Step 705: Based on the first sample features, the second sample features, and the third sample features, obtain the unprocessed course status features corresponding to the sampled course video.

[0082] In this embodiment of the application, the extracted first sample features, second sample features and third sample features are fused to obtain the unprocessed course status features corresponding to the sampled course video.

[0083] Step 706: Based on the status features of the courses to be processed corresponding to at least two sampled course videos in the sample data corresponding to the sampled course videos, obtain the predicted score output by the evaluation model to be trained.

[0084] In this embodiment, the evaluation model to be trained analyzes and processes the course status features corresponding to at least two sampled course videos in the sample data corresponding to the sampled course videos. Based on the model's internal learning mechanism and parameter settings, it outputs a predicted score. For example, the course status features corresponding to at least two sampled course videos can be concatenated according to the chronological order of the at least two sampled course videos to obtain a cross-course sequence to be processed. The evaluation model to be trained then analyzes this cross-course sequence to obtain a predicted score.

[0085] In one possible implementation, any sample data may also include sampled break-time videos of the sampled objects within a preset sampling duration. These break-time videos are input into the break-time expert network to be trained and the course network corresponding to the sampled objects to be trained, resulting in fourth and fifth sample features. The fourth and fifth sample features are then concatenated to obtain the break-time state features to be processed. Following a chronological order, the course state features corresponding to at least two sampled course videos and the break-time state features corresponding to the break-time videos are concatenated to obtain the cross-course sequence to be processed. The cross-course sequence to be processed is then analyzed using the evaluation model to be trained to obtain a predicted score.

[0086] Step 707: Based on the predicted score and the learning state self-score corresponding to the sampled object, adjust the parameters of the evaluation model to be trained.

[0087] In this embodiment, the predicted score output by the evaluation model to be trained is compared with the self-assessment of the learning state corresponding to the sampled object. For example, the difference between the predicted score and the self-assessment of the learning state corresponding to the sampled object can be calculated using a loss function: mean squared error (MSE). Based on the difference between the two, a suitable optimization algorithm (such as stochastic gradient descent) is used to adjust the parameters of the model. By continuously adjusting the parameters of the evaluation model to be trained, the similarity between the predicted score output by the evaluation model to be trained and the self-assessment of the learning state of the sample data is made greater than a first similarity threshold. For example, optimization algorithms such as stochastic gradient descent (SGD) and batch gradient descent (BGD) can be used to adjust the parameters of the evaluation model to be trained. The self-assessment of the learning state can be obtained by conducting a questionnaire survey on the sampled object regarding their self-assessment of their learning state of the course.

[0088] Step 708: If the stopping condition is met, the evaluation model to be trained is determined as the target evaluation model.

[0089] In this embodiment, the stopping condition may include conditions such as the loss value of the model to be trained and evaluated reaching a preset threshold, or the number of training epochs of the model to be trained and evaluated reaching a preset threshold. When the stopping condition is met, it indicates that the model has basically learned how to accurately evaluate the learning state, and at this time, the current model to be trained and evaluated is determined as the target evaluation model.

[0090] In one possible implementation, to encourage sparsity in the routing mechanism, load balancing losses can be added to prevent the overuse of a few expert networks.

[0091] In this embodiment of the application, by training the evaluation model to be trained, the evaluation model to be trained can learn the feature extraction capabilities of different courses, different grades and cross courses during the training process, and thus learn the evaluation capabilities of diversified learning states, so as to better evaluate the learning state of the evaluation object.

[0092] For example, Figure 2 A schematic diagram of the architecture of a target evaluation model is shown, such as... Figure 2 As shown, the target assessment model can be divided into two parts: a "single-course layer," which demonstrates the processing flow of data for a single course, including feature extraction, three routing mechanisms, and an expert pool; and a "cross-course layer," which demonstrates the fusion of multi-course representations and the final prediction. Assuming that at least two course videos corresponding to the object to be assessed within a preset assessment interval include math videos, language arts videos, etc., in the single-course layer, for any course video, the course video is input into the multimodal feature extraction module to obtain the multimodal feature encoding corresponding to the course video. Then, at least two dimensions of learning state features are obtained through a single-course hybrid expert model, and these at least two dimensions of learning state features are concatenated and fused to obtain the refined representation R corresponding to each course video, i.e., the course state features. The course state features corresponding to at least two course videos are concatenated in chronological order to obtain a cross-course feature sequence, which is processed by a cross-course encoder to capture the dependencies between different courses. Finally, the output of the Transformer encoder is passed through a fully connected regression layer to obtain the final rejection score, i.e., the learning state score corresponding to the object to be assessed. In this way, a Transformer-based multimodal student rejection behavior assessment model is used to design specific expert networks for different course types. Experts are dynamically selected through a routing mechanism, enabling the model to adaptively combine the knowledge of different experts according to the course context. This allows the student rejection behavior assessment scheme to dynamically adjust the assessment strategy according to the course context, thereby improving the accuracy and interpretability of the assessment.

[0093] For example, Figure 3 A schematic diagram of the architecture of a single-course hybrid expert model is shown, such as... Figure 3As shown, the diagram includes three routing mechanisms and their connection to an expert pool, as well as a fusion module. It obtains the multimodal feature sequence of the object to be evaluated corresponding to the course video, and then encodes it using a feature encoder to obtain a preliminary contextual representation c, i.e., multimodal feature encoding. The multimodal feature encoding is input into the three routing mechanisms: a specific course routing mechanism, a shared course routing mechanism, and a grade-level routing mechanism. In the specific course routing mechanism, two target specific course expert networks are selected based on the matching degree. The results of the two target specific course expert networks are fused and concatenated to obtain the specific course expert output E_s, i.e., the first learning state feature. A target shared course expert network is selected based on the matching degree to obtain the shared course expert output E_c, i.e., the second learning state feature. Based on the grade label of the object to be evaluated, a grade-level expert network corresponding to the grade label is selected to obtain E_g, i.e., the third learning state feature. The first, second, and third learning state features are fused and concatenated to obtain the refined representation R corresponding to the course video, i.e., the course state feature.

[0094] It should be noted that the execution entity of the school refusal behavior assessment method provided in this application embodiment can be a school refusal behavior assessment device, or a control module in the school refusal behavior assessment device for executing the school refusal behavior assessment method. This application embodiment uses the school refusal behavior assessment device executing the school refusal behavior assessment method as an example to illustrate the school refusal behavior assessment device provided in this application embodiment.

[0095] Figure 4 This is a structural diagram of an assessment device for school refusal behavior provided in an embodiment of this application, with reference to... Figure 4 The device may include: The first acquisition module 801 is used to acquire at least two course videos corresponding to the object to be evaluated within a preset evaluation interval, and input the at least two course videos into the target evaluation model. The second acquisition module 802 is used to acquire the multimodal feature sequence corresponding to the object to be evaluated for any one of the at least two course videos; The first analysis module 803 is used to perform state analysis on the multimodal feature encoding corresponding to the multimodal feature sequence through an expert network of at least two dimensions in the target evaluation model, so as to obtain learning state features of at least two dimensions. The third acquisition module 804 is used to acquire the learning status score of the object to be evaluated, output by the target evaluation model, based on the learning status features corresponding to the at least two dimensions of the at least two course videos.

[0096] Optionally, the third acquisition module 804 includes: The first determining module is used to determine the course status features corresponding to each course video based on the learning status features of at least two dimensions corresponding to each of the at least two course videos; The first splicing module is used to splice the course status features corresponding to the at least two course videos to obtain a cross-course feature sequence; The first acquisition submodule is used to acquire the learning status score corresponding to the object to be evaluated based on the cross-course feature sequence.

[0097] Optionally, the device further includes: The fourth acquisition module is used to acquire the break-time videos corresponding to the objects to be evaluated within the preset evaluation interval, and to acquire the break-time status features corresponding to each break-time video. The first splicing module includes: The first splicing submodule is used to splice the course status features corresponding to the at least two course videos and the break-time status features corresponding to each of the break-time videos in chronological order to obtain a cross-course feature sequence.

[0098] Optionally, the second acquisition module 802 includes: The first extraction module is used to extract images from the course video at preset time intervals to obtain a set of human images corresponding to the object to be evaluated; the set of human images contains human images corresponding to multiple acquisition times. The second extraction module is used to extract features from each human image contained in the human image set to obtain the facial features corresponding to each human image. The first generation module is used to generate pose features corresponding to each of the human body images based on the pose description information corresponding to each of the human body images. The second stitching module is used to stitch together the facial features and pose features of the human body image corresponding to any acquisition time to obtain the multimodal features corresponding to the acquisition time. The third splicing module is used to splice the multimodal features corresponding to each acquisition time in chronological order to obtain the multimodal feature sequence corresponding to the object to be evaluated.

[0099] Optionally, the expert network with at least two dimensions includes a course-specific expert network, a shared course expert network, and a grade-level expert network; the first analysis module 803 includes: The first encoding module is used to encode the multimodal feature sequence based on the feature encoder in the target evaluation model to obtain the multimodal feature code; The second acquisition submodule is used to input the multimodal feature encoding into a target-specific course expert network that matches the course video, and acquire the first learning state features output by the target-specific course expert network. The third acquisition submodule is used to input the multimodal feature encoding into a target shared course expert network that matches the course video, and acquire the second learning state features output by the target shared course expert network; The fourth acquisition submodule is used to input the multimodal feature encoding into a target grade expert network that matches the object to be evaluated, and to acquire the third learning feature output by the target grade expert network.

[0100] Optionally, the device further includes: The fifth acquisition module is used to obtain the query vector of the specific course network corresponding to each subject; The first calculation module is used to calculate the matching degree between the query vector of the specific course network corresponding to any subject and the multimodal feature code corresponding to the course video for any specific course network; The second determining module is used to determine the specific course expert network corresponding to a preset number of subjects with the highest matching degree as the target specific course expert network that matches the course video, according to the matching degree from high to low.

[0101] Optionally, the target evaluation model is trained through the following steps: The first input module is used to input the sample dataset into the model to be trained and evaluated; the sample dataset contains multiple sets of sample data, and each set of sample data includes at least two sampled course videos corresponding to the sampled object within a preset sampling time and a self-score of learning status. The third extraction module is used to extract features from the sample data for any sampled course video based on the subject-specific expert network to be trained for the subject of the sampled course video, so as to obtain the first sample features. The fourth extraction module is used to extract features from the sample data based on the subject-specific shared course expert network corresponding to the sampled course videos, to obtain the second sample features; The fifth extraction module is used to extract features from the sample data based on the course network of the grade to be trained corresponding to the sampled object, so as to obtain the third sample features; The sixth acquisition module is used to acquire the unprocessed course status features corresponding to the sampled course video based on the first sample features, the second sample features, and the third sample features; The seventh acquisition module is used to acquire the predicted score output by the evaluation model to be trained based on the status features of the course to be processed corresponding to at least two sampled course videos in the sample data corresponding to the sampled course videos. The first adjustment module is used to adjust the parameters of the evaluation model to be trained based on the predicted score and the learning state self-score corresponding to the sampled object. The third determining module is used to determine the evaluation model to be trained as the target evaluation model when the stopping condition is met.

[0102] The assessment device for school refusal behavior in this application embodiment can be a device, or a component, integrated circuit, or chip in a terminal. The device can be a mobile electronic device or a non-mobile electronic device. For example, mobile electronic devices can be mobile phones, tablets, laptops, PDAs, in-vehicle electronic devices, wearable devices, ultra-mobile personal computers (UMPCs), netbooks, or personal digital assistants (PDAs), etc., while non-mobile electronic devices can be servers, network attached storage (NAS), personal computers (PCs), televisions (TVs), ATMs, or self-service machines, etc. This application embodiment does not impose specific limitations.

[0103] The assessment device for school refusal behavior in this application embodiment can be a device with an operating system. The operating system can be Android, iOS, or other possible operating systems; this application embodiment does not specifically limit the specific operating system used.

[0104] Optional, such as Figure 5 As shown, this application embodiment also provides an electronic device, including a processor 910, a memory 909, and a program or instructions stored in the memory 909 and executable on the processor 910. When the program or instructions are executed by the processor 910, they implement the various processes of the above-described assessment method embodiment for school refusal behavior and achieve the same technical effect. To avoid repetition, they will not be described again here.

[0105] It should be noted that the electronic devices in the embodiments of this application include the mobile electronic devices and non-mobile electronic devices described above.

[0106] Figure 6 A schematic diagram of the hardware structure of an electronic device to implement an embodiment of this application.

[0107] The electronic device 900 includes, but is not limited to, components such as: radio frequency unit 901, network module 902, audio output unit 903, input unit 904, sensor 905, display unit 906, user input unit 907, interface unit 908, memory 909, and processor 910.

[0108] Those skilled in the art will understand that the electronic device 900 may also include a power supply (such as a battery) for supplying power to various components. The power supply may be logically connected to the processor 910 through a power management system, thereby enabling functions such as managing charging, discharging, and power consumption through the power management system. Figure 6 The electronic device structure shown does not constitute a limitation on the electronic device. The electronic device may include more or fewer components than shown, or combine certain components, or have different component arrangements, which will not be elaborated here.

[0109] It should be understood that, in this embodiment, the input unit 904 may include a graphics processing unit (GPU) 9041 and a microphone 9042. The GPU 9041 processes image data of still images or videos obtained by an image capture device (such as a camera) in video capture mode or image capture mode. The display unit 906 may include a display panel 9061, which may be configured in the form of a liquid crystal display, an organic light-emitting diode, or the like. The user input unit 907 includes at least one of a touch panel 9071 and other input devices 9072. The touch panel 9071 is also called a touch screen. The touch panel 9071 may include a touch detection device and a touch controller. Other input devices 9072 may include, but are not limited to, physical keyboards, function keys (such as volume control buttons, power buttons, etc.), trackballs, mice, and joysticks, which will not be described in detail here.

[0110] The memory 909 can be used to store software programs and various data. The memory 909 may primarily include a first storage area for storing programs or instructions and a second storage area for storing data. The first storage area may store the operating system, application programs or instructions required for at least one function (such as sound playback, image playback, etc.). Furthermore, the memory 909 may include volatile memory or non-volatile memory, or both. The non-volatile memory may be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory can be random access memory (RAM), static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct memory bus RAM (DRRAM). The memory 909 in the embodiments of this application includes, but is not limited to, these and any other suitable types of memory.

[0111] Processor 910 may include one or more processing units; optionally, processor 910 integrates an application processor and a modem processor, wherein the application processor mainly handles operations involving the operating system, user interface, and applications, and the modem processor mainly handles wireless communication signals, such as a baseband processor. It is understood that the aforementioned modem processor may also not be integrated into processor 910.

[0112] This application also provides a readable storage medium storing a program or instructions. When the program or instructions are executed by a processor, they implement the various processes of the above-described method for evaluating school refusal behavior and achieve the same technical effect. To avoid repetition, they will not be described again here.

[0113] The processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk.

[0114] This application embodiment also provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is used to run programs or instructions to implement the various processes of the above-described assessment method embodiment for school refusal behavior, and can achieve the same technical effect. To avoid repetition, it will not be described again here.

[0115] It should be understood that the chip mentioned in the embodiments of this application may also be referred to as a system-on-a-chip, system chip, chip system, or system-on-a-chip, etc.

[0116] This application provides a computer program product stored in a storage medium. The program product is executed by at least one processor to implement the various processes of the above-described method for assessing school refusal behavior, and can achieve the same technical effect. To avoid repetition, it will not be described again here.

[0117] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, it should be noted that the scope of the methods and apparatuses in the embodiments of this application is not limited to performing functions in the order shown or discussed, but may also include performing functions substantially simultaneously or in the reverse order, depending on the functions involved. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.

[0118] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of this application.

[0119] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.

Claims

1. A method for assessing school refusal behavior, characterized in that, The method includes: Obtain at least two course videos corresponding to the object to be evaluated within a preset evaluation interval, and input the at least two course videos into the target evaluation model; For any one of the at least two course videos, obtain the multimodal feature sequence corresponding to the object to be evaluated; The multimodal feature encoding corresponding to the multimodal feature sequence is analyzed by using an expert network with at least two dimensions in the target evaluation model to obtain learning state features with at least two dimensions. Based on the learning status features corresponding to the at least two dimensions of the at least two course videos, the learning status score of the object to be evaluated, output by the target evaluation model, is obtained.

2. The method according to claim 1, characterized in that, The step of obtaining the learning status score corresponding to the object to be evaluated, output by the target evaluation model, based on the learning status features corresponding to the at least two dimensions of the at least two course videos, includes: Based on the learning state features of at least two dimensions corresponding to each of the at least two course videos, determine the course state features corresponding to each course video; The course status features corresponding to at least two course videos are concatenated to obtain a cross-course feature sequence. Based on the cross-course feature sequence, the learning status score corresponding to the object to be evaluated is obtained.

3. The method according to claim 2, characterized in that, The method further includes: Obtain the break-time videos corresponding to the objects to be evaluated within the preset evaluation interval, and obtain the break-time status features corresponding to each break-time video; The step of concatenating the course status features corresponding to the at least two course videos to obtain a cross-course feature sequence includes: The course status features corresponding to the at least two course videos and the break-time status features corresponding to each of the break-time videos are concatenated in chronological order to obtain a cross-course feature sequence.

4. The method according to claim 1, characterized in that, The step of obtaining the multimodal feature sequence corresponding to the object to be evaluated includes: According to a preset time interval, images are extracted from the course video to obtain a set of human images corresponding to the object to be evaluated; the set of human images contains human images corresponding to multiple acquisition times; Feature extraction is performed on each human image in the human image set to obtain the facial features corresponding to each human image; Based on the pose description information corresponding to each of the human body images, generate pose features corresponding to each of the human body images; For any acquisition time, the facial features and pose features corresponding to the human image at the acquisition time are stitched together to obtain the multimodal features corresponding to the acquisition time. The multimodal features corresponding to each acquisition time are spliced ​​together in chronological order to obtain the multimodal feature sequence corresponding to the object to be evaluated.

5. The method according to claim 1, characterized in that, The at least two-dimensional expert network includes a course-specific expert network, a shared course expert network, and a grade-level expert network; the state analysis of the multimodal feature encoding corresponding to the multimodal feature sequence through the at least two-dimensional expert network in the target evaluation model yields at least two-dimensional learning state features, including: Based on the feature encoder in the target evaluation model, the multimodal feature sequence is encoded to obtain the multimodal feature code; The multimodal feature encoding is input into a target-specific course expert network that matches the course video, and the first learning state feature output by the target-specific course expert network is obtained. The multimodal feature encoding is input into a target shared course expert network that matches the course video, and the second learning state feature output by the target shared course expert network is obtained; The multimodal feature encoding is input into a target grade expert network that matches the object to be evaluated, and the third learning feature output by the target grade expert network is obtained.

6. The method according to claim 5, characterized in that, The method further includes: Obtain the query vector of the specific course network corresponding to each subject; For any given subject and its corresponding specific course network, calculate the matching degree between the query vector of the subject and the multimodal feature code corresponding to the course video; Based on the order of matching degree from high to low, the specific course networks corresponding to the preset number of subjects with the highest matching degree are determined as the target specific course expert networks that match the course videos.

7. The method according to claim 1, characterized in that, The target evaluation model is trained through the following steps: Input the sample dataset into the model to be trained and evaluated; the sample dataset contains multiple sets of sample data, and each set of sample data includes at least two sampled course videos corresponding to the sampled object within a preset sampling time and a self-score of learning status. For any sampled course video, based on the subject-specific expert network to be trained for the sampled course video, feature extraction is performed on the sample data to obtain the first sample feature; Based on the subject-specific expert network for the shared courses to be trained from the sampled course videos, feature extraction is performed on the sample data to obtain the second sample features; Based on the course network of the grade to be trained corresponding to the sampling object, feature extraction is performed on the sample data to obtain the third sample feature; Based on the first sample features, the second sample features, and the third sample features, obtain the unprocessed course status features corresponding to the sampled course video; Based on the status features of the courses to be processed corresponding to at least two sampled course videos in the sample data corresponding to the sampled course videos, the predicted score output by the evaluation model to be trained is obtained. Based on the predicted score and the learning state self-score corresponding to the sampled object, the parameters of the evaluation model to be trained are adjusted. If the stopping condition is met, the evaluation model to be trained is determined as the target evaluation model.

8. A device for assessing school refusal behavior, characterized in that, The device includes: The first acquisition module is used to acquire at least two course videos corresponding to the object to be evaluated within a preset evaluation interval, and input the at least two course videos into the target evaluation model. The second acquisition module is used to acquire the multimodal feature sequence corresponding to the object to be evaluated for any one of the at least two course videos; The first analysis module is used to perform state analysis on the multimodal feature encoding corresponding to the multimodal feature sequence through an expert network of at least two dimensions in the target evaluation model, so as to obtain learning state features of at least two dimensions. The third acquisition module is used to acquire the learning status score of the object to be evaluated, output by the target evaluation model, based on the learning status features corresponding to the at least two dimensions of the at least two course videos.

9. An electronic device, characterized in that, It includes a processor, a memory, and a program or instructions stored in the memory and executable on the processor, wherein the program or instructions, when executed by the processor, implement the steps of the assessment method for school refusal behavior as described in any one of claims 1-7.

10. A readable storage medium, characterized in that, The readable storage medium stores a program or instructions that, when executed by a processor, implement the steps of the method for evaluating school refusal behavior as described in any one of claims 1-7.