Learner cooperation role interpretable recognition method and device based on multi-dimensional features

By collecting and processing multi-dimensional feature data in collaborative learning and combining SHAP algorithm for interpretation and analysis, the lack of collaborative role classification methods in the existing technology in the analysis dimension, transparency and automation degree is solved, and more accurate and transparent collaborative role recognition is achieved.

CN120180295APending Publication Date: 2025-06-20BEIJING NORMAL UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510203770.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-24
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

The existing collaborative role classification methods have shortcomings in the analysis dimension, transparency and automation level, which limits the application of classification results in educational design and practice.

Method used

A learner collaborative role interpretation recognition method based on multi-dimensional features is adopted. By collecting and processing collaborative learning data, multiple target feature matrices are constructed, and they are fused and standardized, and interpreted and analyzed in combination with SHAP algorithms to identify and interpret the collaborative roles of learners in the collaborative learning process.

Benefits of technology

By integrating feature matrices of multiple dimensions, a comprehensive collaborative role analysis framework is built, which significantly improves the accuracy and transparency of collaborative role recognition, enhances the credibility of classification results, and provides support for educators to understand learner behavior and role distribution.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120180295A_ABST
    Figure CN120180295A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence, and discloses a learner cooperation role interpretable recognition method and device based on multi-dimensional features, and the method comprises the steps: processing collected cooperation learning data, and obtaining a plurality of target feature matrixes; fusing the plurality of target feature matrixes and performing standardization processing to obtain standardized learner feature data; marking cooperation roles participated by the learners and fusing the cooperation roles with the standardized learner feature data to form a model training data set; inputting the model training data set into a target training model, and training the characteristics of the learner to obtain a cooperative role of the learner; and interpreting and analyzing the cooperation role of the learner based on an SHAP algorithm, and determining the characteristics influencing the cooperation role of the learner. According to the method, the contribution of each dimension feature in role classification is clearly displayed, the working mechanism of the target training model is transparent, the credibility of the classification result is improved, and powerful support is provided for an educator to understand learner behaviors and role distribution.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of artificial intelligence, and particularly to a method and device for interpretable recognition of learner collaboration roles based on multi-dimensional features. Background Art

[0002] As an important teaching method in the field of education, collaborative learning emphasizes promoting social cognitive processes through the joint activities of group members. In collaborative learning, individuals play different roles, which reflect their task division and functional positioning in the group. Existing collaborative role classification methods mainly rely on scripted roles and generative roles, but such methods have deficiencies in analysis dimensions, transparency, and automation, thus limiting the application of classification results in educational design and practice. Summary of the Invention

[0003] The technical problem to be solved by the present invention is to provide a method for interpretable recognition of learner collaboration roles based on multi-dimensional features, which effectively solves the technical problems of deficiencies in analysis dimensions, transparency, and automation.

[0004] The above technical problem is solved by the following technical solutions:

[0005] A method for interpretable recognition of learner collaboration roles based on multi-dimensional features, the method comprising:

[0006] Collecting collaborative learning data generated by learners during the collaborative learning process;

[0007] Processing the collaborative learning data to obtain a plurality of target feature matrices;

[0008] Fusing a plurality of the target feature matrices and performing normalization processing on the fused feature matrix to obtain normalized learner feature data;

[0009] Annotating the collaboration roles participated by the learners and fusing the collaboration roles with the normalized learner feature data to form a model training data set;

[0010] Inputting the model training data set into a target training model to train the learner feature data to obtain the collaboration roles of the learners during the collaborative learning process;

[0011] Interpretively analyzing the collaboration roles of the learners during the collaborative learning process based on the SHAP algorithm to determine the features affecting the collaboration roles of the learners.

[0012] Compared with the background art, the method for interpretable recognition of learners' collaborative roles based on multi-dimensional features according to the present invention has the following beneficial effects: By collecting the collaborative learning data generated by learners during the collaborative learning process and processing the collaborative learning data, a plurality of target feature matrices are obtained; then, the plurality of target feature matrices are fused, and the fused feature matrix is standardized to obtain standardized learner feature data. By fusing the target feature matrices of multiple dimensions, an all-round collaborative role analysis framework is constructed, so as to be able to finely depict the performance of learners during the collaborative learning process, providing a scientific basis for the recognition of collaborative roles.

[0013] Furthermore, the collaborative roles participated by learners are labeled, and the collaborative roles are fused with the standardized learner feature data to form a model training data set; the model training data set is input into a target training model to train the learner feature data, and the collaborative roles of learners during the collaborative learning process are obtained. Through the powerful analysis ability of the target training model to analyze and train the learner feature data, the accuracy of collaborative role recognition is significantly improved, and the role characteristics in complex collaborative learning scenarios can be effectively captured. Then, based on the SHAP algorithm, the collaborative roles of learners during the collaborative learning process are explained and analyzed to determine the features that affect the collaborative roles of learners. By using the SHAP algorithm to explain and analyze the collaborative roles of learners during the collaborative learning process, the contribution of each dimension feature in role classification can be clearly displayed, making the working mechanism of the target training model transparent. This not only improves the credibility of the classification results but also provides strong support for educators to understand the behavior and role distribution of learners.

[0014] In one embodiment, after collecting the collaborative learning data generated by learners during the collaborative learning process, it further includes:

[0015] Taking the number of times of collaborative texts in the collaborative learning process as the time unit, dividing the collaborative learning data, and sorting the collaborative texts;

[0016] Taking the serial number of the collaborative text as the unit of the sliding window size, traversing each sliding window size and calculating the coverage rate of the current window size; the coverage rate is the coverage degree of the text content contained in the window for a specific target content in the collaborative text under the current window size;

[0017] In response to the coverage rate meeting the target range value, determining the current window size as the optimal window.

[0018] In one embodiment, the calculation formula for the coverage rate of the current window size is:

[0019]

[0020] Among them, Coverage(W) is the coverage rate of the current window size; UniqueLearners(1, W) is the number of unique learners within the current window range (from serial number 1 to W); TotalLearners is the total number of all unique learners in the collaborative learning data.

[0021] In one embodiment, in the step of fusing the multiple target feature matrices and performing normalization processing on the fused feature matrix to obtain the normalized learner feature data, the calculation formula is:

[0022]

[0023] The mean μ is the average value of the normalized learner feature data, and the calculation formula is:

[0024]

[0025] The calculation formula for the standard deviation σ is:

[0026]

[0027] Among them, X′ is the normalized learner feature data; X is the original learner feature data; μ is the mean; σ is the standard deviation; X i is the i-th sample value of the original learner feature data; n is the total number of learner feature data samples.

[0028] In one embodiment, in the step of inputting the model training data set into the target training model to train the learner feature data to obtain the role of the learner in the collaborative learning process, the loss function used in the training process is:

[0029]

[0030] Among them, C is the total number of role categories; y i,c is the one-hot encoding of the true label of the i-th sample; p i,c is the probability that the i-th sample predicted by the model belongs to category c; Ω(f k ) is the L2 regularization term.

[0031] In one embodiment, the calculation formula for explaining and analyzing the collaborative role of the learner in the collaborative learning process based on the SHAP algorithm to determine the features affecting the learner's collaborative role is:

[0032]

[0033] Among them, f(x) is the trained classifier; g(x) is the adder used to explain the evaluation result of the target training model; n is the total number of learner feature data samples; φ0 is the prediction benchmark value of the target training model, that is, the mean of all sample status levels; φ i is the SHAP value of the target training model, used to quantify the feature χ i 's influence on the output of the target training model; That is, the feature subset that does not contain x i ; |S| is the number of elements in set S; f x (S ∪ {χ i}) and f x (S) are the predicted values including the feature x i and not including the feature x i respectively.

[0034] In one embodiment, the collaborative learning data includes behavior log data, collaborative text data, and collaborative social relationship data; the target feature matrix includes a behavior feature matrix, a cognitive feature matrix, an emotional feature matrix, and a social network feature matrix; the processing of the collaborative learning data to obtain multiple target feature matrices includes:

[0035] Construct a group social interaction network based on the collaborative social relationship data, and calculate social interaction network indicators to form the social network feature matrix. The social interaction network indicators include out-degree, in-degree, effective size, and degree.

[0036] In one embodiment, the processing of the collaborative learning data to obtain multiple target feature matrices further includes:

[0037] Extract the behavior features of the collaborative role within a unit time from the behavior log data based on a statistical algorithm. The behavior features include the number of posts by the learner participating in the collaboration, the number of replies by the learner participating in the collaboration, the posting speed of the learner participating in the collaboration, the reply speed of the learner participating in the collaboration, the time variance of the posting moment when the collaboration group responds to the collaboration, and the time variance of the reply interval when the collaboration group responds to the collaboration;

[0038] Construct the behavior feature matrix based on the behavior features.

[0039] In one embodiment, the processing of the collaborative learning data to obtain multiple target feature matrices further includes:

[0040] Construct the dataset of the learner's cognitive performance based on the collaborative text data. The dataset includes a training set, a validation set, and a test set;

[0041] Input the training set into the text classification model to train the text classification model;

[0042] Evaluate and verify the reliability of the trained text classification model based on the test set, and correct the trained text classification model based on the evaluation and verification results of the reliability to improve the trained text classification model;

[0043] Extract the cognitive features of the collaborative text data based on the improved text classification model to obtain the probability distribution of each cognitive performance dimension, so as to form the cognitive feature matrix.

[0044] On the other hand, the present invention also provides a learner collaborative role interpretable recognition device based on multi-dimensional features, including:

[0045] An acquisition module for acquiring collaborative learning data generated by learners during collaborative learning;

[0046] A processing module for processing the collaborative learning data to obtain a plurality of target feature matrices;

[0047] A fusion module for fusing a plurality of the target feature matrices and performing standardization processing on the fused feature matrix to obtain standardized learner feature data;

[0048] A labeling module for labeling the collaborative roles participated by the learners and fusing the collaborative roles with the standardized learner feature data to form a model training data set;

[0049] A training module for inputting the model training data set into a target training model to train the learner feature data to obtain the collaborative roles of the learners during collaborative learning;

[0050] A determination module for analyzing and interpreting the collaborative roles of the learners during collaborative learning based on the SHAP algorithm to determine the features affecting the collaborative roles of the learners. Description of the Drawings

[0051] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for use in the description of the specific embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0052] Figure 1 It is a schematic flow chart of a method for interpretable recognition of learner collaborative roles based on multi-dimensional features according to an embodiment of the present invention;

[0053] Figure 2A social network schematic diagram drawn according to example data in a method for interpretable recognition of learner collaboration roles based on multi-dimensional features according to an embodiment of the present invention;

[0054] Figure 3 A flowchart of a method for interpretable recognition of learner collaboration roles based on multi-dimensional features according to another embodiment of the present invention;

[0055] Figure 4 A schematic diagram of a global interpretation example in a method for interpretable recognition of learner collaboration roles based on multi-dimensional features according to an embodiment of the present invention;

[0056] Figure 5 A schematic diagram of a local interpretation example in a method for interpretable recognition of learner collaboration roles based on multi-dimensional features according to an embodiment of the present invention;

[0057] Figure 6 A schematic diagram of the structure of a device for interpretable recognition of learner collaboration roles based on multi-dimensional features according to an embodiment of the present invention;

[0058] Figure 7 A schematic diagram of the structure of an electronic device according to an embodiment of the present invention. Detailed implementation manners

[0059] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0060] In the description of the present invention, it should be understood that the terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the present invention, unless otherwise stated, the meaning of "a plurality" is two or more. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific situations.

[0061] Collaborative learning is a learning method that triggers social cognitive processes through the joint activities of group members, involving the cycle of individual and social knowledge construction. Collaborative roles reflect the task division and functional positioning of individuals in the group and are an important indicator reflecting the group interaction pattern. The existing collaborative role classifications are mainly divided into scripted roles and generative roles. Scripted roles are pre-assigned by teachers or instructional designers and have clear responsibilities, such as facilitators, summarizers, etc., to coordinate group interaction and improve collaborative effects; generative roles are spontaneously formed and evolved by learners during the collaborative process, such as leaders, promoters, marginalizers, etc., which reflect the internal coordination results of group members on task division and resource allocation.

[0062] Existing methods mainly classify collaborative roles from the cognitive dimension (based on conversation or posting content) and the social dimension (based on social network indicators). However, single-dimensional analysis is difficult to reveal the real role-playing situation of learners. Secondly, most current research on collaborative role feature extraction and classification is based on subjective judgment, self-report, manual coding based on text content, and post hoc statistics of collaborative data, and rarely involves the automated analysis of learners' actual role-playing and execution; a small number of studies that attempt to use social network indicators for automated role identification can only achieve the identification of simple roles such as highly influential and marginalized ones. Finally, most machine learning models are essentially black-box models, and it is difficult to intuitively understand how they obtain classification results. This lack of transparency may lead to difficulties for researchers in explaining why a certain learner is classified as a certain role (e.g., highly influential, encourager, etc.), thus limiting the application of classification results in educational design and practice.

[0063] In summary, there are still many deficiencies in the current collaborative role recognition methods in terms of analysis dimensions, method transparency, and automation level.

[0064] The present invention discloses an interpretable recognition method and device for learners' collaborative roles based on multi-dimensional features, which can combine multi-dimensional features (such as cognitive, social, behavioral, and emotional) to explore a more comprehensive and objective role classification method, while improving the interpretability and automation level of the classification process to better support educational design and the optimization of collaborative learning practice.

[0065] To better describe the interpretable recognition method for learners' collaborative roles based on multi-dimensional features disclosed in the present invention, a specific illustration will be given in combination with an actual scenario: Taking the "circuit design" theme in a fifth-grade primary school science class as an example, in a fifth-grade science class, students are divided into groups of 5 - 6 to complete two inquiry tasks based on virtual experiments. The task content includes designing a circuit to meet specified conditions, such as powering a series of electrical appliances. The learning process is supported by a group collaboration record form and an online learning platform, while collecting students' behavioral, communication, and social interaction data.

[0066] According to an embodiment of the present invention, as Figure 1 shown, a method for interpretable recognition of learners' collaborative roles based on multi-dimensional features is provided, including the following steps:

[0067] Step S100: Collect the collaborative learning data generated by learners during the collaborative learning process;

[0068] Step S200: Process the collaborative learning data to obtain multiple target feature matrices;

[0069] Step S300: Fusion the multiple target feature matrices, and perform normalization processing on the fused feature matrix to obtain normalized learner feature data;

[0070] Step S400: Label the collaborative roles participated by learners, and fuse the collaborative roles with the normalized learner feature data to form a model training data set;

[0071] Step S500: Input the model training data set into the target training model to train the learner feature data, and obtain the collaborative roles of learners during the collaborative learning process;

[0072] Step S600: Based on the SHAP algorithm, interpret and analyze the collaborative roles of learners during the collaborative learning process, and determine the features that affect the collaborative roles of learners.

[0073] In this embodiment, by collecting the collaborative learning data generated by learners during the collaborative learning process, and processing the collaborative learning data to obtain multiple target feature matrices; then fusing the multiple target feature matrices, and performing normalization processing on the fused feature matrix to obtain normalized learner feature data. By fusing the target feature matrices of multiple dimensions, an all-round collaborative role analysis framework is constructed, so as to be able to depict the performance of learners during the collaborative learning process in detail, and provide a scientific basis for the recognition of collaborative roles.

[0074] Further annotate the collaborative roles participated by learners, and fuse the collaborative roles with the standardized learner characteristic data to form a model training data set. Input the model training data set into the target training model to train the learner characteristic data, and obtain the collaborative roles of learners in the collaborative learning process. Through the powerful analysis ability of the target training model to analyze and train the learner characteristic data, the accuracy of collaborative role recognition is significantly improved, and the role characteristics in complex collaborative learning scenarios can be effectively captured. Then, based on the SHAP algorithm, explain and analyze the collaborative roles of learners in the collaborative learning process to determine the characteristics that affect the collaborative roles of learners. By using the SHAP algorithm to explain and analyze the collaborative roles of learners in the collaborative learning process, the contribution of each dimension feature in role classification can be clearly displayed, making the working mechanism of the target training model transparent. This not only improves the credibility of the classification results but also provides strong support for educators to understand the learner behavior and role distribution.

[0075] In step S100, the collaborative learning data includes collaborative behavior log data, collaborative text data, and collaborative social relationship data. The collaborative learning data includes information in four dimensions: behavior, cognition, emotion, and society. For example, the log data records the interaction characteristics of students, such as task submission, tool use, and posting frequency. The collaborative text data contains the emotional expressions of students during collaboration, such as joy and confusion, and also includes the emotional interaction of students, such as encouragement. In addition, it also contains the cognitive investment level of learners in the conceptualization, exploration, and summary stages of collaborative participation. The collaborative social relationship data contains the social network structure information between group members, such as in-degree and out-degree.

[0076] Among them, the collaborative behavior log data can be directly obtained from the database of the online learning platform. The collaborative text data can directly record the collaborative process of learners and transcribe it into a video script, or it can be the text data in the online learning platform, usually the forum posts for certain collaborative topics. The collaborative social relationship data can be extracted from the behavior logs of the online learning platform or obtained through questionnaires. For example, obtaining the collaborative behavior log data specifically includes that on a certain online learning platform, students are grouped to participate in a project-based collaborative learning task, such as completing a scientific experiment design. The collaborative behavior log data can be login and active time, task submission records, interaction frequency, and task operation logs.

[0077] Obtaining collaborative text data specifically includes students discussing the details of scientific experiment design in discussion boards or online meetings. The collaborative text data can be discussion content, that is, the message content posted by students in the online discussion board, such as "We should measure the current first and then adjust the resistance value." It can also be Q&A records, recording the questions and answers between students, such as "How to adjust the voltage of this circuit?" "It can be achieved by changing the battery voltage." Or comments and feedback: the evaluations of posts or submitted tasks among group members, such as "This hypothesis is a bit incomplete. Experimental data can be supplemented."

[0078] Obtaining collaborative social relationship data specifically includes recording and analyzing the social interaction patterns of students in the team through online questionnaires or learning platforms. The collaborative social relationship data can be social network data, where the initiating relationship: the number of times a student initiates a question or interaction with another student, such as "Student A asks Student B 3 times." The response relationship: Student B's 2 replies to Student A's posts or messages. It can be questionnaire feedback, where students fill out questionnaires and report which classmates they interact with more. For example, "In this project, I (Student D) mainly cooperate with A and C, and they provided a lot of help."

[0079] As Figure 2 shown, it intuitively illustrates how to construct a collaborative social network from collaborative social relationship data. Taking the in-degree of Student D as an example, the in-degree represents the number of social interactions initiated by other students towards him during the collaborative learning process. It can be seen from the figure that the in-degree of Student D is 2 (interactions initiated by Student A and Student C respectively). Through a similar method, other social network indicators (such as out-degree, centrality, etc.) can also be calculated to further analyze the social status and interaction patterns of students in collaboration. Here, only the in-degree of Student D is taken as an example, and the rest of the indicators will not be elaborated.

[0080] In one embodiment, after step S100, it further includes:

[0081] Dividing the collaborative learning data with the number of collaborative texts during the collaborative learning process as the time unit, and sorting the collaborative texts;

[0082] Taking the serial number of the collaborative text as the unit of the sliding window size, traversing each sliding window size and calculating the coverage rate of the current window size; the coverage rate is the degree of coverage of the text content contained in the window for the specific target content in the collaborative text under the current window size;

[0083] In response to the coverage rate meeting the target range value, determining the current window size as the optimal window.

[0084] In this embodiment, the collaborative learning data is divided with the number of collaborative texts in the collaborative learning process as the time unit to adapt to the diversity of different collaborative tasks. For example, in the collaboration in the offline face-to-face or online video conferencing scenario, the conversations between students are timely, while in the collaborative activities in the asynchronous community, the conversations between students are lagged. Dividing the time unit based on the number of collaborative texts can effectively improve the applicable scope of the present invention.

[0085] Take the serial number of the collaborative text as the unit of the sliding window size W, traverse each sliding window size W and calculate the coverage rate of the current window size; when the coverage rate meets the target coverage rate range value, the current window size can be used as the optimal window. Further, the calculation formula for calculating the coverage rate of the current window size is:

[0086]

[0087] Among them, Coverage(W) is the coverage rate of the current window size; UniqueLearners(1,W) is the number of unique learners within the current window range (from serial number 1 to W); TotalLearners is the total number of all unique learners in the collaborative learning data.

[0088] To better determine the optimal window size, first initialize the window size and the step size; calculate the coverage rate of the current window in each iteration and compare the difference between it and the target coverage rate; dynamically adjust the window size according to the difference in the coverage rate. If the coverage rate is lower than the target coverage rate, increase the window, otherwise decrease the window; when the coverage rate meets the target coverage rate range value, determine the current window size as the optimal window. The calculation formula is as follows:

[0089]

[0090] where W * is the optimal window size; W is to traverse each sliding window size.

[0091] In step S200, the process of processing the collaborative learning data includes generating a behavior feature matrix, generating a cognitive feature matrix, generating an emotional feature matrix, and generating a social network feature matrix. Among them, the behavior feature matrix can be obtained by statistical algorithms. The social network feature matrix can be obtained by network analysis methods. The cognitive feature matrix and the emotional feature matrix can adopt the probability representation method and obtain their probability distributions through machine learning algorithms. Regarding the cognitive feature matrix and the emotional feature matrix, compared with the traditional single-category representation method, the probability representation method has higher interpretability and can more specifically represent the specific distribution of cognitive performance characteristics. The following gives examples of four-dimensional features respectively. Table 1 is the behavior feature matrix, Table 2 is the cognitive feature matrix, Table 3 is the emotional feature matrix, and Table 4 is the social feature matrix:

[0092] Table 1

[0093] Stu B1 B2 B3 B4 B5 B6 AG -0.832 0.424 0.250 0.458 0.610 27.204 AM 0.178 0.559 0.578 0.571 0.615 31.171 AW -0.276 0.580 0.631 0.543 0.593 35.486 … … … … … … …

[0094] Table 2

[0095] Stu C1 C2 C3 C4 AG 0.133 0.864 0.003 0.001 AM 0.289 0.602 0.067 0.042 AW 0.221 0.633 0.068 0.078 … … … … …

[0096] Table 3

[0097] Stu E1 E2 E3 E4 E5 E6 E7 E8 AG 0.121 0.000 0.000 0.000 0.126 0.250 0.000 0.000 AM 0.230 0.032 0.032 0.062 0.199 0.022 0.000 0.108 AW 0.085 0.000 0.000 0.054 0.128 0.095 0.027 0.133 … … … … … … … … …

[0098] Table 4

[0099] Stu S1 S2 S3 S4 AG 0.500 0.000 1.000 0.903 AM 0.500 0.750 1.900 0.813 AW 0.750 0.500 1.300 0.872 … … … … …

[0100] In Tables 1 to 4, Stu is the primary key, and each row represents the characteristic data of a student.

[0101] In one embodiment, step S200 includes: constructing a group social interaction network based on the collaborative social relationship data, and calculating social interaction network metrics to form a social network feature matrix. The social interaction network metrics include out-degree, in-degree, effective size, and hierarchical degree. Specifically, refer to Table 5:

[0102] Table 5 Attributes of the Social Performance Dimension of the Collaborative Role per Unit Time

[0103]

[0104]

[0105] In one embodiment, step S200 further includes: extracting the behavioral characteristics of the collaborative role per unit time from the behavior log data based on a statistical algorithm. The behavioral characteristics include the number of posts by the learner participating in the collaboration, the number of replies by the learner participating in the collaboration, the posting speed of the learner participating in the collaboration, the reply speed of the learner participating in the collaboration, the time variance of the posting moment when the collaboration group responds to the collaboration, and the time variance of the reply interval when the collaboration group responds to the collaboration; constructing a behavioral feature matrix based on the behavioral characteristics. Specifically, refer to Table 6:

[0106] Table 6 Attributes of the Behavioral Performance Dimension of the Collaborative Role per Unit Time

[0107]

[0108]

[0109] As Figure 3 shown, in one embodiment, step S200 further includes:

[0110] Step S210: Construct a dataset of learners' cognitive performance based on collaborative text data. The dataset includes a training set, a validation set, and a test set.

[0111] In step S210, annotate the students' collaborative texts according to the dimensional attributes of the students' collaborative role cognitive performance, as shown in Tables 7 and 8 specifically:

[0112] Table 7 Collaborative text sequence within a unit of time

[0113]

[0114] Table 8 Dimensional attributes of collaborative role cognitive performance within a unit of time

[0115]

[0116]

[0117] Further perform data preprocessing operations on the students' collaborative texts, including text cleaning, text tokenization, and word vectorization. Text cleaning can remove meaningless words (such as stop words) and symbols in the collaborative texts. Text tokenization can tokenize sentences to generate a word sequence. Word vectorization then converts each word into a vector representation using pre-trained word embeddings (such as Word2Vec).

[0118] Finally, obtain the dataset of learners' cognitive performance, including the vector representation of the collaborative texts and the cognitive performance dimension labels. The dataset of learners' cognitive performance is divided into a training set, a validation set, and a test set, with proportions of 60%, 20%, and 20% respectively.

[0119] Step S220: Input the training set into the text classification model to train the text classification model.

[0120] Step S230: Evaluate and verify the reliability of the trained text classification model based on the test set, and correct the trained text classification model based on the evaluation and verification results of the reliability to improve the trained text classification model.

[0121] In steps S220 - S230, the text classification model can be a TextCNN model. The structure of the TextCNN model includes an input, a convolutional layer, a pooling layer, and a fully connected layer. The input can be the training set and is represented in the form of word embeddings. The convolutional layer extracts local features in the collaborative texts through multiple convolutional kernels of different sizes (such as 2-gram, 3-gram). The pooling layer performs max pooling on the features output by the convolutional layer to capture the most significant features. The fully connected layer passes the pooled feature vectors to the fully connected layer, and the output is the probability distribution of each cognitive performance dimension.

[0122] During the training process of the text classification model, the Huber loss function is intended to be used as the criterion for judging the training degree of the text classification model. The Huber loss function combines the advantages of the mean squared error loss (L2 Loss) function and the absolute value loss (L1 Loss) function, enabling the gradient descent curve in the training of the text classification model to be smoother. The Stochastic Gradient Descent (SGD) algorithm can make the training process of the text classification model faster. Therefore, the SGD algorithm is used for iterative calculation. After several iterations, the value of the loss function is finally stabilized, completing the preliminary fitting of the text classification model. When the text classification model is trained, the reliability of the trained text classification model is evaluated and verified according to the test set. The commonly used P value (accuracy), R value (recall rate), and F-score in deep learning are selected as evaluation indicators. The text classification model is corrected according to the results of the reliability evaluation and verification of the text classification model, and the text classification model is iteratively improved until the performance of the text classification model is stable, thereby obtaining a text classification model with stable performance.

[0123] Step S240: Extract the cognitive features of the collaborative text data based on the improved text classification model to obtain the probability distribution of each cognitive performance dimension, so as to form a cognitive feature matrix.

[0124] In one embodiment, the TextCNN algorithm can also be used to extract the emotional probability distribution of the learner from the collaborative text. The method and principle are the same as those of the cognitive feature matrix and will not be elaborated here.

[0125] Among them, the attributes of the emotional performance dimension of the collaborative role per unit time are shown in Table 9:

[0126] Table 9

[0127]

[0128] In steps S300 - S400, fusion can be performed according to the student ID. Using the student ID as the primary key, an inner join is performed on the behavior, cognitive, emotional, and social feature matrices to ensure that all feature data of each student is merged into the same row to obtain the fused feature matrix, as shown in Table 10:

[0129] Table 10

[0130]

[0131]

[0132] Further, the fused feature matrix is standardized to obtain the standardized learner feature data. The calculation formula is:

[0133]

[0134] The mean μ is the average of the standardized learner characteristic data, and the calculation formula is:

[0135]

[0136] The calculation formula for the standard deviation σ is:

[0137]

[0138] Among them, X′ is the standardized learner characteristic data; X is the original learner characteristic data; μ is the mean; σ is the standard deviation; X i is the i-th sample value of the original learner characteristic data; n is the total number of learner characteristic data samples.

[0139] Then, manually label the collaborative roles participated by learners during the collaborative learning process to determine the types of collaborative roles, as shown in Table 11 specifically:

[0140] Table 11

[0141] Role Coding Role Type Functional Performance R1 Coordinator Lead the task direction, conduct division of labor and coordination, and supervise the discussion process R2 Explorer Responsible for specific tasks, such as collecting data, retrieving information, etc. R3 Assistant Provide auxiliary work such as putting forward ideas or suggestions, and reflect on the task process R4 Marginalizer In a marginal position, with low cognitive participation and relatively negative emotions

[0142] Fuse the collaborative roles with the standardized learner characteristic data to form a model training data set, so as to provide basic data for training the model with the training target.

[0143] In step S500, the target training model can be an XGBoost model. The XGBoost model has a high training speed and accuracy and is suitable for data analysis with complex features. Further, the loss function used during training is:

[0144]

[0145] Among them, C is the total number of role categories; y i,c is the one-hot encoding of the true label of the i-th sample; p i,c is the probability that the model predicts the i-th sample belongs to category c; Ω(f k ) is the L2 regularization term, which is used to prevent overfitting.

[0146] Further, considering the dynamic change characteristics of the collaborative roles of students in the collaborative learning scenario, the present invention makes two aspects of adaptive adjustments to the XGBoost model to better meet the requirements of the collaborative scenario:

[0147] First, introduce sliding window dynamic training data generation. Introduce the time sliding window mechanism into the training data generation process to generate training data based on time windows, providing richer time series samples for the XGBoost model. This improves the original static processing method of the XGBoost model for training data, enabling it to dynamically adjust training samples. For example, for each time window t, use the features of the previous n time windows as input and the label of the current window as output. Generate samples by sliding over time to dynamically train the XGBoost model. This can capture the performance changes of students at different stages of collaborative tasks and reflect the dynamic evolution of their roles. For example, a student may initially appear as a "marginalizer", but as the task progresses, their characteristics gradually shift towards a "coordinator". Through time window division and the sliding window mechanism, this transformation process can be clearly captured.

[0148] Second, introduce the Softmax function for probability distribution conversion. In the present invention, the output of the XGBoost model is converted into a probability distribution through the Softmax function. This not only improves the interpretability of the output of the XGBoost model but also can more comprehensively reflect the diversity of students' roles, especially suitable for the multi-faceted role characteristics of students in collaborative learning. For example, a student may have tendencies towards multiple roles within a specific time window, such as a 70% probability of being a "coordinator", a 20% probability of being an "inquirer", and a 10% probability of being a "marginalizer". This probability distribution provides more flexible and fine-grained role analysis support for teachers, no longer limited to the static classification of a single role.

[0149] In one embodiment, the calculation formula in step S600 is:

[0150]

[0151] where f(x) is the trained classifier; g(x) is the additive used to explain the evaluation results of the target training model; n is the total number of learner feature data samples; φ0 is the prediction benchmark value of the target training model, that is, the mean of all sample status levels; φ i is the SHAP value of the target training model, used to quantify the influence of feature χ i on the output of the target training model; that is, the feature subset without x i ; |S| is the number of elements in set S; f x (S ∪ {χ i}) and f x (S) are the predicted values including feature x i and without feature x i respectively.

[0152] Since f x (S ∪ {χi}) represents the predicted value of the status level containing feature x i , and f x (S) represents the predicted value of the status level without feature x i . Therefore, the difference between the two can represent the contribution degree of feature x i to the evaluation result of the interaction level. Then, the average contribution degree of feature x i is calculated through the above formula.

[0153] As Figure 4 , Figure 5 shown, in step S600, the output result of the trained XGBoost model is analyzed interpretably by using the SHAP (Shapley Additive Explanations) algorithm, that is, the collaborative role of the learner in the collaborative learning process is analyzed to obtain the important features affecting the learner's collaborative role.

[0154] Figure 4 For the global interpretation example, the global interpretation identifies 8 indicators that respectively have the greatest impact on the four roles, as shown in Table 12:

[0155] Table 12 Important features affecting role recognition

[0156] Ranking Coordinator Explorer Assistant Marginalizer 1 E5 B1 E7 S2 2 S2 B5 E8 B3 3 C2 C2 E3 C3 4 C3 B2 E4 E5 5 B3 B4 B5 B2 6 S1 C1 S3 E8 7 B2 E6 C2 E4 8 E8 E4 B2 C2

[0157] As Figure 5 shown, further predict the role of the random sample student AM.

[0158] Among them, the model prediction results include: Coordinator: f(x) = 2.92; Explorer: f(x) = -1.02; Assistant: f(x) = -2.22; Marginal: f(x) = -2.37; Convert the predicted values into probability distributions through the Softmax formula:

[0159]

[0160] Coordinator: Explorer: Assistant: Marginal:

[0161] In Figure 5Among them, the student's role is predicted as "coordinator" (97.05%); Positive contribution features (red area): For example, out-degree (S2 = 0.75, Mean = 0.500), indicating that the frequency of the student's initiating behavior has a greater impact on the "coordinator" role classification. Effective size (S3 = 2.5, Mean = 1.662), an efficient social interaction structure increases the probability of this role classification. Conceptualization level (C1 = 0.231, Mean = 0.228), indicating that the student contributed more knowledge points in the discussion. Negative contribution features (blue area): For example, in-degree (S1 = 0.500, Mean = 0.467), indicating that although the student has certain passive interaction behaviors, the relatively low in-degree weakens the certainty of the "coordinator" identity. Comprehensive impact: The positive contribution in the red area is much greater than the negative contribution in the blue area, making the model prediction result tend to be "coordinator".

[0162] Predict the student's role as other roles. For example, the low probability (0.49%) of the "marginalizer" role; Positive contribution (red area): Conclusiveness (C3 = 0.081, Mean = 0.072), a slight increase slightly higher than the mean. Negative contribution (blue area): Out-degree (S2 = 0.75, Mean = 0.500), indicating that the student's active interaction frequency does not conform to the characteristics of the "marginalizer". Comprehensive impact: The negative contribution in the blue area is much greater than the positive contribution in the red area, making the model prediction result tend to be "marginalizer".

[0163] On the other hand, as Figure 6 shown, an embodiment of the present invention further provides a multi-dimensional feature-based learner collaboration role interpretable recognition device, including:

[0164] A collection module 100 for collecting collaborative learning data generated by learners during the collaborative learning process.

[0165] A processing module 200 for processing the collaborative learning data to obtain multiple target feature matrices.

[0166] A fusion module 300 for fusing multiple target feature matrices and performing standardization processing on the fused feature matrix to obtain standardized learner feature data; The calculation formula is:

[0167]

[0168] The mean μ is the average value of the standardized learner feature data, and the calculation formula is:

[0169]

[0170] The calculation formula of the standard deviation σ is:

[0171]

[0172] Among them, X′ is the standardized learner characteristic data; X is the original learner characteristic data; μ is the mean; σ is the standard deviation; X i is the i-th sample value of the original learner characteristic data; n is the total number of learner characteristic data samples.

[0173] The annotation module 400 is used to annotate the collaborative roles participated by learners and fuse the collaborative roles with the standardized learner characteristic data to form a model training data set.

[0174] The training module 500 is used to input the model training data set into the target training model, train the learner characteristic data, and obtain the collaborative roles of learners in the collaborative learning process; the loss function used in the training process is:

[0175]

[0176] Among them, C is the total number of role categories; y i,c is the one-hot encoding of the true label of the i-th sample; p i,c is the probability that the i-th sample predicted by the model belongs to category c; Ω(f k ) is the L2 regularization term.

[0177] The first determination module 600 is used to analyze the collaborative roles of learners in the collaborative learning process based on the SHAP algorithm and determine the characteristics that affect the collaborative roles of learners; the calculation formula is:

[0178]

[0179] Among them, f(x) is the trained classifier; g(x) is the additive used to explain the evaluation result of the target training model; n is the total number of learner characteristic data samples; φ0 is the prediction benchmark value of the target training model, that is, the mean of all sample status levels; φ i is the SHAP value of the target training model, used to quantify the influence of the feature χ i on the output of the target training model; that is, the feature subset that does not include x i ; |S| is the number of elements in the set S; f x (S∪{χ i}) and f x (S) are the predicted values including the feature x i and not including the feature x i respectively.

[0180] In one of the embodiments, it further includes:

[0181] A sorting module, which is used to divide collaborative learning data with the number of times of collaborative text in the collaborative learning process as the time unit, and sort the collaborative text;

[0182] A calculation module, which is used to use the serial number of the collaborative text as the unit of the sliding window size, traverse each sliding window size and calculate the coverage rate of the current window size; the coverage rate is the coverage degree of the text content contained in the window for the specific target content in the collaborative text under the current window size; the calculation formula for calculating the coverage rate of the current window size is:

[0183]

[0184] where Coverage(W) is the coverage rate of the current window size; UniqueLearners(1,W) is the number of unique learners within the current window range (from serial number 1 to W); TotalLearners is the total number of all unique learners in the collaborative learning data;

[0185] A second determination module, which is used to determine the current window size as the optimal window in response to the coverage rate meeting the target range value.

[0186] In one embodiment, the collaborative learning data includes behavior log data, collaborative text data, and collaborative social relationship data; the target feature matrix includes a behavior feature matrix, a cognitive feature matrix, an emotional feature matrix, and a social network feature matrix; the processing module 200 includes:

[0187] A first construction unit, which is used to construct a group social interaction network based on the collaborative social relationship data and calculate social interaction network metrics to form a social network feature matrix. The social interaction network metrics include out-degree, in-degree, effective size, and hierarchical degree.

[0188] In one embodiment, the processing module 200 further includes:

[0189] An extraction unit, which is used to extract the behavior features of the collaborative role within a unit time from the behavior log data based on a statistical algorithm. The behavior features include the number of posts of the learner participating in the collaboration, the number of replies of the learner participating in the collaboration, the posting speed of the learner participating in the collaboration, the reply speed of the learner participating in the collaboration, the time variance of the posting moment when the collaboration group responds to the collaboration, and the time variance of the reply interval when the collaboration group responds to the collaboration;

[0190] A second construction unit, which is used to construct a behavior feature matrix based on the behavior features.

[0191] In one embodiment, the processing module 200 further includes:

[0192] A third construction unit for constructing a dataset of learners' cognitive performance based on collaborative text data, where the dataset includes a training set, a validation set, and a test set;

[0193] A training unit for inputting the training set into a text classification model to train the text classification model;

[0194] An evaluation unit for evaluating and validating the reliability of the trained text classification model based on the test set, and correcting the trained text classification model based on the evaluation and validation results of the reliability to improve the trained text classification model;

[0195] A fourth construction unit for extracting the cognitive features of the collaborative text data based on the improved text classification model to obtain the probability distributions of each cognitive performance dimension, so as to form a cognitive feature matrix.

[0196] Figure 7 The structural schematic diagram of the embodiment of the electronic device provided by the embodiment of the present invention is shown. The specific implementation of the electronic device in the specific embodiment of the present invention is not limited.

[0197] As Figure 7 shown, the electronic device may include: a processor 502, a communication interface 504, a memory 506, and a communication bus 508.

[0198] Wherein: the processor 502, the communication interface 504, and the memory 506 communicate with each other through the communication bus 508. The communication interface 504 is used to communicate with network elements of other devices such as clients or other servers. The processor 502 is used to execute the program 510, and specifically can execute the relevant steps in the above-mentioned embodiment of the method for interpretable recognition of learners' collaborative roles based on multi-dimensional features.

[0199] Specifically, the program 510 may include program code, and the program code includes computer-executable instructions.

[0200] The processor 502 may be a central processing unit CPU, or a specific integrated circuit ASIC (Application Specific Integrated Circuit), or one or more integrated circuits configured to implement the embodiments of the present invention. One or more processors included in the electronic device may be of the same type of processor, such as one or more CPUs; or may be of different types of processors, such as one or more CPUs and one or more ASICs.

[0201] A memory 506 for storing a program 510. The memory 506 may include high-speed RAM memory and may also include non-volatile memory, such as at least one disk memory.

[0202] Specifically, the program 510 can be called by the processor 502 to cause the electronic device to execute the relevant steps in the above-mentioned embodiments of the method for interpretable recognition of the collaborative role of learners based on multi-dimensional features.

[0203] Those of ordinary skill in the art can understand that Figure 7 The structure shown is only illustrative and does not limit the structure of the above-mentioned device. For example, the electronic device may further include more or fewer components than those shown in Figure 7 or have a different configuration from that shown in Figure 7 the figure.

[0204] The embodiments of the present invention also provide a computer-readable storage medium. The method according to the embodiments of the present invention can be implemented in hardware, firmware, or be implemented as computer code that can be recorded on a storage medium, or be implemented by downloading through a network and originally stored in a remote storage medium or a non-transitory machine-readable storage medium and will be stored in a local storage medium, so that the method described herein can be processed by such software stored on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only memory, a random access memory, a flash memory, a hard disk, or a solid-state drive, etc.; further, the storage medium can also include a combination of the above-mentioned types of memories. It can be understood that a computer, a processor, a microprocessor controller, or programmable hardware includes a storage component that can store or receive software or computer code, and when the software or computer code is accessed and executed by the computer, the processor, or the hardware, the method shown in the above embodiments is implemented.

[0205] In the specific content of the above specific embodiments, the technical features can be combined arbitrarily without contradiction. For the sake of brevity of description, not all possible combinations of the above technical features are described. However, as long as the combinations of these technical features do not conflict, they should all be considered as the scope described in this specification.

[0206] The specific content of the above specific embodiments only expresses several embodiments of the present invention, and its description is relatively specific and detailed, but it should not be understood as a limitation to the scope of the patent of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several deformations and improvements can be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the patent of the present invention should be subject to the appended claims.

Claims

1. A method for interpretable identification of learner collaboration roles based on multi-dimensional features, characterized in that: The method comprises: Collect collaborative learning data generated by learners in the collaborative learning process; Processing the collaborative learning data to obtain multiple target feature matrices; Fusing the plurality of target feature matrices, and performing standardization processing on the fused feature matrices to obtain standardized learner feature data; Annotating the collaborative roles in which the learners participate, and fusing the collaborative roles with the standardized learner feature data to form a model training data set; Inputting the model training data set into a target training model, training the learner feature data, and obtaining the collaborative role of the learner in the collaborative learning process; The collaborative roles of the learners in the collaborative learning process are explained and analyzed based on the SHAP algorithm, and the characteristics that affect the collaborative roles of the learners are determined.

2. The interpretable identification method of learner collaboration roles based on multi-dimensional features according to claim 1 is characterized in that: After collecting the collaborative learning data generated by the learners in the collaborative learning process, the method further includes: Taking the number of collaborative texts in the collaborative learning process as a time unit, dividing the collaborative learning data, and sorting the collaborative texts; Taking the serial number of the collaborative text as the unit of the sliding window size, traversing each sliding window size and calculating the coverage rate of the current window size; the coverage rate is the coverage degree of the text content contained in the window for the specific target content in the collaborative text under the current window size; In response to the coverage satisfying a target range value, the current window size is determined to be an optimal window.

3. The interpretable identification method of learner collaboration roles based on multi-dimensional features according to claim 2 is characterized in that: The calculation formula for calculating the coverage of the current window size is: Among them, Coverage(W) is the coverage of the current window size; UniqueLearners(1,W) is the number of unique learners in the current window range (from sequence number 1 to W); TotalLearners is the total number of all unique learners in the collaborative learning data.

4. The interpretable identification method of learner collaboration roles based on multi-dimensional features according to claim 1 is characterized in that: The target feature matrices are fused, and the fused feature matrices are standardized to obtain standardized learner feature data. The calculation formula is: The mean μ is the average value of the standardized learner feature data, and the calculation formula is: The calculation formula of standard deviation σ is: Among them, X′ is the standardized learner characteristic data; X is the original learner characteristic data; μ is the mean; σ is the standard deviation; X i is the i-th sample value of the original learner feature data; n is the total number of learner feature data samples.

5. The method for interpretable identification of learner collaboration roles based on multi-dimensional features according to claim 1 is characterized in that: The model training data set is input into the target training model, the learner feature data is trained, and the role of the learner in the collaborative learning process is obtained. The loss function used in the training process is: Where C is the total number of role categories; y i,c is the one-hot encoding of the true label of the i-th sample; p i,c is the probability that the i-th sample belongs to category c predicted by the model; Ω(f k ) is the L2 regularization term.

6. The method for interpretable identification of learner collaboration roles based on multi-dimensional features according to claim 1 is characterized in that: The calculation formula for analyzing the collaborative role of the learner in the collaborative learning process based on the SHAP algorithm and determining the characteristics that affect the collaborative role of the learner is: Among them, f(x) is the trained classifier; g(x) is the additive function used to explain the evaluation results of the target training model; n is the total number of learner feature data samples; φ0 is the predicted benchmark value of the target training model, that is, the mean of the state levels of all samples; φ i SHAP value of the target training model, used to quantify the feature χ i Impact on the output of the target training model; That is, it does not contain x i The characteristic subset of |S| is the number of elements in set S; f x (S∪{X i }) and f x (S) respectively contain the feature x i and does not contain feature x i The predicted value of .

7. The method for interpretable identification of learner collaboration roles based on multi-dimensional features according to claim 1 is characterized in that: The collaborative learning data includes behavior log data, collaborative text data and collaborative social relationship data; the target feature matrix includes a behavior feature matrix, a cognitive feature matrix, an emotional feature matrix and a social network feature matrix; The collaborative learning data is processed to obtain a plurality of target feature matrices, including: A group social interaction network is constructed based on the collaborative social relationship data, and social interaction network indicators are calculated to form the social network feature matrix, wherein the social interaction network indicators include out-degree, in-degree, effective scale and hierarchical degree.

8. The method for interpretable identification of learner collaboration roles based on multi-dimensional features according to claim 7 is characterized in that: The processing of the collaborative learning data to obtain a plurality of target feature matrices also includes: Extracting behavioral characteristics of collaborative roles per unit time from the behavioral log data based on a statistical algorithm, the behavioral characteristics include the number of posts by learners participating in the collaboration, the number of replies by learners participating in the collaboration, the posting speed by learners participating in the collaboration, the reply speed by learners participating in the collaboration, the time variance of the posting time of the collaborative group responding to the collaboration, and the time variance of the reply interval of the collaborative group responding to the collaboration; The behavior feature matrix is ​​constructed based on the behavior features.

9. The method for interpretable identification of learner collaboration roles based on multi-dimensional features according to claim 7 is characterized in that: The processing of the collaborative learning data to obtain a plurality of target feature matrices also includes: Constructing a dataset of the learner's cognitive performance based on the collaborative text data, wherein the dataset includes a training set, a validation set, and a test set; Inputting the training set into a text classification model to train the text classification model; Performing reliability evaluation and verification on the trained text classification model based on the test set, and revising the trained text classification model based on the reliability evaluation and verification result to improve the trained text classification model; The cognitive features of the collaborative text data are extracted based on the improved text classification model to obtain the probability distribution of each cognitive performance dimension to form the cognitive feature matrix.

10. A learner collaborative role interpretable identification device based on multi-dimensional features, characterized in that: include: The collection module is used to collect the collaborative learning data generated by learners in the collaborative learning process; A processing module, used for processing the collaborative learning data to obtain a plurality of target feature matrices; A fusion module, used for fusing the plurality of target feature matrices and performing standardization processing on the fused feature matrices to obtain standardized learner feature data; A labeling module, used for labeling the collaborative roles in which the learners participate, and fusing the collaborative roles with the standardized learner feature data to form a model training data set; A training module, used for inputting the model training data set into a target training model, training the learner feature data, and obtaining the collaborative role of the learner in the collaborative learning process; The determination module is used to interpret and analyze the collaborative role of the learner in the collaborative learning process based on the SHAP algorithm, and determine the characteristics that affect the collaborative role of the learner.