Text semantic calculation and cognitive input evaluation method based on conceptualized representation
By constructing an educational concept dictionary and F-A model, combining BERT model and correlation analysis, the coarse-grained and insufficient targeted representation of educational concepts in the existing technology is solved, and an accurate assessment of the semantic complexity and cognitive investment level of online discussion texts is achieved.
Patent Information
- Application Number
- CN202411922510.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-25
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2044-12-25
AI Technical Summary
The existing technology relies on a common corpus in the field of educational concept representation, and the text complexity analysis is coarse-grained, and the semantic information of educational concepts is not deeply explored, resulting in weak targeting and unexplainable.
Using a conceptual representation-based method, an educational concept dictionary and an F-A model were constructed, and the semantic correlation between text and educational concepts was calculated through the BERT model, the familiarity and abstraction of educational concepts were quantified, the semantic complexity of online discussion text was calculated, and the level of cognitive investment was analyzed in correlation.
It realizes the accurate representation and understanding of educational concepts, improves the semantic complexity calculation accuracy of online discussion texts, deeply understands the learners' cognitive investment status, and automatically evaluates the cognitive investment level.
Smart Images

Figure CN120012780A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of text information processing, and in particular to a text semantic computing and cognitive input evaluation method based on conceptual representation. Background Art
[0002] In large-scale online learning platforms, the textual semantic complexity of posts published by learners is a predictor of learners' cognitive engagement level. When learners publish posts in discussions that are more complex and involve multiple viewpoints and domain concepts, it indicates that they have conducted in-depth thinking and exploration, that is, their cognitive engagement level is high. On the contrary, if the content of their posts is relatively simple and lacks depth, it indicates that the learners' cognitive engagement is low. From this point of view, the textual semantic complexity of online discussion posts can be used as an indicator to measure learners' cognitive engagement.
[0003] In the field of education, understanding and characterizing the cognitive characteristics of educational concepts is crucial for effective teaching and learning. Educational concepts are the basic units in the field of education and an important part of the subject knowledge system. Understanding and characterizing the cognitive characteristics of educational concepts helps teachers and learners to deeply understand the essence of the subject and grasp the knowledge structure, thereby improving the effectiveness of teaching and learning.
[0004] Current research in the field of educational concept representation mostly relies on general corpora, and text complexity analysis is mostly coarse-grained analysis. Therefore, there is a lack of in-depth analysis specifically targeting educational concepts, and the semantic information of educational concepts has not been fully explored, resulting in weak targeting and inability to be explained. Summary of the invention
[0005] In view of this, an embodiment of the present invention provides a text semantic computing and cognitive investment evaluation method based on conceptual representation.
[0006] The first aspect of the present invention provides a text semantic computing and cognitive input evaluation method based on conceptual representation, comprising the following steps:
[0007] Constructing an education concept dictionary, using the education concept dictionary to match the education concepts in the online discussion text, and obtaining candidate education concepts in the online discussion text;
[0008] Using the BERT model to calculate the semantic relevance between the online discussion text and the candidate educational concepts, and filtering the candidate educational concepts whose semantic relevance is lower than a preset relevance threshold;
[0009] constructing an educational concept FA model according to the educational concept dictionary, and obtaining the complexity and abstractness of the candidate educational concept from the educational concept FA model;
[0010] Calculating the semantic complexity of the online discussion text according to the complexity and abstractness of the candidate educational concepts;
[0011] Conduct correlation analysis on the cognitive engagement level and semantic complexity of online discussion texts; train models to automatically identify the cognitive engagement level of online discussion texts.
[0012] Furthermore, the construction of the educational concept dictionary specifically includes the following steps:
[0013] Extract several educational concepts from the preset corpus;
[0014] The educational concepts are stored and organized using a Trie tree data structure to form an educational concept dictionary.
[0015] Furthermore, the method of using the education concept dictionary to match the education concepts in the online discussion text to obtain candidate education concepts of the online discussion text specifically includes the following steps:
[0016] Perform word segmentation on online discussion texts;
[0017] The edit distance between the online discussion text after word segmentation and each educational concept in the educational concept dictionary is calculated, and the educational concepts whose edit distance is lower than a preset distance threshold are taken as candidate educational concepts for the online discussion text.
[0018] Furthermore, the use of the BERT model to calculate the semantic relevance between the online discussion text and the candidate educational concept specifically includes the following steps:
[0019] Use the BERT model to construct a vector space of online discussion texts and the candidate educational concepts;
[0020] The semantic relevance of each candidate educational concept to the online discussion text is calculated using the following formula:
[0021]
[0022] Where V P The vector space representing the online discussion text, V Ci Represents the vector space of each candidate educational concept, i∈[1,j], j is the total number of candidate educational concepts.
[0023] Furthermore, filtering the candidate educational concepts whose semantic relevance is lower than a preset relevance threshold specifically includes the following steps:
[0024] Constructing a candidate educational concept set according to the candidate educational concepts;
[0025] The preset relevance threshold is determined by determining the mean value and standard deviation of the candidate educational concept set.
[0026] Furthermore, constructing an educational concept FA model according to the educational concept dictionary and obtaining the complexity and abstractness of the candidate educational concept from the educational concept FA model specifically includes the following steps:
[0027] Performing feature vectorization on the educational concepts in the educational concept dictionary to construct a language feature vector of the educational concepts;
[0028] Performing regression analysis based on the language feature vectors of the plurality of educational concepts, determining the familiarity and abstractness of each educational concept, and establishing an educational concept FA model;
[0029] The complexity and abstractness of the candidate educational concepts are obtained in the educational concept FA model.
[0030] Furthermore, the calculating of the semantic complexity of the online discussion text according to the complexity and abstractness of the candidate educational concepts specifically includes the following steps:
[0031] representing the online discussion text as a set of candidate educational concepts;
[0032] The semantic complexity of online discussion text is calculated using the following formula:
[0033]
[0034] Where N represents the total number of candidate educational concepts, i represents each candidate educational concept, i∈[1,N]; W F represents the familiarity weight of the candidate educational concept, F(C i ) represents the familiarity of the candidate educational concept in the educational concept FA model; W A represents the abstractness weight of the candidate educational concept, A(C i ) represents the familiarity of the candidate educational concept in the educational concept FA model.
[0035] Furthermore, the familiarity weight and abstractness weight of the candidate educational concepts are calculated by the quasi-Newton method.
[0036] Furthermore, the correlation analysis of the cognitive engagement level and semantic complexity of the online discussion text specifically includes the following steps:
[0037] Coding online discussion texts for cognitive engagement;
[0038] A correlation analysis tool is used to perform a correlation description on the online discussion text to obtain the correlation between the online discussion text and the semantic complexity and cognitive investment level.
[0039] Further, the training model automatically identifies the cognitive engagement level of the online discussion text, specifically including the following steps: determining the target machine learning model for training;
[0040] The training dataset is constructed with semantic complexity as input feature vector and cognitive investment level as model output;
[0041] The model is trained using the training data set to obtain a machine learning model that can automatically identify the cognitive input level of online discussion texts.
[0042] The embodiment of the present invention also discloses a computer program product or a computer program, which includes a computer instruction stored in a computer-readable storage medium. A processor of a computer device can read the computer instruction from the computer-readable storage medium, and the processor executes the computer instruction, so that the computer device executes the above method.
[0043] The embodiments of the present invention have the following beneficial effects: The text semantic calculation and cognitive engagement evaluation method based on conceptual representation of the present invention starts from the perspective of psycholinguistics, takes the language characteristics of educational concepts as the basis, formally describes the two important dimensions of familiarity and abstractness of educational concepts, quantifies and calculates the attribute values of familiarity and abstractness of educational concepts, and constructs a two-dimensional attribute model to better understand and apply educational concepts. The constructed "familiarity-abstraction (FA)" model for educational concepts effectively maps educational concepts into two-dimensional space. By quantifying and calculating the psycholinguistic attribute values of educational concepts, the conceptualization of online discussion texts is further realized. With the help of this model, the text semantic complexity of online discussion texts can be accurately calculated, so as to have a deeper understanding of the learner's cognitive engagement state and cognitive processing process, and the cognitive engagement level of online discussion texts can be further automatically evaluated.
[0044] Additional aspects and advantages of the present invention will be given in part in the following description, and in part will be obvious from the following description, or will be learned through practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0046] Figure 1 It is a basic implementation flow chart of a text semantic calculation and cognitive input evaluation method based on conceptual representation of the present invention.
[0047] Figure 2 It is a schematic diagram of the effect of constructing an educational concept dictionary of a Trie tree data structure according to the present invention.
[0048] Figure 3 This is a schematic diagram of the BERT model structure used in the present invention.
[0049] Figure 4 It is a schematic diagram of the effect of the FA model constructed by the present invention.
[0050] Figure 5 It is a schematic diagram of the FA model attribute control constructed by the present invention.
[0051] Figure 6 is an error bar chart of the cognitive engagement level and semantic complexity of the present invention.
[0052] Figure 7 It is a schematic diagram of the effect of using a training data set to train a machine learning model in the present invention. DETAILED DESCRIPTION
[0053] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0054] In order to define and describe the learner discussion text more objectively and accurately, the embodiment of the present invention provides a text semantic calculation and cognitive investment evaluation method based on conceptual representation. Figure 1 As shown, the method of the embodiment of the present invention includes the following steps:
[0055] S1. Construct an education concept dictionary, use the education concept dictionary to match the education concepts in the online discussion text, and obtain candidate education concepts in the online discussion text.
[0056] In step S1, the embodiment of the present invention constructs the education concept dictionary into a Trie tree structure, matches the segmented online discussion text with the education concept dictionary Trie tree, and derives related education concepts from the online discussion text to form a candidate concept set.
[0057] S2. Use the BERT model to calculate the semantic relevance between online discussion texts and candidate educational concepts, and filter out candidate educational concepts whose semantic relevance is lower than the preset relevance threshold.
[0058] In step S2, the embodiment of the present invention uses the BERT pre-training model to construct a vector space of online discussion texts and all candidate educational concepts, and calculates the semantic relevance between each candidate educational concept and the online discussion text to obtain a list of semantic relevance between the online discussion text and the educational concept. In addition, the central limit theorem is used to determine the threshold of the relevance list, and the candidate educational concepts whose relevance is lower than the threshold are filtered out.
[0059] S3. Construct an educational concept FA model based on the educational concept dictionary, and obtain the complexity and abstractness of the candidate educational concepts from the educational concept FA model.
[0060] In step S3, the embodiment of the present invention constructs an FA model of educational concepts to achieve a vector space formal representation of the psycholinguistic attributes of educational concepts. The FA model projects educational concepts into a two-dimensional space, thereby characterizing the psycholinguistic attributes of educational concepts, so as to better understand and diagnose learners' cognitive input state and cognitive processing of concept understanding.
[0061] S4. Calculate the semantic complexity of online discussion texts based on the complexity and abstractness of candidate educational concepts.
[0062] In step S4, the embodiment of the present invention calculates the semantic complexity of the online discussion text by using the familiarity and abstractness of the educational concepts contained in the online discussion text.
[0063] S5. Conduct correlation analysis on the cognitive engagement level and semantic complexity of online discussion texts; train the model to automatically identify the cognitive engagement level of online discussion texts.
[0064] In step S5, the embodiment of the present invention obtains the cognitive engagement level of the online discussion text through manual coding based on the ICAP framework of cognitive engagement, with a total of 4 levels from low to high. The cognitive engagement of the text is correlated with the semantic complexity to explore the semantic complexity characteristics of online discussion texts with different cognitive engagements, and select appropriate feature vectors to train a model for evaluating the cognitive engagement level of online discussion texts.
[0065] The method of the present invention effectively maps educational concepts into two-dimensional space by constructing a "familiarity-abstraction (FA)" model for educational concepts. By quantitatively calculating the psycholinguistic attribute values of educational concepts, the conceptualization of online discussion texts is further realized, so that learners' discussion texts can be defined and described more objectively and accurately. With the help of this model, the text semantic complexity of online discussion texts can be accurately calculated, thereby gaining a deeper understanding of learners' cognitive engagement state and cognitive processing, and the cognitive engagement level of online discussion texts can be further automatically evaluated.
[0066] As a preferred embodiment, the implementation process of each step of the method of the present invention is specifically discussed below:
[0067] S1. Construct an education concept dictionary, use the education concept dictionary to match the education concepts in the online discussion text, and obtain candidate education concepts in the online discussion text.
[0068] In step S1, constructing an education concept dictionary specifically includes the following steps:
[0069] S1-1. Extract several educational concepts from the preset corpus.
[0070] The embodiment of the present invention crawls corpus from MOOCCube (a large corpus covering different fields), MOOC (massive open online courses), educational document libraries (such as HowNet, major document libraries, etc.) and other channels to obtain learners' online discussion texts. Educational concepts are extracted from the online discussion texts to form a dictionary containing m educational concepts, which is defined as D = {C1, C2, C3, ..., C m-1 ,C m}.
[0071] S1-2. Use the Trie tree data structure to store and organize educational concepts to form an educational concept dictionary.
[0072] Trie tree data structure is as follows Figure 2 As shown. The dictionary trie starts with a root node, which is used as the starting point for all words; by inserting the first word of the educational concept into this root node; then the subsequent words of the educational concept are added to the dictionary trie one by one until the last character of the educational concept is reached. If the corresponding word already exists in the dictionary trie, it is transferred to the child node and the next word is processed. When the end of the educational concept is reached, a special keyword (EOT) is inserted to mark the end of the term.
[0073] In step S1, the educational concepts in the online discussion text are matched using the educational concept dictionary to obtain candidate educational concepts of the online discussion text, which specifically includes the following steps:
[0074] S1-3. Perform word segmentation on online discussion text.
[0075] The embodiment of the present invention uses existing natural language processing tools such as jieba, ltp, IKAnalyzer, Pangu word segmentation, etc. to segment the crawled online discussion text P.
[0076] S1-4. Calculate the edit distance between the online discussion text after word segmentation and each educational concept in the educational concept dictionary, and use the educational concepts whose edit distance is lower than a preset distance threshold as candidate educational concepts for the online discussion text.
[0077] In the embodiment of the present invention, the edit distance (Levenshtein distance) is used to generate a candidate concept set.
[0078] The edit distance refers to the minimum number of edit operations required to convert one string into another, which can measure the orthogonal similarity between two target objects. The online discussion text P after word segmentation and the educational concepts in the educational concept dictionary Trie tree are measured, and the maximum edit distance is set to 1. The relevant educational concepts C in the online discussion text P are derived, and the candidate concept set of the online discussion text P is generated, that is, P = {C1, C2, C3, ..., C j-1 ,C j}, j is the number of candidate concepts. Examples are shown in Table 1.
[0079] Table 1. Example of concept matching of online discussion texts
[0080]
[0081] S2. Use the BERT model to calculate the semantic relevance between online discussion texts and candidate educational concepts, and filter out candidate educational concepts whose semantic relevance is lower than the preset relevance threshold.
[0082] In step S2, the semantic relevance between the online discussion text and the candidate educational concepts is calculated using the BERT model, which specifically includes the following steps:
[0083] S2-1. Use the BERT model to construct a vector space of online discussion texts and candidate educational concepts;
[0084] like Figure 3 As shown in the figure, the full name of the BERT model is Bidirectional Encoder Representations from Transformer. The goal of the BERT model is to obtain the semantic representation of text using large-scale unlabeled corpus training, and then fine-tune the semantic representation of the text in a specific NLP (natural language processing) task, and finally apply it to the NLP task. In the NLP method based on deep neural networks, the characters / words in the text are usually represented by one-dimensional vectors; on this basis, the neural network will take the one-dimensional word vector of each character or word in the text as input, and after a series of complex transformations, output a one-dimensional word vector as the semantic representation of the text. In particular, we usually hope that the distance between characters / words with similar semantics in the feature vector space is also relatively close, so that the text vector converted from the character / word vector can also contain more accurate semantic information.
[0085] The embodiment of the present invention uses the BERT pre-training model to measure the semantic relevance between candidate educational concepts and online discussion texts, and filters out candidate educational concepts whose relevance is lower than a threshold. For example, given an online discussion text P, its corresponding candidate educational concept set is {C1, C2, C3, …, C j-1 ,C j}. Use the BERT pre-training model to construct the vector space of post P and all candidate educational concepts, and obtain the corresponding vector V P and
[0086] In some embodiments, in addition to the BERT model, other derivative models, such as word2vec, Tencent AI Lab, and ERNIE models, can also be used to extract semantic vector representations.
[0087] S2-2. Calculate the semantic relevance of each candidate educational concept to the online discussion text using the following formula:
[0088]
[0089] Where V P The vector space representing the online discussion text, V Ci Represents the vector space of each candidate educational concept, i∈[1,j], j is the total number of candidate educational concepts.
[0090] By using the semantic relevance formula, we can calculate the following list consisting of j semantic relevances R = {R1, R2, R3, ..., R j-1 ,R j};
[0091]
[0092] In step S2, candidate educational concepts whose semantic relevance is lower than a preset relevance threshold are filtered, which specifically includes the following steps:
[0093] S2-3. Construct a set of candidate educational concepts based on the candidate educational concepts.
[0094] S2-4. Determine the mean and standard deviation of the candidate educational concept set to determine the preset relevance threshold.
[0095] In the embodiment of the present invention, according to the average value of the correlation list R The lower limit threshold is determined by using the sum of the standard deviation σ, and the candidate educational concepts with relevance below the threshold are filtered out. That is, each post can be composed of N educational concepts with high semantic relevance to the post, P = {C1, C2, C3, …, C N-1 ,C N}.
[0096] S3. Construct an educational concept FA model based on the educational concept dictionary, and obtain the complexity and abstractness of the candidate educational concepts from the educational concept FA model.
[0097] In step S3, an educational concept FA model is constructed according to the educational concept dictionary, and the complexity and abstractness of the candidate educational concepts are obtained from the educational concept FA model, which specifically includes the following steps:
[0098] S3-1. Perform feature vectorization on the educational concepts in the educational concept dictionary and construct the language feature vectors of the educational concepts.
[0099] In the embodiment of the present invention, the educational concept dictionary D = {C1, C2, C3, ..., C m-1 ,C m The educational concepts in} are represented by feature vectorization based on word length, pinyin, strokes, part of speech, and semantics, namely Many studies have shown that the linguistic features of Chinese vocabulary, such as word length, pinyin, strokes and semantics, are closely related to its psycholinguistic properties. Quantifying the features of educational concepts can better infer their psycholinguistic properties.
[0100] S3-2. Perform regression analysis based on the language feature vectors of multiple educational concepts, determine the familiarity and abstractness of each educational concept, and establish an educational concept FA model.
[0101] In the embodiment of the present invention, the psycholinguistic attributes of the educational concept C are formally described as a two-dimensional attribute value vector L(C)=(F,A), where F and A represent the values of familiarity and abstractness, respectively, and the value range is between 1 and 7 (when the attribute value is 1, it means very unfamiliar or very specific; when the attribute value is 7, it means very familiar or very abstract).
[0102] Familiarity (F) and abstractness (A) are both psycholinguistic attributes of text components such as vocabulary, concepts, and sentences. They are important factors that affect human cognitive processing and modeling of text components such as vocabulary, domain concepts, sentences, and discourse. Familiarity refers to the learner's familiarity with a concept, including the understanding of the concept's definition, instances, attributes, and relevance. Familiarity can be measured by the learner's prior knowledge, experience, and familiarity with the concept. High familiarity means that the learner has more prior knowledge and experience of the concept, while low familiarity means that the learner is relatively unfamiliar with the concept or lacks sufficient prior knowledge. Abstractness refers to the depth and breadth of thinking involved in the concept, which can be started from the connotation and extension of the concept. Connotation refers to the essential attributes and core characteristics of the concept, while extension refers to the specific instances and situations covered by the concept. Concepts with high abstractness usually have fewer real instances and situations, while concepts with low abstractness are the opposite.
[0103] The embodiment of the present invention utilizes the language feature vector of the educational concept C The familiarity and abstractness of educational concepts are calculated using the psycholinguistic attribute model L(C) = (F, A). The established FA model is as follows: Figure 4 shown.
[0104] Specifically, the embodiment of the present invention first labels the familiarity F and abstractness A of a part of educational concepts, and then maps the educational concept language feature vector FV to the predicted values of the psychological language attributes familiarity F and abstractness A respectively through a regression analysis model (such as Catboost, LGBT and other integrated learning models), determines the familiarity F value and abstractness A value of all educational concepts, and finally obtains L(C)=(F,A).
[0105] S3-3. Obtain the complexity and abstractness of candidate educational concepts in the educational concept FA model.
[0106] In the embodiment of the present invention, the familiarity F and abstractness A attribute values of educational concepts are divided into two parts from low to high, forming a "two-dimensional four-quadrant" psycholinguistic attribute space. Each educational concept C can find the corresponding coordinates in this space, as shown in the following example: Figure 5 shown. Figure 5 The tree control shown provides a basis for in-depth analysis and calculation of psycholinguistic attribute values of educational texts (such as online discussion texts, student essays, teacher comments, etc.). Characterizing the psycholinguistic attributes of educational concepts can better understand and diagnose learners' cognitive input state and cognitive processing process for concept understanding.
[0107] S4. Calculate the semantic complexity of online discussion texts based on the complexity and abstractness of candidate educational concepts.
[0108] In step S4, the semantic complexity of the online discussion text is calculated according to the complexity and abstractness of the candidate educational concepts, which specifically includes the following steps:
[0109] S4-1. Represent online discussion text as a set of candidate educational concepts.
[0110] In the embodiment of the present invention, the online discussion text P is represented as a set of N highly relevant educational concepts after word segmentation, educational concept matching and relevance detection and filtering, that is, P = {C1, C2, C3, ..., C N-1 ,C N}.
[0111] At the same time, since the familiarity and abstractness of domain terms are important indicators that affect the semantic complexity of text, the familiarity and abstractness of educational concepts contained in online discussion texts can be used to estimate their text semantic complexity.
[0112] S4-2. Use the following formula to calculate the semantic complexity of online discussion text:
[0113]
[0114] Where N represents the total number of candidate educational concepts, i represents each candidate educational concept, i∈[1,N]; W F represents the familiarity weight of the candidate educational concept, F(C i ) represents the familiarity of the candidate educational concept in the educational concept FA model; W A represents the abstractness weight of the candidate educational concept, A(C i ) represents the familiarity of the candidate educational concept in the educational concept FA model.
[0115] In the embodiment of the present invention, through the concept C in the educational concept dictionary i The corresponding psycholinguistic attribute value can be found by looking up F(C) in step S3 L(C) = (F, A) i ) and A(C i ) corresponding value. F , W A 、F(C i )、A(C i )Enter the formula In the example above, the semantic complexity results of online discussion texts can be output.
[0116] As a preferred embodiment, the familiarity weight W of the candidate educational concept in the semantic complexity formula is F and the abstractness weight W A , calculated using the Quasi-Newton Methods.
[0117] S5. Conduct correlation analysis on the cognitive engagement level and semantic complexity of online discussion texts; train the model to automatically identify the cognitive engagement level of online discussion texts.
[0118] In step S5, correlation analysis is performed on the cognitive investment level and semantic complexity of the online discussion text, which specifically includes the following steps:
[0119] S5-1. Coding cognitive engagement of online discussion texts.
[0120] In the embodiment of the present invention, the cognitive engagement coding of the online discussion text is mainly based on the ICAP cognitive engagement framework. The ICAP framework is a theoretical model for understanding and analyzing the cognitive engagement level of learners in the learning process. ICAP is an abbreviation of four English words, which represent the following four cognitive engagement levels: (1) Passive: In this state, learners mainly receive information without any active processing. For example, when listening to a lecture or reading a textbook, learners may just passively receive the content without deep thinking or interaction. (2) Active: Learners actively participate in the learning process and process information at this stage. This may include taking notes, participating in discussions, or solving problems. Active learning helps to deepen the understanding of the material. (3) Constructive: In the constructive stage, learners not only actively participate, but also integrate new knowledge with existing knowledge to construct their own understanding. This may involve creatively applying knowledge, analyzing and synthesizing to form new ideas or models. (4) Interactive: The highest level of cognitive engagement, learners interact with others and jointly construct knowledge. Through group discussions, collaborative learning and other methods, learners deepen their understanding and challenge each other's views in communication.
[0121] In the embodiment of the present invention, the cognitive investment of online discussion texts is divided into four levels (passive (1) - active (2) - constructive (3) - interactive (4)) from low to high through the ICAP cognitive investment framework.
[0122] S5-2. Use the correlation analysis tool to describe the correlation of online discussion texts and obtain the correlation between online discussion texts and semantic complexity and cognitive investment level.
[0123] In the embodiment of the present invention, according to the results of cognitive engagement and semantic complexity of the online discussion text calculated in step S4, a correlation analysis tool is used to describe the correlation, and the semantic complexity characteristics of online discussion texts with different cognitive engagement levels based on psycholinguistic attributes are obtained. The correlation analysis tool used in the embodiment of the present invention is SPSS Modeler. SPSS Modeler is a tool for predictive analysis, data mining and text analysis. It allows users to create complex data flows through an intuitive graphical interface, and perform data preprocessing, modeling, evaluation and deployment.
[0124] After obtaining the semantic complexity characteristics of online discussion texts with different cognitive input levels based on psycholinguistic attributes, the embodiment of the present invention draws an error bar graph of cognitive input level (X-axis) and semantic complexity (Y-axis) to analyze the correlation between online discussion texts with different cognitive input levels and their semantic complexity. Figure 6 As shown in the figure, it can be seen that as cognitive investment increases, its semantic complexity also shows an overall upward trend. The highest cognitive investment level is interactive. In addition to considering its semantic complexity, this type of post also needs to consider the interactivity in the discussion, such as the number of likes and comments.
[0125] In some embodiments, the ANOVA (Analysis of Variance) tool may also be used to test the association description, and the test results are shown in Table 2:
[0126] Table 2 ANOVA test results
[0127]
[0128]
[0129] In step S5, the training model automatically identifies the cognitive investment level of the online discussion text, which specifically includes the following steps:
[0130] S5-3. Determine the target machine learning model for training;
[0131] S5-4. Use semantic complexity as input feature vector and cognitive input level as model output to construct training data set;
[0132] S5-5. Use the training data set to train the model to obtain a machine learning model that can automatically identify the cognitive input level of online discussion texts.
[0133] like Figure 7As shown, the embodiment of the present invention conceptualizes and represents the online discussion text P, and then quantifies the psychological language attributes such as familiarity F and abstractness A, as well as the calculated semantic complexity, number of comment likes, etc. as feature vectors (X feature vectors), and uses cognitive engagement as the prediction result (Y prediction result) to construct a data set. The best_model function in the pycaret library is called to compare multiple machine learning models to obtain the best model for automatically identifying cognitive engagement in online discussion texts. (In the embodiment of the present invention, it is a random Forest model with a recognition accuracy of 0.8416, a precision of 0.8366, a recall of 0.8416, and an F1 score of 0.8386).
[0134] The present invention can further calculate the semantic complexity of online discussion texts and improve the analysis accuracy and automatic evaluation level of online discussion texts by constructing a "familiarity-abstraction (FA)" conceptual representation model based on psycholinguistic attributes for educational concepts. Compared with the prior art, the model can objectively and accurately define and describe educational concepts, avoiding the lack of pertinence caused by corpus dependence in traditional methods. By quantitatively calculating psycholinguistic attributes, the present invention not only achieves an in-depth understanding of the learner's cognitive engagement state, but also effectively improves the calculation accuracy of semantic complexity, as well as the effectiveness and interpretability of automatic evaluation of cognitive engagement. This provides a more reliable tool for research and application in the field of education.
[0135] In addition, the multi-stage processing flow of the present invention combines the advanced BERT model to ensure the semantic relevance of educational concepts and texts, making subsequent analysis more refined. Based on the "Familiarity-Abstractness (FA)" conceptual representation model of psycholinguistic attributes and the cognitive engagement of online discussion texts represented by the ICAP framework, the semantic complexity of online discussion texts and the training model are used to evaluate their cognitive engagement. This method not only improves interpretability and reduces the "black box" effect, but also provides clear quantitative standards for the cognitive processing of educational concepts, which helps to better realize the automated evaluation and analysis of online discussion texts.
[0136] This embodiment also provides a computer program product. When the computer program product runs on a computer, it enables the computer to execute the above-mentioned related steps to implement the text semantic calculation and cognitive investment evaluation method based on conceptual representation provided by the above embodiment.
[0137] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in the embodiments of the present invention are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data need to comply with the relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.
[0138] Those skilled in the art will appreciate that the modules in the devices in the embodiments of the present invention may be adaptively changed and arranged in one or more devices different from the embodiments. The modules or units or components in the embodiments of the present invention may be combined into one module or unit or component, and in addition they may be divided into a plurality of submodules or subunits or subcomponents. Except that at least some of such features and / or processes or units are mutually exclusive, all features disclosed in this specification (including corresponding claims, abstracts and drawings) and all processes or units of any method or device disclosed in this manner may be combined in any combination. Unless otherwise expressly stated, each feature disclosed in this specification (including corresponding claims, abstracts and drawings) may be replaced by an alternative feature providing the same, equivalent or similar purpose.
[0139] Through the description of the above implementation methods, technical personnel in the relevant field can understand that for the convenience and simplicity of description, only the division of the above-mentioned functional modules is used as an example. In actual applications, the above-mentioned functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above.
[0140] It should be noted that the logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, device or apparatus (such as a computer-based system, a system including a processor, or other system that can fetch instructions from an instruction execution system, device or apparatus and execute instructions), or in combination with these instruction execution systems, devices or apparatuses. For the purposes of this specification, "computer-readable medium" can be any device that can contain, store, communicate, propagate or transmit a program for use by an instruction execution system, device or apparatus, or in combination with these instruction execution systems, devices or apparatuses. More specific examples of computer-readable media (a non-exhaustive list) include the following: an electrical connection portion with one or more wirings (electronic device), a portable computer disk box (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable and programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disk read-only memory (CDROM). In addition, the computer-readable medium may even be paper or other suitable medium on which the program is printed, since the program may be obtained electronically, for example, by optically scanning the paper or other medium and then editing, interpreting or processing in other suitable ways if necessary, and then stored in a computer memory.
[0141] In addition, each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other. In particular, for embodiments such as devices and equipment, since they are basically similar to method embodiments, the relevant parts can refer to the partial description of the method embodiments. The embodiments of the devices and equipment described above are merely schematic, wherein the modules, units, etc. described as separate components may or may not be physically separated, that is, they may be located in one place, or they may be distributed in multiple places, such as nodes of a system network. Specifically, some or all of the modules and units may be selected according to actual needs to achieve the purpose of the above-mentioned embodiment scheme. Those skilled in the art can understand and implement it without paying creative labor.
[0142] It should be understood that the various parts of the present invention can be implemented by hardware, software, firmware or a combination thereof. In the above-mentioned embodiments, a plurality of steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, it can be implemented by any one of the following technologies known in the art or their combination: a discrete logic circuit having a logic gate circuit for implementing a logic function for a data signal, a dedicated integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.
[0143] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "examples", "specific examples", or "some examples" means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representation of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner.
[0144] In addition, the terms "first", "second", etc. used in the embodiments of the present invention are only used for descriptive purposes and should not be understood as indicating or implying relative importance, or implicitly indicating the number of technical features indicated in the present embodiment. Therefore, the features defined by the terms "first", "second", etc. in the embodiments of the present invention can explicitly or implicitly indicate that the embodiment includes at least one of the features. In the description of the present invention, the word "multiple" means at least two or two or more, such as two, three, four, etc., unless otherwise clearly and specifically defined in the embodiments.
[0145] In the embodiments of the present invention, the terms "include", "comprises" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the sentence "includes a ..." does not exclude the existence of other identical elements in the process, method, article or device including the element. In addition, components, features, and elements with the same name in different embodiments of the present invention may have the same meaning or different meanings, and their specific meanings need to be determined by their explanation in the specific embodiment or further combined with the context in the specific embodiment.
[0146] Although embodiments of the present invention have been shown and described above, it is to be understood that the above embodiments are exemplary and cannot be construed as limitations of the present invention, and those of ordinary skill in the art may change, modify, replace and modify the above embodiments within the scope of the present invention. Those skilled in the art will readily come to think of other embodiments of the present invention after considering the specification and practicing the present invention. This application is intended to cover any variations, uses or adaptations of the present invention, which follow the general principles of the present invention and include common knowledge or customary technical means in the art that are not disclosed in the present invention. The specification and embodiments are intended to be exemplary only, and the true scope and spirit of the present invention are indicated by the following claims.
Claims
1. A text semantic computing and cognitive investment evaluation method based on conceptual representation, characterized in that: The following steps are involved: Constructing an education concept dictionary, using the education concept dictionary to match the education concepts in the online discussion text, and obtaining candidate education concepts in the online discussion text; Using the BERT model to calculate the semantic relevance between the online discussion text and the candidate educational concepts, and filtering the candidate educational concepts whose semantic relevance is lower than a preset relevance threshold; constructing an educational concept FA model according to the educational concept dictionary, and obtaining the complexity and abstractness of the candidate educational concept from the educational concept FA model; Calculating the semantic complexity of the online discussion text according to the complexity and abstractness of the candidate educational concepts; Conduct correlation analysis on the cognitive engagement level and semantic complexity of online discussion texts; train models to automatically identify the cognitive engagement level of online discussion texts.
2. The text semantic computing and cognitive input evaluation method based on conceptual representation according to claim 1 is characterized in that: The construction of the educational concept dictionary specifically includes the following steps: Extract several educational concepts from the preset corpus; The educational concepts are stored and organized using a Trie tree data structure to form an educational concept dictionary.
3. The text semantic computing and cognitive input evaluation method based on conceptual representation according to claim 1 is characterized in that: The method of using the education concept dictionary to match the education concepts in the online discussion text to obtain candidate education concepts of the online discussion text specifically includes the following steps: Perform word segmentation on online discussion texts; The edit distance between the online discussion text after word segmentation and each educational concept in the educational concept dictionary is calculated, and the educational concepts whose edit distance is lower than a preset distance threshold are taken as candidate educational concepts for the online discussion text.
4. The text semantic computing and cognitive input evaluation method based on conceptual representation according to claim 1 is characterized in that: The method of using the BERT model to calculate the semantic relevance between the online discussion text and the candidate educational concept specifically includes the following steps: Use the BERT model to construct a vector space of online discussion texts and the candidate educational concepts; The semantic relevance of each candidate educational concept to the online discussion text is calculated using the following formula: Where V P The vector space representing the online discussion text, V Ci Represents the vector space of each candidate educational concept, i∈[1,j], j is the total number of candidate educational concepts.
5. The text semantic computing and cognitive input evaluation method based on conceptual representation according to claim 1 is characterized in that: The filtering of candidate educational concepts whose semantic relevance is lower than a preset relevance threshold specifically includes the following steps: Constructing a candidate educational concept set according to the candidate educational concepts; Determine the mean value and standard deviation of the candidate educational concept set to determine a preset relevance threshold.
6. The text semantic computing and cognitive input evaluation method based on conceptual representation according to claim 1 is characterized in that: The step of constructing an educational concept FA model according to the educational concept dictionary and obtaining the complexity and abstractness of the candidate educational concept from the educational concept FA model specifically includes the following steps: Performing feature vectorization on the educational concepts in the educational concept dictionary to construct a language feature vector of the educational concepts; Performing regression analysis based on the language feature vectors of the plurality of educational concepts, determining the familiarity and abstractness of each educational concept, and establishing an educational concept FA model; The complexity and abstractness of the candidate educational concepts are obtained in the educational concept FA model.
7. The text semantic computing and cognitive input evaluation method based on conceptual representation according to claim 1 is characterized in that: The step of calculating the semantic complexity of the online discussion text according to the complexity and abstractness of the candidate educational concepts specifically includes the following steps: representing the online discussion text as a set of candidate educational concepts; The semantic complexity of online discussion text is calculated using the following formula: Where N represents the total number of candidate educational concepts, i represents each candidate educational concept, i∈[1,N]; W F represents the familiarity weight of the candidate educational concept, F(C i ) represents the familiarity of the candidate educational concept in the educational concept FA model; W A represents the abstractness weight of the candidate educational concept, A(C i ) represents the familiarity of the candidate educational concept in the educational concept FA model.
8. The text semantic computing and cognitive input evaluation method based on conceptual representation according to claim 7 is characterized in that: The familiarity weight and abstractness weight of the candidate educational concepts are calculated by the quasi-Newton method.
9. The text semantic computing and cognitive input evaluation method based on conceptual representation according to claim 1 is characterized in that: The correlation analysis of the cognitive investment level and semantic complexity of the online discussion text specifically includes the following steps: Coding online discussion texts for cognitive engagement; A correlation analysis tool is used to perform a correlation description on the online discussion text to obtain the correlation between the online discussion text and the semantic complexity and cognitive investment level.
10. The text semantic computing and cognitive input evaluation method based on conceptual representation according to claim 1 is characterized in that: The training model automatically identifies the cognitive engagement level of online discussion texts, specifically including the following steps Determine the target machine learning model for training; The training dataset is constructed with semantic complexity as input feature vector and cognitive investment level as model output; The model is trained using the training data set to obtain a machine learning model that can automatically identify the cognitive input level of online discussion texts.
Citation Information
Patent Citations
Cognitive input tracking method based on double features and semi-supervised learning
CN114936647A
Cognitive support quality evaluation method and system for online teaching feedback of teachers
CN116258390A
Interaction analysis method and system based on cognition and social network mining
CN118733994A