An Automatic Construction Method and System for a Mathematics Course Knowledge Graph Based on Teaching Materials
Through textbooks, the math course knowledge graph system is automatically constructed, and the open source knowledge base and deep learning model are used to solve the problem of building high-quality knowledge graphs under low resource conditions, and efficient and automated map construction is realized, reducing manual workload.
Patent Information
- Application Number
- CN202210574577.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-25
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2042-05-25
AI Technical Summary
In low-resource scenarios, how to quickly build a high-quality mathematics course knowledge graph to reduce labor and time costs.
Through the automatic construction system of math course knowledge graph based on textbooks, including preprocessing modules, course term set construction modules, co-occurrence relationship extraction modules, training data generation modules and relationship set construction modules, the course knowledge graph is automatically constructed using open source knowledge bases and deep learning models.
While ensuring the quality of the map, it reduces manual workload and realizes efficient and automatic construction of math course knowledge graphs under low resource conditions.
Smart Images

Figure CN114969365B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of artificial intelligence and knowledge engineering, and particularly to an automatic construction method and system for a knowledge graph of mathematics courses based on textbooks. Background Art
[0002] As a data structure that can reveal the relationships between knowledge, knowledge graphs have been widely used in various fields in recent years. In the field of education, knowledge graphs can clearly display the internal logical structure of knowledge, eliminating the shortcoming of the single linear arrangement of knowledge according to textbooks in traditional teaching, thereby helping teachers teach better and assisting students in grasping the knowledge context.
[0003] However, constructing a knowledge graph with high quality and large scale often requires huge human and time costs. How to balance accuracy and efficiency and quickly construct a high-quality domain knowledge graph in a low-resource scenario is an important challenge in the field of knowledge engineering. Summary of the Invention
[0004] The present invention is proposed in view of the problem of how to achieve efficient automatic construction of a knowledge graph of mathematics courses based on textbooks in a low-resource scenario.
[0005] To solve the above technical problems, the present invention provides the following technical solutions:
[0006] On the one hand, the present invention provides an automatic construction system for a knowledge graph of mathematics courses based on textbooks, which is applied to implement an automatic construction method for a knowledge graph of mathematics courses based on textbooks. The system includes a preprocessing module, a course term set construction module, a co-occurrence relationship extraction module, a training data generation module, a relationship set construction module, and a course knowledge graph construction module:
[0007] Among them, the preprocessing module is used to preprocess the textbook.
[0008] The course term set construction module is used to construct a course term set according to the preprocessed textbook.
[0009] The co-occurrence relationship extraction module is used to extract the co-occurrence relationships between terms in the textbook based on the course term set.
[0010] The training data generation module is used to generate training data for basic relationship categories according to the co-occurrence relationships between terms and the predefined basic relationship categories between terms in mathematics courses.
[0011] A relationship set construction module, which is used to construct a relationship set according to the training data of basic relationship categories and the co-occurrence relationships of the remaining unlabeled basic relationship categories; wherein, the co-occurrence relationships of the remaining unlabeled basic relationship categories are the co-occurrence relationships between terms excluding the co-occurrence relationships of the basic relationship category training data.
[0012] A course knowledge graph construction module, which is used to construct a course knowledge graph according to the relationship set.
[0013] Optionally, the preprocessing module includes a redundant information deletion module, a table of contents hierarchical structure acquisition module, and a sentence segmentation module.
[0014] Optionally, the redundant information deletion module is used to delete examples, exercises, pictures, and tables in the textbook text according to regular expressions.
[0015] The table of contents hierarchical structure acquisition module is used to obtain the hierarchical structure of the table of contents in the textbook according to regular expressions.
[0016] The sentence segmentation module is used to segment the textbook text output by the redundant information deletion module; automatically annotate the segmented results according to the table of contents structure output by the table of contents hierarchical structure acquisition module, and automatically annotate the chapter to which the segmented results belong.
[0017] Optionally, the course term set construction module includes a crawler programming module, a word segmentation module, and a course term set output module.
[0018] Optionally, the crawler programming module is used to crawl the terms related to the course of the knowledge graph to be constructed on the open source knowledge base, and take the union of the terms with the term list in the textbook appendix to obtain a term reference set.
[0019] The word segmentation module is used to preprocess the term reference set output by the crawler programming module, and segment the preprocessed textbook text according to the Chinese word segmentation method.
[0020] The course term set output module is used to determine whether the word segmentation results output by the word segmentation module are terms related to the course of the knowledge graph to be constructed according to the trained deep learning model, count the terms related to the course of the knowledge graph to be constructed, and obtain a course term set according to the statistical results and the terms related to the course of the knowledge graph to be constructed.
[0021] Optionally, the training data generation module includes a relationship type predefined module, a clustering module, and a basic relationship category training data generation module.
[0022] Optionally, the relationship type predefined module is used to predefine multiple basic relationship types between terms in mathematics courses, and define seed data for each basic relationship type.
[0023] The clustering module is used to take the seed data as the clustering centers in the high-dimensional space, map the co-occurrence relationships into the space of the dimension of the clustering centers, select the k co-occurrence relationships closest to the clustering centers, and label the relationship category labels for the k co-occurrence relationships to obtain the labeled data.
[0024] The basic relationship category training data generation module is used to automatically generate the labeled data of the relationship types between terms according to the idea of distant supervision and the labeled data to obtain the basic relationship category training data.
[0025] Optionally, the relationship set construction module includes a data preprocessing module, a vectorization representation module, a sentence vectorization representation module, a relationship vector concatenation module, a training module, and a relationship set module.
[0026] Optionally, the curriculum knowledge graph construction module is further used to construct a curriculum knowledge graph according to the curriculum term set, the relationship set, and the statistical results.
[0027] On the other hand, the present invention provides a method for automatically constructing a mathematics curriculum knowledge graph based on textbooks, which is implemented by a system for automatically constructing a mathematics curriculum knowledge graph based on textbooks. The system includes a preprocessing module, a curriculum term set construction module, a co-occurrence relationship extraction module, a training data generation module, a relationship set construction module, and a curriculum knowledge graph construction module. The method includes:
[0028] S1. Preprocess the textbook based on the preprocessing module.
[0029] S2. Construct a curriculum term set based on the curriculum term set construction module and the preprocessed textbook.
[0030] S3. Extract the co-occurrence relationships between terms in the textbook based on the co-occurrence relationship extraction module and the curriculum term set.
[0031] S4. Generate basic relationship category training data based on the training data generation module, the co-occurrence relationships between terms, and the predefined basic relationship categories between mathematics curriculum terms.
[0032] S5. Construct a relationship set based on the relationship set construction module, the basic relationship category training data, and the co-occurrence relationships of the remaining unlabeled basic relationship categories; wherein, the co-occurrence relationships of the remaining unlabeled basic relationship categories are the co-occurrence relationships between terms excluding the co-occurrence relationships of the basic relationship category training data.
[0033] S6. Construct a curriculum knowledge graph based on the curriculum knowledge graph construction module and the relationship set.
[0034] Optionally, the preprocessing module includes a redundant information deletion module, a directory hierarchy structure acquisition module, and a sentence and clause segmentation module.
[0035] Optionally, a redundant information deletion module is used to delete examples, exercises, pictures, and tables in the textbook text according to regular expressions.
[0036] A table of contents hierarchical structure acquisition module is used to obtain the hierarchical structure of the table of contents in the textbook according to regular expressions.
[0037] A sentence segmentation and clause separation module is used to separate clauses from the textbook text output by the redundant information deletion module; automatically annotate the clause separation results according to the table of contents structure output by the table of contents hierarchical structure acquisition module, and automatically annotate the chapter to which the clause separation results belong.
[0038] Optionally, the course term set construction module includes a crawler program design module, a word segmentation module, and a course term set output module.
[0039] Optionally, the crawler program design module is used to crawl terms related to the course of the knowledge graph to be constructed on the open source knowledge base, and take the union of the terms with the glossary in the textbook appendix to obtain a term reference set.
[0040] The word segmentation module is used to preprocess the term reference set output by the crawler program design module and segment the preprocessed textbook text according to the Chinese word segmentation method.
[0041] The course term set output module is used to determine whether the word segmentation results output by the word segmentation module are terms related to the course of the knowledge graph to be constructed according to the trained deep learning model, count the terms related to the course of the knowledge graph to be constructed, and obtain the course term set according to the statistical results and the terms related to the course of the knowledge graph to be constructed.
[0042] Optionally, the training data generation module includes a relationship type predefined module, a clustering module, and a basic relationship category training data generation module.
[0043] Optionally, the relationship type predefined module is used to predefine multiple basic relationship types between course terms in mathematics courses and define seed data for each basic relationship type.
[0044] The clustering module is used to use the seed data as the clustering center in the high-dimensional space, map the co-occurrence relationship to the space of the dimension of the clustering center, select the k co-occurrence relationships closest to the clustering center, and label the relationship category label for the k co-occurrence relationships to obtain labeled data.
[0045] The basic relationship category training data generation module is used to automatically generate term relationship type labeled data according to the idea of distant supervision and the labeled data to obtain basic relationship category training data.
[0046] Optionally, the relation set construction module includes a data preprocessing module, a vectorization representation module, a sentence vectorization representation module, a relation vector concatenation module, a training module, and a relation set module.
[0047] Optionally, the course knowledge graph construction module is further configured to construct a course knowledge graph according to the course term set, the relation set, and the statistical results.
[0048] On the one hand, an electronic device is provided, which includes a processor and a memory. At least one instruction is stored in the memory, and the at least one instruction is loaded and executed by the processor to implement the above-mentioned automatic construction method of the mathematics course knowledge graph based on textbooks.
[0049] On the one hand, a computer-readable storage medium is provided. At least one instruction is stored in the storage medium, and the at least one instruction is loaded and executed by a processor to implement the above-mentioned automatic construction method of the mathematics course knowledge graph based on textbooks.
[0050] The beneficial effects brought by the technical solutions provided in the embodiments of the present invention at least include:
[0051] In the above solution, the open-source knowledge base is fully utilized to automatically construct a mathematics course knowledge graph in an effective manner, reducing the manual workload while ensuring the quality of the graph. The input of this method is the original textbook text and the term list in the textbook appendix, and the output is a course knowledge graph composed of terms and the relationships between terms. Description of the Drawings
[0052] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings required for the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0053] Figure 1 is a block diagram of the automatic construction system of the mathematics course knowledge graph based on textbooks provided by the embodiments of the present invention;
[0054] Figure 2 is a schematic diagram of the term extraction process provided by the embodiments of the present invention;
[0055] Figure 3 is a schematic diagram of the knowledge graph construction result provided by the embodiments of the present invention;
[0056] Figure 4 is a schematic diagram of the process of the automatic construction method of the mathematics course knowledge graph based on textbooks provided by the embodiments of the present invention;
[0057] Figure 5It is a schematic structural diagram of an electronic device provided by an embodiment of the present invention. Detailed implementation manners
[0058] To make the technical problems, technical solutions and advantages to be solved by the present invention clearer, the following will be described in detail with reference to the accompanying drawings and specific embodiments.
[0059] As Figure 1 shown, an embodiment of the present invention provides a system for automatically constructing a knowledge graph of mathematics courses based on textbooks. The system includes a preprocessing module, a course term set construction module, a co-occurrence relationship extraction module, a training data generation module, a relationship set construction module, and a course knowledge graph construction module:
[0060] Among them, the preprocessing module is used to preprocess textbooks.
[0061] Optionally, the preprocessing module includes a redundant information deletion module, a directory hierarchical structure acquisition module, and a sentence segmentation and clause annotation module.
[0062] Optionally, the redundant information deletion module is used to delete examples, exercises, pictures, and tables in the textbook text according to regular expressions.
[0063] The directory hierarchical structure acquisition module is used to obtain the hierarchical structure of the table of contents in the textbook according to regular expressions.
[0064] The sentence segmentation and clause annotation module is used to segment the textbook text output by the redundant information deletion module into clauses; automatically annotate the clause results according to the directory structure output by the directory hierarchical structure acquisition module, and automatically annotate the chapter to which the clause results belong.
[0065] In a feasible implementation manner, the steps for the preprocessing module to preprocess the textbook may include:
[0066] S101. Use the redundant information deletion module to delete redundant information. Since there are many examples, exercises, and table pictures in mathematics courses, and these information contain less key information, first use regular expressions to delete examples, exercises, pictures, tables, etc. in the textbook text.
[0067] S102. The directory hierarchical structure acquisition module uses regular expressions to obtain the hierarchical structure of the table of contents from the textbook, cuts the textbook according to the table of contents, and outputs the results separately to multiple files.
[0068] S103. The sentence segmentation and clause annotation module segments the textbook text processed in S101 into clauses: first, segment the text according to Chinese punctuation marks such as full stops and question marks, and set the maximum length L max of the sentence and the minimum length L min . In this embodiment, L max is set to 128, Lmin Set to 16. If the length of a clause exceeds L max then further divide the sentence according to semicolons, colons, etc. in the sentence. If the length of the divided clause is less than L min then merge the two clauses. Automatically label the chapter to which the clause result belongs according to the directory structure obtained in S102, and output the clause result and the belonging chapter to a file.
[0069] A course term set construction module, which is used to construct a course term set according to the preprocessed teaching materials.
[0070] Optionally, the course term set construction module includes a crawler programming module, a word segmentation module, and a course term set output module.
[0071] Optionally, the crawler programming module is used to crawl terms related to the course of the knowledge graph to be constructed on the open source knowledge base, and take the union of the terms with the glossary in the appendix of the teaching materials to obtain a term reference set.
[0072] The word segmentation module is used to preprocess the term reference set output by the crawler programming module, and segment the preprocessed teaching material text according to the Chinese word segmentation method.
[0073] The course term set output module is used to judge whether the word segmentation result output by the word segmentation module is a term related to the course of the knowledge graph to be constructed according to the trained deep learning model, count the terms related to the course of the knowledge graph to be constructed, and obtain the course term set according to the statistical results and the terms related to the course of the knowledge graph to be constructed.
[0074] In a feasible implementation, as Figure 2 shown, the course term set construction module starts from the teaching materials and constructs the course term set by using web crawler technology, Chinese word segmentation technology and deep learning model; the steps of this part can include:
[0075] S201. Design a crawler program through the crawler programming module, crawl terms that may be related to this course on the open source knowledge base, and take the union with the glossary in the appendix of the teaching materials to obtain a term reference set for this course field.
[0076] S202. Based on the term reference set for this course field constructed in S201, the word segmentation module uses the Chinese word segmentation method to segment the preprocessed teaching material text. Using the term reference set for the course field as a word segmentation dictionary can solve part of the ambiguity problem and long-tail problem, and output the word segmentation result to a file.
[0077] S203. The course term set output module uses the term reference set for this course field constructed in S201 to train a deep learning model to judge whether a certain word is a term of this course.
[0078] Specifically, word2vec can be used to vectorize the input words. After average pooling, a fully connected layer and a softmax classifier are used for classification. The prediction results and the true results are trained using the cross-entropy loss function, and the model is persisted to the hard disk. Traverse each word segmentation result obtained in step S202, determine whether it is a term of this course through the model, and at the same time count the full-text word frequency of the term, the chapter where it first appears, and the chapter where it appears the most. Output the statistical results and terms to a file to obtain the course term set.
[0079] The co-occurrence relationship extraction module is used to extract the co-occurrence relationships between terms in the teaching material based on the course term set.
[0080] In a feasible implementation, based on the processed course term set, extract the co-occurrence relationships between terms in the teaching material; that is, if any two terms appear in the same clause in the teaching material, it is determined that these two terms have a co-occurrence relationship, and output all co-occurrence relationships to a file.
[0081] The training data generation module is used to generate basic relationship category training data according to the co-occurrence relationships between terms and the predefined basic relationship categories between terms in mathematics courses.
[0082] Optionally, the training data generation module includes a predefined module for basic relationship types between terms in mathematics courses, a clustering module, and a basic relationship category training data generation module.
[0083] Optionally, the relationship type predefined module is used to predefine multiple basic relationship types between terms in mathematics courses and define seed data for each basic relationship type.
[0084] The clustering module is used to use the seed data as the clustering centers in the high-dimensional space, map the co-occurrence relationships to the space of the dimension of the clustering centers, select the k co-occurrence relationships closest to the clustering centers, and label the relationship category labels for the k co-occurrence relationships to obtain labeled data.
[0085] The basic relationship category training data generation module is used to automatically generate term relationship type labeled data according to the idea of distant supervision and the labeled data to obtain basic relationship category training data.
[0086] In a feasible implementation, the training data generation module is based on the predefined basic relationship categories between terms in mathematics courses and seed samples, and combines clustering and distant supervision methods to automatically label and generate basic relationship category training data in the co-occurrence relationships between terms; the steps in this part can include:
[0087] S401. Relationship type predefined module: Due to the characteristics of rigorous specification and strong logic in mathematics curriculum textbooks, we have predefined seven logically basic relationship types among mathematics curriculum terms and defined some seed data for each relationship type in advance. The following are each relationship type and a corresponding example:
[0088] (i) Synonymy:
[0089] If A takes the value of true under all its assignments, then A is called <e1>tautology< / e1> or <e2>tautology< / e2> ;
[0090] (ii) Coordination:
[0091] In any directed graph, the sum of all nodes <e1>in-degree< / e1> is equal to the sum of all nodes <e2>out-degree< / e2> sum.
[0092] (iii) Positive derivation:
[0093] with a truth value of true <e1>proposition< / e1> is called <e2>true proposition< / e2> .
[0094] (iv) Negative derivation:
[0095] <e1>undirected graph< / e1> and the directed graph are collectively called <e2>graph< / e2> .
[0096] (v) Causality:
[0097] Because there is <e1>true (false) assignment< / e1> , so this formula is <e2>satisfiable formula< / e2> .
[0098] (vi) Irrelevance:
[0099] A proposition always has a definite true <e1>or< / e1> false "value", which is called <e2>truth value< / e2> .
[0100] (vii) Other:
[0101] Generally, in the whole individual domain, for <e1>universal quantifier< / e1> , the characteristic predicate is often used as the antecedent of an implication <e2>antecedent< / e2> .
[0102] S402. The clustering module selects a clustering method, such as K-Means (K-Means clustering algorithm), and designates the seed data in S401 as the clustering center C of each relationship category in the high-dimensional space. Here, the high-dimensional space refers to the space where the set of n-dimensional vectors obtained by vectorizing all Chinese texts is located. Map the co-occurrence relationships obtained by the co-occurrence relationship extraction module into the high-dimensional space of the same dimension, and select the k co-occurrence relationships closest to C and label them with relationship category tags. In this embodiment, k is set to 10.
[0103] S403. Based on the idea of distant supervision, the basic relationship category training data generation module assumes that if there is a certain relationship between two terms in the already labeled data, then the clauses containing these two data can all represent this relationship. Based on the data labeled in S402, automatically generate a certain amount of labeled data of the relationship types between terms as training data.
[0104] In this embodiment, a total of 2388 training data are obtained through the above method.
[0105] The relationship set construction module is used to construct a relationship set according to the basic relationship category training data and the co-occurrence relationships of the remaining unlabeled basic relationship categories.
[0106] Among them, the co-occurrence relationships of the remaining unlabeled basic relationship categories are the co-occurrence relationships between terms excluding the co-occurrence relationships in the basic relationship category training data.
[0107] Optionally, the relationship set construction module includes a data preprocessing module, a vectorization representation module, a sentence vectorization representation module, a relationship vector splicing module, a training module, and a relationship set module.
[0108] In a feasible implementation manner, the relationship set construction module trains a deep learning model to predict the remaining unlabeled co-occurrence relationships; the steps in this part may include:
[0109] S501. The data preprocessing module performs data preprocessing, adds the start flag [CLS] and the end flag [SEP] at the beginning and end of the sentence, and uses # and $ to mark the head entity (Entity1) and the tail entity (Entity2).
[0110] S502. The vectorization representation module uses the Bert (Bidirectional Encoder Representations from Transformer) model to extract the vectorization representation of the text.
[0111] S503. The sentence vector representation module performs average pooling on the vectors corresponding to Entity1 and Entity2 and uses a fully connected layer to obtain the vector representation of the entity. For the special token [CLS] at the beginning of the sentence, a fully connected layer is used to obtain the vector representation of the sentence.
[0112] S504. The relation vector concatenation module concatenates the entity vector and the sentence vector to obtain the relation vector.
[0113] S505. The training module uses a softmax classifier for classification and trains the prediction result and the true result using the cross-entropy loss function.
[0114] S506. The relation set module persists the trained relation classification model to the hard disk and uses the saved model to predict the relation categories of unlabeled data to obtain the relation set.
[0115] The curriculum knowledge graph construction module is used to construct a curriculum knowledge graph according to the relation set.
[0116] Optionally, the curriculum knowledge graph construction module is further used to construct a curriculum knowledge graph according to the curriculum term set, the relation set, and the statistical results.
[0117] In a feasible implementation manner, the curriculum knowledge graph construction module constructs a curriculum knowledge graph, uses the term set as the nodes of the graph, and the relation set as the edges of the graph; and uses the full-text word frequency, the first occurrence chapter, and the most-occurring chapter statistically in S203 as the attributes of the terms; uses the chapter and the specific clauses to which each relation belongs as the attributes of the edges to construct the curriculum knowledge graph.
[0118] Specifically, as Figure 3 shown, the automatic construction of the mathematics curriculum knowledge graph is carried out by using web crawler technology, Chinese word segmentation technology, and deep learning-based relation extraction technology. Specifically, the teaching materials are preprocessed, including deleting redundant information, obtaining the directory hierarchy structure, and clause segmentation; starting from the teaching materials, the curriculum term set is constructed by using web crawler technology, Chinese word segmentation technology, and deep learning models; based on the processed curriculum term set, the co-occurrence relationships between terms in the teaching materials are extracted; based on the predefined basic relation categories and seed samples between mathematics curriculum terms, combined with clustering and distant supervision methods, in the co-occurrence relationships between terms, the basic relation category training data are automatically labeled and generated; the deep learning model is trained to predict the remaining unlabeled co-occurrence relationships; the curriculum knowledge graph is constructed, using the term set as the nodes of the graph, the relation set as the edges of the graph, and adding the corresponding attributes.
[0119] In the embodiments of the present invention, an open-source knowledge base is fully utilized to automatically construct a knowledge graph for mathematics courses in an effective manner, reducing the manual workload while ensuring the quality of the graph. The input of this method is the original text of the textbook and the glossary in the textbook appendix, and the output is a course knowledge graph composed of terms and the relationships between terms.
[0120] As Figure 4 shown, the embodiments of the present invention provide a method for automatically constructing a knowledge graph for mathematics courses based on textbooks, and this method can be implemented by an electronic device. As Figure 1 shown in the flowchart of the method for automatically constructing a knowledge graph for mathematics courses based on textbooks, the processing flow of this method may include the following steps:
[0121] S1. Preprocess the textbook based on the preprocessing module.
[0122] S2. Construct a course term set based on the course term set construction module and the preprocessed textbook.
[0123] S3. Extract the co-occurrence relationships between terms in the textbook based on the co-occurrence relationship extraction module and the course term set.
[0124] S4. Generate training data for basic relationship categories based on the training data generation module, the co-occurrence relationships between terms, and the predefined basic relationship category types between terms in mathematics courses.
[0125] S5. Construct a relationship set based on the relationship set construction module, the training data for basic relationship categories, and the co-occurrence relationships of the remaining unlabeled basic relationship categories; wherein, the co-occurrence relationships of the remaining unlabeled basic relationship categories are the co-occurrence relationships in the co-occurrence relationships between terms excluding the co-occurrence relationships of the training data for basic relationship categories.
[0126] S6. Construct a course knowledge graph based on the course knowledge graph construction module and the relationship set.
[0127] Optionally, the preprocessing module includes a redundant information deletion module, a table of contents hierarchical structure acquisition module, and a sentence segmentation and clause separation module.
[0128] Optionally, the redundant information deletion module is used to delete examples, exercises, pictures, and tables in the textbook text according to regular expressions.
[0129] The table of contents hierarchical structure acquisition module is used to obtain the hierarchical structure of the table of contents in the textbook according to regular expressions.
[0130] The sentence segmentation and clause separation module is used to separate the clauses of the textbook text output by the redundant information deletion module; automatically annotate the clause separation results according to the table of contents structure output by the table of contents hierarchical structure acquisition module, and automatically annotate the chapter to which the clause separation results belong.
[0131] Optionally, the course term set construction module includes a crawler program design module, a word segmentation module, and a course term set output module.
[0132] Optionally, the crawler program design module is used to crawl the terms related to the course of the knowledge graph to be constructed on the open source knowledge base, and take the union of the terms with the glossary in the textbook appendix to obtain a term reference set.
[0133] The word segmentation module is used to preprocess the term reference set output by the crawler program design module, and segment the preprocessed textbook text according to the Chinese word segmentation method.
[0134] The course term set output module is used to determine whether the word segmentation results output by the word segmentation module are terms related to the course of the knowledge graph to be constructed according to the trained deep learning model, count the terms related to the course of the knowledge graph to be constructed, and obtain a course term set according to the statistical results and the terms related to the course of the knowledge graph to be constructed.
[0135] Optionally, the training data generation module includes a relationship type predefined module, a clustering module, and a basic relationship category training data generation module.
[0136] Optionally, the relationship type predefined module is used to predefine multiple basic relationship types between the terms of mathematics courses, and define seed data for each basic relationship type.
[0137] The clustering module is used to use the seed data as the clustering center in the high-dimensional space, map the co-occurrence relationship to the space of the dimension of the clustering center, select the k co-occurrence relationships closest to the clustering center, and label the relationship category label for the k co-occurrence relationships to obtain labeled data.
[0138] The basic relationship category training data generation module is used to automatically generate the relationship type labeled data between terms according to the idea of distant supervision and the labeled data to obtain the basic relationship category training data.
[0139] Optionally, the relationship set construction module includes a data preprocessing module, a vector representation module, a sentence vector representation module, a relationship vector splicing module, a training module, and a relationship set module.
[0140] Optionally, the course knowledge graph construction module is further used to construct a course knowledge graph according to the course term set, the relationship set, and the statistical results.
[0141] In the embodiment of the present invention, the open source knowledge base is fully utilized to automatically construct a mathematics course knowledge graph in an effective manner, reducing the manual workload while ensuring the quality of the graph. The input of this method is the original textbook text and the glossary in the textbook appendix, and the output is a course knowledge graph composed of terms and the relationships between terms.
[0142] Figure 5 FIG. is a schematic structural diagram of an electronic device 500 provided by an embodiment of the present invention. The electronic device 500 may vary greatly due to different configurations or performances, and may include one or more central processing units (CPUs) 501 and one or more memories 502. Among them, at least one instruction is stored in the memory 502, and the at least one instruction is loaded and executed by the processor 501 to implement the following method for automatically constructing a knowledge graph of mathematics courses based on textbooks:
[0143] S1. Preprocess the textbook based on a preprocessing module.
[0144] S2. Construct a course term set based on a course term set construction module and the preprocessed textbook.
[0145] S3. Extract the co-occurrence relationships between terms in the textbook based on a co-occurrence relationship extraction module and the course term set.
[0146] S4. Generate training data for basic relationship categories based on a training data generation module, the co-occurrence relationships between terms, and predefined basic relationship categories between terms in mathematics courses.
[0147] S5. Construct a relationship set based on a relationship set construction module, the training data for basic relationship categories, and the co-occurrence relationships of the remaining unlabeled basic relationship categories; wherein, the co-occurrence relationships of the remaining unlabeled basic relationship categories are the co-occurrence relationships in the co-occurrence relationships between terms excluding the co-occurrence relationships of the training data for basic relationship categories.
[0148] S6. Construct a course knowledge graph based on a course knowledge graph construction module and the relationship set.
[0149] In an exemplary embodiment, a computer-readable storage medium is also provided, such as a memory including instructions, and the above instructions can be executed by a processor in a terminal to complete the above method for automatically constructing a knowledge graph of mathematics courses based on textbooks. For example, the computer-readable storage medium may be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device, etc.
[0150] Those of ordinary skill in the art can understand that all or part of the steps of implementing the above embodiments can be completed by hardware, or can be completed by a program instructing relevant hardware. The described program can be stored in a computer-readable storage medium, and the above-mentioned storage medium can be a read-only memory, a magnetic disk, or an optical disc, etc.
[0151] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. An automatic construction system for a knowledge graph of mathematics courses based on textbooks, characterized in that, The system includes a preprocessing module, a course term set construction module, a co-occurrence relationship extraction module, a training data generation module, a relationship set construction module, and a course knowledge graph construction module: Among them, the preprocessing module is used to preprocess teaching materials; The course term set construction module is used to construct a course term set according to the preprocessed teaching materials; The co-occurrence relationship extraction module is used to extract the co-occurrence relationships between terms in the teaching materials based on the course term set; The training data generation module is used to generate basic relationship category training data according to the co-occurrence relationships between the terms and the predefined basic relationship categories between terms in mathematics courses; The relationship set construction module is used to construct a relationship set according to the basic relationship category training data and the co-occurrence relationships of the remaining unlabeled basic relationship categories; among them, the co-occurrence relationships of the remaining unlabeled basic relationship categories are the co-occurrence relationships between the terms excluding the co-occurrence relationships of the basic relationship category training data; The course knowledge graph construction module is used to construct a course knowledge graph according to the relationship set; the preprocessing module includes a redundant information deletion module, a table of contents hierarchical structure acquisition module, and a sentence and clause segmentation module; The course term set construction module includes a crawler programming module, a word segmentation module, and a course term set output module; The training data generation module includes a relationship type predefined module, a clustering module, and a basic relationship category training data generation module; The relationship type predefined module is used to predefine multiple basic relationship types between terms in mathematics courses and define seed data for each basic relationship type; among them, the multiple basic relationship types include: synonymy, juxtaposition, positive derivation, negative derivation, causality, irrelevance, and others; The clustering module is used to use the seed data as the clustering centers in the high-dimensional space, map the co-occurrence relationships into the space of the dimension of the clustering centers, select the k co-occurrence relationships closest to the clustering centers, and label the relationship category labels for the k co-occurrence relationships to obtain labeled data; The basic relationship category training data generation module is used to automatically generate term relationship type labeled data according to the idea of distant supervision and the labeled data to obtain basic relationship category training data; The relationship set construction module includes a data preprocessing module, a vector representation module, a sentence vector representation module, a relationship vector splicing module, a training module, and a relationship set module; The data preprocessing module is used to add the start flag [CLS] and the end flag [SEP] at the beginning and end of the sentence, and use # and $ to mark the head entity and the tail entity; The vector representation module is used to extract the vector representation of the text using the Bert model; The sentence vector representation module is used to perform average pooling on the vectors corresponding to the head entity and the tail entity and use a fully connected layer to obtain the vector representation of the entity, and use a fully connected layer for the special marker [CLS] at the beginning of the sentence to obtain the vector representation of the sentence; The relationship vector splicing module is used to splice the entity vector and the sentence vector to obtain a relationship vector; The training module is used to perform classification using a softmax classifier and train the prediction results and the true results using a cross-entropy loss function; The relationship set module is used to persist the trained relationship classification model to the hard disk, and use the saved model to predict the relationship categories of unlabeled data to obtain a relationship set; The course knowledge graph construction module is further used to construct a course knowledge graph according to the course term set, the relationship set, and the statistical results, including: Using the term set as the nodes of the graph and the relationship set as the edges of the graph; and using the statistically calculated full-text word frequency, the chapter where the term first appears, and the chapter where the term appears most frequently as the attributes of the term; using the chapter and the specific clause to which each relationship belongs as the attributes of the edge to construct a course knowledge graph.
2. The system according to claim 1, characterized in that, The redundant information deletion module is used to delete examples, exercises, pictures, and tables in the textbook text according to regular expressions; The table of contents hierarchical structure acquisition module is used to obtain the hierarchical structure of the table of contents in the textbook according to regular expressions; The sentence segmentation module is used to segment the textbook text output by the redundant information deletion module; Automatically label the sentence segmentation results according to the table of contents structure output by the table of contents hierarchical structure acquisition module, and automatically label the chapter to which the sentence segmentation results belong.
3. The system according to claim 1, wherein The crawler program design module is used to crawl the terms related to the course of the knowledge graph to be constructed on the open-source knowledge base, and take the union of the terms and the glossary in the textbook appendix to obtain a term reference set; The word segmentation module is used to preprocess the term reference set output by the crawler program design module, and segment the preprocessed textbook text according to the Chinese word segmentation method; The course term set output module is used to determine whether the word segmentation results output by the word segmentation module are terms related to the course of the knowledge graph to be constructed according to the trained deep learning model, count the terms related to the course of the knowledge graph to be constructed, and obtain a course term set according to the statistical results and the terms related to the course of the knowledge graph to be constructed.
4. A method for automatically constructing a knowledge graph of mathematics courses based on textbooks, characterized in that, The method is implemented by a system for automatically constructing a mathematics course knowledge graph based on textbooks. The system includes a preprocessing module, a course term set construction module, a co-occurrence relationship extraction module, a training data generation module, a relationship set construction module, and a course knowledge graph construction module: The method includes: S1. Preprocess the textbook based on the preprocessing module; S2. Construct a course term set based on the course term set construction module and the preprocessed textbook; S3. Extract the co-occurrence relationships between terms in the textbook based on the co-occurrence relationship extraction module and the course term set; S4. Generate basic relationship category training data based on the training data generation module, the co-occurrence relationships between terms, and the predefined basic relationship categories between terms in mathematics courses; S5. Construct a relationship set based on the relationship set construction module, the basic relationship category training data, and the co-occurrence relationships of the remaining unlabeled basic relationship categories; where the co-occurrence relationships of the remaining unlabeled basic relationship categories are the co-occurrence relationships between terms excluding the co-occurrence relationships of the basic relationship category training data; S6. Construct a curriculum knowledge graph based on the curriculum knowledge graph construction module and the relationship set; The preprocessing module includes a redundant information deletion module, a directory hierarchical structure acquisition module, and a sentence segmentation and clause separation module; The curriculum term set construction module includes a crawler program design module, a word segmentation module, and a curriculum term set output module; The training data generation module includes a relationship type predefined module, a clustering module, and a basic relationship category training data generation module; The relationship type predefined module is used to predefine multiple basic relationship types among mathematics curriculum terms and define seed data for each basic relationship type; among them, the multiple basic relationship types include: synonymy, juxtaposition, positive derivation, negative derivation, causality, irrelevance, and others; The clustering module is used to use the seed data as the clustering center in the high-dimensional space, map the co-occurrence relationship into the space of the dimension of the clustering center, select the k co-occurrence relationships closest to the clustering center, and label the relationship category labels for the k co-occurrence relationships to obtain labeled data; The basic relationship category training data generation module is used to automatically generate relationship type labeled data between terms according to the idea of distant supervision and the labeled data to obtain basic relationship category training data; The relationship set construction module includes a data preprocessing module, a vectorized representation module, a sentence vectorized representation module, a relationship vector concatenation module, a training module, and a relationship set module; The data preprocessing module is used to add the start flag [CLS] and the end flag [SEP] at the beginning and end of the sentence, and use # and $ to mark the head entity and the tail entity; The vectorized representation module is used to extract the vectorized representation of the text using the Bert model; The sentence vectorized representation module is used to perform average pooling on the vectors corresponding to the head entity and the tail entity and use a fully connected layer to obtain the vectorized representation of the entity, and use a fully connected layer for the special marker [CLS] at the beginning of the sentence to obtain the vectorized representation of the sentence; The relationship vector concatenation module is used to concatenate the entity vector and the sentence vector to obtain a relationship vector; The training module is used to perform classification using a softmax classifier and train the prediction result and the true result using a cross-entropy loss function; The relationship set module is used to persist the trained relationship classification model to the hard disk and use the saved model to predict the relationship category of the unlabeled data to obtain a relationship set; The curriculum knowledge graph construction module is further used to construct a curriculum knowledge graph according to the curriculum term set, the relationship set, and the statistical results, including: Use the term set as the nodes of the graph and the relationship set as the edges of the graph; and use the statistically calculated full-text word frequency, the first occurrence chapter, and the most-occurring chapter as the attributes of the terms; use the chapter and the specific clause to which each relationship belongs as the attributes of the edges to construct a curriculum knowledge graph.
Citation Information
Patent Citations
Relation extraction method and system of medical health domain knowledge map
CN109145120A
Knowledge graph construction method for mathematical tutoring question-answering system, and system thereof
CN111475629A