Course knowledge graph construction system based on digitization
Through a digital course knowledge graph construction system, combined with automation and manual scoring, the priorities of knowledge points are dynamically adjusted, and the problems of low efficiency and insufficient accuracy of course knowledge graph construction in the existing technology are solved, achieving efficient and accurate knowledge graph construction and learning efficiency improvement.
Patent Information
- Application Number
- CN202510655634.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-21
- Publication Date
- 2025-08-12
AI Technical Summary
The existing curriculum knowledge graph construction methods are inefficient and prone to errors. Automatic construction cannot guarantee quality and accuracy, resulting in large errors in the scoring results and cannot meet the efficiency and accuracy requirements of educational resource processing.
A digital-based course knowledge graph construction system is adopted, including data collection, data processing, data analysis and analysis construction modules. Through the document frequency and weight of knowledge points, weighted results are formed by combining expert scores, and the priorities of knowledge points are dynamically adjusted to form a directed weighted graph and visualize it.
It realizes efficient and automated extraction of important knowledge points in the course, and optimizes automatic scoring with manual scoring to ensure the accuracy and reliability of the knowledge graph, adapt to the personalized needs of different learners, and improves learning efficiency and the integrity of the knowledge graph.
Smart Images

Figure CN120471148A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of knowledge graph technology, and specifically to a system for constructing a digital course knowledge graph. Background Art
[0002] As a data structure that can reveal the relationship between knowledge, knowledge graphs have been widely used in various fields in recent years. In the process of education, knowledge graphs can clearly display the internal logical structure of knowledge, thereby to a certain extent eliminating the shortcomings of traditional teaching in which knowledge is arranged in a single linear manner according to textbooks, helping teachers to teach better and assisting students to grasp the context of knowledge. For courses, which involve more complex and tedious knowledge points, an efficiently cleaned knowledge graph is particularly important.
[0003] However, the current existing graph construction methods include manual construction and automated construction. Manual construction may generally require the subject knowledge in the education field to be manually extracted according to rules or purely manually, so the operation efficiency is low. Especially when a large amount of educational resources need to be processed, manual construction may not only be time-consuming but also prone to human errors. Although automated construction reduces complicated operations, it may not guarantee the quality and accuracy of the constructed educational knowledge graph, and may result in a superficial understanding of knowledge points. That is, automatic scoring may ignore some special concepts or professional terms in the course, resulting in large errors in the scoring results. Therefore, the existing graph construction method may not be effective and efficient. Summary of the Invention
[0004] The purpose of the present invention is to provide a digital-based course knowledge graph construction system to solve the problems raised in the above-mentioned background technology.
[0005] To achieve the above objectives, the present invention provides the following technical solutions: a system for constructing a digital course knowledge graph, comprising:
[0006] Data collection module: collects text content in the course through the data collection module and stores the text content in the database;
[0007] Data processing module: The text content in the database is input into the data processing module, which first cleans the data and then extracts the course knowledge points Ki from the cleaned text content;
[0008] Data analysis module: The knowledge points Ki of the course are input into the data analysis module, which outputs the preliminary score of the knowledge point Ki, the comprehensive evaluation value of the knowledge point Ki, and the priority of the knowledge point Ki;
[0009] Analysis and construction module: The preliminary score of the knowledge point Ki, the comprehensive evaluation value of the knowledge point Ki and the priority of the knowledge point ki are input into the analysis and construction module. The analysis and construction module connects each knowledge point Ki and its priority to form a directed weighted graph, and then visualizes the directed weighted graph. Finally, these knowledge points Ki are stored in a knowledge graph to construct a knowledge graph.
[0010] Optionally, the data analysis module includes: a knowledge review submodule, a score fusion analysis submodule and a knowledge point priority submodule.
[0011] Optionally, the calculation formula of the knowledge review molecule module is as follows:
[0012]
[0013] in:
[0014] SSD ki Refers to the preliminary score of the knowledge point Ki, Ki refers to the knowledge point of the course, i refers to the index of the knowledge point, t refers to the term that appears in the knowledge point, Tk refers to the term set of the knowledge point Ki, SSDT t,ki The frequency of the term t in the knowledge point Ki and its prevalence in the document collection, SSDW t Refers to the weight of term t in the entire course, SSDF t refers to the document frequency of term t, T∈Tk refers to the set of terms related to knowledge point Ki;
[0015] ∑ t∈TK (SSDT t,ki ×SSDW t ) 2 Refers to calculating the product of the prevalence and weight of all relevant terms t and squaring them. The purpose of the sum is to accumulate the impact of each term on the importance of the knowledge point. Squaring makes the impact more significant, especially for low-frequency but important terms;
[0016] 1+∑ t∈TK SSDF t Refers to the degree of influence of document frequency;
[0017] The processing process of the knowledge review molecular module is as follows: the knowledge point Ki of the course is input into the knowledge review molecular module, and the preliminary score SSD of the knowledge point Ki is output based on the term t ki .
[0018] Optionally, the calculation formula of the scoring fusion analysis submodule is as follows:
[0019] SSF ki =[SA×SSD ki +(1-SA×BAki )];
[0020] in:
[0021] SSF ki Refers to the comprehensive evaluation value of knowledge point Ki, SA refers to the weight coefficient of the preliminary score, BA ki Refers to the manual scoring of knowledge points Ki by course education experts;
[0022] The processing process of the score fusion analysis submodule is as follows: the initial score SSD of the knowledge point Ki is converted into ki Input to the scoring fusion analysis submodule, and output the comprehensive evaluation value SSF of the knowledge point Ki based on the weight coefficient SA of the preliminary score ki .
[0023] Optionally, the calculation formula of the knowledge point priority submodule is as follows:
[0024]
[0025] in:
[0026] PPQ ki Refers to the priority of knowledge point ki, PA refers to PS ki,kj The weight factor is: kj refers to the knowledge point similar to ki, Ni refers to the set of kj, PS ki,kj Refers to the similarity between knowledge points ki and kj, PP refers to PS ki,kj The weight factor is two;
[0027] (∑ j∈Ni PS ki,kj ) 2 Reference increases the influence between similar knowledge points by the square of similarity, making the priority differences between similar knowledge points more significant;
[0028] The processing process of the knowledge point priority submodule is as follows: the frequency of term t in knowledge point Ki and its prevalence in the document set SSDT t,ki And the comprehensive evaluation value SSF of knowledge point Ki ki Input to the knowledge point priority submodule, the knowledge point priority submodule outputs the priority PPQ of the knowledge point ki ki .
[0029] Optionally, the similarity PS between the knowledge points ki and kj ki,kj The calculation formula is as follows:
[0030] PS ki,kj =|Tk∩TT|÷|Tk∪TT|;
[0031] Where: Tk and TT are the term sets in knowledge points ki and kj respectively, the intersection represents the common terms, and the union represents all terms;
[0032] PP refers to PS ki,kj The weight factor is two;
[0033] The similarity PS between the knowledge points ki and kj ki,kj The processing process is as follows: Based on the term set T of knowledge point ki and knowledge point kj ki and T kj , output the similarity PS between knowledge points ki and kj through the intersection and union of the two ki,kj .
[0034] Optionally, the data processing module first cleans the input text content to delete irrelevant information, and then uses the Chinese natural language processing tool jieba to extract nouns, verbs and proper nouns in the course from the text content for part-of-speech tagging, and then extracts the knowledge points Ki of the course through TF-IDF technology and text clustering algorithm.
[0035] Optionally, the analysis construction module uses the Neo4j database to store the course knowledge point Ki, the knowledge point kj similar to ki, and the priority PPQki of the knowledge point ki to form a directed weighted graph, and then uses Gephi, GraphXR tools and the Neo4j database to visualize the directed weighted graph to display information such as the network structure and priority of the knowledge points.
[0036] Compared with the prior art, the present invention has the following beneficial effects:
[0037] 1. The present invention outputs the preliminary score SSD of the knowledge point Ki through the knowledge review molecular module ki This sub-module automatically evaluates the importance of each knowledge point in the course resources by processing the terms in the course documents and combining the word weights and document frequencies, providing a preliminary digital and automated scoring basis for the construction of the knowledge graph. The automatic extraction and scoring system can efficiently screen out the most core knowledge points in the course and reflect the relationship between the knowledge points. The high degree of automation in this sub-module saves a lot of manual extraction time, and can effectively process large-scale document data and is suitable for the automated processing of large-scale course content. Automated scoring can provide important input data for the construction of the course knowledge graph. In the early stage of knowledge graph construction, the knowledge review molecular module can efficiently extract highly important knowledge points from large-scale documents and give these knowledge points preliminary scores. The calculation results of this sub-module enable the system to screen out the most relevant parts in the huge knowledge base, laying the foundation for the knowledge graph.
[0038] 2. The present invention outputs the comprehensive evaluation value SSF of the knowledge point Ki through the scoring fusion analysis submodule ki This submodule combines preliminary scoring with manual scoring by experts, and manual scoring plays a supplementary role. By adjusting the weight ratio of automatic scoring and manual scoring, the balance between automation and manual work in the knowledge graph construction process is ensured, the accuracy of scoring is enhanced, and the personalized needs of different learners are adapted to avoid the distortion problems that may occur when relying solely on automatic scoring. The constructed knowledge graph is more accurate through the weighting of expert scoring, and manual scoring compensates for the errors that may be caused by automatic scoring. In courses with strong professional content, manual scoring can effectively improve accuracy. At the same time, the optimization of automatic scoring by the scoring fusion analysis submodule can further increase the weight of knowledge points, ensuring that the core content in the knowledge graph is highlighted.
[0039] 3. The present invention outputs the priority PPQ of the knowledge point ki through the knowledge point priority submodule ki , and then according to the learners' mastery of the knowledge points and the similarity between the knowledge points, the priority of each knowledge point can be dynamically adjusted to effectively improve the learning efficiency and avoid learners wasting time on unnecessary repeated learning. Through multiple sub-modules, it can directly facilitate the construction of the knowledge graph to improve the integrity and accuracy of the knowledge graph construction, and has the advantages of automated construction and manual construction.
[0040] 4. The present invention uses the comprehensive evaluation value SSF of the knowledge point Ki ki To influence the weight of the iterative term t in the entire course SSDW t This iteration adjusts the weight and score of each knowledge point in the knowledge graph. The iterative process not only optimizes the score of the knowledge point, but also adjusts the weight of the term in the entire course based on the comprehensive evaluation value of the knowledge point, making the model construction more accurate and the overall system more efficient. Through iteration, the comprehensive evaluation value of the knowledge point and the weight of the term in the entire course are kept consistent and dynamically adjusted. The comprehensive evaluation value of the knowledge point can be used to iterate the weight of the term in the entire course, thereby optimizing the score of the knowledge point and ensuring the accuracy of the knowledge graph construction. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] Figure 1 A flowchart of the method steps for building a digital course knowledge graph system;
[0042] Figure 2 This is a schematic diagram of the overall structure of the system for building a digital course knowledge graph;
[0043] Figure 3 This is a structural diagram of the data analysis module in the digital course knowledge graph construction system. DETAILED DESCRIPTION
[0044] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0045] This digital-based course knowledge graph construction system is different from existing graph construction systems. Existing graph construction methods include manual construction and automated construction. Manual construction generally requires manual extraction of subject knowledge in the education field according to rules, or pure manual operation is inefficient. Especially when a large amount of educational resources needs to be processed, manual construction is not only time-consuming but also prone to human errors. Although automated construction reduces complicated operations, it cannot guarantee the quality and accuracy of the constructed educational knowledge graph, and may result in a superficial understanding of knowledge points. In other words, automatic scoring may ignore some special concepts or professional terms in the course, resulting in large errors in the scoring results. Therefore, the existing graph construction method may be ineffective and inefficient.
[0046] The module of this graph construction system combines the document frequency of knowledge point terms with the weight of the terms to quickly and automatically extract important knowledge points in the course and score them, providing efficient automated support for the construction of the knowledge graph. The digital automation saves a lot of manual processing time and ensures the accuracy of knowledge point extraction. It can also form a weighted result in combination with expert scoring, thereby effectively solving the errors that may exist in simple automatic scoring, making the final constructed knowledge graph more reliable and avoiding information bias. Therefore, this construction system can achieve a balance between automation and manual scoring, making full use of the advantages of automation in efficiently processing large amounts of data, while optimizing the automatic scoring results through expert manual scoring to ensure the accuracy and reliability of the knowledge graph.
[0047] Example 1: Please refer to Figures 1 to 3 ,This implementation provides a digital-based course knowledge graph construction system, including:
[0048] Data collection module: collects text content in the course through the data collection module and stores the text content in the database;
[0049] Data processing module: The text content in the database is input into the data processing module, which first cleans the data and then extracts the course knowledge points Ki from the cleaned text content;
[0050] Data analysis module: Input the knowledge points Ki of the course into the data analysis module, which outputs the preliminary score of the knowledge point Ki, the comprehensive evaluation value of the knowledge point Ki, and the priority of the knowledge point Ki;
[0051] Analysis and construction module: The preliminary score of the knowledge point Ki, the comprehensive evaluation value of the knowledge point Ki, and the priority of the knowledge point Ki are input into the analysis and construction module. The analysis and construction module connects each knowledge point Ki and its priority to form a directed weighted graph, then visualizes the directed weighted graph, and finally uses these knowledge points Ki to store in a knowledge graph to construct a knowledge graph;
[0052] The data analysis module includes: knowledge review molecular module, scoring fusion analysis submodule and knowledge point priority submodule.
[0053] In this embodiment: This knowledge review molecule module automatically evaluates the importance of each knowledge point in the course resources by performing TF-IDF processing on the terms in the course documents, combining word weights and document frequencies. This provides a preliminary digital and automated scoring basis for the construction of the knowledge graph. Through automatic extraction and scoring, the system can more efficiently screen out the most core knowledge points in the course and reflect the relationships between knowledge points. This submodule has a high degree of automation, saving a lot of manual extraction time; it can effectively process large-scale document data and is suitable for the automated processing of large-scale course content. Automated scoring provides important input data for the construction of the course knowledge graph. In the early stages of knowledge graph construction, the knowledge review molecule module can efficiently extract highly important knowledge points from large-scale documents and assign preliminary scores to these knowledge points. The calculation results enable the system to screen out the most relevant parts in the huge knowledge base, laying the foundation for the knowledge graph.
[0054] The scoring fusion analysis submodule combines preliminary scoring with expert manual scoring, with manual scoring playing a complementary role. By adjusting the weight ratio of automatic scoring and manual scoring, scoring becomes more accurate and reliable, thus ensuring a balance between automation and manual work in the knowledge graph construction process, enhancing scoring accuracy, adapting to the personalized needs of different learners, and avoiding distortion problems that may arise from relying solely on automated scoring. By weighting expert scoring, the constructed knowledge graph becomes more accurate, and manual scoring compensates for the errors that may be caused by automated scoring, especially in highly professional course content. Manual scoring can effectively improve accuracy. At the same time, the optimization of automatic scoring by the scoring fusion analysis submodule can further increase the weight of knowledge points, ensuring that the core content in the knowledge graph is highlighted.
[0055] The knowledge point priority submodule dynamically adjusts the priority of each knowledge point based on the learner's mastery of the knowledge points and the similarity between the knowledge points, thereby effectively improving learning efficiency and avoiding learners wasting time on unnecessary repetitive learning. The above three submodules can directly facilitate the construction of the knowledge graph to improve the completeness and accuracy of the knowledge graph construction, and have the advantages of both automated and manual construction.
[0056] See also Figures 1 to 3 ,The processing process of the knowledge review molecule module is as follows:
[0057]
[0058] in:
[0059] SSD ki Refers to the preliminary score of the knowledge point Ki, with a value range of [0,1];
[0060] Ki refers to the knowledge point of the course, which can be a concept, theme or module in the course. Ki is a specific topic or content block that needs to be mastered during the learning process. Each knowledge point represents a core concept or information unit in the course.
[0061] i refers to the index of the knowledge point;
[0062] t refers to the terms that appear in the knowledge point. They are used to describe and define the core content of the knowledge point ki. The term t is the basic element used to extract and describe knowledge when constructing the knowledge graph. It is usually a concrete representation of the knowledge point.
[0063] Tk refers to the term set of knowledge point Ki. The term set Tk includes all terms and words related to knowledge point Ki. It is a set of related words obtained by segmenting the text of the course.
[0064] Decompose the words in each text into related terms through natural language processing (NLP) techniques such as word segmentation, part-of-speech tagging, and named entity recognition;
[0065] SSDT t,ki Refers to the frequency of term t in knowledge point Ki and its prevalence in the document collection. This value is calculated in the data acquisition module and data processing module based on the text content input using the term frequency (TF) and inverse document frequency (IDF) techniques. The combination of the two is a common technical means for information retrieval and data mining. TF is used to calculate the number of times term t appears in a document divided by the total number of words in the document, and IDF is used to calculate the logarithm of the total number of documents divided by the number of documents containing term t.
[0066] SSDW tRefers to the weight of term t in the entire course;
[0067] SSDF t The document frequency of term t measures the number of documents in which term t appears. Frequently appearing terms may not be highly representative and are therefore assigned a lower weight. This value can be identified using existing technology during the data collection and processing stages and is a well-known prior art approach.
[0068] ∑ t∈TK (SSDT t,ki ×SSDW t ) 2 Refers to calculating the product of the prevalence and weight of all relevant terms t and squaring them. The purpose of the sum is to accumulate the impact of each term on the importance of the knowledge point. Squaring makes the impact more significant, especially for low-frequency but important terms;
[0069] 1+∑ t∈TK SSDF t Refers to the degree of influence of document frequency, plus 1 is to avoid division by zero error, SSDF t The summation of measures the impact of document frequency and is used to reduce the weight of common terms and reduce their contribution to the score;
[0070] The processing process of the knowledge review molecule module is as follows: the knowledge point Ki of the course is input into the knowledge review molecule module, and the preliminary score SSD of the knowledge point Ki is output based on the term t ki .
[0071] In this embodiment: This module provides a preliminary automated digital mechanism for knowledge point scoring, using TF-IDF and SSDW t The analysis gives each knowledge point ki an automatic score SSD ki This score represents the importance or relevance of the knowledge point in the knowledge graph and is a key step in automated construction. It measures the information importance of terms through TF-IDF and combines the term weight to further weight the contribution of each term to the knowledge point and provide a preliminary knowledge point score. This submodule provides a preliminary scoring framework for the automatic construction of the knowledge graph and provides a basis for the subsequent fusion of manual scoring and final scoring. By combining TF-IDF and term weight, this submodule provides a more accurate evaluation method for the scoring of knowledge points and has automated characteristics, which significantly improves efficiency compared to manual construction.
[0072] See also Figures 1 to 3 ,The processing process of the scoring fusion analysis submodule is as follows:
[0073] SSF ki =[SA×SSD ki+(1-SA×BA ki )];
[0074] in:
[0075] SSF ki Refers to the comprehensive evaluation value of the knowledge point Ki, and SA refers to the weight coefficient of the preliminary score;
[0076] BA ki Refers to the manual scoring of knowledge points Ki by course education experts;
[0077] It should be noted that the manual scoring of knowledge point Ki needs further processing. The manual scoring of knowledge point Ki needs to be normalized. Because the manual scoring is subjectively assessed by people, it is usually a discrete value with a limited range of 0-100 points. Through normalization, its value range is consistent with the value range of the preliminary score of knowledge point Ki, that is, [0,1]. The common normalization method is linear normalization, that is, BA ki =(BA ki,p -BA ki,min )÷(BA ki,min -BA ki,min ), the normalized formula is a well-known technical formula, in which BA ki,p The result of manual scoring of knowledge point ki, with a scoring range of 0 to 100 points, BA ki,min and BA ki,min are the minimum and maximum values of manual scoring, such as 0 and 100, so as to ki Normalization is done to avoid the occurrence of different dimensions;
[0078] The processing process of the score fusion analysis submodule is as follows: the initial score SSD of the knowledge point Ki ki Input to the scoring fusion analysis submodule, and output the comprehensive evaluation value SSF of the knowledge point Ki based on the weight coefficient SA of the preliminary score ki .
[0079] In this embodiment: This submodule provides a correction mechanism for automatic scoring by introducing manual scoring. By integrating manual feedback, the submodule can optimize and correct the preliminary scoring of the knowledge point Ki in the knowledge commentary molecular module to make it more in line with teaching practice and educational goals. This submodule embodies the combination of manual and automatic, makes up for the deviations that may occur in the automatic construction process, ensures the quality of knowledge graph construction, and improves its accuracy. The automatic scoring is adjusted through manual feedback to ensure the accuracy and applicability of knowledge point scoring. The introduction of manual scoring allows course experts to correct automatic scoring according to educational goals and actual conditions, so that the final knowledge graph is more in line with educational needs. On the basis of traditional automated construction methods, this submodule introduces manual scoring, creatively realizes the combination of manual and automatic, makes up for the complexity that is difficult for the automated system to handle, and improves the accuracy and applicability of the knowledge graph.
[0080] See also Figures 1 to 3 ,The processing process of the knowledge point priority submodule is as follows:
[0081]
[0082] in:
[0083] PPQ ki Refers to the priority of knowledge point ki, calculates the priority of each knowledge point, and helps recommend learning paths;
[0084] PA refers to PS ki,kj The weight factor of is 1;
[0085] kj refers to knowledge points that are similar to ki. When calculating priorities, the similarity between knowledge points may affect their importance in the learning process;
[0086] Ni refers to the set of kj, that is, the set of other knowledge points associated with knowledge point ki, indicating the similarity between knowledge points. Ni is constructed through semantic analysis and topic modeling such as LDA and Latent Dirichlet Allocation to determine which knowledge points belong to similar themes or sub-topics in the course content. Furthermore, based on learner behavior data, by analyzing learners' learning paths, such as the path and duration of learning on learning platforms or in online courses, it is determined which knowledge points learners frequently study together.
[0087] Furthermore, metrics such as cosine similarity, Jaccard similarity, or Euclidean distance can be used to determine the similarity between two knowledge points. The higher the similarity, the more likely the two knowledge points are to be considered related.
[0088] PS ki,kjRefers to the similarity between knowledge points ki and kj. If two knowledge points have high similarity, they may be close in the knowledge graph, and the learner may master both knowledge points at the same time;
[0089] PP refers to PS ki,kj The weight factor is two;
[0090] PS ki,kj =|Tk∩TT|÷|Tk∪TT|;
[0091] Where: Tk and TT are the term sets in knowledge points ki and kj respectively, the intersection represents the common terms, and the union represents all terms;
[0092] PP refers to PS ki,kj The weight factor is two;
[0093] Similarity PS between knowledge points ki and kj ki,kj The processing process is as follows: Based on the term set T of knowledge point ki and knowledge point kj ki and T kj , output the similarity PS between knowledge points ki and kj through the intersection and union of the two ki,kj ;
[0094] (∑ j∈Ni PS ki,kj ) 2 Reference increases the influence between similar knowledge points by the square of similarity, making the priority differences between similar knowledge points more significant;
[0095] The processing process of the knowledge point priority submodule is as follows: the frequency of term t in knowledge point Ki and its prevalence in the document set SSDT t,ki And the comprehensive evaluation value SSF of knowledge point Ki ki Input to the knowledge point priority submodule, the knowledge point priority submodule outputs the priority PPQ of the knowledge point ki ki .
[0096] In this embodiment: In the traditional teaching model, the teaching content of the course is often fixed, and all students learn the same knowledge points at the same pace. However, different learners have different levels of mastery, interests, and learning needs for the same knowledge points. Therefore, based on the priority calculated by this sub-module, the system can automatically recommend knowledge points suitable for students' current learning situation based on their learning progress, existing knowledge reserves, and the relevance of course content. Through priority sorting, the knowledge points in the knowledge graph can be flexibly adjusted according to the learners' needs and learning progress. Learners no longer need to learn all knowledge points in a fixed order, but can dynamically adjust the learning order based on their own mastery. Through the calculation of knowledge point priorities, students can more efficiently concentrate on learning the most important and difficult content for their current learning stage, reducing the time wasted on unimportant knowledge points, thereby significantly improving learning efficiency. When constructing a course knowledge graph, for each knowledge point, the priority calculated by this sub-module will help these nodes obtain different priorities in the graph. The weights of the knowledge points reflect their relative importance in course learning. In this way, the knowledge graph will be more in line with actual learning needs, rather than just a static knowledge network. For example, if a knowledge point ki has a high similarity with other important knowledge points kj, the system will think that there is a strong correlation between the two knowledge points, and may arrange them to be learned at similar time nodes in the learning path, thereby helping students build a more coherent knowledge system. This submodule gives priority to recommending knowledge points with high relevance based on the learner's current mastery and learning path. Knowledge points with high priority can help students concentrate on mastering the most important and difficult content according to the learner's learning stage and progress. The teaching system can allocate learning resources according to the priority of the knowledge points, ensuring that learners can use their time efficiently in the learning process and not waste it on relatively unimportant knowledge points. According to the calculated priority, the node weights and relationships in the knowledge graph are dynamically adjusted, making the knowledge graph more in line with actual teaching needs and having higher applicability and flexibility.
[0097] The overall construction system has excellent intelligence. Based on the automatic evaluation of learner behavior data, the system can track learners' learning behavior to assess their mastery of knowledge points. For example, the system can record the choices made by learners during the learning process, the course content browsed, the homework submitted, the test questions completed, etc. If learners spend a lot of time on a certain knowledge point and frequently review or study it repeatedly, it indicates that they may not have a solid grasp of this knowledge point. In addition, some course tests can be used to determine whether learners have mastered the corresponding knowledge points.
[0098] In this submodule, PS ki,kjIt reflects the similarity between knowledge points. For knowledge points that have been mastered, the system will calculate the similarity between the knowledge point and other knowledge points. When the learner has mastered certain knowledge points, the system will appropriately lower the priority of other similar knowledge points in the subsequent knowledge graph recommendation. This is indirectly controlled by calculating the impact of similarity on priority. Assuming that the learner has already mastered a certain knowledge point, then the two have a high similarity in the priority calculation of the knowledge point, then the priority will be suppressed to avoid recommending content that has already been mastered. Similarity calculation will play a significant role in priority calculation, ensuring that the priority of the knowledge points that have been mastered is lower, thereby reducing the repeated push of these contents.
[0099] It is worth noting that the comprehensive evaluation value SSF of knowledge point Ki ki Further operations are used to influence the weight of term t in the knowledge review molecule module in the entire course SSDW t , based on the preliminary score SSD of the knowledge point Ki ki , the comprehensive evaluation value SSF of knowledge point Ki ki and the priority PPQ of knowledge point ki ki Continuous optimization, the specific processing process is as follows:
[0100] First up: SSDW t,new =SSDW t,old ×[1+WA×(SSF ki - SSD ki )];
[0101] Second: Set the iteration termination condition:
[0102] Termination condition 1: The number of iterations is 100;
[0103] Termination Condition 2: SSF ki,new -SSF ki,old <0.001
[0104] in:
[0105] SSDW t,new Refers to the weight of term t in the entire course after iteration, SSDW t,old refers to the weight of the term t in the entire course before iteration, and WA refers to the adjustment factor used to control SSDW t Adjustment strength, SSF ki,ne w refers to the comprehensive evaluation value of the knowledge point Ki after iteration, SSF ki,old Refers to the comprehensive evaluation value of the knowledge point Ki before iteration.
[0106] In this embodiment: This iterative form can fine-tune the weight and score of each knowledge point Ki in the knowledge graph. Through iterative feedback, the introduction of manual scoring can correct the deviation in automatic scoring based on the professional judgment of education experts, thereby improving the accuracy of knowledge point scoring. Through the iterative process, we not only optimize the scoring of knowledge points, but also can improve the comprehensive evaluation value SSF of knowledge point Ki. ki To adjust the weight of term t in the whole course SSDW t The recommended learning materials are more targeted. In the course, learners have different foundations. Some students may have a higher degree of mastery of knowledge points. Through the iteration system, the weights of knowledge points and terms can be automatically adjusted to strengthen learning for each student's weak links, thereby providing a personalized learning path. Through the feedback mechanism, the scoring and term weights are continuously optimized, which can make the recommendation of learning materials more in line with the learner's cognitive structure, and through the iteration of the comprehensive evaluation value SSF of the knowledge point Ki ki and the weight of term t in the whole course SSDW t Maintain consistency and dynamic adjustment between them, through the comprehensive evaluation value SSF of knowledge point Ki ki The weight SSDW of term t in the entire course can be iterated t , thereby optimizing the scoring of knowledge points and ensuring the accuracy of knowledge graph construction. When the comprehensive evaluation value of knowledge point Ki is significantly different from the preliminary score of knowledge point Ki, it means that the model has room for adjustment. For example, if the automatic score of some knowledge points is lower than the manual score, it means that the actual importance of the knowledge point in the course is underestimated. The system will adjust the weight SSDW based on this feedback. t After being dynamically adjusted, SSD ki It will change as the weights are adjusted, and eventually converge to a more accurate score through an iterative process, thereby optimizing the quality of the knowledge graph.
[0107] In the specific implementation process, the multiple sub-modules in this method are used to construct the graph construction system. The knowledge points Ki of the course are input into the knowledge review molecular module, and the preliminary score SSD of the knowledge point Ki is output based on the term t. kiThe knowledge review molecule module automatically evaluates the importance of each knowledge point in the course resources by performing TF-IDF processing on the terms in the course documents and combining the word weight and document frequency. This provides a preliminary digital and automated scoring basis for the construction of the knowledge graph. The automatic extraction and scoring system can efficiently screen out the most core knowledge points in the course and reflect the relationship between the knowledge points. The high degree of automation of this submodule saves a lot of manual extraction time and can effectively process large-scale document data. It is suitable for the automated processing of large-scale course content. Automated scoring provides important input data for the construction of the course knowledge graph. In the early stage of knowledge graph construction, the knowledge review molecule module can efficiently extract knowledge points with high importance from large-scale documents and give these knowledge points a preliminary score. The calculation results of this submodule enable the system to screen out the most relevant parts in the huge knowledge base, laying the foundation for the knowledge graph.
[0108] By taking the initial score SSD of the knowledge point Ki ki Input to the scoring fusion analysis submodule, and output the comprehensive evaluation value SSF of the knowledge point Ki based on the weight coefficient SA of the preliminary score ki The scoring fusion analysis submodule combines preliminary scoring with expert manual scoring. Manual scoring can play a complementary role. By adjusting the weight ratio of automatic scoring and manual scoring, the balance between automation and manual work in the knowledge graph construction process is ensured, the accuracy of scoring is enhanced, and the personalized needs of different learners are adapted to avoid the distortion problem that may occur when relying solely on automatic scoring. The constructed knowledge graph is more accurate through the weighting of expert scoring, and manual scoring makes up for the errors that may be caused by automatic scoring. In courses with strong professional content, manual scoring can effectively improve accuracy. At the same time, the optimization of automatic scoring by the scoring fusion analysis submodule can further increase the weight of knowledge points, ensuring that the core content in the knowledge graph is highlighted.
[0109] By comparing the frequency of term t in knowledge point Ki with its prevalence in the document collection SSDT t,ki And the comprehensive evaluation value SSF of knowledge point Ki ki Input to the knowledge point priority submodule, the knowledge point priority submodule outputs the priority PPQ of the knowledge point ki ki The knowledge point priority submodule dynamically adjusts the priority of each knowledge point according to the learner's mastery of the knowledge points and the similarity between the knowledge points, thereby effectively improving learning efficiency and avoiding learners wasting time on unnecessary repeated learning. Furthermore, through multiple submodules, it can directly facilitate the construction of the knowledge graph, thereby improving the integrity and accuracy of the knowledge graph construction, and has the advantages of both automated and manual construction;
[0110] Comprehensive evaluation value SSF based on knowledge point Ki kiFurther operations are used to influence the weight of term t in the knowledge review molecule module in the entire course SSDW t This iterative form fine-tunes the weight and score of each knowledge point Ki in the knowledge graph. The introduction of manual scoring through iterative feedback can correct the deviation in automatic scoring based on the professional judgment of education experts, thereby improving the accuracy of knowledge point scoring. Through the iterative process, not only the score of the knowledge point is optimized, but also the comprehensive evaluation value SSF of the knowledge point Ki can be improved. ki To adjust the weight of term t in the whole course SSDW t Make the model construction more accurate and the overall system more efficient by iterating the comprehensive evaluation value SSF of the knowledge point Ki ki and the weight of term t in the whole course SSDW t Maintain consistency and dynamic adjustment between them, through the comprehensive evaluation value SSF of knowledge point Ki ki Iterable after the weight of term t in the entire course SSDW t , and then optimize the scoring of knowledge points to ensure the accuracy of knowledge graph construction;
[0111] This allows the various sub-modules to cooperate with each other in calculations, and to perform overall cycles and iterations, so that the overall system has the effect of automatic optimization and updating, and thus better adaptability.
[0112] Example 2: Please refer to Figure 1 、 Figure 2 and Figure 3 ,The data processing module first cleans the input text content to remove irrelevant information, and then uses the Chinese natural language processing tool jieba to extract the nouns, verbs and proper nouns in the course from the text content for part-of-speech tagging.,Afterwards, the knowledge points Ki of the course are extracted through TF-IDF technology and text clustering algorithm;
[0113] The analysis and construction module uses the Neo4j database to store the course knowledge points Ki, the knowledge points kj similar to ki, and the priority PPQki of the knowledge point ki to form a directed weighted graph. Then, Gephi, GraphXR tools and the Neo4j database are used to visualize the directed weighted graph to display information such as the network structure and priority of the knowledge points.
[0114] In this embodiment: the input text content data is extracted by the data processing module to extract the knowledge points ki, and then the knowledge points ki extracted from the course are helpful for the subsequent graph construction. Through the analysis and construction module and based on the calculation and analysis of the data analysis module, a directed weighted graph is constructed according to the relationship between the knowledge points. The nodes represent the knowledge points and the edges represent the relationship between the knowledge points. The constructed directed weighted graph data is then imported into the knowledge graph storage Neo4j database, thereby digitally and intelligently constructing the course knowledge graph.
[0115] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.
Claims
1. A digital course knowledge graph construction system, characterized by: include: Data collection module: collects text content in the course through the data collection module and stores the text content in the database; Data processing module: The text content in the database is input into the data processing module, which first cleans the data and then extracts the course knowledge points Ki from the cleaned text content; Data analysis module: The knowledge points Ki of the course are input into the data analysis module, which outputs the preliminary score of the knowledge point Ki, the comprehensive evaluation value of the knowledge point Ki, and the priority of the knowledge point Ki; Analysis and construction module: The preliminary score of the knowledge point Ki, the comprehensive evaluation value of the knowledge point Ki and the priority of the knowledge point ki are input into the analysis and construction module. The analysis and construction module connects each knowledge point Ki and its priority to form a directed weighted graph, and then visualizes the directed weighted graph. Finally, these knowledge points Ki are stored in a knowledge graph to construct a knowledge graph.
2. The digital course knowledge graph construction system according to claim 1 is characterized by: The data analysis module includes: a knowledge review submodule, a score fusion analysis submodule and a knowledge point priority submodule.
3. The digital course knowledge graph construction system according to claim 2 is characterized by: The calculation formula of the knowledge review molecular module is as follows: in: SSD ki Refers to the preliminary score of the knowledge point Ki, Ki refers to the knowledge point of the course, i refers to the index of the knowledge point, t refers to the term that appears in the knowledge point, Tk refers to the term set of the knowledge point Ki, SSDT t,ki The frequency of the term t in the knowledge point Ki and its prevalence in the document collection, SSDW t Refers to the weight of term t in the entire course, SSDF t refers to the document frequency of term t, T∈Tk refers to the set of terms related to knowledge point Ki; ∑ t∈TK (SSDT t,ki ×SSDW t ) 2 Refers to calculating the product of the prevalence and weight of all relevant terms t and squaring them. The purpose of the sum is to accumulate the impact of each term on the importance of the knowledge point. Squaring makes the impact more significant, especially for low-frequency but important terms; 1+∑ t∈TK SSDF t Refers to the degree of influence of document frequency; The processing process of the knowledge review molecular module is as follows: the knowledge point Ki of the course is input into the knowledge review molecular module, and the preliminary score SSD of the knowledge point Ki is output based on the term t ki .
4. The digital course knowledge graph construction system according to claim 3 is characterized by: The calculation formula of the scoring fusion analysis submodule is as follows: SSF ki =[SA×SSD ki +(1-SA×BA ki )]; in: SSF ki Refers to the comprehensive evaluation value of knowledge point Ki, SA refers to the weight coefficient of the preliminary score, BA ki Refers to the manual scoring of knowledge points Ki by course education experts; The processing process of the score fusion analysis submodule is as follows: the initial score SSD of the knowledge point Ki is converted into ki Input to the scoring fusion analysis submodule, and output the comprehensive evaluation value SSF of the knowledge point Ki based on the weight coefficient SA of the preliminary score ki .
5. The digital course knowledge graph construction system according to claim 4 is characterized by: The calculation formula of the knowledge point priority submodule is as follows: in: PPQ ki Refers to the priority of knowledge point ki, PA refers to PS ki,kj The weight factor is: kj refers to the knowledge point similar to ki, Ni refers to the set of kj, PS ki,kj Refers to the similarity between knowledge points ki and kj, PP refers to PS ki,kj The weight factor is two; (∑ j∈Ni PS ki,kj ) 2 Reference increases the influence between similar knowledge points by the square of similarity, making the priority differences between similar knowledge points more significant; The processing process of the knowledge point priority submodule is as follows: the frequency of term t in knowledge point Ki and its prevalence in the document set SSDT t,ki And the comprehensive evaluation value SSF of knowledge point Ki ki Input to the knowledge point priority submodule, the knowledge point priority submodule outputs the priority PPQ of the knowledge point ki ki .
6. The digital course knowledge graph construction system according to claim 5 is characterized by: The similarity PS between the knowledge points ki and kj ki,kj The calculation formula is as follows: PS ki,kj =|Tk∩TT|÷|Tk∪TT|; Where: Tk and TT are the term sets in knowledge points ki and kj respectively, the intersection represents the common terms, and the union represents all terms; PP refers to PS ki,kj The weight factor is two; The similarity PS between the knowledge points ki and kj ki,kj The processing process is as follows: Based on the term set T of knowledge point ki and knowledge point kj ki and T kj , output the similarity PS between knowledge points ki and kj through the intersection and union of the two ki,kj .
7. The digital course knowledge graph construction system according to claim 1 is characterized by: The data processing module first cleans the input text content to remove irrelevant information, then uses the Chinese natural language processing tool Jieba to extract nouns, verbs, and proper nouns in the course from the text content for part-of-speech tagging, and then extracts the course knowledge points Ki through TF-IDF technology and text clustering algorithm.
8. The digital course knowledge graph construction system according to claim 1 is characterized by: The analysis construction module uses the Neo4j database to store the course knowledge point Ki, the knowledge point kj similar to ki, and the priority PPQki of the knowledge point ki to form a directed weighted graph, and then uses Gephi, GraphXR tools and the Neo4j database to visualize the directed weighted graph to display information such as the network structure and priority of the knowledge points.
Citation Information
Cited By
Archive description and knowledge graph construction method and system based on deep learning
CN121387896A