A personalized learning recommendation method based on a personalized knowledge graph

By using personalized knowledge graphs and reinforcement learning DQN networks, learning data is analyzed in real time and learning content is automatically adjusted, solving the problem of learning content mismatch in existing technologies and improving learning effectiveness and the efficiency of knowledge graphs.

CN119808919BActive Publication Date: 2026-05-29SUZHOU YANTU EDUCATION TECH CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SUZHOU YANTU EDUCATION TECH CO LTD
Filing Date
2024-12-26
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

Existing technologies cannot accurately capture the complex relationships between knowledge points in personalized learning recommendations, resulting in a mismatch between learning content and the user's actual situation. Furthermore, the dynamic updating and maintenance of knowledge graphs are difficult, affecting learning outcomes.

Method used

By constructing a personalized knowledge graph, combining learner profiling models and reinforcement learning DQN networks with knowledge graph technology and learning recommendation algorithms, we can analyze learning data in real time, automatically identify and adjust learning content, and provide personalized learning paths.

Benefits of technology

It improves the matching degree of learning content and learning effectiveness, reduces human intervention, improves the efficiency of knowledge graph construction and updating, and enhances the interpretability of learning paths.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119808919B_ABST
    Figure CN119808919B_ABST
Patent Text Reader

Abstract

A personalized learning recommendation method based on a personalized knowledge graph, including knowledge graph construction, a learner filling in a learning style self-description and learning goal expectation content; and through an entrance test module, a preliminary investigation and evaluation is conducted on the current learning situation of the learner; an initial learner portrait model is established, the learner enters a database for learning, and a learning evaluation module conducts a periodic evaluation on the learning situation of the learner; the learner portrait model is updated, a learning path and learning resources suitable for the learner are generated by using a knowledge graph technology and a reinforcement learning DQN network; a learning path database is updated; and a new learning path is recommended to the learner; the learning data of the student is analyzed to determine the basic situation of the student, the learner portrait model is constructed, the personalized learning characteristics are adapted to the learning knowledge element sequence by using the knowledge graph technology and the learning recommendation algorithm, the personalized learning path recommendation mode is constructed, and the learning effect is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of data processing technology, and specifically relates to a personalized learning recommendation method based on personalized knowledge graphs. Background Technology

[0002] With lifelong learning gaining popularity, most people are constantly improving their abilities through continuous learning. To help users learn better, there are many training courses on the market designed for various exams. However, most of these courses are designed for a broad range of users. Since each user's learning interests, cognitive level, and learning ability differ, their comprehension and learning progress vary, resulting in inconsistent learning outcomes and impacting learning efficiency. This also reflects that the recommended courses or teaching content in such learning models fail to meet the actual needs of users, thus affecting their learning effectiveness.

[0003] Even if some platforms can offer personalized learning courses, their inability to adjust them in a timely manner according to the learning situation of customers at different times will lead to the problem that the teaching content does not match the actual situation of the user, which will seriously affect the learning effect.

[0004] However, in knowledge management systems across various platforms, the extraction and organization of knowledge points often rely on manual operations, resulting in a large workload, low efficiency, and susceptibility to subjective judgment during processing. With the development of artificial intelligence technology, large AI models such as BERT and GPT have made significant progress in natural language processing, enabling automated knowledge extraction. While some models, trained through deep learning algorithms, can understand and process large amounts of text data, identifying entities, concepts, attributes, and relationships within the text, existing technical solutions suffer from the following problems:

[0005] 1. Currently, most knowledge management systems use keyword matching and expert systems to extract and organize knowledge points.

[0006] 2. Keyword matching methods rely on a predefined list of keywords to identify relevant knowledge points by searching for keywords in documents. However, this method often fails to understand the context and easily overlooks important information that does not contain keywords.

[0007] 3. Expert systems rely on the knowledge and rules of domain experts to extract and organize knowledge points through a series of customized rules. However, this method requires a lot of human intervention and is difficult to adapt to constantly changing knowledge systems.

[0008] 4. Another approach is to use traditional relational databases to store knowledge points. However, this approach has limitations in handling complex relationships and dynamically updating knowledge graphs because it cannot effectively represent complex relationships and attributes between entities.

[0009] At the same time, existing technologies still have some limitations in data processing:

[0010] Existing technologies often require significant human intervention when processing large amounts of unstructured data, which not only increases costs but also reduces efficiency.

[0011] Existing technologies struggle to accurately capture the complex relationships between knowledge points, especially in cross-domain and cross-document scenarios. The construction and maintenance of knowledge graphs often require specialized knowledge and skills, limiting their application in a wider range of fields. Existing solutions also perform poorly in dynamically updating knowledge graphs, failing to adapt to rapidly changing information environments. In summary, existing technologies heavily rely on manual labor for knowledge point extraction during data processing, resulting in low efficiency and high costs. Furthermore, the difficulty in accurately capturing the relationships between knowledge points further complicates extraction and processing. Even when knowledge graphs are used in some existing technologies, their dynamic updates and maintenance remain challenging. Therefore, existing technologies suffer from inaccurate knowledge point extraction and time-consuming, difficult-to-maintain knowledge graph construction. Summary of the Invention

[0012] Purpose of the Invention: To overcome the above shortcomings, the purpose of this invention is to provide a personalized learning recommendation method based on personalized knowledge graphs. This method constructs a learner profile model based on the user's current learning situation and set goals, combined with learning ability, learning level, learning interest, and learning style. Learning process data is recorded through a learning process database, and a learning assessment module conducts timely assessments. Learning content is adjusted accordingly based on the assessment results.

[0013] By analyzing the learning data generated by students during their learning process on the platform, we can clarify the basic information of each student and build a learner profile model. By combining subject knowledge graphs, we can use knowledge graph technology and learning recommendation algorithms to match personalized learning features with learning knowledge element sequences, search and optimize the sequence combination of subject knowledge elements, and build a learning push path recommendation mode. This will enable us to provide different personalized learning recommendations to different students and help them improve their learning outcomes.

[0014] Technical Solution: To achieve the above objectives, this invention provides a personalized learning recommendation method based on personalized knowledge graphs, comprising:

[0015] S1): To construct a knowledge graph, the data source is first preprocessed. Then, a dedicated knowledge point extraction and fine-tuning model is used to automatically identify entities in the text and extract the relationships between entities to obtain preprocessed knowledge points and relationships. Next, disambiguation processing is performed on the knowledge points and relationships. A graph database model is established, entity nodes are created, and relationship edges are formed between entity nodes. Finally, the disambiguated entity nodes and relationship edges are stored in the graph database model to form a dynamic and scalable knowledge graph.

[0016] S2): When learners first enter the personalized learning platform through the learner interaction interface, they must first provide basic personal information and complete the registration.

[0017] Afterwards, under the guidance of the personalized learning platform's wizard, learners enter the learning style assessment module and the learning goal setting module to complete their learning style assessment and learning goal setting, including selectively completing the learner's learning style questionnaire, filling in a self-description of learning style and expected content of learning goals; and through the entrance test module, an initial assessment of the learner's current learning situation is conducted.

[0018] S3): Initial learner profile model establishment, that is, the personalized learning recommendation platform builds an initial learner profile model based on the collected user information through the learner profile model establishment module; the learner profile model establishment module processes the collected information and stores it in the user model database;

[0019] S4): Learners enter the database to learn. The learning process recording module records the data of the learner's entire learning process, including the sequence of pages accessed, learning materials, and access time information in real time, and updates the user's learning record database in real time.

[0020] S5): The learning assessment module conducts phased assessments of learners' learning progress. That is, after each knowledge unit is completed, the learning assessment module evaluates the learner's learning progress recorded by the learning process recording module and sends the assessment results to the learner profile model and the personalized learning recommendation engine.

[0021] S6): Update the learner profile model, that is, update the learner profile model according to the assessment results of the learning assessment module, and store the updated learner model in the learner model database.

[0022] S7): Based on the updated learner profile model and learning process records, the personalized learning recommendation engine calculates the learner's mastery of knowledge elements in real time, generates a learner profile accordingly, calculates the user's comprehensive feature vector, and then uses knowledge graph technology and reinforcement learning DQN network to adaptively and self-organize the learning path and learning resources that are suitable for the learner based on the calculated user comprehensive feature vector, so as to provide users with accurate personalized learning guidance services.

[0023] S8): Update the learning path database and store the newly generated learning paths for the learners in S7) into the learning path database;

[0024] S9): Personalized learning content recommendation uses the learning path database to recommend new learning paths to learners;

[0025] S10): Then the learner repeats steps 4) to 9) until all the knowledge points to be learned are completed.

[0026] The personalized learning recommendation method based on personalized knowledge graphs described in this invention, specifically the method for constructing the knowledge graph in step S1) is as follows:

[0027] S11): Data preprocessing, which is to process the data source of the knowledge points to be extracted. The data source of the knowledge points to be extracted includes structured data and unstructured data. Structured data is directly input into the mapper, and the text data in the unstructured data is cleaned.

[0028] S12): Extract knowledge points and relationships from the data source. First, build a dedicated knowledge point extraction fine-tuning model. Input the pre-processed data into the knowledge point extraction fine-tuning model. The knowledge point extraction fine-tuning model automatically identifies entities in the text and extracts the relationships between entities to obtain pre-processed knowledge points and relationships.

[0029] S13): Knowledge point and relationship disambiguation, i.e., the identification and mapping of different descriptions of the same knowledge point; the specific disambiguation process is as follows:

[0030] S131): Establish a standard knowledge point base, that is, first construct a standard set of knowledge point entities. The established standard set of knowledge point entities serves as the target set for mapping and is used to disambiguate the knowledge points extracted from the text.

[0031] S132): Similarity calculation and mapping, that is, using text embedding technology in natural language processing to convert knowledge points and related texts into vectors; then performing similarity calculation, by calculating the cosine similarity or Euclidean distance between the text embedding vectors to determine whether the extracted different expressions point to the same knowledge point. Cosine similarity is used to measure the angle between vectors to determine the similarity between the two.

[0032] S133): Semantic disambiguation based on context: In order to ensure the accuracy of the disambiguation process, the sliding window technique is used to take contextual information into consideration, so that in the process of identifying knowledge points, not only the current sentence is considered, but also the sentences before and after it are referenced, and the relationship between knowledge points is extracted from the preprocessed knowledge points.

[0033] S134): Iterative verification and feedback optimization: After disambiguation, the knowledge points and relationships extracted in the previous step will be mapped to the knowledge point entity standard set, and the two will be matched to find possible corresponding relationships. The extracted relationships will be further verified, and if they do not match, the entity will be deleted.

[0034] S14): Knowledge Graph Construction: First, establish a graph database model, create entity nodes, and create relationship edges between entity nodes. Then, store the disambiguated entity nodes and relationship edges in the graph database model to form a dynamic and scalable knowledge graph. In later use, use an AI model to reanalyze text data regularly and update the knowledge graph to ensure that it reflects the latest knowledge points and their changes in a timely manner.

[0035] The personalized learning recommendation method based on personalized knowledge graphs described in this invention includes the following specific method for data cleaning of textual data in unstructured data in step S11):

[0036] S111): Text cleaning: First, the text data in the unstructured data is cleaned. A deep learning model is used to classify different text regions in the text. By using the context and semantic features of the text, irrelevant content is distinguished from the main text to automatically identify the main text and irrelevant content. Irrelevant content is removed by regular expressions. Noise in the text data is eliminated by deleting redundant spaces, special characters, irrelevant HTML tags, and duplicate text.

[0037] S112): Sentence segmentation: In order to provide a clear context for subsequent knowledge point extraction, the cleaned text needs to be segmented into sentences;

[0038] S113): Duplicate content extraction. This involves extracting duplicate content from text data by detecting duplicate sentences and paragraphs. Specifically:

[0039] To detect duplicate sentences, a text similarity detection algorithm is used to identify and merge similar or duplicate sentences; to detect duplicate paragraphs, text comparison and clustering algorithms are used to detect and merge duplicate paragraphs in longer texts.

[0040] S114): Contextual Preservation. In order to preserve the original context of the text, the following measures are taken:

[0041] Maintain paragraph structure: While dividing the text into sentences, preserve the original paragraph structure so that complete paragraph information can be referenced when extracting knowledge points;

[0042] Using a sliding window: In order to capture a wider range of contextual information, a sliding window technique can be used when extracting knowledge points, which allows the model to consider information from a sentence and the sentences before and after it at the same time.

[0043] The personalized learning recommendation method based on personalized knowledge graphs described in this invention, specifically the sentence segmentation process in step S112), is as follows:

[0044] S1121): Punctuation-based: using periods, question marks, exclamation marks, and other punctuation marks as sentence separators;

[0045] S1122): Based on natural language processing tools, use AI models or natural language processing libraries to segment sentences. These tools can more accurately identify sentence boundaries, even in the absence of obvious punctuation marks.

[0046] S1123): Dependency-based parsing: By analyzing the dependency relationships in a sentence, the main clause and subordinate structure of the sentence can be determined, thereby more accurately dividing the sentence; its core idea is to decompose the sentence into a "dependency relationship" graph between words, thereby determining the predicate verb, subject, object elements in the sentence and their interrelationships.

[0047] The personalized learning recommendation method based on personalized knowledge graphs described in this invention includes the following specific analysis process for dependency parsing in step S1123):

[0048] S11231): Input preprocessing, which involves word segmentation, noise reduction, and normalization of the input text; specifically as follows:

[0049] Text cleaning: First, the input text data needs to be cleaned and preprocessed, specifically by removing noise, symbols, irrelevant HTML tags, and normalizing complex text, such as converting it to lowercase and correcting typos.

[0050] Word segmentation: Before parsing the syntax, word segmentation tools are used to divide the sentence into individual words or phrases;

[0051] S11232): Part-of-speech tagging

[0052] First, natural language processing tools are used to tag the part of speech of each word. That is, a part-of-speech tagger is used to tag each word or phrase after word segmentation to determine the grammatical role of each word or phrase in the sentence, such as noun, verb, or adjective. Part-of-speech tagging provides basic information for subsequent dependency analysis.

[0053] Further simplify the parts of speech: simplify the parts of speech according to the specific application scenario of the words or phrases, such as simplifying the detailed verb form into a single "verb" category;

[0054] S11233): Dependency relation parsing, using a dependency parsing model to generate a dependency relation tree, as detailed below:

[0055] First, sentence component analysis is performed: through dependency parsing, the dependency relationships between words in the sentence are analyzed to obtain a dependency tree. The goal of dependency parsing is to find the subordinate relationship between the "core word" and other words in the sentence.

[0056] Next, identify the core components of the sentence: analyze the predicate verb and its related subject, object, and adverbial components. For example, `subject` depends on `verb`, and `object` also depends on `verb`. These components constitute the basic structure of the sentence.

[0057] Model selection: Choose an appropriate dependency syntax model based on the actual application and construct dependency relations;

[0058] S11234): Core structure extraction, which involves analyzing the subject, verb, object, modifier, and complement in sentences within the text to identify the core components of the sentence;

[0059] S11235): The construction of the dependency tree, using visualization tools to display dependency relationships, allows for a more intuitive analysis of sentence structure, as detailed below:

[0060] Dependency tree structured representation: Each sentence generates a dependency tree. The root node of the tree is the predicate verb of the sentence. Other words are attached to the trunk of the dependency tree in sequence according to their dependencies. Through the dependency tree, it is easy to see the main and subordinate relationships of the components in the sentence, thereby realizing grammatical analysis and sentence reorganization.

[0061] Visualization of dependency trees: Use the `displaCy` tool of `spaCy` or other dependency tree visualization tools to represent the parsed dependency relationships in a graphical form, making it easier to understand the sentence structure more intuitively;

[0062] S11236): Syntax rule validation and optimization

[0063] Dependency tree trunk and subordinate structure identification: By analyzing the dependency tree, the trunk structure of the sentence can be determined, namely the predicate verb and its main modifiers. The subject and object directly depend on the predicate verb, while the subordinate components of attributive and adverbial modifiers depend on the core components.

[0064] Syntactic rule optimization: Based on the needs of the domain, specific syntactic rules are set to correct dependency relations. For example, for complex sentence structures such as compound sentences and coordinate sentences, conjunctions can be processed through dependency syntax to clarify the position and function of conjunctions in the sentence structure and ensure the accuracy of semantic dependency relations.

[0065] The personalized learning recommendation method based on personalized knowledge graph described in this invention uses a context embedding model in the knowledge point and relation disambiguation in S13), which converts the extracted knowledge points and their contexts into high-dimensional vectors, thereby achieving more accurate semantic matching. Compared with traditional word vectors, context embedding can better capture the different meanings of the same word in different contexts and improve the robustness of disambiguation.

[0066] Specifically as follows:

[0067] For the input text t, the context embedding model BERT is used to transform the input text into a high-dimensional vector representation v;

[0068] v t =BERT(t)

[0069] Among them, v t It is the context embedding vector of text t.

[0070] Let cosine similarity calculate the angle between two vectors, defined as:

[0071]

[0072] Here, v1 and v2 are the context embedding vector representations of the two texts;

[0073] In dynamic weighted cosine similarity, considering that different dimensions contribute differently to the similarity, a weight w is assigned to each dimension. i This allows similarity calculations to more flexibly reflect important information in the text:

[0074]

[0075] Where d is the dimension, and d is usually the hidden layer size of the BERT model;

[0076] Among them, w i These are the weights used in weighted cosine similarity. They are dynamically calculated based on context or category, and are used to measure the contribution of different dimensions to the similarity. If two knowledge points are consistent in a certain category, the weight of dimensions related to that category can be increased. i ;

[0077] The w i Based on the following factors

[0078] Term frequency (TF): The frequency with which a word appears in a text, reflecting its importance;

[0079] Inverse Document Frequency (IDF): The rarity of a word in the entire corpus;

[0080] w i =TFi·IDFi

[0081] It can also be calculated using semantic similarity, as follows:

[0082] Suppose we have two sentences s1 and s2, whose semantic vector representations are respectively and

[0083] SBERT uses cosine similarity to measure the similarity between two sentences, s1 and s2:

[0084]

[0085] Compared to traditional sentence representation methods, SBERT can capture more fine-grained semantic information, making the semantic similarity calculation of two sentences more accurate.

[0086] Multi-level similarity:

[0087] To achieve more precise semantic matching, the similarity of knowledge points at multiple levels can be calculated; specifically as follows:

[0088] First, calculate word-level similarity, then calculate sentence-level or document-level similarity, and finally, calculate the weighted average of the similarities at each level.

[0089] total_similarity(x1, x2)

[0090] =α·word_similarity(x1,x2)+β·sentence_similarity(x1,x2)+γ

[0091] ·document_similarity(1, x2)

[0092] Where α, β, and γ are the weights of similarity at different levels, which are adjusted according to the specific application scenario;

[0093] In the processing, if the same entity is represented by multiple expressions, the large model is fine-tuned to determine whether they are the same knowledge point, and then mapping and recognition are performed to obtain the final knowledge point and relationship.

[0094] Similarity calculations are used to map knowledge point entity standard sets.

[0095] It also includes mapping methods, and introduces context-aware semantic matching in the matching process to capture the context of knowledge points in the text; context analysis, combined with the contextual information of surrounding words and sentences, enhances the ability to distinguish synonyms or related knowledge points.

[0096] The personalized learning recommendation method based on personalized knowledge graphs described in this invention, specifically the construction process of the knowledge point extraction and fine-tuning model in step S12) is as follows:

[0097] (1) Data preparation, that is, preparing enough domain-related text data, including the annotation of knowledge points. For different expressions of knowledge points, a standardized set of knowledge points needs to be prepared to facilitate subsequent mapping and verification. The data source can be manually annotated text or existing knowledge bases, such as Wiki or domain datasets annotated by experts.

[0098] (2) Model selection

[0099] Choose a pre-trained large language model (such as BERT, GPT, etc.) as the pre-trained model. This model has been pre-trained with massive amounts of general data and has good language understanding ability. For the task of knowledge point extraction, you can choose a model based on the Transformer architecture. These models perform well in natural language processing tasks.

[0100] (3) Data annotation and cleaning of the text.

[0101] The data is preprocessed, including text cleaning (removing noise and irrelevant content) and sentence segmentation. Then, using annotation tools or writing rules, entities and relationships in the text are annotated. This step can use tools such as SpaCy or NLTK for automatic sentence segmentation and part-of-speech tagging.

[0102] (4) Fine-tune the pre-trained model to obtain the fine-tuned knowledge point extraction model.

[0103] For knowledge point extraction tasks, pre-trained models such as BERT or GPT are fine-tuned to obtain the fine-tuned model. The training process of the pre-trained model is completed through the following stages:

[0104] Entity recognition: Automatically identify entities in text, such as concepts, facts, and principles, using pre-trained models;

[0105] Relationship extraction: Extracting relationships between text entities automatically identified by the pre-trained model, such as inclusion, dependency, and leader;

[0106] Disambiguation: For different descriptions of the same knowledge point, vectorization techniques are used to calculate similarity, mapping similar expressions to the same entity;

[0107] (5) Model Validation and Optimization

[0108] The fine-tuned model is tested on a validation set to evaluate its accuracy in knowledge point recognition and relationship extraction tasks. A similarity algorithm is used to verify the accuracy and consistency of the extracted entities and relationships. If the accuracy and consistency meet the requirements, the model can meet the needs.

[0109] The personalized learning recommendation method based on personalized knowledge graphs described in this invention, in step S2), involves the following process for constructing the learner profile model:

[0110] S21): Data collection, collecting various data of learners, including learning behavior data, learning resource interaction data, and test scores, and using association rule algorithms and data mining techniques to mine user learning process data;

[0111] S22): Preprocess the collected data, including cleaning up outliers and handling missing values ​​to ensure data quality. Through attribute matching and feature extraction, combined with learners' personal attributes and learning styles, integrate various learning profile modules using statistical methods to obtain learners' learning abilities, cognitive levels, and learning goal profiles.

[0112] S23): From the three dimensions of learning ability, cognitive level and learning goals, the learner’s personalized learning characteristics are extracted. The learning feature matrix of the learner is formed by quantifying the feature labels of learning ability, cognitive level and learning goals through AprioriAll. Each row represents a learner and each column corresponds to a feature dimension.

[0113] S24): Based on the learning feature matrix, machine learning algorithms are used to group or classify learners in order to build learner profile models.

[0114] The personalized learning recommendation method based on personalized knowledge graph described in this invention uses a reinforcement learning DQN network in the personalized learning recommendation engine to make decisions through the interaction between the agent and the environment. These decisions need to bring as many rewards as possible. Here, the agent is the personalized learning recommendation engine, the environment is the user's state and knowledge graph, and the reward is the benefit of the learning effect.

[0115] The specific work process is as follows:

[0116] S61): First, the data is processed to obtain user profile features.

[0117] The user profile features consist of a learning ability vector c, a cognitive level vector l, and a learning target vector g. The learning ability vector c, cognitive level vector l, and learning target vector g are calculated separately, and then the user comprehensive feature vector is processed. The specific process is as follows:

[0118] The normalization formulas used below are all:

[0119]

[0120] Where X min For the minimum value, X max The maximum value is used to map the data range to the interval [0, 1].

[0121] The learning ability mentioned includes the dimensions of memory, logical reasoning ability, learning speed, and learning focus:

[0122] The memory ability dimension is measured by the score rate M, which is calculated from the scores of memory-based questions and the total score.

[0123]

[0124] Normalization:

[0125]

[0126] The logical reasoning ability dimension is measured by the score rate L of reasoning-type questions;

[0127]

[0128] Normalization:

[0129]

[0130] Learning speed dimension: This reflects a user's ability to effectively complete learning tasks per unit of time, demonstrating how quickly a student can accept and master new knowledge; it includes average learning time and learning focus.

[0131] Average study time:

[0132]

[0133] Learning speed is inversely proportional to average learning time. kp

[0134]

[0135] Normalization:

[0136]

[0137] Learning focus:

[0138] The learning focus refers to a student's ability to maintain concentration and avoid distraction during the learning process, reflecting the degree of engagement in learning;

[0139] The main indicators of learning focus are as follows:

[0140] Continuous learning period, number of learning interruptions, and other abnormal operations;

[0141] Learning focus index

[0142]

[0143] Average study time

[0144]

[0145] Formula for final score of learning focus

[0146]

[0147] A min and A max These are the minimum and maximum values ​​of the average continuous learning time for users, respectively.

[0148] D min and D max These are the minimum and maximum values ​​of the distraction metric for all users, respectively.

[0149] Where ω1 and ω2 are both weights, with a value of 0.5;

[0150] Learning ability matrix representation:

[0151] C = [M] norm L norm L kp_norm A score ]

[0152] The cognitive levels in the cognitive level vector l are as follows:

[0153]

[0154] Among them, A total M represents the user's total score across all tests within the subject. total is the sum of the maximum scores for all tests within the subject; l is the ratio of the user's total score for all tests within the subject to the sum of the maximum scores for all tests within the subject.

[0155] Normalization:

[0156]

[0157] The learning objective in the learning objective vector g:

[0158] Based on the user's target university and major, knowledge points are selected from the knowledge graph, and then the proficiency level of each knowledge point is calculated based on the target grades.

[0159] Assume the knowledge points obtained based on the target universities and majors are as follows:

[0160] k = [k1 k2 ... k] n ]

[0161] The weight of each knowledge point is:

[0162] ω(k i ), where i is 1...n;

[0163] Weight normalization for each knowledge point

[0164]

[0165] The target total score S is calculated based on the normalized weights. goal Assign weights to each knowledge point and adjust the base values ​​to ensure the weights are greater than 1.

[0166] S target (k i )=ω′(k i )×(S goal )+1

[0167] The matrix of the user's learning target vector g is as follows:

[0168] g = [S target (k1) S target (k2) ... S target (k i )]

[0169] The above features are combined to form the user's comprehensive feature vector.

[0170] F student =[c, l, g]

[0171] Where c is the learning ability vector, l is the cognitive level vector, and g is the learning target vector;

[0172] Using the comprehensive feature vector F student As a clustering feature, students are divided into K groups, and each student is assigned a cluster label G. j (j = 1, 2, ..., K), representing the group to which a student belongs, and assigning cluster labels G to the students. j As a feature, it is incorporated into the state representation of subsequent models;

[0173] S62): Knowledge graph embedding, which involves collecting all knowledge points and constructing a knowledge point set K = k1, k2, ..., k N ;

[0174] The Node2Vec algorithm is used to embed knowledge points from the knowledge graph into a low-dimensional vector space, resulting in the embedding vector e of the i-th knowledge point. i ;

[0175] Then, by adding the user's comprehensive feature vector, the knowledge point embedding vector, and the student's clustering label, the corresponding user state vectors are obtained, as follows:

[0176] The user state vector S, which incorporates the user's comprehensive feature vector, is represented as follows:

[0177] s = [s1, s2, ..., s N c, l, g]

[0178] Where c is the learning ability vector, l is the cognitive level vector, and g is the learning target vector;

[0179] Student clustering labels G were added j and the embedding vector e of knowledge points i ;

[0180] The method for calculating knowledge point vectors is as follows:

[0181]

[0182] in,

[0183] To what extent students have mastered the knowledge point k: i The degree of mastery, ranging from [0,1];

[0184] Learning frequency: i.e., the student's learning of knowledge point k i The number of times;

[0185] For testing performance: that is, students' performance on knowledge point k i Average test score on;

[0186] The most recent learning time interval: that is, the time since the last time knowledge point k was learned. i The time interval since then;

[0187] e i Embedding vectors of knowledge points;

[0188] f i Features of frequent patterns;

[0189] Knowledge point k i The support level is:

[0190]

[0191] Add student cluster labels G j The user state vector S after the change is represented as follows:

[0192] s = [s1, s2, ..., s N c, l, g, G j ]

[0193] If a user has not fully grasped the knowledge points, the following action options are available:

[0194] A′(s)={k i ∈A(s)|g i >0}

[0195] A(s) represents the set of knowledge points that satisfy the prerequisite relationships but are not fully mastered, g i >0 indicates knowledge point k i These are the students' learning goals;

[0196] Reward learners for achieving their learning objectives. The specific reward function is as follows:

[0197] r t =w(k i )×Δs i ×f(c,l)×G(k i )×PM(k i )×AR(k i )×LF(k i )×FC(k i )

[0198] Among them, w(k) i ) represents knowledge point k i Importance weight

[0199] Δs i For students to understand knowledge point k i Improved mastery

[0200] f(c, l) is the adjustment function for students' learning ability and cognitive level, as follows:

[0201] f(c, l) = δ1·c + δ2·l, where δ1 and δ2 are weight parameters with values ​​of 0.7 and 0.3, respectively;

[0202] G(k i ) is the weighting function for the learning objectives.

[0203]

[0204] Among them, g i Let i be the i-th element of the student's learning objective vector, representing the student's understanding of knowledge point k. iThe degree of learning need; γ is a constant greater than 1, which is the reward bonus coefficient for the learning target knowledge point;

[0205] When g i >0 (Knowledge point k) i When G(k) is the student's learning objective, i =γ, the reward is amplified by γ times (usually 1.2-2.0), and the expected reward for submitting the selection of this knowledge point is more likely to be recommended by the personalized learning recommendation engine;

[0206] When g i =0, G(k) i With a reward of 1, the personalized learning recommendation engine will not give special preference to recommending this knowledge point.

[0207] PM(k i The function is a frequent pattern weighting function, which adjusts the reward based on the importance of the knowledge point in the frequent itemset, reflecting the importance of the knowledge point in the learning group, as detailed below:

[0208] PM(k i )=1+λ PM ×Support(k i )

[0209] Where, λ PM To adjust the parameter, the degree of influence on the reward is usually a small positive number, such as 0.1 or 0.2;

[0210] Support(k i ) represents knowledge point k i Support, representing k i Frequency of occurrence in frequent itemsets;

[0211] By considering the frequency of knowledge points appearing in frequent learning patterns, the personalized learning recommendation engine is guided to recommend knowledge points that are commonly learned in the group.

[0212] In the association rule algorithm, AR(k) i The weighting function for association rules is:

[0213] AR(k i )=1+λ AR ×Conf(k i )

[0214] Where, λ AR To adjust the parameters and control the degree of influence of association rules on rewards, the values ​​are usually taken as positive numbers;

[0215] Conf(k i Knowledge Point k iThe highest confidence level in an association rule represents k. i The strength of the connection with previously learned knowledge points;

[0216] Utilizing previously learned knowledge: By considering k i By associating knowledge points with what students have already learned, the model is guided to recommend knowledge points closely related to the learned content, thus promoting coherent learning.

[0217] LF(k i The learning frequency adjustment function is:

[0218]

[0219] Where, λ LF Take w(k) i The reciprocal of the number of knowledge points; the more important the knowledge point, the smaller the reduction in reward.

[0220] As the learning frequency increases, the function value will gradually decrease, reducing its contribution to the reward and preventing students from excessively and repeatedly learning familiar knowledge points.

[0221] FC(k i Forgetting Curve Adjustment Function

[0222]

[0223] Where, λ FC Take a direct proportional function of the knowledge point weights, with a value range of 0.5-1; when Increase the reward, decrease the reward, and prompt the student to review the knowledge point.

[0224] S63): The reinforcement learning DQN network updates its network parameters as follows:

[0225] Input layer: User state vector s

[0226] Output layer: Q-values ​​for each optional action (Q-value is the mathematical expectation of the sum of rewards from all future states);

[0227] The DQN algorithm network training steps are as follows:

[0228] A. State Acquisition: Acquire the current state s t ;

[0229] B. Action selection: Select an action from the set of available actions according to the greedy strategy.

[0230] C. Execution Action: Students learn the knowledge points, update their level of mastery, and obtain a new state. t+1 ;

[0231] D. Reward Calculation: Calculate the immediate reward based on the reward function;

[0232] E. Experience storage: (s) t a t r t s t+1 Stored in the experience replay pool;

[0233] F. Network Update: Randomly draw a small batch of samples (s) from the experience pool. j a j r j s j+1 )

[0234] Train and update the network parameters;

[0235] The real-time recommendation process using a recommendation engine is as follows:

[0236] Status Acquisition: Retrieve the student's current user status S;

[0237] Action selection: Based on the greedy strategy, select the optimal knowledge point from A′(s);

[0238] Knowledge Point Recommendations: Recommending knowledge points to students and providing corresponding learning resources.

[0239] S64): Feedback and updates on learners' progress are collected through user feedback, as detailed below:

[0240] Learning feedback collection: Record students' learning outcomes for recommended knowledge points, and update their mastery level, learning frequency, and recent learning time;

[0241] User Status Update: Update the student's user status S based on learning feedback.

[0242] Reinforcement Learning DQN Network Update: Regularly train the reinforcement learning DQN network, incorporate new learning data, and improve model performance;

[0243] Add learners' feedback to the experience replay pool

[0244] In reinforcement learning, experience samples are typically represented as quadruples:

[0245] (s t a t r t s t+1 )

[0246] To incorporate student feedback into the experience sample, the structure of the experience sample needs to be expanded to include feedback information:

[0247] (s t a t r t′ s t+1 ft )

[0248] r t′ The adjusted instant rewards take student feedback into account.

[0249] f t Students perform action a at time step t t Feedback

[0250] If the student's feedback is positive, provide an immediate reward.

[0251] r t′ =r t +δr

[0252] If the result is negative, reduce the reward.

[0253] r t′ =r t -δr.

[0254] The personalized learning recommendation method based on personalized knowledge graphs described in this invention, specifically the process of constructing the learning path recommendation pattern in step S7) is as follows:

[0255] S601): The learning feature representation is based on the subject knowledge graph. It uses the association rule algorithm to extract subject knowledge elements and related learning resources, which together form the learning elements of the knowledge graph, represented by the triple L=(Ki,Kj,r).

[0256] Where r represents the logical relationship between knowledge element Ki and learning resource Kj;

[0257] Then, knowledge graph technology is used to serialize and label the attributes of subject knowledge elements and related learning resources, thereby characterizing and describing the basic features of learning elements;

[0258] Based on the learners' learning needs, the learning elements of the learning activity process are formed into a learning element sequence (s1, s2, ..., s...). N );

[0259] At this time, the user state S is represented as follows:

[0260] s[=s1 s2 ... s N ]

[0261] S602): Learning path recommendation. Based on the user's current learning state matrix and the priori relationship of knowledge points in the knowledge graph, recommend the next knowledge point that will receive the maximum reward and the corresponding learning materials.

[0262] S603): Learning Path Evaluation and Optimization: During the learning process, learners complete tasks based on the learning paths and resources recommended by personalized learning and submit feedback on their learning experience and effectiveness. The personalized learning recommendation engine collects and analyzes user feedback, combines it with the learner's goal completion status, compares expert paths with recommended paths, and evaluates the accuracy and practicality of the recommendations.

[0263] Based on evaluation results and user feedback, we continuously optimize the personalized learning recommendation engine to improve the effectiveness of personalized recommendations and user satisfaction.

[0264] As can be seen from the above technical solution, the present invention has the following beneficial effects:

[0265] 1. The personalized learning recommendation method based on personalized knowledge graphs described in this invention constructs a learner profile model based on the user's current actual learning situation and set goals, combined with learning ability, learning level, learning interest, and learning style. Learning process data is recorded through a learning process database, and timely assessments are conducted through a learning assessment module. Learning content is adjusted in a timely manner based on the assessment results.

[0266] By analyzing the learning data generated by students during their learning process on the platform, we can clarify the basic information of each student and build a learner profile model. By combining subject knowledge graphs, we can use knowledge graph technology and learning recommendation algorithms to match personalized learning features with learning knowledge element sequences, search and optimize the sequence combination of subject knowledge elements, and build a learning push path recommendation mode. This will enable us to provide different personalized learning recommendations to different students and help them improve their learning outcomes.

[0267] 2. The knowledge graph construction in this invention starts with preprocessing the data to obtain processed data. Then, by using a knowledge point extraction fine-tuning model, knowledge points are automatically extracted to determine entities and relational edges. Various disambiguation methods are used to further identify and optimize the extracted knowledge points, identify and construct complex relationships between knowledge points, effectively improving the efficiency and accuracy of data processing. Finally, the disambiguated data is stored in a graph database model to construct a knowledge graph, realizing the automated construction and updating of the knowledge graph, effectively solving the problems of large workload and time consumption in data processing.

[0268] 3. This invention's knowledge graph provides a visual and structured way to present knowledge and decision-making processes. By embedding the knowledge graph into the framework of deep reinforcement learning, the interpretability of the system's decisions can be enhanced, enabling learners and system administrators to understand the logic and basis behind the recommended learning paths and strategy selections. The knowledge graph can also provide a global explanation of the overall strategy learning process; that is, the system can use the knowledge graph to display the relationships between all knowledge points and explain the recommended strategies for the entire learning path.

[0269] 4. The data preprocessing process effectively improves the accuracy of data processing, removes invalid data, and improves the efficiency of data processing by means of text cleaning, sentence segmentation, duplicate content extraction, and context preservation.

[0270] 5. Extracting knowledge points and relationships from data sources: By constructing a specialized extraction and fine-tuning model, the system accurately identifies entities and the relationships between them, effectively solving the problem of difficulty in identifying the relationships between entities in current data processing.

[0271] 6. This invention utilizes a large AI model combined with vector embedding, semantic similarity calculation, and symbolic reasoning to efficiently identify and disambiguate different representations of the same knowledge point. The dynamic updating of the standard set and the integration of domain knowledge further enhance the accuracy and flexibility of knowledge point disambiguation. This method is particularly effective in constructing and maintaining large-scale knowledge graphs, ensuring their high accuracy and robustness.

[0272] 7. This invention can dynamically adjust the difficulty and pace of learning materials according to the learner's performance, thereby improving learning efficiency and ensuring that learners can master the necessary knowledge and skills.

[0273] 8. In this invention, machine learning can automatically adjust recommendation strategies through algorithms, continuously improving recommendation quality over time; it supports the modeling of nonlinear relationships and can handle complex learning behavior data; it can also automatically identify which factors have the greatest impact on learning effects and optimize recommendations accordingly to better meet user needs.

[0274] 9. The data processing in this invention supports the storage and rapid querying of massive amounts of data, ensuring the response speed of the recommendation system; by analyzing large-scale datasets, potential learning trends and patterns can be discovered. Attached Figure Description

[0275] Figure 1 This is a schematic diagram illustrating the steps of the personalized learning recommendation method based on personalized knowledge graphs as described in this invention.

[0276] Figure 2 This is a framework diagram for constructing the knowledge graph in this invention;

[0277] Figure 3 This is a flowchart for disambiguating knowledge points and relationships in this invention;

[0278] Figure 4 This is a schematic diagram of the learner profile model in this invention;

[0279] Figure 5 This is a framework diagram of the knowledge graph-based personalized learning recommendation platform in this invention;

[0280] Figure 6 This is a schematic diagram of the structure of the knowledge graph-based personalized learning recommendation platform described in this invention;

[0281] Figure 7 This is a flowchart of the reinforcement learning DQN network algorithm in this invention. Detailed Implementation

[0282] The present invention will be further explained below with reference to the accompanying drawings and specific embodiments.

[0283] Example

[0284] like Figure 1 The illustrated personalized learning recommendation method based on a personalized knowledge graph includes:

[0285] S1): To construct a knowledge graph, the data source is first preprocessed. Then, a dedicated knowledge point extraction and fine-tuning model is used to automatically identify entities in the text and extract the relationships between entities to obtain preprocessed knowledge points and relationships. Next, disambiguation processing is performed on the knowledge points and relationships. A graph database model is established, entity nodes are created, and relationship edges are formed between entity nodes. Finally, the disambiguated entity nodes and relationship edges are stored in the graph database model to form a dynamic and scalable knowledge graph.

[0286] S2): When learners first enter the personalized learning platform through the learner interaction interface, they must first provide basic personal information and complete the registration.

[0287] Afterwards, under the guidance of the personalized learning platform's wizard, learners enter the learning style assessment module and the learning goal setting module to complete their learning style assessment and learning goal setting, including selectively completing the learner's learning style questionnaire, filling in a self-description of learning style and expected content of learning goals; and through the entrance test module, an initial assessment of the learner's current learning situation is conducted.

[0288] S3): Initial learner profile model establishment, that is, the personalized learning recommendation platform builds an initial learner profile model based on the collected user information through the learner profile model establishment module; the learner profile model establishment module processes the collected information and stores it in the user model database;

[0289] S4): Learners enter the database to learn. The learning process recording module records the data of the learner's entire learning process, including the sequence of pages accessed, learning materials, and access time information in real time, and updates the user's learning record database in real time.

[0290] S5): The learning assessment module conducts phased assessments of learners' learning progress. That is, after each knowledge unit is completed, the learning assessment module evaluates the learner's learning progress recorded by the learning process recording module and sends the assessment results to the learner profile model and the personalized learning recommendation engine.

[0291] S6): Update the learner profile model, that is, update the learner profile model according to the assessment results of the learning assessment module, and store the updated learner model in the learner model database.

[0292] S7): Based on the updated learner profile model and learning process records, the personalized learning recommendation engine calculates the learner's mastery of knowledge elements in real time, generates a learner profile accordingly, calculates the user's comprehensive feature vector, and then uses knowledge graph technology and reinforcement learning DQN network to adaptively and self-organize the learning path and learning resources that are suitable for the learner based on the calculated user comprehensive feature vector, so as to provide users with accurate personalized learning guidance services.

[0293] S8): Update the learning path database and store the newly generated learning paths for the learners in S7) into the learning path database;

[0294] S9): Personalized learning content recommendation uses the learning path database to recommend new learning paths to learners;

[0295] S10): Then the learner repeats steps 4) to 9) until all the knowledge points to be learned are completed.

[0296] The personalized learning recommendation method based on personalized knowledge graphs described in this embodiment, specifically the method for constructing the knowledge graph in S1) is as follows:

[0297] S11): Data preprocessing, which is to process the data source of the knowledge points to be extracted. The data source of the knowledge points to be extracted includes structured data and unstructured data. Structured data is directly input into the mapper, and the text data in the unstructured data is cleaned.

[0298] S12): Extract knowledge points and relationships from the data source. First, build a dedicated knowledge point extraction fine-tuning model. Input the pre-processed data into the knowledge point extraction fine-tuning model. The knowledge point extraction fine-tuning model automatically identifies entities in the text and extracts the relationships between entities to obtain pre-processed knowledge points and relationships.

[0299] S13): Knowledge point and relationship disambiguation, i.e., the identification and mapping of different descriptions of the same knowledge point; the specific disambiguation process is as follows:

[0300] S131): Establish a standard knowledge point base, that is, first construct a standard set of knowledge point entities. The established standard set of knowledge point entities serves as the target set for mapping and is used to disambiguate the knowledge points extracted from the text.

[0301] S132): Similarity calculation and mapping, that is, using text embedding technology in natural language processing to convert knowledge points and related texts into vectors; then performing similarity calculation, by calculating the cosine similarity or Euclidean distance between the text embedding vectors to determine whether the extracted different expressions point to the same knowledge point. Cosine similarity is used to measure the angle between vectors to determine the similarity between the two.

[0302] S133): Semantic disambiguation based on context: In order to ensure the accuracy of the disambiguation process, the sliding window technique is used to take contextual information into consideration, so that in the process of identifying knowledge points, not only the current sentence is considered, but also the sentences before and after it are referenced, and the relationship between knowledge points is extracted from the preprocessed knowledge points.

[0303] S134): Iterative verification and feedback optimization: After disambiguation, the knowledge points and relationships extracted in the previous step will be mapped to the knowledge point entity standard set, and the two will be matched to find possible corresponding relationships. The extracted relationships will be further verified, and if they do not match, the entity will be deleted.

[0304] S14): Knowledge Graph Construction: First, establish a graph database model, create entity nodes, and create relationship edges between entity nodes. Then, store the disambiguated entity nodes and relationship edges in the graph database model to form a dynamic and scalable knowledge graph. In later use, use an AI model to reanalyze text data regularly and update the knowledge graph to ensure that it reflects the latest knowledge points and their changes in a timely manner.

[0305] The personalized learning recommendation method based on personalized knowledge graphs described in this embodiment includes the following specific method for data cleaning of text data in unstructured data in step S11):

[0306] S111): Text cleaning: First, the text data in the unstructured data is cleaned. A deep learning model is used to classify different text regions in the text. By using the context and semantic features of the text, irrelevant content is distinguished from the main text to automatically identify the main text and irrelevant content. Irrelevant content is removed by regular expressions. Noise in the text data is eliminated by deleting redundant spaces, special characters, irrelevant HTML tags, and duplicate text.

[0307] S112): Sentence segmentation: In order to provide a clear context for subsequent knowledge point extraction, the cleaned text needs to be segmented into sentences;

[0308] S113): Duplicate content extraction. This involves extracting duplicate content from text data by detecting duplicate sentences and paragraphs. Specifically:

[0309] To detect duplicate sentences, a text similarity detection algorithm is used to identify and merge similar or duplicate sentences; to detect duplicate paragraphs, text comparison and clustering algorithms are used to detect and merge duplicate paragraphs in longer texts.

[0310] S114): Contextual Preservation. In order to preserve the original context of the text, the following measures are taken:

[0311] Maintain paragraph structure: While dividing the text into sentences, preserve the original paragraph structure so that complete paragraph information can be referenced when extracting knowledge points;

[0312] Using a sliding window: In order to capture a wider range of contextual information, a sliding window technique can be used when extracting knowledge points, which allows the model to consider information from a sentence and the sentences before and after it at the same time.

[0313] The personalized learning recommendation method based on personalized knowledge graphs described in this embodiment, specifically the sentence segmentation process in step S112), is as follows:

[0314] S1121): Punctuation-based: using periods, question marks, exclamation marks, and other punctuation marks as sentence separators;

[0315] S1122): Based on natural language processing tools, use AI models or natural language processing libraries to segment sentences. These tools can more accurately identify sentence boundaries, even in the absence of obvious punctuation marks.

[0316] S1123): Dependency-based parsing: By analyzing the dependency relationships in a sentence, the main clause and subordinate structure of the sentence can be determined, thereby more accurately dividing the sentence; its core idea is to decompose the sentence into a "dependency relationship" graph between words, thereby determining the predicate verb, subject, object elements in the sentence and their interrelationships.

[0317] The personalized learning recommendation method based on personalized knowledge graphs described in this embodiment, specifically the dependency parsing process in S1123) is as follows:

[0318] S11231): Input preprocessing, which involves word segmentation, noise reduction, and normalization of the input text; specifically as follows:

[0319] Text cleaning: First, the input text data needs to be cleaned and preprocessed, specifically by removing noise, symbols, irrelevant HTML tags, and normalizing complex text, such as converting it to lowercase and correcting typos.

[0320] Word segmentation: Before parsing the syntax, word segmentation tools are used to divide the sentence into individual words or phrases;

[0321] S11232): Part-of-speech tagging

[0322] First, natural language processing tools are used to tag the part of speech of each word. That is, a part-of-speech tagger is used to tag each word or phrase after word segmentation to determine the grammatical role of each word or phrase in the sentence, such as noun, verb, or adjective. Part-of-speech tagging provides basic information for subsequent dependency analysis.

[0323] Further simplify the parts of speech: simplify the parts of speech according to the specific application scenario of the words or phrases, such as simplifying the detailed verb form into a single "verb" category;

[0324] S11233): Dependency relation parsing, using a dependency parsing model to generate a dependency relation tree, as detailed below:

[0325] First, sentence component analysis is performed: through dependency parsing, the dependency relationships between words in the sentence are analyzed to obtain a dependency tree. The goal of dependency parsing is to find the subordinate relationship between the "core word" and other words in the sentence.

[0326] Next, identify the core components of the sentence: analyze the predicate verb and its related subject, object, and adverbial components. For example, `subject` depends on `verb`, and `object` also depends on `verb`. These components constitute the basic structure of the sentence.

[0327] Model selection: Choose an appropriate dependency syntax model based on the actual application and construct dependency relations;

[0328] S11234): Core structure extraction, which involves analyzing the subject, verb, object, modifier, and complement in sentences within the text to identify the core components of the sentence;

[0329] S11235): The construction of the dependency tree, using visualization tools to display dependency relationships, allows for a more intuitive analysis of sentence structure, as detailed below:

[0330] Dependency tree structured representation: Each sentence generates a dependency tree. The root node of the tree is the predicate verb of the sentence. Other words are attached to the trunk of the dependency tree in sequence according to their dependencies. Through the dependency tree, it is easy to see the main and subordinate relationships of the components in the sentence, thereby realizing grammatical analysis and sentence reorganization.

[0331] Visualization of dependency trees: Use the `displaCy` tool of `spaCy` or other dependency tree visualization tools to represent the parsed dependency relationships in a graphical form, making it easier to understand the sentence structure more intuitively;

[0332] S11236): Syntax rule validation and optimization

[0333] Dependency tree trunk and subordinate structure identification: By analyzing the dependency tree, the trunk structure of the sentence can be determined, namely the predicate verb and its main modifiers. The subject and object directly depend on the predicate verb, while the subordinate components of attributive and adverbial modifiers depend on the core components.

[0334] Syntactic rule optimization: Based on the needs of the domain, specific syntactic rules are set to correct dependency relations. For example, for complex sentence structures such as compound sentences and coordinate sentences, conjunctions can be processed through dependency syntax to clarify the position and function of conjunctions in the sentence structure and ensure the accuracy of semantic dependency relations.

[0335] It should be noted that the subordinate relationship between the core word and other words mainly follows these principles:

[0336] 1. Centered on the predicate verb: The predicate verb is the root node of the dependency tree, and other words establish dependency relationships around it.

[0337] 2. Dependency relationship of main components: The subject and object directly depend on the predicate verb, while subordinate components such as attributives and adverbs depend on the core components.

[0338] 3. Hierarchical structure: By using a dependency tree, the hierarchical relationship between the components in a sentence is clearly displayed, which facilitates grammatical analysis and sentence restructuring.

[0339] 4. Syntax rules: Specific syntax rules will be set according to the needs of specific domains, especially when dealing with complex sentence structures such as compound sentences and coordinate sentences, to ensure the accuracy of semantic dependency relationships.

[0340] The word association in S1235 is based on grammatical relations and hierarchical structure, and the association process follows the following principles:

[0341] - First, the predicate verb is used as the root node and trunk of the dependency tree.

[0342] - The main grammatical components (such as subject and object) depend directly on the predicate verb, while other modifying components (such as attributive and adverbial modifiers) depend on the core components they modify.

[0343] - The whole structure forms a hierarchical structure, clearly showing the main and subordinate relationships among the components of the sentence.

[0344] In the specific process of linking sentences, the main structure of a sentence is identified by analyzing its subject, verb, object, modifier, and complement components. Specific grammatical rules are followed, especially when dealing with complex sentence structures such as compound sentences and coordinate sentences, to ensure the accuracy of semantic dependencies.

[0345] It should be noted that the predicate verb is the core, the root node of the dependency tree, and an important component of the main structure of the entire sentence;

[0346] The core components and their subordinate relationships are as follows:

[0347] The subject and object directly depend on the predicate verb, while modifiers such as attributives and adverbs are subordinate components that depend on the core components.

[0348] 3. The concept of core ingredients:

[0349] - "Core words" specifically refer to the words that govern a sentence, usually the predicate verb.

[0350] - "Sentence core components" is a broader concept, including the predicate verb and the sentence components directly related to it, such as the subject, object, and adverbial.

[0351] This hierarchical structure clearly displays the hierarchical relationships between the components in a sentence through a dependency tree.

[0352] Application scenarios

[0353] Machine translation: Dependency parsing can help machine translation systems better understand the structure of complex sentences and ensure translation accuracy.

[0354] -Text summarization: By extracting the core structure of sentences, dependency parsing can be used to generate concise text summaries.

[0355] - Question answering systems: Dependency parsing can help question answering systems understand the core components of user questions, thereby generating more accurate answers.

[0356] Dependency parsing, as an effective method for analyzing the internal structure of sentences, has wide applications in text analysis, information extraction, and other fields.

[0357] In this embodiment, during the extraction of knowledge points and relationships from the data source in S12), an AI large-scale model is used to identify entities and concepts in the text as preprocessed knowledge points. The application of the large-scale model primarily involves extracting knowledge points through Prompt and mounting a finite knowledge set. It should be noted that the AI ​​large-scale model uses the existing pre-trained model, Alibaba's qwen-max, but other models that meet the requirements can also be selected according to actual needs. Mounting a finite knowledge set means specifying a limited set of knowledge points within a vertical scope. This content comes from external knowledge bases, such as Wikipedia, or annotations from domain experts within the company, such as computer science teachers defining 408 subject exam knowledge points and custom rule sets: based on specific needs, a set of finite knowledge points is defined as the constraints for entity and relationship extraction.

[0358] The following is a Prompt text:

[0359] You are a knowledge graph expert in the field of online education, specializing in entity extraction and relationship extraction. Extract entities from text and the relationships between them.

[0360] The types of entities are limited to the following range:

[0361] Facts: Information describing observed facts, knowledge about specific things, such as capital cities, dates, formulas, etc.

[0362] Composition: refers to understanding the components of something, such as human anatomy, engine parts, computer hardware, etc.

[0363] Concept: refers to an abstract and generalized definition of things, such as democracy, evolution, force, etc.

[0364] Principle: refers to the understanding of how things work, such as Newton's laws of motion, circuit principles, programming languages, etc., including subcategories such as rules and formulas.

[0365] Procedure: refers to the steps and methods for completing a task, such as solving math problems, conducting experiments, or writing essays.

[0366] Core courses: In an education system, these are foundational and crucial courses designed to achieve specific educational goals. Core courses typically cover a broad range of basic knowledge and skills. For example, "Advanced Mathematics: A required foundational course for university science and engineering students."

[0367] "

[0368] Entity relationships are limited to the following scope:

[0369] The word "contains" means that one entity contains another entity, that is, a smaller part or subset is contained within a larger entity. For example, "The mathematics curriculum includes algebra, geometry, and calculus."

[0370] Subordinate: "belong to" indicates that knowledge point B is a part of knowledge point A. For example, "Hardware is a part of a computer system."

[0371] Precedence: One entity is a preceding state or stage of another entity, usually an earlier stage in a process. For example, "Algebra is a prerequisite course for learning calculus."

[0372] Synonyms: Two entities that are identical or very similar in meaning and can be used interchangeably. For example, "differential and derivative are synonymous in some contexts."

[0373] Reference: a related relationship between two entities, such as "the instruction word length depends on the length of the opcode, the length and number of operand addresses".

[0374] Application: `applyto` indicates that knowledge point A is a skill, method, or approach to applying knowledge point B. For example: "Through learning and practice, theoretical knowledge can be transformed into practical skills."

[0375] Correlation: There is some kind of connection or relationship between two entities, but it is not necessarily a causal relationship. For example, "There is a correlation between study time and exam scores."

[0376] Follow these steps:

[0377] Step 1: High-precision entity recognition and entity extraction.

[0378] - Gain a deep understanding of entity context and accurately identify entity attributes and meanings.

[0379] - The type of an entity is limited to the provided types; new categories cannot be arbitrarily generated.

[0380] Step 2: Parse complex relationships and extract the relationships between entities.

[0381] - Ensure that the parsing logic is rigorous and conforms to the principles and standards for knowledge graph construction.

[0382] - Entity relationships are limited to the pre-provided range of relationships and new relationships cannot be generated arbitrarily.

[0383] Step 3: The final output data format is as follows:

[0384] "entities":[

[0385] {

[0386] "name":"Entity Name",

[0387] "type":"Entity Type"

[0388] }

[0389] ],

[0390] "relations":[

[0391] {

[0392] "head":"Relationship Header Entity Name",

[0393] "tail":"Relationship tail entity name",

[0394] "type":"Relationship type"

[0395] } ]

[0397] }

[0398] It should be noted that the JSON data format includes array keys such as "entities".

[0399] For example:

[0400] Content: Recursion and iteration are two common methods for solving data structure problems. Recursion breaks down a problem into smaller subproblems, while iteration repeatedly executes a set of instructions until a condition is met.

[0401] result:

[0402] {

[0403] "entities":[

[0404] {

[0405] "name":"Semiconductor random access memory",

[0406] "type": "concept"

[0407] },

[0408] {

[0409] "name":"Computer Performance Metrics",

[0410] "type": "concept"

[0411] },

[0412] {

[0413] "name":"Basic Principles of Multiplication / Division Operations",

[0414] "type":"Principle"

[0415] },

[0416] {

[0417] "name":"Operation Method",

[0418] "type":"Principle"

[0419] },

[0420] {

[0421] "name":"Instruction",

[0422] "type": "concept"

[0423] },

[0424] {

[0425] "name":"Basic Principles of Cache",

[0426] "type":"Principle"

[0427] },

[0428] {

[0429] "name":"The algorithm for replacing main memory blocks in the cache",

[0430] "type":"Principle"

[0431] },

[0432] {

[0433] "name":"Instruction Format",

[0434] "type": "concept"

[0435] },

[0436] {

[0437] "name":"Data Structure",

[0438] "type":"Core Courses"

[0439] }

[0440] ],

[0441] "relations":[

[0442] {

[0443] "head":"Calculation Method",

[0444] "tail":"Basic principles of multiplication / division operations",

[0445] "type": "contains"

[0446] },

[0447] {

[0448] "head":"command",

[0449] "tail":"command format",

[0450] "type": "contains"

[0451] },

[0452] {

[0453] "head":"The algorithm for replacing main memory blocks in the cache",

[0454] "tail":"Basic Principles of Cache",

[0455] "type":"Related"

[0456] } ]

[0458] }

[0459] Below is an example scenario:

[0460] enter:

[0461] Advanced mathematics is a foundational course for university science and engineering students, encompassing calculus, linear algebra, and probability theory. Calculus serves as a prerequisite for further study in physics, while linear algebra is closely related to engineering mathematics. Probability theory is primarily applied in the field of data analysis.

[0462] Output:

[0463]

[0464]

[0465]

[0466] The aforementioned knowledge points and relationships are disambiguated using cosine similarity to measure the angle between two vectors, thereby determining whether they represent the same concept. The formula for calculating cosine similarity is:

[0467] Here, a and b are two vectors, and ||a|| and ||b|| are their magnitudes.

[0468] In this way, similar knowledge points can be mapped to the same entity, thus constructing accurate entities and relationships in the knowledge graph. This method can improve the accuracy and robustness of the knowledge graph, enabling it to better reflect the knowledge and information in textual materials.

[0469] The knowledge entity standard set is the foundation of knowledge disambiguation. It typically comes from authoritative public resources (such as Wikipedia and DBpedia), specialized knowledge bases annotated by domain experts, or internal enterprise knowledge bases. Ensuring the accuracy and coverage of the standard set is crucial in this process.

[0470] Authoritative public resources: such as the structured data provided by WikiData, which can cover basic concepts and entities in multiple fields.

[0471] Expert-annotated specialized data sources: Combining knowledge from specific fields, enterprises can build custom sets of standard knowledge points according to their own business needs, which is particularly suitable for applications in vertical scenarios (such as mathematics, education, etc.).

[0472] Dynamic update mechanism: The knowledge point entity standard set is updated regularly from the latest literature and databases during use, and new knowledge points are continuously introduced to ensure the real-time nature of the knowledge in the system.

[0473] Standardization and naming rules:

[0474] Each knowledge point should be named and standardized to ensure consistent naming. For example, a course can be consistently named "Linear Algebra" instead of variations like "Linear Algebra" or "Linear Algebra Course". Furthermore, these standardized knowledge points need to include a unique identifier (such as a knowledge point ID) for subsequent mapping and querying.

[0475] The personalized learning recommendation method based on personalized knowledge graph described in this embodiment uses a context embedding model in the knowledge point and relation disambiguation in S13), which converts the extracted knowledge points and their contexts into high-dimensional vectors, thereby achieving more accurate semantic matching. Compared with traditional word vectors, context embedding can better capture the different meanings of the same word in different contexts and improve the robustness of disambiguation.

[0476] Specifically as follows:

[0477] For the input text t, the context embedding model BERT is used to transform the input text into a high-dimensional vector.

[0478] The quantity is represented by v;

[0479] v t =BERT(t)

[0480] Among them, v t It is the context embedding vector of text t.

[0481] Let cosine similarity calculate the angle between two vectors, defined as:

[0482]

[0483] Here, v1 and v2 are the context embedding vector representations of the two texts;

[0484] In dynamic weighted cosine similarity, considering that different dimensions contribute differently to the similarity, a weight w is assigned to each dimension. i This allows similarity calculations to more flexibly reflect important information in the text:

[0485]

[0486] Where d is the dimension, and d is usually the hidden layer size of the BERT model;

[0487] Among them, w i These are the weights used in weighted cosine similarity. They are dynamically calculated based on context or category, and are used to measure the contribution of different dimensions to the similarity. If two knowledge points are consistent in a certain category, the weight of dimensions related to that category can be increased. i ;

[0488] The w i Based on the following factors

[0489] Term frequency (TF): The frequency with which a word appears in a text, reflecting its importance;

[0490] Inverse Document Frequency (IDF): The rarity of a word in the entire corpus;

[0491] w i =TFi·IDFi

[0492] It can also be calculated using semantic similarity, as follows:

[0493] Suppose we have two sentences s1 and s2, whose semantic vector representations are respectively and

[0494] SBERT uses cosine similarity to measure the similarity between two sentences, s1 and s2:

[0495]

[0496] Compared to traditional sentence representation methods, SBERT can capture more fine-grained semantic information, making the semantic similarity calculation of two sentences more accurate.

[0497] Multi-level similarity:

[0498] To achieve more precise semantic matching, the similarity of knowledge points at multiple levels can be calculated; specifically as follows:

[0499] First, calculate word-level similarity, then calculate sentence-level or document-level similarity, and finally, calculate the weighted average of the similarities at each level.

[0500] total_similarity(x1,x2)

[0501] =α·word_similarity(x1,x2)+β·sentence_similarity(x1,x2)+γ

[0502] ·document_similarity(x1,x2)

[0503] Here, α, β, and γ are the weights of similarity at different levels, which are adjusted according to the specific application scenario. In the processing, if multiple representations are used for the same entity, the large model is fine-tuned to determine whether they are the same knowledge point, and then mapping and recognition are performed to obtain the final knowledge point and relationship.

[0504] Similarity calculations are used to map knowledge point entity standard sets.

[0505] It also includes mapping methods, and introduces context-aware semantic matching during the matching process by capturing the context of knowledge points in the text; context analysis, combined with the contextual information of surrounding words and sentences, enhances the ability to distinguish synonyms or related knowledge points. For example, if "quantum mechanics" is mentioned in the context of physics, then it can be inferred that the statement is consistent with related knowledge points in that field.

[0506] It's worth noting that dynamic weighted cosine similarity is suitable for scenarios where the importance of different dimensions needs to be considered. By assigning weights w_i to different dimensions, the calculation results more flexibly reflect important information. Especially when two knowledge points are consistent in a certain category, the weight of the relevant dimensions within that category can be increased.

[0507] Semantic Hierarchical Similarity (SBERT): Suitable for scenarios requiring finer-grained semantic analysis. Compared to traditional sentence representation methods, SBERT can capture more granular semantic information, making the similarity calculation between sentences more accurate. Multi-level Similarity: Suitable for scenarios requiring more comprehensive semantic matching; this method calculates similarity at multiple levels simultaneously.

[0508] The personalized learning recommendation method based on personalized knowledge graphs described in this embodiment, specifically the construction process of the knowledge point extraction and fine-tuning model in S12) is as follows:

[0509] (1) Data preparation, that is, preparing enough domain-related text data, including the annotation of knowledge points. For different expressions of knowledge points, a standardized set of knowledge points needs to be prepared to facilitate subsequent mapping and verification. The data source can be manually annotated text or existing knowledge bases, such as Wiki or domain datasets annotated by experts.

[0510] (2) Model selection

[0511] Choose a pre-trained large language model (such as BERT, GPT, etc.) as the pre-trained model. This model has been pre-trained with massive amounts of general data and has good language understanding ability. For the task of knowledge point extraction, you can choose a model based on the Transformer architecture. These models perform well in natural language processing tasks.

[0512] (3) Data annotation and cleaning of the text.

[0513] The data is preprocessed, including text cleaning (removing noise and irrelevant content) and sentence segmentation. Then, using annotation tools or writing rules, entities and relationships in the text are annotated. This step can use tools such as SpaCy or NLTK for automatic sentence segmentation and part-of-speech tagging.

[0514] (4) Fine-tune the pre-trained model to obtain the fine-tuned knowledge point extraction model.

[0515] For knowledge point extraction tasks, pre-trained models such as BERT or GPT are fine-tuned to obtain the fine-tuned model. The training process of the pre-trained model is completed through the following stages:

[0516] Entity recognition: Automatically identify entities in text, such as concepts, facts, and principles, using pre-trained models;

[0517] Relationship extraction: Extracting relationships between text entities automatically identified by the pre-trained model, such as inclusion, dependency, and leader;

[0518] Disambiguation: For different descriptions of the same knowledge point, vectorization techniques are used to calculate similarity, mapping similar expressions to the same entity;

[0519] (5) Model Validation and Optimization

[0520] The fine-tuned model is tested on a validation set to evaluate its accuracy in knowledge point recognition and relationship extraction tasks. A similarity algorithm is used to verify the accuracy and consistency of the extracted entities and relationships. If the accuracy and consistency meet the requirements, the model can meet the needs.

[0521] The specific process of fine-tuning the pre-trained model to obtain the fine-tuned model is as follows:

[0522] (401) First, determine the task objectives: that is, determine the task objectives of entity recognition, relation extraction and disambiguation respectively;

[0523] (402) Prepare the dataset: Label entities and relations (BIO labels, entity pairs and relation categories);

[0524] (403) Model selection: Based on actual needs, select the appropriate model from "BERT-type models (suitable for classification and sequence labeling), GPT-type models (suitable for generation tasks), and graph neural networks (suitable for complex relationship modeling)";

[0525] (404) Design fine-tuning tasks: that is, set fine-tuning tasks for the selected model. Specifically, for entity recognition, perform sequence labeling (BIO) and fine-tune using linear layers; for relation extraction, extract relations from entities and context, extract and output relations through a classifier; for disambiguation, perform entity vectorization description and train a similarity classifier or clustering model for fine-tuning.

[0526] (405): The fine-tuned model is trained and optimized using the AdamW optimizer with a learning rate of 1e. -5 up to 5e -5 ;

[0527] (406): The fine-tuned training model is evaluated, specifically: the entity recognition extraction is evaluated by the F1 score; the accuracy of relation extraction is evaluated by the classification accuracy of the text and the F1 score; and the disambiguation effect is evaluated by the Top-k accuracy and similarity evaluation.

[0528] The personalized learning recommendation method based on personalized knowledge graphs described in this embodiment includes the following specific methods for model validation and optimization:

[0529] Adaptive threshold adjustment: During the similarity calculation process, the similarity threshold is dynamically adjusted by continuously learning the accuracy of historical data;

[0530] The specific process is as follows:

[0531] Data initialization: Collect initial historical data for similarity calculation, including the results of each similarity calculation and the correctness of actual verification;

[0532] Initial threshold setting: Set the similarity threshold based on common empirical values, such as the default threshold for cosine similarity being 0.8;

[0533] Similarity calculation: For each knowledge point to be processed, a preliminary similarity score is obtained based on cosine similarity, SBERT or other similarity algorithms.

[0534] Feedback collection: Based on the results of preprocessing, the system collects feedback data from users or experts and marks whether each disambiguation is correct.

[0535] Adaptive learning: Threshold adjustment is triggered by accumulated feedback data and the error of disambiguation results. The system uses machine learning algorithms to dynamically adjust the similarity threshold. The adjustment range is usually between 0.05 and 0.10. The threshold can be adaptively changed according to different domains, text length, semantic complexity and other factors.

[0536] Apply dynamic thresholds: In the new round of similarity calculation, dynamically adjusted thresholds are used to determine whether two knowledge points are considered the same.

[0537] For example, when the system disambiguates “quantum physics” and “quantum mechanics”, it can adaptively adjust the threshold based on similar data processed previously to ensure that these statements are considered the same knowledge point.

[0538] In this embodiment, the extracted knowledge points and relationships are stored in a graph database. The specific process of forming a knowledge graph is as follows: 1. Insert knowledge points and relationships into the graph database.

[0539] Graph databases (such as Neo4j, JanusGraph, ArangoDB, etc.) store data in the form of nodes and edges. Nodes represent entities (knowledge points), and edges represent relationships between entities. The process of building a knowledge graph mainly involves the following steps:

[0540] (1) Design of graph data model

[0541] Entity Nodes: Each extracted knowledge point is inserted into the graph database as a node. Nodes typically have multiple attributes, such as entity name, type, and unique ID.

[0542] Edges: The relationship between any two entities is represented as an edge. The attributes of an edge include the relationship type (such as "contains", "related", "leader") and the direction of the relationship.

[0543] (2) Insertion process using Neo4j as an example

[0544] Connecting to a graph database: Use a Neo4j Python client (such as `neo4j-driver`) to connect to a Neo4j database.

[0545] Python

[0546] from neo4j import GraphDatabase

[0547] uri="bolt: / / localhost:7687"

[0548] driver=GraphDatabase.driver(uri,auth=("username","password"))

[0549] Creating nodes and relationships:

[0550] Create entity nodes: Each extracted knowledge point is inserted into the graph database as a node.

[0551] ```cypher

[0552] CREATE(n:Entity{name:'Advanced Mathematics',type:'Course'})

[0553] ```

[0554] You can insert nodes in batches using the Python API:

[0555]

[0556] 2. Create relationship edges: Create relationships between entity nodes.

[0557] ```cypher

[0558] MATCH(a:Entity{name:'Advanced Mathematics'}),(b:Entity{name:'Calculus'})

[0559] CREATE(a)-[:contains]->(b)

[0560] ```

[0561] In Python, relationships can be created in batches:

[0562]

[0563] (3) Insertion process using JanusGraph

[0564] For distributed graph databases like JanusGraph, Gremlin is typically used as the query language, similar to Neo4j's Cypher. Batch insertion of nodes and relationships is achieved through the Gremlin interface.

[0565] ```gremlin

[0566] g.addV('Entity').property('name','Advanced Mathematics').property('type','Course')

[0567] gV().has('name','Advanced Mathematics').addE('Includes').to(gV().has('name','Calculus')).

[0568] The personalized learning recommendation method based on personalized knowledge graphs described in this embodiment can incorporate symbolic reasoning during the disambiguation process. Symbolic reasoning can further correct the disambiguation results using known rules and logic. The specific process is as follows:

[0569] (1): Knowledge point extraction: Extract potential similar knowledge points from the original data in the literature, and obtain preliminary semantic similarity through deep learning models;

[0570] (2): Semantic disambiguation: Using deep learning methods to perform preliminary similarity calculations, those with similarity values ​​lower than the target value will be deleted. This involves some basic knowledge points of disambiguation.

[0571] (3): Application of symbolic reasoning rules: For knowledge points that fail to disambiguate or have insufficient similarity, symbolic reasoning is combined with the defined rule set or logical reasoning to further correct the results based on known knowledge in the domain;

[0572] (4): Inference optimization: The results of symbolic inference in the previous step are fed back to the deep learning model to further adjust the weights or thresholds of the deep learning model in order to reduce errors;

[0573] (5): Output the final result: Combine the final result of symbolic reasoning and deep learning disambiguation to determine whether the two knowledge points are the same.

[0574] The known rules include logical rules based on logical expressions (mutual exclusivity, transitivity), ontology rules for defining the hierarchy, attributes and relationships of concepts, data consistency rules for verifying data consistency, semantic rules for reasoning through context or defined semantic relationships, mathematical or physical laws and priority rules.

[0575] The deep learning model (such as BERT or SBERT) obtains preliminary semantic similarity and generates fixed word vectors, which can be used to calculate semantic similarity, as follows:

[0576] Word2Vec: Generates word vectors based on a context window and uses the cosine similarity of the vectors to compare semantic similarity;

[0577] GloVe: Generates word embeddings based on global statistical information, suitable for semantic similarity calculation.

[0578] FastText extends Word2Vec to support subword representation, enabling better handling of word form variations. The defined rule set, which is the core of symbolic reasoning, typically requires domain knowledge in its definition during implementation. Below are some general principles and examples for defining rule sets:

[0579] The rule set is structured as follows:

[0580] (1) Static rules, which are deterministic rules based on known facts and domain constraints. For example:

[0581] Classification rule: If a knowledge point belongs to a certain category, it may have the attributes common to that category.

[0582] Association rule: Two knowledge points are equivalent or related under certain conditions.

[0583] (2) Dynamic rules, which involve data-driven patterns and can often be combined with statistical analysis or frequent itemset discovery:

[0584] According to the conditional probability rule, if knowledge points A and B appear frequently at the same time, it can be inferred that they may be the same or related.

[0585] Context-sensitive rules mean that the meaning of a knowledge point changes with the context.

[0586] The following is an example of a specific rule set:

[0587] (1) Domain: Mathematical knowledge points

[0588] Equivalence rule: If the description of a knowledge point contains "equivalent to" or "also called", they may be the same concept.

[0589] TextContains(x,"also known as") → Equivalent(x,y)

[0590] If two formulas have the same transformation form, they may be equivalent, such as...

[0591] FormulaSimilar(A,B)∧Transform(A,B)→Equivalent(A,B)

[0592] Derivation rule: If one formula can be derived from another formula through simple algebraic transformations, they may be based on the same knowledge point;

[0593] AlgebraicTransform(Equation1,Equation2)→SamePoint(Equation1,Equation2)

[0594] (2) Field: Biological terminology

[0595] Synonym rule: If two terms differ only in spelling (such as British and American spelling), they may be synonyms.

[0596] Hypernym rule: If one term is a hypernym of another term, and a certain condition is met, they may be partially related:

[0597] (3) Cross-domain rules

[0598] Time relevance: If knowledge points A and B are frequently cited within the same time period, they may be related.

[0599] 4. The rule set is constructed as follows:

[0600] Based on expert knowledge: Domain experts summarize the relationships between concepts and manually define rules.

[0601] Data mining-based: Extracting latent patterns from corpora to generate rules.

[0602] Use tools such as the Apriori algorithm to generate frequent itemsets.

[0603] Use knowledge graphs to mine terminological relationships.

[0604] Automatic generation and optimization: Combining deep learning and reinforcement learning, new rules are generated and optimized through verification feedback.

[0605] 5. Rule Validation and Optimization

[0606] Validation: Test the defined rules and use known data to verify the correctness of the reasoning.

[0607] Optimization: Adjust rule weights based on feedback from actual applications, or design meta-rules to prioritize conflicting rules.

[0608] By combining these rule sets, symbolic reasoning can effectively correct ambiguous semantics during the disambiguation process and further improve the accuracy and interpretability of the final results.

[0609] The specific logic of reasoning can be divided into the following key steps:

[0610] 1. Formalizing Input Data: To enable symbolic reasoning to handle problems, the input data (knowledge points, semantic information) must first be formalized into a logical language. Common forms include:

[0611] Predicate logic: describing the relationships between knowledge points through logical predicates.

[0612] Term("Optical diffraction") → Related("Optical interference")

[0613] Graph structure: Knowledge points can be represented by a knowledge graph, where nodes are knowledge points and edges are relationships.

[0614] 2. Application of knowledge point matching rules: Using a predefined set of rules, pattern matching is performed on knowledge points and their contextual information to extract potential logical relationships.

[0615] The rule matching method is as follows:

[0616] 2.1 Static matching: Directly matching the conditions in the rules and comparing them with known knowledge points or context;

[0617] For example, the rule:

[0618] Synonym(a,b)→Equivalent(a,b)

[0619] If Synonym("optical diffraction", "diffraction") is true, then the following can be derived:

[0620] Equivalent ("optical diffraction", "diffraction")

[0621] 2.2 Dynamic condition derivation: For uncertain knowledge points, conditional judgments are made in conjunction with the logical rules of the context;

[0622] Context("optics")∧Related("light waves","interference")→Related("optics","diffraction")

[0623] 2.3 Fuzzy Rule Processing: For knowledge points with insufficient semantic similarity, fuzzy logic is used to define a credibility threshold.

[0624] For example: Rules:

[0625] Similarity(a,b)>0.7→PossibleEquivalent(a,b)

[0626] If Similarity("Optical Diffraction","Diffraction") = 0.75, then it can be inferred that they may be equivalent.

[0627] 3. Application of the inference engine: By executing the rule set using the inference engine and combining it with logical deduction of the relationships between knowledge points, the recommendation engine can achieve the following inference:

[0628] 3.1 Forward reasoning, that is, starting from the known conditions, triggering rules in sequence, and deriving new knowledge point relationships;

[0629] Given:

[0630] Synonym("Optical Diffraction", "Diffraction")

[0631] rule:

[0632] Equivalent(a,b)→Related(a,b)

[0633] Derivation:

[0634] Related("Optical diffraction", "Diffraction").

[0635] 3.2 Backward reasoning: Starting from the goal, we search backward for the conditions that satisfy the goal;

[0636] 3.3 Fuzzy logic reasoning, which introduces a confidence level between the conditions and the conclusion.

[0637] 4. Conflict resolution and prioritization: Multiple rule conflicts may occur during symbolic reasoning, and their priorities need to be clearly defined.

[0638] 4.1. Weight-based priority: Define weights for rules, and execute rules with higher weights first.

[0639] 4.2. Conflict resolution based on meta-rules: Conflict resolution strategies are defined through meta-rules.

[0640] During the reasoning process, the text is matched using the rule Synonym(a,b), the context is checked for support, and finally the reasoning results are merged to draw a conclusion.

[0641] If RuleA∧RuleB conflict→Prefer(RuleA)

[0642] 5. Example Process

[0643] Question: Are "optical diffraction" and "interference" equivalent concepts?

[0644] Given:

[0645] rule:

[0646] Synonym(a,b)→Equivalent(a,b)

[0647] Context("Optics")∧Related(a,b)→Equivalent(a,b)

[0648] data:

[0649] Synonym("Optical Diffraction", "Diffraction")

[0650] Context("optics")

[0651] Related("light waves", "interference")

[0652] Reasoning process:

[0653] 1. Match using the rule Synonym(a,b):

[0654] Synonym("Optical Diffraction", "Diffraction") → Equivalent("Optical Diffraction", "Diffraction")

[0655] 2. Check if the context supports it:

[0656] Context("optics")∧Related("light waves","interference")→Equivalent("optical diffraction","interference")

[0657] 3. Merging the inference results:

[0658] Equivalent("Optical Diffraction","Interference")∧Equivalent("Optical Diffraction","Diffraction")→True

[0659] The final conclusion is that "optical diffraction" and "interference" are equivalent.

[0660] 5. Incorporate feedback from the deep learning model; that is, after symbolic reasoning is completed, the results can be used as feedback to further adjust the similarity calculation or knowledge representation of the deep learning model. For example:

[0661] Update the semantic similarity model: If the symbolic reasoning result is True, adjust the similarity weights to make them closer to the corrected standard.

[0662] Retrain the semantic embedding model: Add the new disambiguation results to the corpus and update the word vectors or relation graph.

[0663] In this way, symbolic reasoning and deep learning can optimize each other, improving the overall accuracy and robustness of disambiguation.

[0664] The personalized learning recommendation method based on personalized knowledge graph described in this embodiment uses a domain knowledge-based disambiguation mechanism during the disambiguation process. In addition to general semantic matching and similarity calculation, it can also combine expert knowledge or rules in a specific domain to perform disambiguation.

[0665] Implementation process:

[0666] A. Knowledge point semantic matching: Use a deep learning model to perform general semantic matching on the extracted knowledge points;

[0667] In the disambiguation task, the general semantic matching in Part A mainly relies on deep learning models to compare knowledge points at the semantic level. This process involves specific semantic embedding representations, similarity calculations, and matching rules. The following are the specific implementation methods and corresponding matching rules:

[0668] B. Introduction of domain expert knowledge: Introduce expert rules, definitions, vocabularies or knowledge graphs for specific domains, and further verify and correct the semantic matching results through this domain knowledge;

[0669] C. Domain rule disambiguation: Based on the rule set of a specific domain, the matching results are classified or excluded in a more refined manner;

[0670] D. Output disambiguation results: Combine the domain knowledge disambiguation results with the general semantic matching results to output the final disambiguated knowledge point matching.

[0671] In knowledge point disambiguation, the combination method of Part D is not a simple direct superposition. Instead, it integrates the results of domain knowledge disambiguation and general semantic matching through methods such as weight allocation, multi-layer decision logic, or collaborative optimization to improve the accuracy and robustness of the final result. The following is an analysis of the specific combination method, as well as the purpose and function of the combination:

[0672] The core process of knowledge point semantic matching is as follows:

[0673] (1) Semantic representation of knowledge points: knowledge points are transformed into comparable semantic vectors through deep learning models.

[0674] Word vector-based embedding can be done in the following ways:

[0675] Use pre-trained word vector models (such as Word2Vec and GloVe) to transform knowledge points into vectors.

[0676] Use language models (such as BERT, RoBERTa) to generate context-sensitive vector representations.

[0677] If the knowledge points contain multimodal information (such as text, formulas, and images), joint embedding representations can be generated using multimodal models (such as CLIP).

[0678] (2) Semantic similarity calculation, that is, calculating the similarity between the vector representations of two knowledge points.

[0679] (3) Setting the matching threshold: Based on the similarity score, a matching threshold is defined:

[0680] Exact match: If similarity > 0.9

[0681] Possible match: 0.7 ≤ similarity ≤ 0.9

[0682] Mismatch: Similarity < 0.7

[0683] It should be noted that the threshold setting can be adjusted according to the actual field.

[0684] 3. Tools and models for achieving general semantic matching can be adopted as follows:

[0685] (1) Pre-trained language models: Models pre-trained on large-scale corpora generate semantic embeddings, for example:

[0686] BERT and RoBERTa are suitable for text semantic matching; SciBERT is optimized for scientific literature; Sentence-BERT generates sentence-level semantic vectors and is suitable for calculating similarity.

[0687] (2) Knowledge base and dictionary: Use general knowledge bases (such as WordNet and ConceptNet) to expand semantic information and use domain-related dictionaries to supplement terminology.

[0688] (3) Semantic similarity calculation libraries, including Python tools, scikit-learn: provides cosine similarity and Euclidean distance calculation; sentence-transformers directly calculates the similarity of embedded vectors.

[0689] Example to illustrate: The complete process of general semantic matching

[0690] Question: Are "optical diffraction" and "wave interference" the same concept?

[0691] 1. Using BERT to generate embeddings:

[0692] "Optical Diffraction" → [0.32, 0.45, 0.76]

[0693] "Wave interference" → [0.33, 0.44, 0.78]

[0694] 2. Calculate the cosine similarity: CosineSimilarity = 0.98

[0695] 3. Matching rule validation: Validation is performed by checking spell variants.

[0696] If we consider "optical diffraction" and "wave interference"...

[0697] EditDistance("Optical Diffraction", "Wave Interference")>3→NoMatch

[0698] EditDistance(a,b): Represents the edit distance between strings a and b. Edit distance (LevenshteinDistance) is a commonly used metric for measuring the similarity between two strings; it represents the minimum number of edit operations required to transform one string into another. These edit operations typically include:

[0699] Insert a character;

[0700] Delete a character;

[0701] Replace one character.

[0702] >3: This indicates that the edit distance is greater than 3; that is, the number of edit operations between strings a and b is greater than 3.

[0703] NoMatch: This means that if the condition is met, i.e. the edit distance is greater than 3, then a and b are considered to be unmatched (NoMatch).

[0704] Check context overlap: If the context overlap is greater than or equal to 0.7, a match is possible;

[0705] 4. Matching threshold determination: Based on similarity and rule results, determine whether a match is "possible".

[0706] In the process of combining domain knowledge disambiguation results with general semantic matching results, a weighted fusion method is used.

[0707] Different weights are assigned to the domain knowledge disambiguation results and the general semantic matching results, and the two are weighted and calculated according to the importance of the actual task.

[0708] Specifically as follows:

[0709] General semantic similarity score: S1 = 0.8

[0710] Domain rule disambiguation score: S2 = 0.9

[0711] Weighted formula: S final =w1·S1+w2·S2

[0712] Among them, w1 and w2 are weighting coefficients, which can be adjusted experimentally;

[0713] Assume the weights are w1 = 0.4 and w2 = 0.6.

[0714] S final =0.4·0.8 + 0.6·0.9 = 0.86

[0715] The matching can be determined based on the set threshold.

[0716] Using weighted fusion can effectively overcome the limitations of a single method, resulting in a more comprehensive and accurate disambiguation result. Specifically:

[0717] (1) To make up for the shortcomings of a single method

[0718] Limitations of general semantic matching: General semantic matching is weak in handling domain-specific terms or rare knowledge points, which may lead to misjudgments or omissions. For example, a general model may mistakenly consider "optical diffraction" and "wave interference" to be unrelated.

[0719] The drawback of domain rule disambiguation: Domain rules rely heavily on known expert knowledge and may be powerless in the face of unknown or uncovered situations;

[0720] By combining the broad coverage of general semantic matching and the high precision of domain rules, the disambiguation results are made more reliable.

[0721] (2) Handling special cases,

[0722] Some matching pairs may score low in general semantic matching, but domain knowledge shows that they are closely related;

[0723] For example: General matching result: Similarity = 0.6

[0724] Domain rules: Rules indicate that a causal relationship exists between two things;

[0725] After combining: it is determined to be a match;

[0726] (3) Improve confidence and interpretability. The introduction of domain rules provides a clear logical basis, making the matching results more interpretable, while general semantic matching results provide quantitative support.

[0727] (4) Optimize model performance.

[0728] By incorporating feedback from domain rules, the parameters of the semantic model can be dynamically adjusted to gradually adapt to domain requirements.

[0729] Although the domain rule disambiguation in Part C has already categorized and excluded general semantic matching results, the following reasons make its combination even more necessary:

[0730] (1) Improve matching coverage.

[0731] Domain-specific rules have limited coverage and cannot encompass all knowledge points. Combining general semantic matching can handle aspects that are undefined or not covered by the rules.

[0732] (2) Handling low-confidence matching

[0733] In semantic matching and rule disambiguation, some knowledge points may be in boundary states simultaneously:

[0734] For example, the semantic similarity score is close to the threshold or there are contradictions or incompleteness in the domain rules;

[0735] By combining the results of both, these boundary cases can be addressed more comprehensively.

[0736] (3) Verification and calibration results: By combining the results of general semantic matching, the judgment of domain rules can be verified, thereby reducing errors;

[0737] (4) Improve scalability. After combining semantic matching, the system can still output effective results even when the domain rules are insufficient, making the model more general and flexible.

[0738] Examples are given below:

[0739] Determine whether "wave interference" and "optical diffraction" are the same.

[0740] 1. General semantic matching:

[0741] Semantic similarity score: 0.6 (below the matching threshold);

[0742] Preliminary assessment: No match;

[0743] 2. Domain rule disambiguation:

[0744] Rule 1: If the knowledge point contains "wave" and "diffraction", then the relevance is high;

[0745] Rule application result: Relevance score 0.9;

[0746] 3. Weighted fusion:

[0747] Weighting: w1 = 0.3, w2 = 0.7;

[0748] Final score:

[0749] S final =0.3·0.6 + 0.7·0.9 = 0.78

[0750] The results above show that the match exceeds the matching threshold and is therefore considered a match.

[0751] 4. Output Result: Knowledge Point Matching: "Wave Interference" = "Optical Diffraction"

[0752] Therefore, combining general semantic matching and domain rule disambiguation not only improves coverage and confidence, but also provides the model with the ability to dynamically adjust and expand. This combined strategy can output more accurate and reliable disambiguation results when a single method is insufficient.

[0753] For example, while "myocardial infarction" and "heart attack" are similar in general semantics in the medical field, they have more subtle differences in some literature, so they need to be strictly distinguished in accordance with domain rules.

[0754] The personalized learning recommendation method based on personalized knowledge graphs described in this embodiment employs several methods to ensure the accuracy of inserted knowledge points and relationships during the knowledge graph construction process, as detailed below:

[0755] (1) Multi-level verification

[0756] Rule-based validation: Different validation rules are defined for different domains to detect knowledge points and relationships. If data that does not conform to the validation rules is matched, potential errors are marked, and knowledge points and relationships that are not marked as errors will be deleted.

[0757] Contextual consistency verification: By analyzing the context in which the knowledge points are located, ensure that the extracted entities and relationships are consistent with the logic in the text;

[0758] Domain expert review: For knowledge in a specific domain, an expert-annotated knowledge base can be introduced for comparison to ensure that the extracted knowledge points are consistent with the information in the existing knowledge base;

[0759] (2) Semantic similarity matching

[0760] Semantic similarity calculation is used to verify the extracted knowledge points. For example, the extracted knowledge points and standard knowledge points are vectorized using language models such as BERT and SBERT. Then, the similarity is calculated to determine whether the knowledge points have been accurately extracted. If the similarity is lower than the set threshold, it is marked as a potential error and deleted.

[0761] (3) Relational logic verification

[0762] Relationship type restrictions: Ensure that relationship types conform to the logic within the domain. Logical validation of relationship types is performed through the domain rule base. For example, in the education domain, the relationship between "course" and "chapter" should be "contains" rather than "precedes".

[0763] Loop detection: Graph algorithms are used to detect unreasonable loops in the knowledge graph, avoiding incorrect relationships that could lead to an unreasonable knowledge graph structure.

[0764] The personalized learning recommendation method based on personalized knowledge graphs described in this embodiment, specifically the means and process of verifying the accuracy of extracting knowledge points and relationships from the data source in step S2) are as follows:

[0765] Before storing knowledge points and relationships in a graph database, use the following steps to verify the accuracy of the extracted knowledge points and relationships:

[0766] (1): Automated test dataset

[0767] Construct a standard labeled dataset and use it as a benchmark to input into the AI ​​model. Then compare the knowledge points extracted by the knowledge point extraction fine-tuning model with the standard answer. If the extracted knowledge points and relationships are consistent with those in the labeled dataset, the model passes the verification. Otherwise, further debugging of the knowledge point extraction fine-tuning model or adjustment of the algorithm is required.

[0768] Accuracy: Number of correctly extracted knowledge points / Total number of extracted knowledge points;

[0769] Recall rate: Number of correctly extracted knowledge points / Number of knowledge points in the standard set;

[0770] (2): Manual review and expert feedback

[0771] Knowledge points in specific fields undergo manual review by experts to ensure that the extracted content meets industry standards; for example, in the medical field, doctors can review the extracted terminology and relationships.

[0772] (3): Context validation

[0773] After extracting knowledge points, a second verification is performed in conjunction with their context to ensure that the extracted knowledge points are consistent with their meaning in the context. For example, "Java" as a knowledge point of programming language may also represent other concepts in different contexts. Context verification can reduce such errors.

[0774] (4): Model Adaptive Optimization

[0775] By continuously collecting and fine-tuning error cases in the knowledge point extraction model, analyzing and providing feedback on them, and incorporating this feedback into the training data, the model's performance is continuously optimized. This cyclical feedback mechanism can effectively improve the accuracy of extraction.

[0776] In the process of knowledge graph construction, complex knowledge graphs can be built by inserting extracted knowledge points and relationships into a graph database and utilizing tools such as Neo4j and JanusGraph. Simultaneously, the accuracy and logical consistency of the inserted data are ensured through methods such as context consistency checks, rule-based validation, semantic similarity matching, and loop detection. To further verify the accuracy of knowledge points and relationships, various methods such as automated dataset testing, expert review, and contextual validation can be used to ensure the high quality and high accuracy of the knowledge graph.

[0777] The personalized learning recommendation method based on personalized knowledge graphs described in this embodiment includes the following process for constructing the learner profile model in step S2):

[0778] S21): Data collection, collecting various data of learners, including learning behavior data, learning resource interaction data, and test scores, and using association rule algorithms and data mining techniques to mine user learning process data;

[0779] S22): Preprocess the collected data, including cleaning up outliers and handling missing values ​​to ensure data quality. Through attribute matching and feature extraction, combined with learners' personal attributes and learning styles, integrate various learning profile modules using statistical methods to obtain learners' learning abilities, cognitive levels, and learning goal profiles.

[0780] S23): From the three dimensions of learning ability, cognitive level and learning goals, the learner’s personalized learning characteristics are extracted. The learning feature matrix of the learner is formed by quantifying the feature labels of learning ability, cognitive level and learning goals through AprioriAll. Each row represents a learner and each column corresponds to a feature dimension.

[0781] S24): Based on the learning feature matrix, machine learning algorithms are used to group or classify learners in order to build learner profile models.

[0782] The personalized learning recommendation method based on personalized knowledge graph described in this embodiment uses a reinforcement learning DQN network in the personalized learning recommendation engine to make decisions through the interaction between the agent and the environment. These decisions need to bring as many rewards as possible. Here, the agent is the personalized learning recommendation engine, the environment is the user's state and knowledge graph, and the reward is the benefit of the learning effect.

[0783] The specific work process is as follows:

[0784] S61): First, the data is processed to obtain user profile features.

[0785] The user profile features consist of a learning ability vector c, a cognitive level vector l, and a learning target vector g. The learning ability vector c, cognitive level vector l, and learning target vector g are calculated separately, and then the user comprehensive feature vector is processed. The specific process is as follows:

[0786] The normalization formulas used below are all:

[0787]

[0788] Where X min For the minimum value, X max The maximum value is used to map the data range to the interval [0, 1].

[0789] The learning ability mentioned includes the dimensions of memory, logical reasoning ability, learning speed, and learning focus:

[0790] The memory ability dimension is measured by the score rate M, which is calculated from the scores of memory-based questions and the total score.

[0791]

[0792] Normalization:

[0793]

[0794] The logical reasoning ability dimension is measured by the score rate L of reasoning-type questions;

[0795]

[0796] Normalization:

[0797]

[0798] Learning speed dimension: This reflects a user's ability to effectively complete learning tasks per unit of time, demonstrating how quickly a student can accept and master new knowledge; it includes average learning time and learning focus.

[0799] Average study time:

[0800]

[0801] Learning speed is inversely proportional to average learning time. kp

[0802]

[0803] Normalization:

[0804]

[0805] Learning focus:

[0806] The learning focus refers to a student's ability to maintain concentration and avoid distraction during the learning process, reflecting the degree of engagement in learning;

[0807] The main indicators of learning focus are as follows:

[0808] Continuous learning period, number of learning interruptions, and other abnormal operations;

[0809] Learning focus index

[0810]

[0811] Average study time

[0812]

[0813] Formula for final score of learning focus

[0814]

[0815] A min and A max These are the minimum and maximum values ​​of the average continuous learning time for users, respectively.

[0816] D min and D max These are the minimum and maximum values ​​of the distraction metric for all users, respectively.

[0817] Where ω1 and ω2 are both weights, with a value of 0.5;

[0818] Learning ability matrix representation:

[0819] C = [M] norm L norm L kp_norm A score ]

[0820] The cognitive levels in the cognitive level vector l are as follows:

[0821]

[0822] Among them, A total M is the user's total score across all tests within the subject. total is the sum of the maximum scores for all tests within the subject; l is the ratio of the user's total score for all tests within the subject to the sum of the maximum scores for all tests within the subject.

[0823] Normalization:

[0824]

[0825] The learning objective in the learning objective vector g:

[0826] Based on the user's target university and major, knowledge points are selected from the knowledge graph, and then the proficiency level of each knowledge point is calculated based on the target grades.

[0827] Assume the knowledge points obtained based on the target universities and majors are as follows:

[0828] k = [k1 k2 ... k] n ]

[0829] The weight of each knowledge point is:

[0830] ω(k i), where i is 1...n;

[0831] Weight normalization for each knowledge point

[0832]

[0833] The target total score S is calculated based on the normalized weights. goal Assign weights to each knowledge point and adjust the base values ​​to ensure the weights are greater than 1.

[0834] S target (k i )=ω′(k i )×(S goal )+1

[0835] The matrix for the user's learning target vector g is as follows:

[0836] g = [S target (k1) S target (k2) ... S target (ki)] The above features are combined to form the user's comprehensive feature vector.

[0837] F student =[c, l, g]

[0838] Where c is the learning ability vector, l is the cognitive level vector, and g is the learning target vector;

[0839] Using the comprehensive feature vector F student As a clustering feature, students are divided into K groups, and each student is assigned a cluster label G. j (j = 1, 2, ..., K), representing the group to which a student belongs, and assigning cluster labels G to the students. j As a feature, it is incorporated into the state representation of subsequent models;

[0840] S62): Knowledge graph embedding, which involves collecting all knowledge points and constructing a knowledge point set K = k1, k2, ..., k N ;

[0841] The Node2Vec algorithm is used to embed knowledge points from the knowledge graph into a low-dimensional vector space, resulting in the embedding vector e of the i-th knowledge point. i ;

[0842] Then, by adding the user's comprehensive feature vector, the knowledge point embedding vector, and the student's clustering label, the corresponding user state vectors are obtained, as follows:

[0843] The user state vector S, which incorporates the user's comprehensive feature vector, is represented as follows:

[0844] s = [s1, s2, ..., sN c, l, g]

[0845] Where c is the learning ability vector, l is the cognitive level vector, and g is the learning target vector;

[0846] Student clustering labels G were added j and the embedding vector e of knowledge points i ;

[0847] The method for calculating knowledge point vectors is as follows:

[0848]

[0849] in,

[0850] To what extent students have mastered the knowledge point k: i The degree of mastery, ranging from [0,1];

[0851] Learning frequency: i.e., the student's learning of knowledge point k i The number of times;

[0852] For testing performance: that is, students' performance on knowledge point k i Average test score;

[0853] The most recent learning time interval: that is, the time since the last time knowledge point k was learned. i The time interval since then;

[0854] e i Embedding vectors of knowledge points;

[0855] f i Features of frequent patterns;

[0856] Knowledge point k i The support level is:

[0857]

[0858] Add student cluster labels G j The user state vector S after the change is represented as follows:

[0859] s = [s1, s2, ..., s N c, l, g, G j ]

[0860] If a user has not fully grasped the knowledge points, the following action options are available:

[0861] A′(s)={k i ∈A(s)|g i >0}

[0862] A(s) represents the set of knowledge points that satisfy the prerequisite relationships but are not fully mastered, g i >0 indicates knowledge point k i These are the students' learning goals;

[0863] Reward learners for achieving their learning objectives. The specific reward function is as follows:

[0864] r t =w(k i )×Δs i ×f(c,l)×G(k i )×PM(k i )×AE(k i )×LF(k i )×FC(k i )

[0865] Among them, w(k) i ) represents knowledge point k i Importance weight

[0866] Δs i For students to understand knowledge point k i Improved mastery

[0867] f(c, l) is the adjustment function for students' learning ability and cognitive level, as follows:

[0868] f(c, l) = δ1·c + δ2·l, where δ1 and δ2 are weight parameters with values ​​of 0.7 and 0.3, respectively;

[0869] G(k i ) is the weighting function for the learning objectives.

[0870]

[0871] Among them, g i Let i be the i-th element of the student's learning objective vector, representing the student's understanding of knowledge point k. i The degree of learning need; γ is a constant greater than 1, which is the reward bonus coefficient for the learning target knowledge point;

[0872] When g i >0 (Knowledge point k) i When G(k) is the student's learning objective, i =γ, the reward is amplified by γ times (usually 1.2-2.0), and the expected reward for submitting the selection of this knowledge point is more likely to be recommended by the personalized learning recommendation engine;

[0873] When g i =0, G(k) iWith a reward of 1, the personalized learning recommendation engine will not give special preference to recommending this knowledge point.

[0874] PM(k i The function is a frequent pattern weighting function, which adjusts the reward based on the importance of the knowledge point in the frequent itemset, reflecting the importance of the knowledge point in the learning group, as detailed below:

[0875] PM(k i )=1+λ PM ×Support(k i )

[0876] Where, λ PM To adjust the parameter, the degree of influence on the reward is usually a small positive number, such as 0.1 or 0.2;

[0877] Support(k i ) represents knowledge point k i Support, representing k i Frequency of occurrence in frequent itemsets;

[0878] By considering the frequency of knowledge points appearing in frequent learning patterns, the personalized learning recommendation engine is guided to recommend knowledge points that are commonly learned in the group.

[0879] In the association rule algorithm, AR(k) i The weighting function for association rules is:

[0880] AR(k i )=1+λ AR ×Conf(k i )

[0881] Where, λ AR To adjust parameters and control the degree of influence of association rules on rewards, values ​​are typically positive; Conf(k i Knowledge Point k i The highest confidence level in an association rule represents k. i The strength of the connection with previously learned knowledge points;

[0882] Utilizing previously learned knowledge: By considering k i By associating knowledge points with what students have already learned, the model is guided to recommend knowledge points closely related to the learned content, thus promoting coherent learning.

[0883] LF(k i The learning frequency adjustment function is:

[0884]

[0885] Where, λ LF Take w(k)i The reciprocal of the number of knowledge points; the more important the knowledge point, the smaller the reduction in reward.

[0886] As the learning frequency increases, the function value will gradually decrease, reducing its contribution to the reward and preventing students from excessively and repeatedly learning familiar knowledge points.

[0887] FC(k i Forgetting Curve Adjustment Function

[0888]

[0889] Where, λ FC Take a direct proportional function of the knowledge point weights, with a value range of 0.5-1; when Increase the reward, decrease the reward, and prompt the student to review the knowledge point.

[0890] S63): The reinforcement learning DQN network updates its network parameters as follows:

[0891] Input layer: User state vector s

[0892] Output layer: Q-values ​​for each optional action (Q-value is the mathematical expectation of the sum of rewards from all future states);

[0893] The DQN algorithm network training steps are as follows:

[0894] A. State Acquisition: Acquire the current state s t ;

[0895] B. Action selection: Select an action from the set of available actions according to the greedy strategy.

[0896] C. Execution Action: Students learn the knowledge points, update their level of mastery, and obtain a new state. t+1 ;

[0897] D. Reward Calculation: Calculate the immediate reward based on the reward function;

[0898] E. Experience storage: (s) t a t r t s t+1 Stored in the experience replay pool;

[0899] F. Network Update: Randomly draw a small batch of samples (s) from the experience pool. j a j r j s j+1 )

[0900] Train and update the network parameters;

[0901] The real-time recommendation process using a recommendation engine is as follows:

[0902] Status Acquisition: Retrieve the student's current user status S;

[0903] Action selection: Based on the greedy strategy, select the optimal knowledge point from A′(s);

[0904] Knowledge Point Recommendations: Recommending knowledge points to students and providing corresponding learning resources.

[0905] S64): Feedback and updates on learners' progress are collected through user feedback, as detailed below:

[0906] Learning feedback collection: Record students' learning outcomes for recommended knowledge points, and update their mastery level, learning frequency, and recent learning time;

[0907] User Status Update: Update the student's user status S based on learning feedback.

[0908] Reinforcement Learning DQN Network Update: Regularly train the reinforcement learning DQN network, incorporate new learning data, and improve model performance;

[0909] Add learners' feedback to the experience replay pool

[0910] In reinforcement learning, experience samples are typically represented as quadruples:

[0911] (s t a t r t s t+1 )

[0912] To incorporate student feedback into the experience sample, the structure of the experience sample needs to be expanded to include feedback information:

[0913] (s t a t r t′ s t+1 f t )

[0914] r t′ The adjusted instant rewards take student feedback into account.

[0915] f t Students perform action a at time step t t Feedback

[0916] If the student's feedback is positive, provide an immediate reward.

[0917] r t′ =r t +δr

[0918] If the result is negative, reduce the reward.

[0919] r t′ =r t -δr.

[0920] The personalized learning recommendation method based on personalized knowledge graphs described in this embodiment, specifically the process of constructing the learning path recommendation pattern in step S7) is as follows:

[0921] S601): The learning feature representation is based on the subject knowledge graph. It uses the association rule algorithm to extract subject knowledge elements and related learning resources, which together form the learning elements of the knowledge graph, represented by the triple L=(Ki,Kj,r).

[0922] Where r represents the logical relationship between knowledge element Ki and learning resource Kj;

[0923] Then, knowledge graph technology is used to serialize and label the attributes of subject knowledge elements and related learning resources, thereby characterizing and describing the basic features of learning elements;

[0924] Based on the learners' learning needs, the learning elements of the learning activity process are formed into a learning element sequence (s1, s2, ..., s...). N );

[0925] At this time, the user state S is represented as follows:

[0926] s = [s1 s2 ... s N ]

[0927] S602): Learning path recommendation. Based on the user's current learning state matrix and the priori relationship of knowledge points in the knowledge graph, recommend the next knowledge point that will receive the maximum reward and the corresponding learning materials.

[0928] S603): Learning Path Evaluation and Optimization: During the learning process, learners complete tasks based on the learning paths and resources recommended by personalized learning and submit feedback on their learning experience and effectiveness. The personalized learning recommendation engine collects and analyzes user feedback, combines it with the learner's goal completion status, compares expert paths with recommended paths, and evaluates the accuracy and practicality of the recommendations.

[0929] Based on evaluation results and user feedback, we continuously optimize the personalized learning recommendation engine to improve the effectiveness of personalized recommendations and user satisfaction.

[0930] The personalized learning recommendation method based on personalized knowledge graphs described in this embodiment, in step S3), uses knowledge graph technology for knowledge extraction, which follows the following rules:

[0931] S31): Template-based rules, such as regular expression rules, template matching;

[0932] S32): Syntactic and lexical analysis rules: HanLP (Chinese processing library, supports part-of-speech tagging and word segmentation), spaCy (provides fast part-of-speech tagging), Stanford NLP (part-of-speech tagging, supports multiple languages).

[0933] S33): Context dependency rules;

[0934] S34): Statistical learning rules;

[0935] The specific method for establishing the logical connection between subject knowledge elements and related learning resources is as follows: First, the learning resources are classified and labeled, which can be divided into textbooks, exercises, videos, cases, experiments and extracurricular materials. Each resource has clear metadata to describe its connection with knowledge elements.

[0936] Create a tagging system for knowledge elements and learning resources, and ensure the accuracy of logical connections through manual and automatic annotation (natural language processing or large models to analyze text content, extract knowledge elements or entities for tagging), and manual verification.

[0937] Based on the knowledge systems and curriculum standards of various disciplines, knowledge graph technology is used for knowledge acquisition and extraction. Then, through Python code or large-scale models, entity alignment, attribute alignment, relationship alignment, and conflict resolution are performed on the data to complete knowledge fusion. The knowledge fusion process is as follows:

[0938] Entity alignment: Different data sources may use different names to describe the same entity. Entity alignment is needed to identify them as the same object; for example, "Apple Inc." and "Apple Inc." are actually the same entity.

[0939] Attribute alignment: One data source uses "date of birth" and another uses "birthday", but they are actually the same attribute.

[0940] Relation alignment: Aligning relation types to uniformly describe relations across different data sources. For example...

[0941] "is employed by" and "works for" can be aligned to the same relation type.

[0942] Conflict resolution: Attributes about the same entity provided by different data sources may conflict. For example, one data source might state a person's birth year as 1980, while another might state 1981. Conflict resolution techniques are needed to select the most reliable version.

[0943] For example, reward functions not only provide feedback based on learners' task completion or test scores, but can also offer multi-dimensional explanations by incorporating their knowledge mastery within the knowledge graph. By examining the connections between knowledge points in the graph, the system can explain why learners receive rewards or penalties, such as "After completing the recommended task, the learner mastered multiple key knowledge points in the knowledge graph, thus receiving a higher reward."

[0944] The interconnected paths in a knowledge graph can help the system explain which knowledge points played a key role in the overall progress, making the reward mechanism more transparent.

[0945] It also includes a knowledge graph-based personalized learning recommendation platform, characterized by: including:

[0946] The learner's interface is the login portal for users to log in to the personalized learning recommendation platform.

[0947] The backend management interface is used by administrators or teachers to log in to the personalized learning recommendation platform.

[0948] The learning style assessment module is used to test learners' learning styles.

[0949] The learning goal setting module is used by learners to set learning goals. The learner's interactive interface is connected to the learning style assessment module and the learning goal setting module.

[0950] The entrance test module is used to assess learners' current learning progress.

[0951] The learning process recording module is used to record the learner's learning process;

[0952] The learning assessment module is used to assess, analyze, and evaluate the learning results based on the collected learner learning process data. The learning assessment module is bidirectionally connected to the learning process recording module.

[0953] The learner profile model module is used to establish a learner profile model and update the learner profile model based on the learner's subsequent learning progress. The learning style determination module, learning goal setting module, learning process recording module, and learning assessment module are all connected to the learner profile model module.

[0954] A personalized learning recommendation engine processes collected user learning data and learning assessment results to generate personalized learning recommendations for users. The personalized learning recommendation engine is connected to the learning assessment module and the learning process recording module.

[0955] The personalized learning recommendation module is used to provide users with personalized learning content recommendations and services. The personalized learning recommendation content and service module is connected to the personalized learning recommendation engine.

[0956] The learning database is used to provide a data repository for learning-related content. The learning database is connected to the learner profile model building module, the learning process recording module, the personalized learning recommendation content and service module, and the learning assistance tool module. The personalized learning recommendation engine is bidirectionally connected to the learning database.

[0957] The course content backend is a platform used to manage all learning content for learners. The course content backend is connected to the learning content database in the learning database.

[0958] The knowledge point platform is a backend used to provide knowledge points to the course content backend and the knowledge graph in the learning database; it is connected to the knowledge graph in the learning database and the course content backend, respectively.

[0959] During operation, the learner interaction interface, backend management interface, learning style assessment module, learning goal setting module, entrance test module, learning process recording module, learning assessment module, learning database, course content backend, and knowledge point platform form the data foundation. Each time a learner takes a test in the learning assessment module, the learner profile model in the personalized learning recommendation platform is updated in real time. The personalized learning recommendation engine adjusts the recommended learning path and learning content according to the adjusted user profile model.

[0960] In this embodiment, the learning database includes a learner model database, a learning process database, a learning course database, a learning path database, a learning content database, and a knowledge graph. The learner model database, learning process database, learning course database, learning path database, and learning content database are bidirectionally connected. The learner profile model building module is connected to the learner model database, the learning process recording module is connected to the learning process database, the personalized learning recommendation engine is connected to the learning path database, the learning path database is connected to the personalized learning content recommendation module, the learning content database is connected to the course content backend, and the knowledge graph is connected to the knowledge point platform.

[0961] The above description is only a preferred embodiment of the present invention. It should be noted that for those skilled in the art, several improvements can be made without departing from the principle of the present invention, and these improvements should also be considered within the scope of protection of the present invention.

Claims

1. A personalized learning recommendation method based on personalized knowledge graphs, characterized in that: include: S1): To build a knowledge graph, the data source is first preprocessed, and then the model is fine-tuned through a dedicated knowledge point extraction to automatically identify entities in the text and extract the relationships between entities to obtain preprocessed knowledge points and relationships; then the knowledge points and relationships are disambiguated. Establish a graph database model, innovate entity nodes, and create relationship edges between entity nodes. Then, store the disambiguated entity nodes and relationship edges into the graph database model to form a dynamic and scalable knowledge graph. The specific construction method of the knowledge graph in S1) is as follows: S11): Data preprocessing, that is, processing the data source of the knowledge points to be extracted. The data source of the knowledge points to be extracted includes structured data and unstructured data. The structured data is directly input into the mapper, and the text data in the unstructured data is cleaned. S12): Extract knowledge points and relationships from the data source. First, build a dedicated knowledge point extraction fine-tuning model. Input the pre-processed data into the knowledge point extraction fine-tuning model. The knowledge point extraction fine-tuning model automatically identifies entities in the text and extracts the relationships between entities to obtain pre-processed knowledge points and relationships. S13): Knowledge point and relationship disambiguation, i.e., the identification and mapping of different descriptions of the same knowledge point; the specific disambiguation process is as follows: S131): Establish a standard knowledge point base, that is, first construct a standard set of knowledge point entities. The established standard set of knowledge point entities serves as the target set for mapping and is used to disambiguate the knowledge points extracted from the text. S132): Similarity calculation and mapping, that is, using text embedding technology in natural language processing to convert knowledge points and related texts into vectors; then performing similarity calculation, by calculating the cosine similarity or Euclidean distance between the text embedding vectors to determine whether the extracted different expressions point to the same knowledge point. Cosine similarity is used to measure the angle between vectors to determine the similarity between the two. S133): Semantic disambiguation based on context: In order to ensure the accuracy of the disambiguation process, the sliding window technique is used to take contextual information into consideration, so that in the process of identifying knowledge points, not only the current sentence is considered, but also the sentences before and after it are referred to, and the relationship between knowledge points is extracted from the preprocessed knowledge points. S134): Iterative verification and feedback optimization: After disambiguation, the knowledge points and relationships extracted in the previous step will be mapped to the knowledge point entity standard set, the two will be matched, corresponding relationships will be found, and further verification will be performed through the extracted relationships. If there is no match, the entity will be deleted. In S13), the knowledge point and relation disambiguation process uses a context embedding model to convert the extracted knowledge points and their contexts into high-dimensional vectors. S14): Knowledge graph construction: First, establish a graph database model, create entity nodes, and create relationship edges between entity nodes. Then, store the disambiguated entity nodes and relationship edges in the graph database model to form a dynamic and scalable knowledge graph. In the later use process, use AI models to re-analyze text data regularly and update the knowledge graph to ensure that it reflects the latest knowledge points and their changes in a timely manner. Symbolic reasoning can be incorporated into the disambiguation process. Symbolic reasoning can further correct the disambiguation results using known rules and logic. The specific process is as follows: (1): Knowledge point extraction: Extract potential similar knowledge points from the original data from the literature, and obtain preliminary semantic similarity through a deep learning model; (2): Semantic disambiguation: Using deep learning methods to perform preliminary similarity calculations, those with similarity values ​​lower than the target value will be deleted. This involves some basic knowledge points of disambiguation. (3): Application of symbolic reasoning rules: For knowledge points that fail to disambiguate or have insufficient similarity, symbolic reasoning is combined with the defined rule set or logical reasoning to further correct the results based on known knowledge in the domain; (4): Inference optimization: The results of symbolic inference in the previous step are fed back to the deep learning model to further adjust the weights or thresholds of the deep learning model in order to reduce errors; (5): Output the final result: Combine the final result of symbolic reasoning and deep learning disambiguation to determine whether the two knowledge points are the same; S2): When learners first enter the personalized learning platform through the learner interaction interface, they must first provide basic personal information and complete the registration. S3): Establishing the initial learner profile model; S4): Learners access the database to learn; S5): The learning assessment module conducts periodic assessments of learners' learning progress and sends the assessment results to the learner profile model and the personalized learning recommendation engine; S6): Update the learner profile model; S7): Based on the updated learner profile model and learning process records, the personalized learning recommendation engine calculates the learner's mastery of knowledge elements in real time, generates a learner profile accordingly, calculates the user's comprehensive feature vector, and then uses knowledge graph technology and reinforcement learning DQN network to adaptively and self-organize the learning path and learning resources that are suitable for the learner based on the calculated user comprehensive feature vector. S8): Update the learning path database and store the newly generated learning paths of the adapted learners in S7) into the learning path database; S9): Personalized learning content recommendation uses the learning path database to recommend new learning paths to learners; S10): Then the learner repeats steps 4) to 9) until all the knowledge points to be learned are completed.

2. The personalized learning recommendation method based on personalized knowledge graphs according to claim 1, characterized in that: The specific method for data cleaning of text data in unstructured data in S11) is as follows: S111): Text cleaning: First, the text data in the unstructured data is cleaned. A deep learning model is used to classify different text regions in the text. By using the context and semantic features of the text, irrelevant content is distinguished from the main text to automatically identify the main text and irrelevant content. Irrelevant content is removed by regular expressions. Noise in the text data is eliminated by deleting redundant spaces, special characters, irrelevant HTML tags, and duplicate text. S112): Sentence segmentation: In order to provide a clear context for subsequent knowledge point extraction, the cleaned text needs to be segmented into sentences; S113): Duplicate content extraction, which extracts duplicate content from text data by detecting duplicate sentences and duplicate paragraphs, specifically: To detect duplicate sentences, a text similarity detection algorithm is used to identify and merge similar or duplicate sentences. Detecting duplicate paragraphs: Long texts are detected and duplicate paragraphs are merged using text comparison and clustering algorithms; S114): Contextual preservation: In order to preserve the original context of the text, the following measures are taken: Maintain paragraph structure: While dividing the text into sentences, preserve the original paragraph structure so that complete paragraph information can be referenced when extracting knowledge points; Using a sliding window: In order to capture a wider range of contextual information, a sliding window technique can be used when extracting knowledge points, which allows the model to consider information from a sentence and the sentences before and after it at the same time.

3. The personalized learning recommendation method based on personalized knowledge graphs according to claim 2, characterized in that: The specific process of clause segmentation in S112 is as follows: S1121): Based on punctuation: using periods, question marks, exclamation marks, and other punctuation marks as sentence separators; S1122): Based on natural language processing tools, sentence segmentation is performed using AI models or natural language processing libraries. These tools can more accurately identify sentence boundaries, even in the absence of obvious punctuation marks. S1123): Dependency parsing-based analysis: By analyzing the dependency relations in a sentence, the main clause and subordinate structures of the sentence are determined, thus more accurately dividing the sentence; Its core idea is to decompose a sentence into a "dependency relationship" diagram between words, thereby determining the predicate verb, subject, object elements in the sentence and their interrelationships.

4. The personalized learning recommendation method based on personalized knowledge graphs according to claim 3, characterized in that: The specific analysis process of dependency parsing in S1123 is as follows: S11231): Input preprocessing, which involves word segmentation, noise reduction, and normalization of the input text; specifically as follows: Text cleaning: First, the input text data needs to be cleaned and preprocessed, specifically by removing noise, symbols, irrelevant HTML tags, and normalizing complex text, such as converting it to lowercase and correcting typos. Word segmentation: Before parsing the syntax, word segmentation tools are used to divide the sentence into individual words or phrases; S11232): Part-of-speech tagging First, natural language processing tools are used to tag the part of speech of each word. That is, a part-of-speech tagger is used to tag each word or phrase after word segmentation to determine the grammatical role of each word or phrase in the sentence, such as noun, verb, or adjective. Part-of-speech tagging provides basic information for subsequent dependency analysis. Further simplify the parts of speech: simplify the parts of speech according to the specific application scenario of the words or phrases, such as simplifying the detailed verb form into a single "verb" category; S11233): Dependency relation parsing, using a dependency parsing model to generate a dependency relation tree, as detailed below: First, sentence component analysis is performed: through dependency parsing, the dependency relationships between words in the sentence are analyzed to obtain a dependency tree. The goal of dependency parsing is to find the subordinate relationship between the "core word" and other words in the sentence. Next, identify the core components of the sentence: analyze the predicate verb and its related subject, object, and adverbial components. For example, `subject` depends on `verb`, and `object` also depends on `verb`. These components constitute the basic structure of the sentence. Model selection: Choose an appropriate dependency syntax model based on the actual application and construct dependency relations; S11234): Core structure extraction, which involves analyzing the subject-verb-object and attributive-adverbial-complement components in sentences within the text to identify the core components of the sentences; S11235): The dependency tree is constructed, and the dependency relationships are displayed through visualization tools to more intuitively analyze sentence structure, as detailed below: Dependency tree structured representation: Each sentence generates a dependency tree. The root node of the tree is the predicate verb of the sentence. Other words are attached to the trunk of the dependency tree in sequence according to their dependencies. Through the dependency tree, it is easy to see the main and subordinate relationships of the components in the sentence, thereby realizing grammatical analysis and sentence reorganization. Visualization of dependency trees: Use the `displaCy` tool of `spaCy` or other dependency tree visualization tools to represent the parsed dependency relationships in a graphical form, making it easier to understand the sentence structure more intuitively; S11236): Syntax rule verification and optimization Dependency tree trunk and subordinate structure identification: By analyzing the dependency tree, the trunk structure of the sentence can be determined, namely the predicate verb and its main modifiers. The subject and object directly depend on the predicate verb, while the subordinate components of attributive and adverbial modifiers depend on the core components. Syntactic rule optimization: Based on the needs of the domain, specific syntactic rules are set to correct dependency relations. For complex sentence structures such as compound sentences and coordinate sentences, conjunctions can be processed through dependency syntax to clarify the position and function of conjunctions in the sentence structure and ensure the accuracy of semantic dependency relations.