Cross-domain Query Method for Educational System Data Based on Knowledge Graph
By building multi-level ontologies and knowledge graphs in the education system, semantic modeling and dynamic updates of educational data are solved, and the problems of data semantic inconsistency and low query efficiency in the existing technology are significantly improved.
Patent Information
- Application Number
- CN202411922718.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-25
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2044-12-25
AI Technical Summary
The existing education system has significant shortcomings in the semantic unity of data, dynamic updates, cross-domain query efficiency and result optimization, and it is difficult to meet the needs of the development of modern education informatization.
Using a knowledge graph-based method, we use multi-level ontology, partitioned data storage, set up dynamic update mechanisms, receive natural language query requests, perform semantic analysis and inference, generate cross-partition query paths, parallel query and result optimization.
It significantly improves the unity of educational data and query association efficiency, ensures the real-time and semantic integrity of the knowledge graph, and improves the accuracy and comprehensiveness of query results.
Smart Images

Figure CN119862227B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of educational systems, and particularly to a method for cross-domain query of educational system data based on a knowledge graph. Background Art
[0002] With the development of big data and artificial intelligence technologies, the educational field has gradually adopted intelligent means to improve the efficiency and accuracy of teaching, assessment, and resource management. However, the data of educational systems are usually distributed in different subsystems and regions, with diverse forms, including structured, semi-structured, and unstructured data. The characteristics of multi-source heterogeneity make cross-domain query and data integration an important bottleneck in the development of educational informatization, severely restricting the optimal allocation of educational resources and the popularization of intelligent applications.
[0003] Currently, most educational systems use independent databases or data warehouses to manage data. Existing systems usually implement basic query functions through manual configuration or rule definition, and there are the following problems: First, due to the diverse sources of educational data and the lack of a unified semantic standard, existing technologies cannot perform in-depth semantic modeling and standardized expression of data, resulting in the difficulty of effectively utilizing the relevance and context semantics between data; Second, traditional query methods are mostly based on keyword matching or simple field retrieval. When facing the cross-disciplinary, cross-stage, and cross-regional data query requirements in the educational field, the query results are often inaccurate and incomplete, unable to meet the requirements of complex educational scenarios; Third, existing distributed query systems usually focus on data access performance rather than semantic optimization, and there are significant deficiencies in the efficiency and accuracy of queries involving multi-dimensional semantic associations.
[0004] In summary, existing technologies have significant deficiencies in semantic unification, dynamic update, cross-domain query efficiency, and result optimization of educational data, and it is difficult to meet the needs of modern educational informatization development. The defects of existing technologies directly affect the intelligent level and service ability of educational systems, and there is an urgent need for an innovative method to solve the above problems and provide more efficient and accurate technical support for the data management and application of educational systems. Summary of the Invention
[0005] An object of the present invention is to propose a method for cross-domain query of educational system data based on a knowledge graph, which significantly improves the unity and query association efficiency of educational data.
[0006] A method for cross-domain query of educational system data based on a knowledge graph according to an embodiment of the present invention includes the following steps:
[0007] S1. Based on specific requirements in the educational field, define core concepts, attributes, and their hierarchical relationships, and construct a multi-level ontology in the educational field;
[0008] S2. Partition and store the multi-source heterogeneous data in the education field by subject area, education stage, and region based on the constructed multi-level ontology. Establish an independent knowledge graph storage structure for each partition, and record the association relationships between partitions through semantic connection strategies;
[0009] S3. Set up a dynamic update mechanism for each partition in the knowledge graph. Update the graph nodes and relationships within the partition in real time according to the newly added or changed education data. At the same time, use the semantic connection strategy to synchronously update the association information between partitions;
[0010] S4. Receive the user's natural language query request, perform semantic parsing on the query request, identify the query intent and generate a structured query statement. The structured query statement includes the target partition information to be retrieved and the query semantics;
[0011] S5. According to the target partition information and query semantics in the query request, use the query scheduling algorithm to locate the relevant partitions in the distributed storage structure of the knowledge graph, and perform dynamic reasoning on the semantic association relationships between the target partitions through the semantic connection strategy to generate a cross-partition query path;
[0012] S6. Execute parallel query operations on the located knowledge graph partitions based on the distributed computing framework, and integrate the query results of each partition according to the cross-partition query path to form a preliminary query result set;
[0013] S7. Use the semantic reasoning ability of the multi-level ontology to optimize the semantics of the preliminary query result set, eliminate duplicate results, resolve semantic conflicts, and complete implicit associated data, and integrate the optimized query results into the final output result.
[0014] Optionally, the specific steps of S1 are as follows:
[0015] S11. Determine the specific requirements in the education field, define the set C of core concepts in the education system. The core concepts include subject topics, teaching objectives, learning objects, and resource types related to education. Define the set of attributes A(c i ) for each core concept. Each attribute a j in the set of attributes represents the characteristic information of concept c i ;
[0016] S12. Construct the set of hierarchical relationships between concepts by analyzing the logical relationships in the education field:
[0017] R = {r1, r2,..., r k};
[0018] Among them, each relationship r k includes the hyponymy relationship between concepts subordination relationship and semantic association relationship
[0019] S13. Build the ontology Ontology of the subject field based on the core concept set C, the attribute set A(c i ) and the relationship set R Domain , and the ontology of the subject field includes the knowledge model of specific disciplines;
[0020] S14. Divide the education system into primary school, junior high school and senior high school according to the characteristics of the education stage, and construct the ontology Ontology of the education stage Stage , and the ontology of the education stage defines the core concept C of different stages Stage , the attribute set A Stage and the relationship set R Stage ;
[0021] S15. Define the ontology Ontology of the region according to the characteristics of educational resources and standardization requirements in different regions Region , combine educational resources with regional characteristics, and the concept set C in the regional ontology Region includes educational policies and regional resource distributions;
[0022] S16. Associate the ontology of the subject field, the ontology of the education stage and the ontology of the region to form a multi-level ontology set:
[0023] Ontology Multi ={Ontology Domain ,Ontology Stage ,Ontology Region};
[0024] S17. Based on the multi-level ontology set Ontology Multi perform semantic structured modeling on the multi-source heterogeneous data in the education system, and map the fields in the data to the concepts, attributes and relationships in the multi-level ontology to form a unified ontology basic framework.
[0025] Optionally, the S2 specifically includes the following steps:
[0026] S21. Classify the multi-source heterogeneous data set D in the education field based on the multi-level ontology set, and the data classification basis includes the subject field, the education stage and the regional characteristics;
[0027] S22. For the classified subject field set D Domain , the education stage set D Stage and the regional characteristic set D Region , respectively construct knowledge graph partitions:
[0028] Gi ={E i ,R i ,A i}, where \(i\in\{\text{Domain},\text{Stage},\text{Region}\}\);
[0029] Among them, E i represents the entity set of partition \(i\), including core concept entities and attribute entities, R i represents the relationship set of partition \(i\), which is consistent with the relationship set defined in the multi-level ontology, A i represents the attribute value set of partition \(i\), which is used to describe the specific attribute data of entities;
[0030] S23. In each knowledge graph partition \(G\) i , use the semantic rule set \(Q\) of the multi-level ontology i to perform semantic reasoning and consistency checking on the data within the partition, and eliminate redundant relationships and semantic conflicts within the partition;
[0031] S24. Based on the knowledge graph partition set \(G = \{G\) Domain ,G Stage ,G Region}, design a cross-partition semantic connection strategy \(R\) i Connect :
[0032]
[0033] S25. Store the constructed knowledge graph partition set \(G\) in a distributed graph database, and establish a global index and a local index mechanism for each partition. The global index is used to quickly locate the partition category, and the local index is used to query specific entities and relationships within the partition;
[0034] S26. Combine the semantic connection strategy to dynamically synchronize and update the cross-partition association relationships. When the knowledge graph data in a certain partition changes, the semantic connection rules of its related partitions are automatically adjusted.
[0035] Optionally, the specific steps of S3 are as follows:
[0036] S31. Set a dynamic update trigger mechanism for the knowledge graph partition set \(G\), monitor the new education data set \(D\) New and the changed education data set \(D\) Update , and determine the update trigger conditions according to the data type, including new nodes, relationship updates, or attribute modifications;
[0037] S32. Update the knowledge graph structure within each knowledge graph partition \(G\) i according to the semantic content of the new or changed education data set through the following rules:
[0038] If the data involves new entities, then add the new entity e i to the entity set E n ;
[0039] If the data involves relationship changes, then add or modify the relationship r i to the relationship set R u ;
[0040] If the data involves attribute updates, then add or modify the attribute value a i to the attribute set A u ;
[0041] S33. After updating the graph structure within the partition, use the semantic rule set Q i within the partition to perform consistency verification on the newly added or modified content, check the semantic consistency between the newly added entity and the existing entities, check whether the modified relationship conforms to the constraints defined by the ontology, and check whether the attribute value conforms to the field specifications;
[0042] S34. Based on the cross-partition semantic connection strategy perform synchronous adjustment on the updated content associated across partitions:
[0043] If the newly added entity involves the cross-partition super-subordinate relationship, then update the relevant connection rules;
[0044] If the relationship change involves the subordinate relationship, then update the association rules;
[0045] If the attribute modification involves cross-domain semantic association, then update the connection rules;
[0046] S35. After the dynamic synchronous update is completed, update the index mechanism of the partitioned knowledge graph, update the partition location of the newly added or changed entities or relationships within the partition for the global index I Global and update the specific node or attribute address of the newly added or changed content within the partition for the local index I Local,i ;
[0047] S36. Store the updated knowledge graph partition set G new back to the distributed graph database.
[0048] Optionally, the S4 specifically includes the following steps:
[0049] S41. Receive the user's natural language query request Q NL , where the query request includes the query keywords or sentences input by the user;
[0050] S42. Use the semantic parsing algorithm to parse the natural language query request Q NL Perform parsing, tokenize the user's query request, extract the keyword set K in the query, and based on the multi-level ontology set Ontology Multi Perform semantic matching on the keyword set K, identify the concepts, attributes, and relationships corresponding to the keywords, and resolve the polysemy of the keywords in combination with the user's query context or historical behavior to generate the query intention set I Q ;
[0051] S43. According to the query intention set I Q Locate the knowledge graph partition related to the query according to the concepts, attributes, or relationships in it Represent the semantic content corresponding to the matched query intention as the query semantics S Q , where S Q Includes the query target entity set E Q , the query target relationship set R Q And the query target attribute set A Q , combine the target partition information P and the query semantics S Q Generate the structured query statement Q Structured =(P, S Q );
[0052] S44. In the generated structured query statement Q Structured , verify the logical consistency of the target partition information and the query semantics, so that the structured query statement satisfies the following conditions:
[0053] The entities in the query target entity set E Q Can be mapped to the entity set E within the partition i ;
[0054] The relationships in the query target relationship set R Q Can match the relationship set R within the partition i ;
[0055] The attributes in the query target attribute set A Q Can match the attribute set A within the partition i .
[0056] Optionally, the S5 specifically includes the following steps:
[0057] S51. Extract the knowledge graph partition set to be retrieved from the target partition set P according to the structured query statement Q Structured ;
[0058] S52. Lock the target partition according to the partition type through the global index I Global For each target partition G i∈P, determine the specific locations of the target entity set, target relationship set, and target attribute set within the partition through the local index I Local,i Determine the specific locations of the target entity set, target relationship set, and target attribute set within the partition;
[0059] S53. Based on the target partition set P and the query semantics, perform dynamic semantic reasoning by combining the cross-partition semantic connection strategy to generate a cross-partition semantic association path. Perform dynamic semantic reasoning by combining the cross-partition semantic connection strategy to generate a cross-partition semantic association path Path Cross :
[0060] If the target entity set in the query semantics S Q involves the upper and lower level relationships of partitions G i and G j then infer the upper and lower level path across partitions through the semantic connection rule ; Infer the upper and lower level path across partitions
[0061] If the target relationship set in the query semantics S Q involves the subordinate relationship between partitions, then infer the subordinate path across partitions through the semantic connection rule ; Infer the subordinate path across partitions
[0062] If the target attribute set in the query semantics S Q involves cross-domain semantic associations, then infer the semantic association path across partitions through the semantic connection rule ; Infer the semantic association path across partitions
[0063] S54. Use the semantic association path Path Cross to dynamically integrate the query results between the target partitions to generate a cross-partition query path:
[0064] Path Query ={Path1,Path2,…,Path m}.
[0065] Optionally, the S6 specifically includes the following steps:
[0066] S61. Based on the cross-partition query path set, determine the target knowledge graph partition set corresponding to each path and the relevant query semantics;
[0067] S62. In the distributed computing framework, assign computing tasks to each cross-partition query path Path k and assign an independent computing node N i to each partition G in each path i and assign the target query semantics to each computing node N i ;
[0068] S63. In each computing node N iAmong them, a query operation is performed according to the local knowledge graph partition and the query semantics to generate a local partition query result set R Local,i ;
[0069] S64. Based on the cross-partition query path Path k , according to the semantic connection strategy in the path dynamically integrate the local partition query result set R Local,i to generate a path-level query result set R Path,k :
[0070] If the path Path k involves the upper and lower semantic relationships, then connect the partition query results according to ;
[0071] If the path Path k involves the subordinate semantic relationship, then connect the partition query results according to ;
[0072] If the path Path k involves the semantic association relationship, then connect the partition query results according to ;
[0073] S65. Aggregate all path-level query result sets R Path ={R Path,1 , R Path,2 , …, R Path,m}, to generate a preliminary query result set R Initial .
[0074] Optionally, the S7 specifically includes the following steps:
[0075] S71. Analyze the semantic information in the preliminary query result set by using a multi-level ontology set;
[0076] S72. Perform deduplication processing on the preliminary query result set R Initial to identify duplicate entities in the result entity set Based on the unique identifier e of the entity i ∈E Q and its corresponding multi-level ontology definition to determine duplicates, and identify duplicate relationships in the result relationship set Determine duplicates according to the upper and lower semantics of the relationship and the semantic connection rules associated in the target partition ;
[0077] S73. Analyze and optimize the semantic conflicts in the preliminary query result set R Initial . If an entity has multiple meanings in different partitions, then combine the query path Path QueryResolve the ambiguity with the context constraints defined by the multi - level ontology. If there are conflicting relationships in the resulting relationship set R Q optimize the conflicting relationships through the semantic inference rule set Q i so that each relationship meets the constraint conditions defined by the ontology;
[0078] S74. Utilize the semantic inference ability of the multi - level ontology to complement the implicit associated data in the preliminary query result set. If there is an entity e Q in the query result entity set E i whose associated attributes or relationships are not fully revealed, then through the semantic inference rule q k ∈Q i complement the missing associated attributes and relationships. If there are indirect associated relationships in the query result relationship set R Q inference and complement the implicit associated relationships through the cross - partition semantic connection strategy R i Connect ;
[0079] S75. Integrate the optimized query result set R Optimized into the final query result set:
[0080] R Final ={E Final ,R Final ,A Final}.
[0081] The beneficial effects of the present invention are:
[0082] (1) The present invention realizes the semantic modeling and standardized expression of multi - source heterogeneous data in the education field by constructing a multi - level ontology set, solves the problem of semantic inconsistency between multiple data sources in traditional methods, introduces a dynamic update mechanism, automatically adjusts the nodes, relationships and attributes in the knowledge graph according to the real - time changes of education data, combines the cross - partition semantic connection strategy to synchronously update the associated information between partitions, and performs consistency verification and implicit association complementation on the optimized data through the semantic inference rule set, ensuring the real - time nature and semantic integrity of the knowledge graph, and significantly improving the unity and query association efficiency of education data.
[0083] (2) The present invention combines a distributed computing framework, quickly locates the target partition through the global index and local index mechanism in cross - partition queries, and uses the semantic connection strategy to dynamically infer and generate cross - partition query paths. During the query process, parallel computing is used to perform local queries on each target partition, and the query results are dynamically integrated according to the predefined cross - partition semantic association rules to generate efficient cross - domain query paths.
[0084] (3) Based on the preliminary query results, the present invention deeply optimizes the results through the semantic reasoning ability of the multi-level ontology, eliminates duplicate results, resolves semantic conflicts, and complements implicit associated data. For the potential polysemy and implicit association relationships in the query results, automatic parsing and complementation are performed through dynamic semantic reasoning rules, thereby ensuring the semantic consistency and integrity of the query results. The mechanism for complementing implicit associated data in multi-domain association queries can effectively improve the comprehensiveness of the query results. BRIEF DESCRIPTION OF THE DRAWINGS
[0085] The accompanying drawings are used to provide a further understanding of the present invention, and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention, and do not constitute a limitation to the present invention. In the drawings:
[0086] Figure 1 is a flowchart of a method for cross-domain query of educational system data based on a knowledge graph proposed by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0087] Now, the present invention will be further described in detail with reference to the accompanying drawings. These drawings are all simplified schematic diagrams, only illustrating the basic structure of the present invention in a schematic manner, so they only show the components related to the present invention.
[0088] Refer to Figure 1 , a method for cross-domain query of educational system data based on a knowledge graph, includes the following steps:
[0089] S1. Based on the specific requirements in the education field, define the core concepts, attributes, and their hierarchical relationships, and construct a multi-level ontology for the education field;
[0090] S2. Based on the constructed multi-level ontology, partition and store the multi-source heterogeneous data in the education field according to subject fields, education stages, and regions. Establish an independent knowledge graph storage structure for each partition, and record the association relationships between partitions through semantic connection strategies;
[0091] S3. Set a dynamic update mechanism for each partition in the knowledge graph, update the graph nodes and relationships in the partition in real time according to the newly added or changed education data, and simultaneously use the semantic connection strategy to synchronously update the association information between partitions;
[0092] S4. Receive the natural language query request from the user, perform semantic parsing on the query request, identify the query intention, and generate a structured query statement. The structured query statement includes the target partition information to be retrieved and the query semantics;
[0093] S5. According to the target partition information and query semantics in the query request, use the query scheduling algorithm to locate relevant partitions in the distributed storage structure of the knowledge graph, and dynamically infer the semantic association relationships between the target partitions through the semantic connection strategy to generate a cross-partition query path;
[0094] S6. Based on the distributed computing framework, perform parallel query operations on the located knowledge graph partitions, and integrate the query results of each partition according to the cross-partition query path to form a preliminary query result set;
[0095] S7. Use the semantic reasoning ability of the multi-level ontology to perform semantic optimization on the preliminary query result set, eliminate duplicate results, resolve semantic conflicts and complete implicit associated data, and integrate the optimized query results into the final output result.
[0096] In this embodiment, S1 specifically includes the following steps:
[0097] S11. Determine the specific requirements in the education field, define the set of core concepts C in the education system, where the core concepts include subject topics, teaching objectives, learning objects, and resource types related to education, and define the set of its attributes A(c i ) for each core concept, and each attribute a j in the set of attributes represents the characteristic information of the concept c i ;
[0098] S12. Construct the set of hierarchical relationships between concepts by analyzing the logical relationships in the education field:
[0099] R = {r1, r2,..., r k};
[0100] where each relationship r k includes the hyponymy relationship subordination relationship and the semantic association relationship r r k elated -to;
[0101] S13. Establish the subject domain ontology Ontology i according to the set of core concepts C, the set of attributes A(c Domain ), and the set of relationships R. The subject domain ontology includes the knowledge models of specific disciplines;
[0102] S14. Divide the education system into primary school, junior high school, and high school according to the characteristics of the education stage, and construct the education stage ontology Ontology Stage . The education stage ontology defines the core concepts C Stage , the set of attributes A Stage , and the set of relationships RStage ;
[0103] S15. Define the regional ontology Ontology according to the characteristics of educational resources and standardization requirements in different regions Region , combine educational resources with regional characteristics, and the concept set C in the regional ontology Region includes educational policies and regional resource distributions;
[0104] S16. Associate the ontology of subject fields, the ontology of educational stages, and the regional ontology to form a multi-level ontology set:
[0105] Ontology Multi = {Ontology Domain , Ontology Stage , Ontology Region};
[0106] S17. Based on the multi-level ontology set Ontology Multi perform semantic structured modeling on the multi-source heterogeneous data in the education system, map the fields in the data to the concepts, attributes, and relationships in the multi-level ontology, and form a unified ontology basic framework.
[0107] In this embodiment, S2 specifically includes the following steps:
[0108] S21. Classify the multi-source heterogeneous data set D in the education field based on the multi-level ontology set. The data classification basis includes subject fields, educational stages, and regional characteristics;
[0109] S22. For the classified subject field set D Domain , educational stage set D Stage and regional characteristic set D Region , respectively construct knowledge graph partitions:
[0110] G i = {E i , R i , A i}, i ∈ {Domain, Stage, Region};
[0111] Among them, E i represents the entity set of partition i, including core concept entities and attribute entities, R i represents the relationship set of partition i, which is consistent with the relationship set defined in the multi-level ontology, and A i represents the attribute value set of partition i, which is used to describe the specific attribute data of the entity;
[0112] S23. In each knowledge graph partition G iAmong them, the semantic rule set Q of the multi - level ontology is used i to perform semantic reasoning and consistency checking on the data within the partition, eliminating redundant relationships and semantic conflicts within the partition;
[0113] S24. Based on the knowledge graph partition set G = {G Domain , G Stage , G Region}, design a cross - partition semantic connection strategy
[0114]
[0115] S25. Store the constructed knowledge graph partition set G in a distributed graph database, and establish a global index and a local index mechanism for each partition. The global index is used to quickly locate the partition category, and the local index is used to query specific entities and relationships within the partition;
[0116] S26. Combine the semantic connection strategy to dynamically synchronize and update the cross - partition association relationships. When the knowledge graph data in a certain partition changes, the semantic connection rules of its related partitions are automatically adjusted.
[0117] In this embodiment, S3 specifically includes the following steps:
[0118] S31. Set a dynamic update trigger mechanism for the knowledge graph partition set G to monitor the newly added education data set D New and the changed education data set D Update , and determine the update trigger conditions according to the data type, including new node addition, relationship update, or attribute modification;
[0119] S32. Within each knowledge graph partition G i , update the knowledge graph structure within the partition according to the semantic content of the newly added or changed education data set through the following rules:
[0120] If the data involves a new entity, add the new entity e i to the entity set E n ;
[0121] If the data involves a relationship change, add or modify the relationship r i in the relationship set R u ;
[0122] If the data involves an attribute update, add or modify the attribute value a i in the attribute set A u ;
[0123] S33. After updating the graph structure within the partition, use the semantic rule set Q within the partition iPerform consistency verification on the newly added or modified content, check the semantic consistency between the newly added entities and the existing entities, check whether the modified relationships conform to the constraints defined by the ontology, and check whether the attribute values conform to the field specifications;
[0124] S34. Based on the cross-partition semantic connection strategy Synchronously adjust the updated content associated with cross-partitions:
[0125] If the newly added entity involves the upper and lower position relationships across partitions, then update the relevant connection rules;
[0126] If the relationship change involves the subordinate relationship, then update the association rules;
[0127] If the attribute modification involves cross-domain semantic association, then update the connection rules;
[0128] S35. After the dynamic synchronous update is completed, update the index mechanism of the partitioned knowledge graph, for the global index I Global Update the partition location of the newly added or changed entities or relationships within the partition, for the local index I Local,i Update the specific node or attribute address of the newly added or changed content within the partition;
[0129] S36. Store the updated partitioned knowledge graph set G new back to the distributed graph database.
[0130] In this embodiment, S4 specifically includes the following steps:
[0131] S41. Receive the user's natural language query request Q NL , and the query request includes the query keywords or sentences input by the user;
[0132] S42. Use the semantic parsing algorithm to parse the natural language query request Q NL Perform word segmentation on the user's query request, extract the keyword set K in the query, perform semantic matching on the keyword set K based on the multi-level ontology set Ontology Multi , identify the concepts, attributes, and relationships corresponding to the keywords, and resolve the polysemy of the keywords in combination with the user's query context or historical behavior to generate the query intention set I Q ;
[0133] S43. Locate the knowledge graph partitions related to the query according to the concepts, attributes, or relationships in the query intention set I Q Represent the semantic content corresponding to the matched query intention as the query semantics S , where S Q Q Including querying the target entity set E Q , querying the target relationship set R Q and querying the target attribute set A Q , combining the target partition information P and the query semantics S Q to generate a structured query statement Q Structured =(P, S Q );
[0134] S44. In the generated structured query statement Q Structured , verify the logical consistency of the target partition information and the query semantics, so that the structured query statement meets the following conditions:
[0135] The entities in the query target entity set E Q can be mapped to the entity set E within the partition i ;
[0136] The relationships in the query target relationship set R Q can match the relationship set R within the partition i ;
[0137] The attributes in the query target attribute set A Q can match the attribute set A within the partition i .
[0138] In this embodiment, S5 specifically includes the following steps:
[0139] S51. Extract the knowledge graph partition set to be retrieved from the target partition set P according to the structured query statement Q Structured ;
[0140] S52. Lock the target partition according to the partition type through the global index I Global , for each target partition G i ∈P, determine the specific positions of the target entity set, the target relationship set and the target attribute set within the partition through the local index I Local,i ;
[0141] S53. Based on the target partition set P and the query semantics, combine the cross-partition semantic connection strategy to perform dynamic semantic reasoning, generate a cross-partition semantic association path, and combine the cross-partition semantic connection strategy to perform dynamic semantic reasoning to generate a cross-partition semantic association path Path Cross :
[0142] If the target entity set in the query semantics S Q involves the hyponymy relationship between partitions G i and G j , then infer the hyponymy path across partitions through the semantic connection rule ;
[0143] If the target relationship set in the query semantics S Q involves a subordination relationship of sub - intervals, then through the semantic connection rule infer the subordination path across partitions;
[0144] If the target attribute set in the query semantics S Q involves cross - domain semantic associations, then through the semantic connection rule infer the semantic association path across partitions;
[0145] S54. Utilize the semantic association path Path Cross to dynamically integrate the query results between target partitions and generate a cross - partition query path:
[0146] Path Query ={Path1, Path2, …, Path m}.
[0147] In this embodiment, S6 specifically includes the following steps:
[0148] S61. Based on the cross - partition query path set, determine the target knowledge graph partition set corresponding to each path and the relevant query semantics;
[0149] S62. In the distributed computing framework, for each cross - partition query path Path k allocate a computing task, and for each partition G in each path i allocate an independent computing node N i , and for each computing node N i allocate the target query semantics;
[0150] S63. In each computing node N i , execute a query operation according to the local knowledge graph partition and the query semantics, and generate a local partition query result set R Local,i ;
[0151] S64. Based on the cross - partition query path Path k , according to the semantic connection strategy in the path dynamically integrate the local partition query result set R Local,i to generate a path - level query result set R Path,k :
[0152] If the path Path k involves a hyponym - hypernym semantic relationship, then connect the partition query results according to ;
[0153] If the path Path k involves a subordination semantic relationship, then according to Connect the partition query results;
[0154] If the path Path k involves semantic association relationships, then based on Connect the partition query results;
[0155] S65. Aggregate all path-level query result sets R Path ={R Path,1 , R Path,2 , …, R Path,m}, and generate the preliminary query result set R Initial .
[0156] In this embodiment, S7 specifically includes the following steps:
[0157] S71. Analyze the semantic information in the preliminary query result set using a multi-level ontology set;
[0158] S72. Deduplicate the preliminary query result set R Initial and identify duplicate entities in the result entity set Based on the unique identifier e of the entity i ∈E Q and its corresponding multi-level ontology definition to determine duplicates, and identify duplicate relationships in the result relationship set Determine duplicates according to the hyponymy semantics of the relationship and the semantic connection rules associated in the target partition ;
[0159] S73. Analyze and optimize the semantic conflicts in the preliminary query result set R Initial . If an entity has multiple meanings in different partitions, then combine the query path Path Query and the context constraints defined in the multi-level ontology to resolve the multiple meanings. If there are conflicting relationships in the result relationship set R Q , then optimize the conflicting relationships through the semantic inference rule set Q i so that each relationship satisfies the constraints defined by the ontology;
[0160] S74. Use the semantic inference ability of the multi-level ontology to complete the implicit association data in the preliminary query result set. If the associated attributes or relationships of an entity e Q in the query result entity set E i are not fully manifested, then complete the missing associated attributes and relationships through the semantic inference rule q k ∈Q i . If there are indirect association relationships in the query result relationship set R Q , then infer and complete the implicit association relationships through the cross-partition semantic connection strategy ;
[0161] S75. Integrate the optimized query result set R Optimized into the final query result set:
[0162] R Final = {E Final , R Final , A Final}.
[0163] Example 1:
[0164] In the construction project of the basic education informatization platform of a certain city's Education Bureau, the city hopes to achieve cross - regional, cross - disciplinary and cross - educational stage data sharing and intelligent query functions to support the comprehensive evaluation of students' academic performance and accurate resource recommendation. However, due to decentralization and heterogeneity, the existing education data management systems cannot efficiently integrate data from different schools, grades and disciplines. Therefore, the city's Education Bureau decides to adopt the method of the present invention to build a cross - domain query platform for education system data based on a knowledge graph to solve the problems of data dispersion, semantic inconsistency and low query efficiency.
[0165] The data sources in this city include students' achievement data, teacher evaluation data and educational resource data of 60 primary and secondary schools within the region. The data is distributed in different subsystems, including school management systems, online course platforms and teaching resource libraries, and the forms include structured achievement records, semi - structured teaching plan files, and unstructured classroom videos and assignment texts. Previous data queries only supported simple retrieval in a single dimension and could not correlate data between disciplines or educational stages. For example, a certain school hopes to query "the list of students with excellent mathematics scores in primary school and rapid improvement in English scores in junior high school". Traditional methods require separate retrieval in primary and junior high school systems and then manual comparison and analysis, which is time - consuming, labor - intensive and error - prone.
[0166] The whole process of applying the method of the present invention in this project is as follows:
[0167] First, aiming at the dispersion and heterogeneity of the city's education data, a multi - level ontology modeling method is adopted to divide the data into three levels: subject area, educational stage and region. The ontology model defines core concepts (such as "student", "achievement", "resource"), attributes (such as "subject name", "time node") and relationships (such as "achievement belongs to student"). Through the modeling of historical data, a multi - level ontology set is formed, and the data of each school is standardized. For example, the primary school mathematics achievement records were originally stored in Excel files, while the junior high school English achievement data was stored in database tables. After modeling, these data are uniformly mapped into the knowledge graph and represented as "student - subject - achievement" triples.
[0168] Then, using the knowledge graph partition storage technology, the data is partitioned and stored according to subject fields, educational stages, and school regions. For example, data related to primary school mathematics is stored in and junior high school English data is stored in Meanwhile, semantic connection rules across partitions are established. For example, for cross-stage data association of "students", the upper and lower position relationships are defined to represent the data association of a certain student in primary school and junior high school.
[0169] In an actual application scenario, a key middle school in a certain city hopes to query "the list of students who ranked in the top 10% in mathematics in primary school and had a significant improvement in English in junior high school in the short term (the recent two years)". Using traditional methods, this query requires retrieving in the primary school and junior high school systems respectively, exporting in tabular form, and then manually comparing and calculating according to the student ID. This process usually takes about 7 hours for an information management staff to complete, and due to data polysemy and inconsistent formats, the accuracy rate is less than 80%.
[0170] After adopting the method of the present invention, the user inputs a query request "find students who ranked in the top 10% in primary school mathematics and had a significant improvement in junior high school English" through natural language. The system first uses semantic parsing technology to segment and analyze the intent of the query request, identifies the key semantics such as "mathematics score", "English score", "top 10%", and "improvement range", and maps them to the corresponding concepts and attributes in the knowledge graph.
[0171] The system quickly locks the target partitions through a distributed indexing mechanism and During the parallel query execution, the primary school mathematics data partition returns the entity set E of students who ranked in the top 10% in scores Top10 , and the junior high school English data partition returns the entity set E of students with significant improvement according to the score change range within two years Improved . Subsequently, the system integrates the two entity sets according to the cross-partition semantic connection rules to generate a list of students meeting the conditions as the preliminary query result.
[0172] In the result optimization stage, the system complements and deduplicates the query results through the semantic reasoning rules of the multi-level ontology. For example, the names of a certain student in primary school and junior high school with slight differences caused by the pinyin input method are misidentified as two entities. After semantic reasoning, the system automatically merges them into the same entity and complements the missing attribute information. Finally, the system outputs a list of 25 students, and at the same time provides the mathematics ranking of each student, the English score change curve, and the recommended learning resources associated.
[0173] To verify the performance of the method of the present invention, data of 15,000 students from three schools in the city were selected for comparative testing. The results showed that the traditional method took 420 minutes with an accuracy rate of 78%; the method of the present invention took only 15 minutes and the accuracy rate was increased to 96%. The specific data are shown in Table 1 below:
[0174] Table 1 Comparative data between the present invention and the traditional method
[0175] Method Query Time (minutes) Query Accuracy Implicit Data Completion Rate Traditional Method 420 78% None Method of the Present Invention 15 96% 92%
[0176] In addition, in the test of the implicit associated data completion ability, the method of the present invention successfully completed 92% of the missing attribute information based on semantic reasoning rules, while the traditional method completely relied on manual analysis and could not complete the completion. This result fully proves the superiority and practical value of the method of the present invention in cross-domain query of educational system data.
[0177] The present invention realizes semantic modeling and standardized expression of multi-source heterogeneous data in the education field by constructing a multi-level ontology set, solves the problem of semantic inconsistency between multiple data sources in the traditional method, introduces a dynamic update mechanism, automatically adjusts the nodes, relationships and attributes in the knowledge graph according to the real-time changes of educational data, combines the cross-partition semantic connection strategy to synchronously update the associated information between partitions, and performs consistency verification and implicit association completion on the optimized data through a set of semantic reasoning rules, ensuring the real-time and semantic integrity of the knowledge graph, and significantly improving the unity and query association efficiency of educational data.
[0178] The present invention combines a distributed computing framework, quickly locates the target partition through a global index and a local index mechanism in cross-partition query, and uses a semantic connection strategy to dynamically infer and generate a cross-partition query path. In the query process, a parallel computing method is used to perform local queries on each target partition, and the query results are dynamically integrated according to predefined cross-partition semantic association rules to generate an efficient cross-domain query path.
[0179] Based on the preliminary query results, the present invention deeply optimizes the results through the semantic reasoning ability of the multi-level ontology, eliminates duplicate results, resolves semantic conflicts and completes implicit associated data. For the potential polysemy and implicit association relationships in the query results, automatic parsing and completion are performed through dynamic semantic reasoning rules, so as to ensure the semantic consistency and integrity of the query results. The mechanism for completing implicit associated data in multi-domain association query can effectively improve the comprehensiveness of the query results.
[0180] The above are only the preferred specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and inventive concept of the present invention, making equivalent substitutions or changes should be covered within the protection scope of the present invention.
Claims
1. A cross-domain query method for education system data based on knowledge graph, characterized in that: The steps include: S1. Based on the specific needs of the education field, define the core concepts, attributes and their hierarchical relationships, and construct a multi-level ontology in the education field; S2. Based on the constructed multi-level ontology, the multi-source heterogeneous data in the field of education are partitioned and stored according to subject areas, educational stages and regions. An independent knowledge graph storage structure is established for each partition, and the association relationship between the partitions is recorded through the semantic connection strategy; S3. Set up a dynamic update mechanism for each partition in the knowledge graph, update the graph nodes and relationships in the partition in real time according to the newly added or changed education data, and use the semantic connection strategy to synchronously update the association information between partitions; S4. Receive a natural language query request from a user, perform semantic analysis on the query request, identify the query intent and generate a structured query statement, the structured query statement includes target partition information to be retrieved and query semantics; S5. Based on the target partition information and query semantics in the query request, the query scheduling algorithm is used to locate the relevant partitions in the distributed storage structure of the knowledge graph, and the semantic association relationship between the target partitions is dynamically inferred through the semantic connection strategy to generate a cross-partition query path; S6. Perform parallel query operations on the located knowledge graph partitions based on the distributed computing framework, integrate the query results of each partition according to the cross-partition query path, and form a preliminary query result set; S7. Use the semantic reasoning capabilities of multi-level ontologies to semantically optimize the preliminary query result set, eliminate duplicate results, resolve semantic conflicts, and complete implicit related data, and integrate the optimized query results into the final output results.
2. According to claim 1, a cross-domain query method for education system data based on knowledge graph is characterized in that: The S1 specifically includes the following steps: S11. Determine the specific needs in the field of education and define the core concept set C in the education system. The core concepts include education-related subject themes, teaching objectives, learning objects, and resource types. For each core concept, define its attribute set A (c i ), each attribute a in the attribute set j Represents concept c i Characteristics information; S12. Construct a hierarchical relationship set between concepts by analyzing the logical relationships in the field of education: R={r1,r2,…,r k}; Among them, each relationship r k Including the hierarchical relationship between concepts Dependencies and semantic association S13. According to the core concept set C and attribute set A (c i ) and the relationship set R to establish the subject domain ontology Domain ,Discipline domain ontology includes the knowledge model of a specific discipline; S14. Divide the education system into primary school, junior high school and high school according to the characteristics of the education stage, and construct the education stage ontology Stage , the education stage ontology defines the core concepts of different stages C Stage , attribute set A Stage and the relation set R Stage ; S15. Define regional ontology based on the characteristics of educational resources and standardization requirements in different regions Region , combining educational resources with regional characteristics, the concept set C in the regional ontology Region including education policies and regional resource distribution; S16. Associate the subject domain ontology, education stage ontology and region ontology to form a multi-level ontology set: Ontology Multi ={Ontology Domain ,Ontology Stage ,Ontology Region }; S17. Based on multi-level ontology collection Ontology Multi Perform semantic structured modeling on multi-source heterogeneous data in the education system, map the fields in the data with the concepts, attributes and relationships in the multi-level ontology, and form a unified ontology basic framework.
3. According to claim 1, a cross-domain query method for education system data based on knowledge graph is characterized in that: The S2 specifically includes the following steps: S21. Classify the multi-source heterogeneous data set D in the field of education based on a multi-level ontology set, where the data classification is based on subject areas, educational stages, and regional characteristics; S22. For the classified subject area set D Domain , Education stage set D Stage and the set of regional characteristics D Region , respectively construct knowledge graph partitions: G i ={E i ,R i ,A i },i∈{Domain,Stage,Region}; Among them, E i Represents the entity set of partition i, including core concept entities and attribute entities, R i Represents the relationship set of partition i, which is consistent with the relationship set defined in the multi-level ontology. i Represents the attribute value set of partition i, which is used to describe the specific attribute data of the entity; S23. In each knowledge graph partition G i In the above example, we use the semantic rule set Q of the multi-level ontology i Perform semantic reasoning and consistency checks on the data within the partition to eliminate redundant relationships and semantic conflicts within the partition; S24. Based on the knowledge graph partition set G = {G Domain ,G Stage ,G Region Designing semantic connection strategies across partitions S25. Store the constructed knowledge graph partition set G in a distributed graph database, and establish a global index and a local index mechanism for each partition. The global index is used to quickly locate the partition category, and the local index is used to query specific entities and relationships within the partition; S26. Combined semantic connection strategy The cross-partition association relationships are dynamically and synchronously updated. When the knowledge graph data in a partition changes, the semantic connection rules of its related partitions are automatically adjusted.
4. According to claim 1, a cross-domain query method for education system data based on knowledge graph is characterized in that: The S3 specifically includes the following steps: S31. Set a dynamic update trigger mechanism for the knowledge graph partition set G to monitor the newly added education data set D New and change the education data set D Update ,determine the update trigger conditions based on the data type, including adding new nodes, relationship updates, or attribute modifications; S32. In each knowledge graph partition G i According to the semantic content of the newly added or changed educational data set, the knowledge graph structure in the partition is updated according to the following rules: If the data involves a new entity, then in the entity set E i Add a new entity to n ; If the data involves a relationship change, then in the relationship set R i Add or modify a relationship in u ; If the data involves attribute updates, then in the attribute set A i Add or modify attribute value a u ; S33. After updating the graph structure in the partition, use the semantic rule set Q in the partition i Perform consistency checks on new or modified content, check whether the new entities are semantically consistent with the existing entities, check whether the modified relationships comply with the constraints defined in the ontology, and check whether the attribute values comply with the field specifications; S34. Semantic connection strategy based on cross-partition Synchronize and adjust the updated content associated across partitions: If the newly added entity involves a cross-partition hierarchical relationship, update The relevant connection rules; If the relationship change involves a subordinate relationship, update Association rules of If the attribute modification involves cross-domain semantic association, update Connection rules; S35. After the dynamic synchronization update is completed, the index mechanism of the partitioned knowledge graph is updated, and the global index I Global Update the partition location of the newly added or changed entity or relationship in the partition, and update the local index I Local,i Update the specific node or attribute address of the newly added or changed content in the partition; S36. Partition the updated knowledge graph into a set G new Stored back to the distributed graph database.
5. According to claim 1, a cross-domain query method for education system data based on knowledge graph is characterized in that: The S4 specifically comprises the following steps: S41. Receive the user's natural language query request Q NL , the query request includes a query keyword or sentence input by a user; S42. Using semantic parsing algorithm to query natural language Q NL Parse and segment the user's query request, extract the keyword set K in the query, and then use the multi-level ontology set Ontology Multi Perform semantic matching on the keyword set K, identify the concepts, attributes and relationships corresponding to the keywords, resolve the ambiguity of keywords based on the user's query context or historical behavior, and generate the query intent set I Q ; S43. According to the query intention set I Q The concepts, attributes or relations in the knowledge graph are located in the partitions related to the query The semantic content corresponding to the matched query intent is represented as query semantics S Q , where S Q Including the query target entity set E Q , query target relation set R Q and query target attribute set A Q , combining the target partition information P and the query semantics S Q Generate structured query statement Q Structured =(P,S Q ); S44. In the generated structured query statement Q Structured In the above example, the logical consistency of the target partition information and the query semantics is verified so that the structured query statement satisfies the following conditions: Query the target entity set E Q The entities in can be mapped to the entity set E in the partition i ; Query target relation set R Q The relations in can match the relation set R in the partition i ; Query target attribute set A Q The attributes in can match the attribute set A in the partition i .
6. According to claim 1, a cross-domain query method for education system data based on knowledge graph is characterized in that: The S5 specifically includes the following steps: S51. According to the structured query statement Q Structured Extract the knowledge graph partition set to be retrieved from the target partition set P; S52. Based on the partition type, use the global index I Global Lock the target partition, for each target partition G i ∈P, through the local index I Local,i Determine the specific locations of the target entity set, target relationship set, and target attribute set within the partition; S53. Based on the target partition set P and query semantics, dynamic semantic reasoning is performed in combination with the cross-partition semantic connection strategy to generate a cross-partition semantic association path. Dynamic semantic reasoning is performed in combination with the cross-partition semantic connection strategy to generate a cross-partition semantic association path Path Cross : If the query semantics S Q The target entity set in partition G i and G j The relationship between the upper and lower levels is determined by the semantic connection rule. Reasoning about superior and subordinate paths across partitions; If the query semantics S Q The target relationship set in involves the subordinate relationship between partitions, so the semantic connection rule Reasoning about dependency paths across partitions; If the query semantics S Q The target attribute set in involves cross-domain semantic associations, so the semantic connection rules are used Manage the semantic association paths across partitions; S54. Using semantic association path Path Cross Dynamically integrate the query results between target partitions to generate a cross-partition query path: Path Query ={Path1,Path2,…,Path m }。 7. According to claim 1, a cross-domain query method for education system data based on knowledge graph is characterized in that: The S6 specifically comprises the following steps: S61. Determine the target knowledge graph partition set and related query semantics corresponding to each path based on the cross-partition query path set; S62. In the distributed computing framework, for each cross-partition query path Path k Assign computing tasks to each partition G in the path i Allocate independent computing nodes N i , for each computing node N i Assign target query semantics; S63. At each computing node N i In the process, query operations are performed based on local knowledge graph partitions and query semantics to generate a local partition query result set R Local,i ; S64. Query path based on cross-partition k , according to the semantic connection strategy in the path Query the local partition result set R Local,i Perform dynamic integration to generate path-level query result set R Path,k : If the path Path k If it involves a hyponymous semantic relationship, then the Connect partition query results; If the path Path k Involving subordinate semantic relations, Connect partition query results; If the path Path k If it involves semantic association, then according to Connect partition query results; S65. Summarize all path-level query result sets R Path = {R Path,1 ,R Path,2 ,…,R Path,m }, generate a preliminary query result set R Initial .
8. According to claim 1, a cross-domain query method for education system data based on knowledge graph is characterized in that: The S7 specifically comprises the following steps: S71. Analyze the semantic information in the preliminary query result set using a multi-level ontology set; S72. Preliminary query result set R Initial Perform deduplication processing to identify duplicate entities in the result entity set Unique identifier based on entity i ∈E Q Determine duplication with its corresponding multi-level ontology definition and identify duplicate relations in the result relation set According to the semantics of the relationship and the semantic connection rules associated in the target partition Determine duplication; S73. Preliminary query result set R Initial , and if the same entity has ambiguity in different partitions, the query path Path Query and the context constraints defined by the multi-level ontology to resolve ambiguity. If the resulting relation set R Q If there is a conflict relationship in i Optimize conflicting relationships so that each relationship satisfies the constraints defined in the ontology; S74. Use the semantic reasoning ability of multi-level ontology to complete the implicit related data in the initial query result set. If the query result entity set E Q Entity i If the associated attributes or relationships are not fully revealed, the semantic reasoning rule q k ∈Q i Complete the missing associated attributes and relationships. If the query result relationship set R Q If there is an indirect association relationship, the cross-partition semantic connection strategy is used Reasoning to complete implicit associations; S75. The optimized query result set R Optimized Integrate into the final query result set: R Final ={E Final ,R Final ,A Final }。
Citation Information
Patent Citations
Training scoring method based on natural language
CN118467985A
Multi-layer semantic understanding large model agent construction and application method
CN119089931A