Mapping knowledge graph-based ideological and educational knowledge recommendation and question and answer method, device and component
By constructing an ideological and educational knowledge graph based on the knowledge graph, using TransH and GNN models for relationship and structure analysis, and combining RAG technology to achieve intelligent question and answer, the problem of low efficiency in information acquisition and question and answer in the existing ideological and educational system is solved, and more accurate knowledge recommendation and question and answer are achieved, thereby improving user learning effects.
Patent Information
- Application Number
- CN202510737960.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-04
- Publication Date
- 2025-09-12
Smart Images

Figure CN120632042A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a knowledge graph-based knowledge recommendation and question-answering method, device, and component. Background Art
[0002] With the continuous development of information technology and the advancement of digital transformation, traditional ideological and educational methods have gradually exposed some shortcomings, such as inefficient information acquisition and processing, inaccurate content recommendations, and limited knowledge question-and-answer capabilities. Existing ideological and educational systems primarily rely on matching retrieval and traditional teaching methods, which are unable to meet the needs of personalization, intelligence, and real-time updates. With the development of artificial intelligence technologies, knowledge graphs, deep learning, and large language models (LLMs) have brought new approaches to the management, understanding, and analysis of ideological and educational knowledge systems.
[0003] In the existing technology, knowledge graphs, as an effective technology for representing and managing large-scale knowledge, have been applied in many fields. However, their application in the field of ideological education is still in its early stages. In particular, they still face great challenges in building high-quality ideological and educational knowledge graphs and graph-based intelligent recommendation systems and question-answering systems. At present, the construction of ideological and educational knowledge graphs relies on basic deep learning models or LLM to extract entities, but these methods still have certain defects, such as: (1) lack of domain recognition capabilities or poor extraction results due to AI hallucinations; (2) existing recommendation systems are mostly based on traditional deep learning recommendation algorithms and lack of knowledge structured representation; (3) existing question-answering systems are also often limited by traditional rule matching models or AI hallucinations, lack of deep reasoning, intelligent analysis, or reduced professionalism due to AI hallucinations. Summary of the Invention
[0004] The embodiments of the present invention provide a method, apparatus, computer equipment and storage medium for recommending and answering ideological and educational knowledge based on a knowledge graph, aiming to improve the accuracy of query recommendation and question answering of ideological and educational knowledge, thereby improving the ideological and educational learning effect of users.
[0005] In a first aspect, an embodiment of the present invention provides a method for recommending and answering educational knowledge based on a knowledge graph, including:
[0006] Obtaining Sijiao resource data, extracting entities from the Sijiao resource data using a pre-built target model structure to obtain corresponding triple data, and constructing a Sijiao knowledge graph based on the triple data; wherein the target model structure is constructed using the BERT model, a bidirectional long short-term memory network, and a conditional random field, and an adapter is used to fine-tune parameters during the construction process;
[0007] A visualization interface is provided for the Sijiao knowledge graph, and in response to a user's search call to the visualization interface, the Sijiao knowledge graph is displayed through an Echarts chart; wherein the visualization interface provides multiple search methods;
[0008] Perform word semantic relationship analysis on the entity relationships in the Sijiao knowledge graph using the TransH model, and perform structural relationship analysis on the nodes and edges in the Sijiao knowledge graph using the GNN model, and systematically recommend entities based on the results of the word semantic relationship analysis and the structural relationship analysis;
[0009] Generate a question-and-answer question based on the triple data, and store the question-and-answer question as a key and the corresponding triple data as a value in a preset vector database;
[0010] In response to the user's query operation on the Sijiao knowledge graph, the vector database is searched and matched based on the RAG technology, and the question and answer results corresponding to the query operation are returned in combination with the AI big model based on the search and matching results.
[0011] In a second aspect, an embodiment of the present invention provides a knowledge graph-based knowledge recommendation and question-answering device, comprising:
[0012] A graph construction unit is used to obtain Sijiao resource data, extract entities from the Sijiao resource data using a pre-built target model structure to obtain corresponding triple data, and construct a Sijiao knowledge graph based on the triple data; wherein the target model structure is obtained by building a BERT model, a bidirectional long short-term memory network, and a conditional random field, and an adapter is used to fine-tune parameters during the construction process;
[0013] A graph visualization unit is used to set a visualization interface for the Sijiao knowledge graph and display the Sijiao knowledge graph through an Echarts chart in response to a user's search call to the visualization interface; wherein the visualization interface provides multiple search methods;
[0014] An entity recommendation unit, configured to perform word semantic relationship analysis on entity relationships in the Sijiao knowledge graph using a TransH model, and perform structural relationship analysis on nodes and edges in the Sijiao knowledge graph using a GNN model, and systematically recommend entities based on the results of the word semantic relationship analysis and the structural relationship analysis;
[0015] A database building unit, configured to generate a question-and-answer question based on the triple data, and store the question-and-answer question as a key and the corresponding triple data as a value in a preset vector database;
[0016] The question and answer matching unit is used to respond to the user's query operation on the Sijiao knowledge graph, search and match the vector database based on RAG technology, and return the question and answer result corresponding to the query operation based on the search and matching results in combination with the AI big model.
[0017] In a third aspect, an embodiment of the present invention provides a computer device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the knowledge graph-based teaching knowledge recommendation and question-answering method as described in the first aspect is implemented.
[0018] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the computer program implements the knowledge graph-based teaching knowledge recommendation and question-answering method as described in the first aspect.
[0019] The embodiment of the present invention extracts entities and constructs a knowledge graph of ideological and political education through an improved Adapter fine-tuning model and algorithm optimization, realizes graph visualization of entity retrieval through a visualization interface, and uses the ideological and political education knowledge graph to realize systematic graph recommendation based on TransH and graph neural network GNN technology. It also uses the constructed ideological and political education knowledge graph as a framework to realize professional knowledge question and answer based on retrieval enhancement generation RAG, which can improve the query recommendation and question and answer accuracy of ideological and political education knowledge, thereby improving the ideological and political education learning effect of users, thereby promoting the digital transformation of ideological and political education, and providing innovative technical support for the intelligentization of ideological and political education. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0021] Figure 1 A flowchart of a method for recommending and answering questions of educational knowledge provided by an embodiment of the present invention;
[0022] Figure 2 A network architecture diagram of a target model structure in a method for recommending educational knowledge and answering questions provided by an embodiment of the present invention;
[0023] Figure 3 A network architecture diagram of an adapter in a method for recommending educational knowledge and answering questions provided by an embodiment of the present invention;
[0024] Figure 4A diagram showing a visualization of a method for recommending and answering questions in accordance with an embodiment of the present invention;
[0025] Figure 5 A system recommendation architecture diagram for a method for recommending educational knowledge and answering questions provided by an embodiment of the present invention;
[0026] Figure 6 A diagram of a knowledge question-answering architecture in a method for recommending and answering knowledge in accordance with an embodiment of the present invention;
[0027] Figure 7 This is a diagram of the overall architecture of a method for recommending and answering questions about educational knowledge provided by an embodiment of the present invention;
[0028] Figure 8 A schematic block diagram of a device for recommending educational knowledge and answering questions provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0029] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0030] It will be understood that when used in this specification and the appended claims, the terms “comprises” and “comprising” indicate the presence of described features, integers, steps, operations, elements and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof.
[0031] It should also be understood that the terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the present invention. As used in the specification and appended claims, the singular forms "a," "an," and "the" are intended to include the plural forms unless the context clearly indicates otherwise.
[0032] It should be further understood that the term "and / or" used in the present description and the appended claims refers to and includes any and all possible combinations of one or more of the associated listed items.
[0033] See below Figure 1 The embodiment of the present invention provides a method for recommending and answering educational knowledge based on a knowledge graph, which specifically includes steps S101 to S105.
[0034] Step S101: Acquire ideological and political education resource data, extract entities from the ideological and political education resource data using a pre-built target model structure to obtain corresponding triple data, and construct an ideological and political education knowledge graph based on the triple data; wherein the target model structure is obtained by building a BERT model, a bidirectional long short-term memory network, and a conditional random field, and an adapter is used to fine-tune parameters during the construction process;
[0035] Step S102: Setting a visualization interface for the Sijiao knowledge graph, and displaying the Sijiao knowledge graph through an Echarts chart in response to a user's search call to the visualization interface; wherein the visualization interface provides multiple search methods;
[0036] Step S103: Perform word semantic relationship analysis on the entity relationships in the Sijiao knowledge graph using the TransH model, and perform structural relationship analysis on the nodes and edges in the Sijiao knowledge graph using the GNN model, and systematically recommend entities based on the results of the word semantic relationship analysis and the structural relationship analysis;
[0037] Step S104: Generate a question and answer based on the triple data, and store the question and answer as a key and the corresponding triple data as a value in a preset vector database;
[0038] Step S105: In response to the user's query operation on the Sijiao knowledge graph, the vector database is searched and matched based on the RAG technology, and the question and answer results corresponding to the query operation are returned in combination with the AI big model based on the search and matching results.
[0039] In this embodiment, the BERT-BiLSTM-CRF target model structure is first used to extract entities from the Sijiao resource data, and the adapter is used to fine-tune the optimization parameters to generate structured triple data, and on this basis, the Sijiao knowledge graph is constructed. Then, a multimodal visualization interface is built through the Echarts chart engine, which supports multiple retrieval methods and provides a dynamic and interactive knowledge graph display. The TransH model is then used to parse the word semantic relationship between entities, and the GNN model is used to mine the graph topology structure features, thereby integrating the word semantic relationship and the graph structure features to achieve systematic knowledge recommendation. At the same time, question and answer pair data is generated based on the RAG framework, and the "question-triplet" key-value pair structure is used to store it in the vector database, providing structured knowledge support for intelligent question and answer. When the user queries and searches, the vector database is searched and matched through retrieval generation technology, and the corresponding context is generated for the search and matching results in combination with the AI large model, so as to return the required question and answer results to the user.
[0040] This embodiment uses an improved Adapter fine-tuning model and algorithm to optimize the extraction of entities and construct a knowledge graph for ideological and political education. At the same time, it realizes the graph visualization of entity retrieval through a visualization interface, and uses the ideological and political education knowledge graph to realize systematic graph recommendation based on TransH and graph neural network GNN technology. It also uses the constructed ideological and political education knowledge graph as a framework to realize professional knowledge question and answer based on retrieval enhancement and generation of RAG. This can improve the query recommendation and question and answer accuracy of ideological and political education knowledge, thereby improving the ideological and political education learning effect of users, thereby promoting the digital transformation of ideological and political education, and providing innovative technical support for the intelligentization of ideological and political education.
[0041] In one embodiment, the acquiring of ideological and educational resource data, extracting entities from the ideological and educational resource data using a pre-built target model structure to obtain corresponding triple data, and constructing an ideological and educational knowledge graph based on the triple data, includes:
[0042] Inputting the thought and teaching resource data into the target model structure;
[0043] Perform contextual semantic encoding on the Sijiao resource data through the BERT model in the target model structure to obtain corresponding word vectors;
[0044] Using a bidirectional long short-term memory network to capture context dependencies on the word vectors, and obtain corresponding sequence features;
[0045] Using a conditional random field to perform sequence labeling on the sequence features to obtain a primary entity label and a secondary entity label;
[0046] The secondary entity labels are mapped to relationships to construct triple data including main entity-relationship-secondary entity, and the triple data is used to construct the Sijiao knowledge graph.
[0047] Combine Figure 2 This example uses the BERT-BiLSTM-CRF target model structure to extract entities from Sijiao resource data. The BERT model uses a Transformer encoder to capture contextual information. After the BERT model outputs the hidden state, the BiLSTM uses a bidirectional approach to capture contextual dependencies within the text, enabling a better understanding of the words in the sequence. Finally, the CRF sequence labeling model post-processes the labels output by the BiLSTM to optimize the dependencies between the labels, ultimately obtaining accurate entities.
[0048] The actual training process begins by acquiring unstructured, semi-structured, and structured online resources. Because unstructured data often contains web page tags or illegal characters, it is necessary to preprocess the data and create a dataset, removing all content except for text and punctuation. Because the model must be trained before testing, only a portion of the text is selected as the training and test sets, with the remaining content used as the validation set. After preprocessing the unstructured data, a BERT-BiLSTM-CRF model fine-tuned using the adapter is trained and tested on the training set. Finally, after validating the model on the validation set, entity extraction is performed on the entire dataset. Secondary entity labels are mapped to relations to extract primary entity-relationship-secondary entity triplets. For example, in the text "Student A once joined Association B," "Student A Deng" will be identified as NAME (personal name) and used as the primary entity, and "Association B" will be identified as ORG (organization) and used as the secondary entity. This allows the extraction of the triple "Student A - Organization - Association B," where the relationship "Organization" is mapped from the secondary entity label ORG. Finally, the semi-structured and structured data are integrated and imported into the Neo4j graph database to store the Sijiao knowledge graph.
[0049] Specifically, the contextual semantic encoding of the teaching resource data by the BERT model in the target model structure to obtain the corresponding word vector includes:
[0050] Embed the adapter into the BERT model, and use the adapter to sequentially perform downsampling processing, nonlinear activation function activation processing, and upsampling processing on the input data of the BERT model to obtain corresponding output results;
[0051] The output of the adapter is input into the Transformer layer through the residual block, and the Transformer layer is used to encode and capture context information.
[0052] In this embodiment, the BERT model is a pre-trained deep learning model specifically designed for language understanding tasks. The BERT model uses the Transformer encoder to capture contextual information and can be pre-trained based on a large amount of corpus. In order to avoid adjusting the parameters of the entire BERT model every time training, this embodiment uses the Adapter fine-tuning method to only add a small number of trainable parameters to some layers of BERT. Specifically, Figure 3As shown, an Adapter network is embedded in each layer of BERT, and each Adapter module consists of a two-layer bottleneck structure. First, the Adapter downsamples the input, then adds a nonlinear activation function, and then performs upsampling to remap the low-dimensional representation back to the original high-dimensional space. Finally, through a residual connection, it is directly added back to the input of the Transformer layer, thereby avoiding affecting the representation ability of the pre-trained model. This embodiment can help the model retain the powerful pre-training ability of BERT through Adapter fine-tuning while adapting to specific tasks with fewer parameter adjustments, significantly improving training efficiency.
[0053] In one embodiment, the knowledge graph-based teaching knowledge recommendation and question-answering method further includes:
[0054] According to the following formula, the initial loss function is set, and the flooding mechanism is used to dynamically control the initial loss function to construct the target loss function:
[0055]
[0056] in, represents the target loss function, L(θ) represents the initial loss function, and b represents the hyperparameter introduced by the flooding mechanism;
[0057] The target model structure is trained and updated using the target loss function.
[0058] During training, this embodiment introduces a flooding mechanism to dynamically control the loss function. During gradient descent or optimization, the flooding mechanism uses a hyperparameter, typically labeled b or flooding value, which represents the maximum allowable gradient value. Exceeding this value triggers "flooding," forcing the gradient value to a reasonable range to control the "amplitude" of updates or adjustments. The following formula is used to add the b value to the training loss optimization process. Enumeration is used to continuously adjust the b value, so that the training loss no longer approaches 0 but approaches b. This adjusts the validation loss to prevent it from gradually increasing.
[0059]
[0060] In the above formula Represents the training loss function after adjustment, L(θ) represents the training loss function before adjustment, and b represents the maximum allowed "gradient value", which is usually set at 1 / 2 of the rising point of the verification loss or a smaller value.
[0061] By constructing a target loss function, this embodiment effectively avoids the risk of overfitting, where the training loss decreases while the validation loss increases, making the model's recognition performance more reliable. In practical applications, this embodiment tested the adapter fine-tuning model, selecting a b value in the range of 0.06 to 0.09, and testing the target model results with and without adapter fine-tuning at a step size of 0.005 (this is because without a b value, the validation loss rises in the range of 0.16 to 0.18). The final test showed that the recognition performance was best when b was 0.065, as shown in Table 1 below:
[0062]
[0063] Table 1
[0064] In Table 1, P, R, F1, and TrainingParameters represent precision, recall, F1-score, and training parameters, respectively. Precision (P) focuses on the percentage of samples correctly predicted as positive by the model. Recall (R) focuses on whether the model can recognize all positive samples. F1-score is the harmonic mean of precision and recall, a comprehensive measure of both. A higher F1 value indicates better model recognition. TrainingParameters indicates the number of training parameters used by BERT during training, expressed in millions (M).
[0065] In one embodiment, when a visualization interface is set for the said teaching knowledge graph and the said teaching knowledge graph is displayed through an Echarts chart in response to a user's search call to the visualization interface, the following is combined with Figure 4 First, define an interface for visualizing the knowledge graph and provide two search methods, namely ordinary search and advanced search. Among them, ordinary search can only input the head entity, and the complete graph related to the head entity is displayed. Advanced search can input the head entity, relationship and tail entity, or only the head entity and relationship. In this case, only the head entity corresponds to the complete graph, and the single data graph related to the relationship is displayed. Then the knowledge graph is displayed through Echarts. When the input data is obtained, the data will be searched in the database according to the corresponding pattern according to the different search methods, and the corresponding triples and relationship categories will be returned. Then, the knowledge graph composed of triples can be displayed through Echarts, and all the information of the retrieved head entity will be displayed on the page for users to understand. For example, relevant graph information and text information can be obtained through head entity retrieval and visualized.
[0066] In one embodiment, the word semantic relationship analysis of the entity relationships in the Sijiao knowledge graph using the TransH model includes:
[0067] Based on the Sijiao knowledge graph, constructing head and tail entity embedding vectors and relationship embedding vectors, and mapping the head and tail entity embedding vectors to the relationship hyperplane;
[0068] The vector distance between the head and tail entity embedding vectors and the relation embedding vector is set as the positive sample score, and after randomly replacing the head entity or the tail entity, the vector distance between the head and tail entity embedding vectors and the relation embedding vector is set as the negative sample score;
[0069] Calculating the cross entropy loss between the positive sample scores and the negative sample scores, and using the cross entropy loss to train a TransH model, and then using the trained TransH model to perform word semantic relationship analysis on the entity relationships in the Sijiao knowledge graph to obtain a first similarity matrix containing multiple head entities;
[0070] The structural relationship analysis of the nodes and edges in the Sijiao knowledge graph using the GNN model includes:
[0071] Build a graph index and convert the index correspondence between the head and tail entities into a graph index representation;
[0072] Randomly initialize node features and combine graph index representation to train the GNN model’s structural representation;
[0073] The GNN model is updated using the node reconstruction loss function, and the updated GNN model is used to perform structural relationship analysis on the nodes and edges in the Sijiao knowledge graph to obtain a second similarity matrix containing multiple head entities.
[0074] Furthermore, the method of systematically recommending entities by combining the results of word semantic relationship analysis and structural relationship analysis includes:
[0075] Performing cosine similarity calculation on multiple head entities in the first similarity matrix to select top k1 head entities with the highest similarity; and performing cosine similarity calculation on multiple head entities in the second similarity matrix to select top k2 head entities with the highest similarity;
[0076] Perform weighted processing on the first k1 head entities and the first k2 head entities according to the preset weights, and select the first K head entities with the highest similarity based on the weighted processing results;
[0077] The top K entities are output as the results of systematic recommendation.
[0078] This embodiment combines the Trans H model and the GNN model to analyze the deep semantic relationship and shallow structural relationship of words in the Sijiao knowledge graph, providing technical support for the systematic recommendation of the main entity in the graph. Figure 5 First, load the triple data in the Sijiao knowledge graph and assign unique IDs to all entities, relations, head entities, tail entities, and triples. Then, for the Trans H model, after constructing the embedding vectors of the head and tail entities and relations, map the head and tail entity vectors to the relational hyperplane, and use the vector distance between the head and tail entities and the relations as the positive sample score. At the same time, the distance obtained after randomly replacing the head entity or the tail entity is used as the negative sample score. Then, the Trans H model is trained by calculating the cross entropy loss between them, so that the embedding vector of the elements in each triple expresses the word relationship between them as much as possible. It should be noted here that this embodiment performs semantic recommendation on the main entity, that is, the head entity. Therefore, when outputting the embedding vector matrix, only the embedding of the head entity is output. For the GNN model, it is first necessary to construct a graph index, convert the index correspondence between the head and tail entities into a graph index representation, and randomly initialize the node features to gradually learn the structural representation between meaningful nodes through the graph structure and the model itself. After that, by constructing the GNN model, the node reconstruction loss function (torch.nn.MSELoss()) is used as the loss function of the model to output the node features corresponding to the head entity.
[0079] The recommendation of a certain entity is based on the similarity sorting to obtain the top K entities with the highest similarity to the entity. Therefore, this embodiment uses cosine similarity to calculate the similarity between each head entity, that is, the cosine similarity is calculated for the similarity matrix output by Trans H and the similarity matrix output by GNN respectively. The specific output process is as follows: After the model training is completed, the node features (embedding vectors) corresponding to each head entity will be obtained and all stored in head_embeddings, which is usually a tensor with a shape of (N, D), where N is the number of head entity samples and D is the dimension of the embedding vector of each entity. Cosine similarity is calculated through F.cosine_similarity and passed into the two variables head_embeddings.unsqueeze(1) and head_embeddings.unsqueeze(0). head_embeddings.unsqueeze(1) will add a new dimension to the second dimension (the dimension with index 1), so that the shape changes from (N, D) to (N, 1, D). head_embeddings.unsqueeze(0) adds a new dimension to the first dimension (the dimension with index 0), changing the shape from (N, D) to (1, N, D). In this way, this embodiment transforms head_embeddings into two shapes suitable for calculating cosine similarity: (N, 1, D) and (1, N, D). F.cosine_similarity is a function used in PyTorch to calculate cosine similarity. Cosine similarity calculates the similarity between two vectors. The embedding vectors of each pair of entities in the two tensors head_embeddings.unsqueeze(1) and head_embeddings.unsqueeze(0) (such as entity 0 and 0, entity 0 and 1...entity 0 and N) will be used to calculate similarity. Therefore, a similarity matrix of size (N, N) can be obtained, which represents the similarity between each pair of head entity embedding vectors. Furthermore, different weights can be set to indicate the degree of skewness of the results, for example, a weight of 0.3 for the Trans H output and a weight of 0.7 for the GNN output. The weighted sums are then combined to obtain the final similarity value for all head entities. Finally, the top K recommended entities for each entity are obtained based on the ranking, thus generating an entity recommendation table for each head entity.
[0080] In this method, the TransH model can analyze the semantic relationship of words in the entity relationship in the knowledge graph, and the GNN model can analyze the structural relationship between nodes and edges in the knowledge graph, thereby systematically recommending entities in two dimensions. It better combines the knowledge graph and semantic understanding. Compared with the traditional recommendation algorithm, this embodiment can make the recommendation function more multidimensional and more in line with the characteristics of the knowledge graph. In the Sijiao knowledge graph retrieval, after searching for an entity and obtaining related entities, the K (for example, 10) recommended entities corresponding to the entity will be output synchronously. At the same time, these recommended entities can jump to the corresponding entity page, making it easier for users to understand entity information similar to the entity.
[0081] In one embodiment, in response to a user's query operation on the Sijiao knowledge graph, the vector database is searched and matched based on the RAG technology, and the question and answer result corresponding to the query operation is returned in combination with the AI big model based on the search and matching results, including:
[0082] Based on the query operation, obtaining a corresponding query embedding vector;
[0083] Calculating similarity between the query embedding vector and each vector data in the vector database;
[0084] Based on the results of the similarity calculation, the top N vector data with the highest similarity are selected and returned as the question and answer results.
[0085] This embodiment uses RAG to implement Sijiao knowledge question answering, which requires a vector database and highly structured Sijiao data to enhance retrieval while reducing the impact of AI-generated hallucinations on the results. Figure 6 Specifically, the constructed Sijiao knowledge graph triples are used to generate questions in various formats based on templates. For example, the entity A-relationship 1-entity B triple can generate questions such as "What has relationship 1 with entity A?" and "What is the relationship between entity A and entity B?" Each generated question is used as a key, and the corresponding triple is stored as a value in a vector database. The lightweight vector database Chroma can meet most requirements and reduce operational complexity. Here, the template can be understood as defining a method for generating a question by concatenating strings. Since this method generates questions in a fixed manner, it is called a template, or rule matching.
[0086] The large language AI model is then used as the output for answer generation. When the user enters a question, the large model's embedding function calculates the embedding vector of the input question and searches for matches in the vector database. A comprehensive prompt word engineering can also be designed, allowing the large AI model to combine the question and answer to output a complete contextual answer. For example, when a user enters the question "What organization has student A joined?", the embedding vector of the input question can be calculated using the embedding function of the large model API or other methods that can calculate text embedding vectors.
[0087] Then, you can define a Chroma client to connect to the vector database. Once connected, it will sequentially calculate the similarity between the input question vector and the vector of each data item in the database. The data item or items with the highest similarity can be selected. For this question, the retrieved answer may be a triple (Student A, Organization, Association B). Since the AI needs to output the prompt word "prompt" and the propmt needs to be passed the user-entered question and answer, it can be roughly written as prompt = "This is the user-entered question {User-entered question} (passed in the parameter 'What organization did Student A join?'), this is the retrieved answer {Retrieved answer} (passed in the parameter '(Student A, Organization, Association B)'). Please generate the context for the answer based on the above information. You can also decide whether to provide additional information." Finally, the context of the AI's answer is output, which may result in the answer "Student A once joined Association B" or the AI may provide more information.
[0088] Based on the above method, knowledge question and answer based on knowledge graph can be better realized, and the answer retrieval and question and answer capabilities can be enhanced through RAG. Compared with ordinary AI and retrieval output answers, the question and answer method enhanced by RAG through knowledge graph can effectively adapt to the specific field knowledge requirements of thinking and education, and reduce the impact of professional problems in answers caused by AI hallucinations.
[0089] In actual application scenarios, this embodiment can set up a corresponding teaching digital system based on the teaching knowledge recommendation and question-answering method, such as Figure 7As shown, the system specifically includes four modules: constructing a knowledge graph for teaching, visualizing and displaying the knowledge graph, recommending teaching entities based on Trans H and GNN, and answering questions about teaching knowledge based on the knowledge graph. First, the system uses adapters to fine-tune the BERT-BiLSTM-CRF model to achieve high-accuracy entity extraction. This, combined with label mapping to relations and triple extraction from text, implements a framework for constructing a knowledge graph for teaching. Secondly, standard or advanced searches of head entities allow for searching the database for relevant data and displaying its knowledge graph and related information. Finally, for entity recommendation, the Trans H model and GNN model systematically recommend graph entities based on both semantic and graph structural relationships. Cosine similarity ranking is used to obtain the top K recommended entities, and an entity recommendation table is generated for each head entity. When searching for a head entity, the K recommended entities corresponding to that entity are displayed according to the head entity recommendation table, and the user can jump to the corresponding entity information interface. In the knowledge question-answering function, the constructed Sijiao knowledge graph is combined with RAG to implement the Sijiao knowledge question-answering strategy. For each triple, a multi-modal question is generated through a template and stored in the vector database as a key-value pair with the corresponding triple. A vector is created with the question key as the index, and the vector is matched with the input question. Finally, the context answer is output by constructing a prompt word project.
[0090] Figure 8 A schematic block diagram of a knowledge graph-based educational knowledge recommendation and question-answering device 800 provided in an embodiment of the present invention, the device 800 includes:
[0091] A graph construction unit 801 is used to obtain ideological and political education resource data, extract entities from the ideological and political education resource data using a pre-built target model structure to obtain corresponding triple data, and construct an ideological and political education knowledge graph based on the triple data; wherein the target model structure is obtained by building a BERT model, a bidirectional long short-term memory network, and a conditional random field, and an adapter is used to fine-tune parameters during the construction process;
[0092] The graph visualization unit 802 is used to set a visualization interface for the Sijiao knowledge graph and display the Sijiao knowledge graph through an Echarts chart in response to a user's search call to the visualization interface; wherein the visualization interface provides multiple search methods;
[0093] The entity recommendation unit 803 is configured to perform word semantic relationship analysis on the entity relationships in the Sijiao knowledge graph using the TransH model, perform structural relationship analysis on the nodes and edges in the Sijiao knowledge graph using the GNN model, and systematically recommend entities based on the results of the word semantic relationship analysis and the structural relationship analysis;
[0094] A database building unit 804 is used to generate a question-answering question based on the triple data, and store the question-answering question as a key and the corresponding triple data as a value in a preset vector database;
[0095] The question and answer matching unit 805 is used to respond to the user's query operation on the Sijiao knowledge graph, search and match the vector database based on RAG technology, and return the question and answer result corresponding to the query operation based on the search and matching results in combination with the AI big model.
[0096] In one embodiment, the map construction unit 801 includes:
[0097] A data input unit, configured to input the thought and teaching resource data into the target model structure;
[0098] A semantic encoding unit, configured to perform contextual semantic encoding on the teaching resource data through the BERT model in the target model structure to obtain corresponding word vectors;
[0099] A dependency capture unit is used to capture the context dependency of the word vector using a bidirectional long short-term memory network to obtain corresponding sequence features;
[0100] A sequence labeling unit, configured to perform sequence labeling on the sequence features using a conditional random field to obtain a primary entity label and a secondary entity label;
[0101] The label mapping unit is used to map the secondary entity labels to relationships, thereby constructing triple data including main entity-relationship-secondary entity, and using the triple data to construct the Sijiao knowledge graph.
[0102] In one embodiment, the semantic encoding unit includes:
[0103] An adaptation and fine-tuning unit, configured to embed the adapter into the BERT model, and use the adapter to sequentially perform downsampling processing, nonlinear activation function activation processing, and upsampling processing on the input data of the BERT model to obtain corresponding output results;
[0104] A context capture unit is used to input the output result of the adapter into the Transformer layer through a residual block, and use the Transformer layer to encode and capture context information.
[0105] In one embodiment, the knowledge graph-based educational knowledge recommendation and question-answering device 800 further includes:
[0106] The loss construction unit is used to set the initial loss function according to the following formula and dynamically control the initial loss function using a flooding mechanism to construct a target loss function:
[0107]
[0108] in, represents the target loss function, L(θ) represents the initial loss function, and b represents the hyperparameter introduced by the flooding mechanism;
[0109] A training update unit is used to train and update the target model structure using the target loss function.
[0110] In one embodiment, the entity recommendation unit 803 includes:
[0111] A vector construction unit, configured to construct a head-tail entity embedding vector and a relationship embedding vector based on the Sijiao knowledge graph, and map the head-tail entity embedding vector to a relationship hyperplane;
[0112] A sample setting unit, configured to set the vector distance between the head and tail entity embedding vectors and the relationship embedding vector as a positive sample score, and to set the vector distance between the head and tail entity embedding vectors and the relationship embedding vector as a negative sample score after randomly replacing the head entity or the tail entity;
[0113] a loss calculation unit, configured to calculate a cross entropy loss between the positive sample scores and the negative sample scores, and train a TransH model using the cross entropy loss, and then use the trained TransH model to perform word semantic relationship analysis on entity relationships in the Sijiao knowledge graph to obtain a first similarity matrix containing multiple head entities;
[0114] The entity recommendation unit 803 further includes:
[0115] A graph index setting unit is used to construct a graph index and convert the index correspondence between the head and tail entities into a graph index representation;
[0116] The structural training unit is used to randomly initialize node features and perform structural representation training on the GNN model in combination with graph index representation;
[0117] The structural analysis unit is used to update the GNN model using a node reconstruction loss function, and use the updated GNN model to perform structural relationship analysis on the nodes and edges in the Sijiao knowledge graph to obtain a second similarity matrix containing multiple head entities.
[0118] In one embodiment, the entity recommendation unit 803 includes:
[0119] a similarity calculation unit, configured to perform cosine similarity calculation on the multiple head entities in the first similarity matrix to select the top k1 head entities with the highest similarity; and perform cosine similarity calculation on the multiple head entities in the second similarity matrix to select the top k2 head entities with the highest similarity;
[0120] A weighted selection unit is used to perform weighted processing on the first k1 head entities and the first k2 head entities according to preset weights, and select the first K head entities with the highest similarity based on the weighted processing results;
[0121] The result output unit is used to output the top K head entities as the results of systematic recommendation.
[0122] The question-answer matching unit 805 includes:
[0123] A vector acquisition unit, configured to acquire a corresponding query embedding vector based on the query operation;
[0124] a vector matching unit, configured to calculate similarity between the query embedding vector and each vector data in the vector database;
[0125] The result return unit is used to select the top N vector data with the highest similarity as the question and answer results based on the similarity calculation results.
[0126] Since the embodiments of the apparatus part correspond to the embodiments of the method part, please refer to the description of the embodiments of the method part for the embodiments of the apparatus part, and they will not be repeated here.
[0127] The present invention also provides a computer-readable storage medium having a computer program stored thereon. When executed, the computer program can implement the steps provided in the above embodiments. The storage medium can include a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk, among other media capable of storing program code.
[0128] The present invention also provides a computer device that may include a memory and a processor. The memory stores a computer program, and when the processor calls the computer program in the memory, the steps provided in the above embodiment can be implemented. Of course, the computer device may also include various network interfaces, a power supply, and other components.
[0129] The various embodiments in the specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same and similar parts between the various embodiments can be referred to each other. For the system disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the method part description. It should be pointed out that for ordinary technicians in this technical field, without departing from the principles of this application, several improvements and modifications can be made to this application, and these improvements and modifications also fall within the scope of protection of the claims of this application.
[0130] It should also be noted that, in this specification, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of additional identical elements in the process, method, article, or apparatus comprising the element.
Claims
1. A method for recommending and answering educational knowledge based on knowledge graph, characterized in that: include: Obtaining Sijiao resource data, extracting entities from the Sijiao resource data using a pre-built target model structure to obtain corresponding triple data, and constructing a Sijiao knowledge graph based on the triple data; wherein the target model structure is constructed using the BERT model, a bidirectional long short-term memory network, and a conditional random field, and an adapter is used to fine-tune parameters during the construction process; A visualization interface is provided for the Sijiao knowledge graph, and in response to a user's search call to the visualization interface, the Sijiao knowledge graph is displayed through an Echarts chart; wherein the visualization interface provides multiple search methods; Perform word semantic relationship analysis on the entity relationships in the Sijiao knowledge graph using the TransH model, and perform structural relationship analysis on the nodes and edges in the Sijiao knowledge graph using the GNN model, and systematically recommend entities based on the results of the word semantic relationship analysis and the structural relationship analysis; Generate a question-and-answer question based on the triple data, and store the question-and-answer question as a key and the corresponding triple data as a value in a preset vector database; In response to the user's query operation on the Sijiao knowledge graph, the vector database is searched and matched based on the RAG technology, and the question and answer results corresponding to the query operation are returned in combination with the AI big model based on the search and matching results.
2. The method for recommending and answering knowledge based on knowledge graph according to claim 1 is characterized in that: The obtaining of ideological and educational resource data, extracting entities from the ideological and educational resource data using a pre-built target model structure to obtain corresponding triple data, and constructing an ideological and educational knowledge graph based on the triple data, includes: Inputting the thought and teaching resource data into the target model structure; Perform contextual semantic encoding on the Sijiao resource data through the BERT model in the target model structure to obtain corresponding word vectors; Using a bidirectional long short-term memory network to capture context dependencies on the word vectors, and obtain corresponding sequence features; Using a conditional random field to perform sequence labeling on the sequence features to obtain a primary entity label and a secondary entity label; The secondary entity labels are mapped to relationships to construct triple data including main entity-relationship-secondary entity, and the triple data is used to construct the Sijiao knowledge graph.
3. The method for recommending and answering knowledge based on knowledge graph according to claim 2 is characterized in that: The contextual semantic encoding of the teaching resource data by the BERT model in the target model structure to obtain the corresponding word vector includes: Embed the adapter into the BERT model, and use the adapter to sequentially perform downsampling processing, nonlinear activation function activation processing, and upsampling processing on the input data of the BERT model to obtain corresponding output results; The output of the adapter is input into the Transformer layer through the residual block, and the Transformer layer is used to encode and capture context information.
4. The method for recommending and answering knowledge based on knowledge graph according to claim 1, characterized in that: Also includes: According to the following formula, the initial loss function is set, and the flooding mechanism is used to dynamically control the initial loss function to construct the target loss function: in, represents the target loss function, L(θ) represents the initial loss function, and b represents the hyperparameter introduced by the flooding mechanism; The target model structure is trained and updated using the target loss function.
5. The method for recommending and answering knowledge based on knowledge graph according to claim 1 is characterized in that: The word semantic relationship analysis of the entity relationships in the Sijiao knowledge graph using the TransH model includes: Based on the Sijiao knowledge graph, constructing head and tail entity embedding vectors and relationship embedding vectors, and mapping the head and tail entity embedding vectors to the relationship hyperplane; The vector distance between the head and tail entity embedding vectors and the relation embedding vector is set as the positive sample score, and after randomly replacing the head entity or the tail entity, the vector distance between the head and tail entity embedding vectors and the relation embedding vector is set as the negative sample score; Calculating the cross entropy loss between the positive sample scores and the negative sample scores, and using the cross entropy loss to train a TransH model, and then using the trained TransH model to perform word semantic relationship analysis on the entity relationships in the Sijiao knowledge graph to obtain a first similarity matrix containing multiple head entities; The structural relationship analysis of the nodes and edges in the Sijiao knowledge graph using the GNN model includes: Build a graph index and convert the index correspondence between the head and tail entities into a graph index representation; Randomly initialize node features and combine graph index representation to train the GNN model’s structural representation; The GNN model is updated using the node reconstruction loss function, and the updated GNN model is used to perform structural relationship analysis on the nodes and edges in the Sijiao knowledge graph to obtain a second similarity matrix containing multiple head entities.
6. The method for recommending and answering knowledge based on knowledge graph according to claim 5 is characterized in that: The method of systematically recommending entities by combining the results of word semantic relationship analysis and structural relationship analysis includes: Performing cosine similarity calculation on multiple head entities in the first similarity matrix to select top k1 head entities with the highest similarity; and performing cosine similarity calculation on multiple head entities in the second similarity matrix to select top k2 head entities with the highest similarity; Perform weighted processing on the first k1 head entities and the first k2 head entities according to the preset weights, and select the first K head entities with the highest similarity based on the weighted processing results; The top K entities are output as the results of systematic recommendation.
7. The method for recommending and answering knowledge based on knowledge graph according to claim 1, characterized in that: In response to the user's query operation on the Sijiao knowledge graph, the vector database is searched and matched based on the RAG technology, and the question and answer result corresponding to the query operation is returned in combination with the AI big model based on the search and matching results, including: Based on the query operation, obtaining a corresponding query embedding vector; Calculating similarity between the query embedding vector and each vector data in the vector database; Based on the results of the similarity calculation, the top N vector data with the highest similarity are selected and returned as the question and answer results.
8. A knowledge graph-based device for recommending and answering educational knowledge, characterized in that: include: A graph construction unit is used to obtain Sijiao resource data, extract entities from the Sijiao resource data using a pre-built target model structure to obtain corresponding triple data, and construct a Sijiao knowledge graph based on the triple data; wherein the target model structure is obtained by building a BERT model, a bidirectional long short-term memory network, and a conditional random field, and an adapter is used to fine-tune parameters during the construction process; A graph visualization unit is used to set a visualization interface for the Sijiao knowledge graph and display the Sijiao knowledge graph through an Echarts chart in response to a user's search call to the visualization interface; wherein the visualization interface provides multiple search methods; An entity recommendation unit, configured to perform word semantic relationship analysis on entity relationships in the Sijiao knowledge graph using a TransH model, and perform structural relationship analysis on nodes and edges in the Sijiao knowledge graph using a GNN model, and systematically recommend entities based on the results of the word semantic relationship analysis and the structural relationship analysis; A database building unit, configured to generate a question-and-answer question based on the triple data, and store the question-and-answer question as a key and the corresponding triple data as a value in a preset vector database; The question and answer matching unit is used to respond to the user's query operation on the Sijiao knowledge graph, search and match the vector database based on RAG technology, and return the question and answer result corresponding to the query operation based on the search and matching results in combination with the AI big model.
9. A computer device, characterized in that: It includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the knowledge graph-based teaching knowledge recommendation and question-answering method as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, which, when executed by a processor, implements the knowledge graph-based thinking and teaching knowledge recommendation and question-answering method as described in any one of claims 1 to 7.