Education knowledge graph construction method and system

Through the fusion of multimodal data and the application of deep learning and graph neural network, the problem of lack of comprehensive analysis in the display and construction of existing educational knowledge graphs is solved, and a higher quality and performance knowledge graph construction is achieved.

CN119990276AInactive Publication Date: 2025-05-13ZHILIN TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510076788.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-17
Publication Date
2025-05-13
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing educational knowledge graph lacks comprehensive analysis of multimodal data during the display and construction process, and the single quality evaluation analysis of the construction method is insufficient, resulting in a decrease in the quality of the knowledge graph.

Method used

By collecting multimodal data in the field of education, including text, pictures, and audio, and using natural language processing, convolutional neural networks and Mel frequency cepspectral coefficients, the features are extracted, and the tandem method is used to fusion features, and a multimodal fusion model is established based on deep learning and graph neural networks, knowledge representation learning and relationship strengthening are carried out, and the quality of the knowledge graph is ensured through multi-layer evaluation indicators.

Benefits of technology

The comprehensive integration of multimodal data and the structural strengthening of the knowledge graph are achieved, the quality and performance of the knowledge graph are improved, and the generalization ability of the model and the reliability of practical applications are ensured.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119990276A_ABST
    Figure CN119990276A_ABST
Patent Text Reader

Abstract

The invention discloses a method and a system for constructing an educational knowledge graph, and relates to the technical field of knowledge graph construction, a series connection method is used for fusing features extracted from different modals together, establishing multi-modal features, performing knowledge representation learning on the fused multi-modal features based on a deep learning model, and on the basis of the fused knowledge representation, constructing a multi-modal knowledge graph. The method comprises the following steps: establishing a relationship between different modals and entities by using a graph neural network, strengthening the structure of a knowledge graph, completing the construction of a multi-modal fusion model, collecting labeled data through a large database, dividing the labeled data into a training set and a test set, training the constructed multi-modal fusion model by using the training set, and constructing a multi-modal fusion model. According to the construction method, the multi-modal data are effectively fused to construct the multi-modal fusion model, so that the constructed multi-modal fusion model is more comprehensive.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of knowledge graph construction, and specifically to a method and system for constructing an educational knowledge graph. Background Art

[0002] Educational knowledge graph is a way of organizing and presenting educational information based on graph theory and artificial intelligence technology. It structures and semantically represents the knowledge elements, concepts, relationships, and connections between disciplines in the field of education, thereby building a comprehensive and visual educational knowledge network.

[0003] The prior art has the following defects:

[0004] 1. Existing educational knowledge graphs usually display text, images, and audio separately. However, in practical applications, since these elements may be related to each other, this separate display method may result in incomplete analysis of the knowledge graph.

[0005] 2. The existing construction method generally performs a single quality assessment and analysis on the constructed education knowledge graph. In actual applications, the constructed education knowledge graph is affected by many factors, which can easily lead to reduced quality. The use of a single quality assessment and analysis method leads to incomplete analysis, which reduces the performance and accuracy of the education knowledge graph. Summary of the invention

[0006] The purpose of the present invention is to provide a method and system for constructing an educational knowledge graph to address the deficiencies in the background technology.

[0007] In order to achieve the above object, the present invention provides the following technical solution: a method for constructing an educational knowledge graph, the construction method comprising the following steps:

[0008] Collect multimodal data in the field of education, including text, images, and audio, and check whether the data has annotation information;

[0009] Perform natural language processing on text data to extract text features, use convolutional neural networks to extract image features from image data, and use Mel-frequency cepstral coefficients to extract acoustic features from audio data;

[0010] The features extracted from different modalities are fused together using a concatenation method to establish multimodal features, and knowledge representation learning is performed on the fused multimodal features based on a deep learning model.

[0011] Based on the fused knowledge representation, the graph neural network is used to establish the relationship between different modalities and entities, strengthen the structure of the knowledge graph, and complete the construction of the multimodal fusion model.

[0012] Collect labeled data through a large database, divide the labeled data into training set and test set, use the training set to train the constructed multimodal fusion model, evaluate the constructed educational knowledge graph through the test set, and complete the construction after the multimodal fusion model is evaluated.

[0013] Preferably, the constructed educational knowledge graph is evaluated by a test set, including the following steps:

[0014] Input the multimodal data in the test set into the constructed multimodal fusion model, perform multimodal feature extraction on each sample in the test set, including text, image, and audio, use the features in the test set to input into the multimodal fusion model for prediction, obtain the output of the model on the test set, obtain the structure and relationship of the knowledge graph, use the selected evaluation indicators to evaluate the performance of the model on the test set, analyze the generalization ability of the model on the test set, and ensure the performance of the model on unseen data.

[0015] Preferably, the multimodal fusion model is constructed after being evaluated as qualified, including the following steps:

[0016] Obtain entity accuracy, graph completeness, embedding quality index, and classification error rate of the multimodal fusion model;

[0017] The model coefficient MXS is obtained by comprehensively calculating the entity accuracy, graph completeness, embedding quality index and classification error rate, and the expression is:

[0018]

[0019] ; Where stz is entity accuracy, twz is graph completeness, qrz is embedding quality index, fcw is classification error rate, fcwi is classification error rate of the i-th data point, qrzi is embedding quality index of the i-th data point, n is the number of data points, k1, k2, k3, k4 are proportional coefficients of entity accuracy, graph completeness, embedding quality index and classification error rate, respectively, and k1, k2, k3, k4 are all greater than 0;

[0020] After obtaining the model coefficient MXS of the multimodal fusion model, the model coefficient is compared with the quality threshold. If the model coefficient is greater than or equal to the quality threshold, the quality of the multimodal fusion model is qualified. If the model coefficient is less than the quality threshold, the quality of the multimodal fusion model is unqualified.

[0021] Preferably, the labeled data is collected through a large database, the labeled data is divided into a training set and a test set, and the constructed multimodal fusion model is trained using the training set, including the following steps:

[0022] Collect annotated multimodal data from a large database and preprocess the collected data, including cleaning and segmenting text data, resizing and normalizing image data, and unifying the sampling rate of audio data. The data set is divided into a training set and a test set. For each sample in the training set, multimodal features are extracted, including text features, image features, and audio features. The constructed multimodal fusion model is trained using the training set.

[0023] Preferably, based on the fused knowledge representation, a graph neural network is used to establish the relationship between different modalities and entities, strengthen the structure of the knowledge graph, and complete the construction of the multimodal fusion model, including the following steps:

[0024] Using the fused multimodal features, we construct a knowledge graph containing entities and relationships, and represent the knowledge graph as a graph structure, where entities are nodes of the graph and relationships are edges of the graph. We use graph neural networks to embed and learn the nodes in the knowledge graph to obtain a low-dimensional representation of each node. We use graph neural networks to model the edges in the graph and learn the representation of relationships. If the relationships in the graph have different importance for the associations between different entities, we introduce a graph attention mechanism to emphasize the influence of important relationships, perform graph convolution operations on the nodes and edges in the graph, and fuse the information of entities and relationships through information transfer. We use the learned node embeddings and relationship representations to complete the task.

[0025] Preferably, the features extracted from different modalities are fused together using a tandem method to establish a multimodal feature, comprising the following steps:

[0026] Features from different modalities are concatenated together to form a feature vector. Dimension alignment is performed before concatenation so that the features of each modality are represented by the same dimension. Through concatenation, a large feature vector containing features from each modality is obtained. Dimensionality reduction technology is used to map the feature vector to a low dimension.

[0027] Preferably, the acoustic features are extracted from the audio data using Mel Frequency Cepstral Coefficients (MFCC), comprising the following steps:

[0028] Load audio data in an audio file. If the audio signal is long, perform segmentation, normalize the audio signal, scale the amplitude range to the standard range, divide the audio signal into short-time windows, use a window of 20-40 milliseconds, apply a window function to each frame, perform discrete Fourier transform on the signal in each window to obtain a spectrum, filter the energy of each spectrum through a Mel filter to obtain the filtered energy, take the logarithm of the filtered energy to obtain a logarithmic energy spectrum, perform discrete cosine transform on the logarithmic energy spectrum to obtain MFCC coefficients, perform differential operations on the MFCC coefficients to obtain first-order differential coefficients, which are used to capture the dynamic characteristics of the audio signal, and concatenate the MFCC coefficients and their first-order differential coefficients to form a feature vector.

[0029] Preferably, collecting multimodal data in the field of education, including text, pictures, audio, etc., and checking whether the data has annotation information includes the following steps:

[0030] Crawling or acquiring data from identified data sources, cleaning the collected data, removing irrelevant or redundant information, performing text cleaning and word segmentation preprocessing operations on text data, denoising and dimensionality reduction on image data and audio data. If there is no labeled information in the data set, using crowdsourcing platforms or professional labeling teams to perform labeling work, labeling work is performed on text, images, audio and other data, and quality control is performed during the labeling process.

[0031] The present invention also provides a system for constructing an educational knowledge graph, including a collection module, a feature extraction module, a fusion module, a structure reinforcement module, a collection module, and an evaluation and testing module;

[0032] Collection module: collects multimodal data in the field of education, including text, pictures, and audio, and checks whether the data has annotation information;

[0033] Feature extraction module: performs natural language processing on text data to extract text features, uses convolutional neural networks to extract image features from image data, and uses Mel-frequency cepstral coefficients to extract acoustic features from audio data;

[0034] Fusion module: uses a concatenation method to fuse the features extracted from different modalities to establish multimodal features;

[0035] Structural reinforcement module: Based on the deep learning model, the fused multimodal features are subjected to knowledge representation learning. On the basis of the fused knowledge representation, the graph neural network is used to establish the relationship between different modalities and entities, strengthen the structure of the knowledge graph, and complete the construction of the multimodal fusion model.

[0036] Acquisition module: collects labeled data through a large database and divides the labeled data into training sets and test sets;

[0037] Evaluation and testing module: Use the training set to train the constructed multimodal fusion model, and use the test set to evaluate the constructed educational knowledge graph. The construction of the multimodal fusion model is completed after the evaluation is qualified.

[0038] In the above technical solution, the technical effects and advantages provided by the present invention are:

[0039] 1. The present invention fuses the features extracted from different modalities through a series method to establish multimodal features, performs knowledge representation learning on the fused multimodal features based on a deep learning model, and uses a graph neural network to establish the relationship between different modalities and entities on the basis of the fused knowledge representation, strengthens the structure of the knowledge graph, and completes the construction of the multimodal fusion model. The labeled data is collected through a large database, and the labeled data is divided into a training set and a test set. The constructed multimodal fusion model is trained using the training set, and the constructed educational knowledge graph is evaluated through the test set. This construction method effectively fuses multimodal data to construct a multimodal fusion model, making the constructed multimodal fusion model more comprehensive.

[0040] 2. The present invention completes the construction after the multimodal fusion model is evaluated as qualified, obtains the entity accuracy, graph completeness, embedding quality index and classification error rate of the multimodal fusion model, and comprehensively calculates the entity accuracy, graph completeness, embedding quality index and classification error rate to obtain the model coefficient MXS. After obtaining the model coefficient MXS of the multimodal fusion model, the model coefficient is compared with the quality threshold. If the model coefficient is greater than or equal to the quality threshold, the quality of the multimodal fusion model is analyzed to be qualified. If the model coefficient is less than the quality threshold, the quality of the multimodal fusion model is analyzed to be unqualified. By performing quality evaluation on the multimodal fusion model, it is possible to effectively determine whether the multimodal fusion model supports use, thereby improving the performance and accuracy of the model. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in the present invention. For ordinary technicians in this field, other drawings can also be obtained based on these drawings.

[0042] Figure 1 The present invention is a flow chart of the method. DETAILED DESCRIPTION

[0043] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0044] Example 1: Please refer to Figure 1 As shown, the method for constructing an educational knowledge graph described in this embodiment includes the following steps:

[0045] Collect multimodal data in the field of education, including text, pictures, audio, etc., and check whether the data has annotation information for subsequent model training and evaluation. Perform natural language processing on text data and extract text features, including keywords, word embeddings or other representations. Use convolutional neural networks (CNN) to extract image features for image data and use Mel-frequency cepstral coefficients (MFCC) to extract acoustic features for audio data. Use the concatenation method to fuse the features extracted from different modalities to establish multimodal features. Perform knowledge representation learning on the fused multimodal features based on the deep learning model. Based on the fused knowledge representation, use graph neural networks to establish the relationship between different modalities and entities, strengthen the structure of the knowledge graph, and complete the construction of the multimodal fusion model. Collect annotated data through a large database, divide the annotated data into training sets and test sets, use the training set to train the constructed multimodal fusion model, and evaluate the constructed education knowledge graph through the test set to verify its accuracy and generalization ability. The construction of the multimodal fusion model is completed after the evaluation is qualified.

[0046] This application fuses the features extracted from different modalities through a concatenation method to establish multimodal features, performs knowledge representation learning on the fused multimodal features based on a deep learning model, and uses a graph neural network to establish the relationship between different modalities and entities on the basis of the fused knowledge representation, strengthen the structure of the knowledge graph, and complete the construction of the multimodal fusion model. The labeled data is collected through a large database, and the labeled data is divided into a training set and a test set. The constructed multimodal fusion model is trained with the training set, and the constructed educational knowledge graph is evaluated with the test set. This construction method effectively fuses multimodal data to construct a multimodal fusion model, making the constructed multimodal fusion model more comprehensive.

[0047] Embodiment 2: Collecting multimodal data in the field of education, including text, pictures, audio, etc., and checking whether the data has annotation information, including the following steps:

[0048] Clearly define data requirements: determine the specific tasks and goals of building an educational knowledge graph, such as subject knowledge association, student learning trajectory analysis, etc.

[0049] Clarify the type of data that needs to be collected based on the task, including text, pictures, audio, etc.

[0050] Determine the source of data: Determine the source of data collection, which may include educational institution textbooks, online learning platforms, student assignments, teaching videos, etc.

[0051] Consider collaborating with educational institutions, teachers, students and other stakeholders to obtain data authorization and support.

[0052] Data crawling and acquisition: Based on the determined data source, appropriate methods are used to crawl or acquire data. This may involve web crawlers, API calls, etc.

[0053] Ensure compliance with relevant regulations and privacy policies during data crawling and obtain legal authorization for data.

[0054] Data cleaning and preprocessing: Clean the collected data to remove irrelevant or redundant information.

[0055] For text data, perform preprocessing operations such as text cleaning and word segmentation.

[0056] Perform denoising and dimensionality reduction on image data and audio data.

[0057] Acquisition of annotation information: Ensure that the dataset contains annotation information, which may be manually annotated text labels, image labels, audio labels, etc.

[0058] If there is no annotation information in the dataset, you can consider using a crowdsourcing platform or a professional annotation team to do the annotation work.

[0059] Establish annotation standards: If multiple people are involved in the annotation work, establish clear annotation standards to ensure consistency and accuracy of annotation.

[0060] For image and audio data, annotation specifications can be established to define the meaning of each label.

[0061] Quality control of annotation work: Annotation work is performed on text, images, audio and other data, and quality control is performed during the annotation process.

[0062] Methods such as cross-validation and consistency check between annotators can be used to ensure the quality of annotation.

[0063] Perform natural language processing on text data to extract text features, which include keywords, word embeddings or other representations, including the following steps:

[0064] Text cleaning: Remove noises such as special characters, punctuation marks, numbers, etc. from the text.

[0065] Perform case normalization and convert the text to lowercase.

[0066] Tokenization: Split the text into individual words or tokens.

[0067] For Chinese, Chinese tokenization tools such as jieba can be used.

[0068] Stop word removal: Remove common stop words, which usually do not contribute much to text analysis.

[0069] Stop words include frequently occurring words such as "de", "shi", "zai", etc.

[0070] Stemming or Lemmatization: Reduce the complexity of the vocabulary by reducing words to their base forms.

[0071] Stemming truncates the word endings, while Lemmatization reduces words to their original roots.

[0072] Build a Bag of Words: Convert the text into a bag-of-words model, which represents the occurrence of words in the text.

[0073] Count the frequency of each word in the text.

[0074] TF-IDF (Term Frequency-Inverse Document Frequency): Calculate the TF-IDF weights of words to measure the importance of words in the text.

[0075] TF represents term frequency, and IDF represents inverse document frequency.

[0076] Word Embeddings: Use pre-trained word embedding models such as Word2Vec, GloVe, FastText, etc. to convert each word into a high-dimensional vector representation.

[0077] Word embedding vectors capture the semantic relationships between words.

[0078] Other text representation methods: In addition to using bag of words, TF-IDF, word embeddings, etc., other text representation methods such as Doc2Vec, BERT, etc. can also be used.

[0079] Feature selection: According to the task requirements and the training complexity of the model, perform feature selection to select the most representative text features.

[0080] Vectorization: Vectorize text features so that they can be input into machine learning models for training and analysis.

[0081] Construct a text feature matrix: Combine the processed text features into a text feature matrix, where each row corresponds to a text sample and each column corresponds to a text feature.

[0082] Using a convolutional neural network (CNN) to extract image features from image data includes the following steps:

[0083] Data Preprocessing: Resizing images: Resize the images to the input size required by the model.

[0084] Normalization: Scale pixel values ​​to a smaller range (usually between 0 and 1).

[0085] Load pre-trained CNN models: Use CNN models pre-trained on large-scale image datasets, such as VGG16, ResNet, Inception, etc.

[0086] The weights of the pre-trained model contain rich feature learning of the image.

[0087] Feature extraction layer: Choose appropriate feature extraction layers, usually the convolutional layers of the model. These layers are more inclined to capture local and global features of the image.

[0088] Build the model: Keep the first few convolutional layers of the CNN model and remove the fully connected layers (or only keep the global average pooling layer).

[0089] The resulting model will output the feature map of the image at the convolutional layer.

[0090] Image feature extraction: Input the image into the constructed model to obtain the feature map at the convolutional layer.

[0091] These feature maps capture the characteristics of the image at different levels of abstraction.

[0092] Global Average Pooling (GAP): Perform global average pooling on each feature map and reduce the size of each feature map to 1x1.

[0093] This will produce a feature vector with the same number as the feature map as the final image features.

[0094] Feature vector representation: The feature vector obtained by global average pooling is used as the final representation of the image.

[0095] This representation is usually a high-dimensional vector that contains the abstract features of the image learned in the convolutional layers.

[0096] Dimensionality reduction (optional): If the dimensionality of the feature vector is too high, you can consider using dimensionality reduction techniques such as principal component analysis (PCA) or t-SNE to project the feature vector to a lower dimension.

[0097] Feature vector application: The obtained feature vector is used for subsequent tasks such as image classification, detection, recognition, etc.

[0098] These features can be fed into a machine learning model for training or directly applied to other tasks.

[0099] The Mel Frequency Cepstral Coefficient (MFCC) is used to extract acoustic features from audio data, including the following steps:

[0100] Audio loading: Load audio data from an audio file, ensuring that parameters such as sampling rate and bit number are correct.

[0101] Preprocessing: If the audio signal is long, it may need to be segmented for subsequent processing.

[0102] Normalize the audio signal to scale the amplitude range to a standard range.

[0103] Frame: Divide the audio signal into short-time windows (frames), usually with a window length of 20-40 milliseconds.

[0104] There may be an overlap between adjacent frames, for example a 50% overlap.

[0105] Windowing: Apply a window function to each frame. Commonly used window functions include Hamming window and Hamming window.

[0106] The window function helps to reduce spectral leakage.

[0107] Fourier transform: Perform discrete Fourier transform (DFT) on the signal in each window to obtain the spectrum.

[0108] The Fast Fourier Transform (FFT) algorithm is usually used.

[0109] Mel filter bank: Design a set of Mel filters that are evenly distributed over the Mel frequencies.

[0110] The energy of each spectrum is filtered through a Mel filter to obtain the filtered energy.

[0111] Logarithmic operation: Take the logarithm of the filtered energy to obtain the logarithmic energy spectrum (Log Mel Spectrum).

[0112] Logarithmic operations help simulate the human ear's perception of sound intensity.

[0113] Discrete Cosine Transform (DCT): Perform discrete cosine transform on the logarithmic energy spectrum to obtain MFCC coefficients.

[0114] Usually the first few coefficients are retained as the final MFCC features.

[0115] Feature Difference (Delta Coefficients, Delta): Perform a differential operation on the MFCC coefficients to obtain the first-order differential coefficients, which are used to capture the dynamic characteristics of the audio signal.

[0116] Feature concatenation (Optional): The MFCC coefficients and their first-order difference coefficients can be concatenated together to form a richer feature vector.

[0117] Feature normalization: Normalize the extracted feature vector to make it numerically stable.

[0118] Feature application: The extracted MFCC features are used for tasks such as audio classification and speech recognition.

[0119] The features extracted from different modalities are fused together using a concatenation method to establish multimodal features, including the following steps:

[0120] Feature concatenation: Features from different modalities are directly concatenated together to form a longer feature vector.

[0121] If the feature dimensions between the modalities are different, it may be necessary to align the dimensions before concatenation to ensure that the features of each modality are represented in the same dimension.

[0122] Feature vector: Through concatenation, a large feature vector containing features from each mode is obtained.

[0123] This vector contains information from different modalities, forming a multimodal feature representation.

[0124] Dimensionality reduction (optional): If the dimension of the merged feature vector is high, you can consider using dimensionality reduction techniques, such as principal component analysis (PCA), to map the feature vector to a lower dimension.

[0125] Feature application: The obtained multimodal features are used for subsequent tasks such as classification, regression, clustering, etc.

[0126] This feature vector can be input into a machine learning model for training or directly applied to other tasks.

[0127] The knowledge representation learning of the fused multimodal features is performed based on the deep learning model, including the following steps:

[0128] Build a deep learning model: Design a deep learning model to perform knowledge representation learning on the fused multimodal features.

[0129] Common models include multi-layer perceptron (MLP), convolutional neural network (CNN), recurrent neural network (RNN), autoencoder, etc.

[0130] Define the network structure: Define the input layer of the network to ensure that it can accept the fused multimodal features.

[0131] Design the hidden layer and output layer, and choose the appropriate structure according to the requirements of the task and the characteristics of the data.

[0132] Choice of loss function: Choose an appropriate loss function to measure the difference between the model output and the true label.

[0133] For unsupervised tasks, such as autoencoders, the reconstruction error can be used as the loss function.

[0134] Training model: Use the fused multimodal features to train the deep learning model.

[0135] The back propagation algorithm is used to optimize the network parameters and reduce the value of the loss function.

[0136] Use appropriate optimization algorithms and learning rates as needed.

[0137] Regularization and avoiding overfitting: Add regularization terms such as Dropout, L1 regularization, L2 regularization, etc. to prevent the model from overfitting the training data.

[0138] The regularization parameter can be tuned using cross-validation.

[0139] Based on the fused knowledge representation, the graph neural network is used to establish the relationship between different modalities and entities, strengthen the structure of the knowledge graph, and complete the construction of the multimodal fusion model, including the following steps:

[0140] Build a knowledge graph: Use the fused multimodal features to build a knowledge graph containing entities and relationships.

[0141] Entities can represent elements of different modalities, and relationships represent the associations between entities.

[0142] Graph representation: The knowledge graph is represented as a graph structure, where entities are nodes and relationships are edges.

[0143] The feature vector of each node can be a fused multi-modal feature.

[0144] Graph neural network design: Select a suitable graph neural network model, such as Graph Convolutional Network (GCN), GraphSAGE, GAT (Graph Attention Network), etc.

[0145] Define the number of layers and structure of the graph neural network.

[0146] Node embedding learning: Use graph neural networks to embed nodes in the knowledge graph and obtain a low-dimensional representation of each node, namely node embedding.

[0147] Node embeddings capture the contextual information of nodes in the graph structure.

[0148] Relationship modeling: Use graph neural networks to model the edges in the graph and learn the representation of relationships.

[0149] Consider information such as the weight, direction, and type of the relationship.

[0150] Graph Attention Mechanism (optional): If the relationships in the graph have different importance for the associations between different entities, you can consider introducing a graph attention mechanism to emphasize the influence of important relationships.

[0151] Graph convolution operation: Perform graph convolution operations on the nodes and edges in the graph to fuse the information of entities and relationships through information transfer.

[0152] Multiple rounds of graph convolution can help the model better capture the complex structure in the graph.

[0153] Prediction and classification: Use the learned node embeddings and relationship representations to complete specific tasks, such as entity relationship prediction, node classification, etc.

[0154] The specific form of the task depends on the application scenario of the knowledge graph.

[0155] Collect labeled data from a large database, divide the labeled data into a training set and a test set, and use the training set to train the constructed multimodal fusion model, including the following steps:

[0156] Data collection: Collect annotated multimodal data from large databases, including text, images, audio, etc.

[0157] Ensure that the data annotation information is relevant and effective for the model training task.

[0158] Data preprocessing: preprocess the collected data, including text data cleaning and word segmentation, image data resizing and normalization, audio data sampling rate unification, etc.

[0159] Make sure the data from different modalities are aligned and in a format ready for input into the model.

[0160] Data partitioning: Divide the dataset into training and test sets, usually using a common partitioning ratio such as 80% training set and 20% test set.

[0161] You can use random partitioning or partitioning according to some strategy (such as by entity, by time).

[0162] Feature extraction: For each sample in the training set, multimodal features are extracted, including text features, image features, audio features, etc.

[0163] Use the related methods introduced before, such as natural language processing, convolutional neural networks, Mel-frequency cepstral coefficients, etc.

[0164] Model training: Use the training set to train the constructed multimodal fusion model.

[0165] Through the back propagation algorithm, the model parameters are optimized and the value of the loss function is reduced.

[0166] The constructed educational knowledge graph is evaluated through the test set to verify its accuracy and generalization ability. The construction of the multimodal fusion model is completed after the evaluation is qualified, including the following steps:

[0167] Test set preparation: Input the multimodal data in the test set into the constructed multimodal fusion model.

[0168] Make sure the examples in your test set cover different situations that your model might encounter in real life.

[0169] Feature extraction: Multimodal feature extraction is performed on each sample in the test set, including text, image, audio, etc.

[0170] Keep the same feature extraction method as used during training to ensure consistency.

[0171] Model prediction: Use the features in the test set to input into the multimodal fusion model for prediction.

[0172] Get the output of the model on the test set, that is, the structure and relations of the knowledge graph.

[0173] Evaluation metric selection: Choose an appropriate evaluation metric to measure the performance of the model on the test set.

[0174] For the knowledge graph construction task, indicators may include the accuracy of entity relationships, the completeness of the graph, the connectivity of the knowledge graph, etc.

[0175] Performance evaluation: Evaluate the performance of the model on the test set using the selected evaluation metric.

[0176] Analyze the prediction accuracy of the model for different types of entities and relationships.

[0177] If there are subtasks, such as classification or prediction, each subtask is evaluated independently.

[0178] Generalization ability analysis: Analyze the generalization ability of the model on the test set to ensure the performance of the model on unseen data.

[0179] Consider the difference in distribution between the test set and the training set, as well as edge cases in the test set.

[0180] Result visualization (optional): Visualize a portion of the knowledge graph to intuitively understand the graph structure built by the model.

[0181] Graphical tools or other visualization methods can be used to demonstrate the learning effect of the model on different modalities.

[0182] Improvement and tuning (optional): If the evaluation results are not ideal, you can consider improving and tuning the model.

[0183] You can adjust model parameters, optimize training strategies, or consider using more complex model structures.

[0184] Model deployment (optional): If the model performs well on the test set, you can consider deploying it to actual application scenarios.

[0185] Further testing and monitoring are carried out in actual scenarios to ensure the stability and performance of the model.

[0186] Documentation and Reporting: Write documentation and reports on the model evaluation, recording the evaluation process, results, and analysis.

[0187] Provide enough detail so that others can understand the capabilities and limitations of the model.

[0188] After the multimodal fusion model is evaluated and qualified, the construction is completed, including the following steps:

[0189] Obtain entity accuracy, graph completeness, embedding quality index, and classification error rate of the multimodal fusion model;

[0190] The model coefficient MXS is obtained by comprehensively calculating the entity accuracy, graph completeness, embedding quality index and classification error rate, and the expression is:

[0191]

[0192] ; Where stz is entity accuracy, twz is graph completeness, qrz is embedding quality index, fcw is classification error rate, fcwi is classification error rate of the i-th data point, qrzi is embedding quality index of the i-th data point, n is the number of data points, k1, k2, k3, k4 are proportional coefficients of entity accuracy, graph completeness, embedding quality index and classification error rate, respectively, and k1, k2, k3, k4 are all greater than 0;

[0193] After obtaining the model coefficient MXS of the multimodal fusion model, the model coefficient is compared with the quality threshold. If the model coefficient is greater than or equal to the quality threshold, the quality of the multimodal fusion model is qualified. If the model coefficient is less than the quality threshold, the quality of the multimodal fusion model is unqualified.

[0194] Entity Accuracy: Entity accuracy is used to measure the accuracy of the model for entity recognition. The model predicts the entities in the test set, compares them with the true labels, and calculates the entity accuracy.

[0195] Knowledge Graph Completeness: Graph completeness is used to evaluate whether the knowledge graph constructed by the model is complete and whether it captures the important entities and relationships in the data. Analyze the number of entities and relationships contained in the knowledge graph and compare them with the number in the actual data.

[0196] Embedding Quality Index: The Embedding Quality Index is used to evaluate the quality of node embeddings and relationship embeddings learned by the model. It uses similarity metrics of embeddings, such as cosine similarity, to evaluate the semantic similarity of nodes and relationships. It analyzes the distance and distribution between categories in the embedding space.

[0197] Classification Error Rate: The classification error rate is used to evaluate the performance of the model in node classification or relationship classification tasks. The model predicts the nodes or relationships in the test set, compares them with the true labels, and calculates the classification error rate.

[0198] Embodiment 3: The system for constructing an educational knowledge graph described in this embodiment includes a collection module, a feature extraction module, a fusion module, a structure enhancement module, a collection module, and an evaluation and testing module;

[0199] Collection module: collects multimodal data in the field of education, including text, pictures, audio, etc., and checks whether the data has annotation information for subsequent model training and evaluation;

[0200] Feature extraction module: Perform natural language processing on text data to extract text features, including keywords, word embedding or other representations. Use convolutional neural network (CNN) to extract image features from image data, and use Mel frequency cepstral coefficient (MFCC) to extract acoustic features from audio data.

[0201] Fusion module: uses a concatenation method to fuse the features extracted from different modalities to establish multimodal features;

[0202] Structural reinforcement module: Based on the deep learning model, the fused multimodal features are subjected to knowledge representation learning. On the basis of the fused knowledge representation, the graph neural network is used to establish the relationship between different modalities and entities, strengthen the structure of the knowledge graph, and complete the construction of the multimodal fusion model.

[0203] Acquisition module: collects labeled data through a large database and divides the labeled data into training sets and test sets;

[0204] Evaluation and testing module: Use the training set to train the constructed multimodal fusion model, and use the test set to evaluate the constructed educational knowledge graph to verify its accuracy and generalization ability. The construction of the multimodal fusion model is completed after the evaluation is qualified.

[0205] The above formulas are all dimensionless and numerical calculations. The formula is a formula for the most recent real situation obtained by collecting a large amount of data and performing software simulation. The preset parameters in the formula are set by technicians in this field according to actual conditions.

[0206] In the description of this specification, the description with reference to the terms "one embodiment", "example", "specific example", etc. means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representation of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.

[0207] The preferred embodiments of the present invention disclosed above are only used to help explain the present invention. The preferred embodiments do not describe all the details in detail, nor do they limit the invention to only specific implementation methods. Obviously, many modifications and changes can be made according to the content of this specification. This specification selects and specifically describes these embodiments in order to better explain the principles and practical applications of the present invention, so that those skilled in the art can understand and use the present invention well. The present invention is limited only by the claims and their full scope and equivalents.

Claims

1. A method for constructing an educational knowledge graph, characterized in that: The construction method comprises the following steps: Collect multimodal data in the field of education, including text, images, and audio, and check whether the data has annotation information; Perform natural language processing on text data to extract text features, use convolutional neural networks to extract image features from image data, and use Mel-frequency cepstral coefficients to extract acoustic features from audio data; The features extracted from different modalities are fused together using a concatenation method to establish multimodal features, and knowledge representation learning is performed on the fused multimodal features based on a deep learning model. Based on the fused knowledge representation, the graph neural network is used to establish the relationship between different modalities and entities, strengthen the structure of the knowledge graph, and complete the construction of the multimodal fusion model. Collect labeled data through a large database, divide the labeled data into training set and test set, use the training set to train the constructed multimodal fusion model, evaluate the constructed educational knowledge graph through the test set, and complete the construction after the multimodal fusion model is evaluated.

2. The method for constructing an educational knowledge graph according to claim 1, characterized in that: The constructed educational knowledge graph is evaluated through the test set, including the following steps: Input the multimodal data in the test set into the constructed multimodal fusion model, perform multimodal feature extraction on each sample in the test set, including text, image, and audio, use the features in the test set to input into the multimodal fusion model for prediction, obtain the output of the model on the test set, obtain the structure and relationship of the knowledge graph, use the selected evaluation indicators to evaluate the performance of the model on the test set, analyze the generalization ability of the model on the test set, and ensure the performance of the model on unseen data.

3. The method for constructing an educational knowledge graph according to claim 2, characterized in that: After the multimodal fusion model is evaluated and qualified, the construction is completed, including the following steps: Obtain entity accuracy, graph completeness, embedding quality index, and classification error rate of the multimodal fusion model; The model coefficient MXS is obtained by comprehensively calculating the entity accuracy, graph completeness, embedding quality index and classification error rate, and the expression is: Where stz is the entity accuracy, twz is the graph completeness, qrz is the embedding quality index, fcw is the classification error rate, fcwi is the classification error rate of the i-th data point, qrzi is the embedding quality index of the i-th data point, n is the number of data points, k1, k2, k3, k4 are the proportional coefficients of entity accuracy, graph completeness, embedding quality index and classification error rate, respectively, and k1, k2, k3, k4 are all greater than 0; After obtaining the model coefficient MXS of the multimodal fusion model, the model coefficient is compared with the quality threshold. If the model coefficient is greater than or equal to the quality threshold, the quality of the multimodal fusion model is qualified. If the model coefficient is less than the quality threshold, the quality of the multimodal fusion model is unqualified.

4. The method for constructing an educational knowledge graph according to claim 3, characterized in that: Collect labeled data from a large database, divide the labeled data into a training set and a test set, and use the training set to train the constructed multimodal fusion model, including the following steps: Collect annotated multimodal data from a large database and preprocess the collected data, including cleaning and segmenting text data, resizing and normalizing image data, and unifying the sampling rate of audio data. The data set is divided into a training set and a test set. For each sample in the training set, multimodal features are extracted, including text features, image features, and audio features. The constructed multimodal fusion model is trained using the training set.

5. The method for constructing an educational knowledge graph according to claim 4, characterized in that: Based on the fused knowledge representation, the graph neural network is used to establish the relationship between different modalities and entities, strengthen the structure of the knowledge graph, and complete the construction of the multimodal fusion model, including the following steps: Using the fused multimodal features, we construct a knowledge graph containing entities and relationships, and represent the knowledge graph as a graph structure, where entities are nodes of the graph and relationships are edges of the graph. We use graph neural networks to embed and learn the nodes in the knowledge graph to obtain a low-dimensional representation of each node. We use graph neural networks to model the edges in the graph and learn the representation of relationships. If the relationships in the graph have different importance for the associations between different entities, we introduce a graph attention mechanism to emphasize the influence of important relationships, perform graph convolution operations on the nodes and edges in the graph, and fuse the information of entities and relationships through information transfer. We use the learned node embeddings and relationship representations to complete the task.

6. The method for constructing an educational knowledge graph according to claim 5, characterized in that: The features extracted from different modalities are fused together using a concatenation method to establish multimodal features, including the following steps: Features from different modalities are concatenated together to form a feature vector. Dimension alignment is performed before concatenation so that the features of each modality are represented by the same dimension. Through concatenation, a large feature vector containing features from each modality is obtained. Dimensionality reduction technology is used to map the feature vector to a low dimension.

7. The method for constructing an educational knowledge graph according to claim 6, characterized in that: The Mel Frequency Cepstral Coefficient (MFCC) is used to extract acoustic features from audio data, including the following steps: Load audio data in an audio file. If the audio signal is long, perform segmentation, normalize the audio signal, scale the amplitude range to the standard range, divide the audio signal into short-time windows, use a window of 20-40 milliseconds, apply a window function to each frame, perform discrete Fourier transform on the signal in each window to obtain a spectrum, filter the energy of each spectrum through a Mel filter to obtain the filtered energy, take the logarithm of the filtered energy to obtain a logarithmic energy spectrum, perform discrete cosine transform on the logarithmic energy spectrum to obtain MFCC coefficients, perform differential operations on the MFCC coefficients to obtain first-order differential coefficients, which are used to capture the dynamic characteristics of the audio signal, and concatenate the MFCC coefficients and their first-order differential coefficients to form a feature vector.

8. The method for constructing an educational knowledge graph according to claim 7, characterized in that: Collect multimodal data in the field of education, including text, pictures, audio, etc., and check whether the data has annotation information, including the following steps: Crawling or acquiring data from identified data sources, cleaning the collected data, removing irrelevant or redundant information, performing text cleaning and word segmentation preprocessing operations on text data, denoising and dimensionality reduction on image data and audio data. If there is no labeled information in the data set, using crowdsourcing platforms or professional labeling teams to perform labeling work, labeling work is performed on text, images, audio and other data, and quality control is performed during the labeling process.

9. A system for constructing an educational knowledge graph, used to implement the construction method according to any one of claims 1 to 8, characterized in that: It includes collection module, feature extraction module, fusion module, structure enhancement module, acquisition module and evaluation and testing module; Collection module: collects multimodal data in the field of education, including text, pictures, and audio, and checks whether the data has annotation information; Feature extraction module: performs natural language processing on text data to extract text features, uses convolutional neural networks to extract image features from image data, and uses Mel-frequency cepstral coefficients to extract acoustic features from audio data; Fusion module: uses a concatenation method to fuse the features extracted from different modalities to establish multimodal features; Structural reinforcement module: Based on the deep learning model, the fused multimodal features are subjected to knowledge representation learning. On the basis of the fused knowledge representation, the graph neural network is used to establish the relationship between different modalities and entities, strengthen the structure of the knowledge graph, and complete the construction of the multimodal fusion model. Acquisition module: collects labeled data through a large database and divides the labeled data into training sets and test sets; Evaluation and testing module: Use the training set to train the constructed multimodal fusion model, and use the test set to evaluate the constructed educational knowledge graph. The construction of the multimodal fusion model is completed after the evaluation is qualified.