Intelligent education question and answer information processing system based on AI teaching assistance

By designing an intelligent educational question-and-answer information processing system based on AI teaching assistants, using multiple collaborative modules to deeply understand educational problems, extract relevant knowledge and generate high-quality answers, the existing system's shortcomings in accuracy, completeness, knowledge correlation, adaptability and real-timeness are solved, and the system's significant performance improvement and application value enhancement are achieved.

CN120067256AInactive Publication Date: 2025-05-30GUANGZHOU COLLEGE OF TECH BUSINESS CO LTD
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510131529.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-06
Publication Date
2025-05-30
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing intelligent education question and answer system has shortcomings in accuracy, completeness, knowledge correlation, adaptability and real-timeness, and it is difficult to provide accurate, comprehensive and personalized answers, and the system is difficult to adapt to the rapid changes in educational content.

Method used

An intelligent education question-and-answer information processing system based on AI teaching assistants was designed. Through the collaborative work of the educational question-and-answer collection module, question division module, text vectorization module, semantic space mapping module, knowledge graph matching module, answer generation module, answer optimization module and other modules, we can deeply understand educational problems, accurately extract relevant knowledge, generate high-quality answers, and continuously improve ourselves through the learning and update module.

Benefits of technology

The system has achieved breakthroughs in many aspects, including the accuracy and completeness of answers, the improvement of knowledge relevance, the enhancement of adaptability and the satisfaction of real-time requirements, which has significantly improved the performance and application value of the intelligent education question and answer system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120067256A_ABST
    Figure CN120067256A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of intelligent education and artificial intelligence, in particular to an intelligent education question and answer information processing system based on AI teaching assistance, which comprises an education question and answer acquisition module, a question division module, a text vectorization module, a semantic space mapping module, a semantic space mapping module and a semantic space mapping module, the semantic space mapping module is in communication connection with the text vectorization module, the knowledge graph matching module is in communication connection with the semantic space mapping module, and the answer generation module is in communication connection with the knowledge graph matching module and is used for receiving a matching result sent by the knowledge graph matching module; the answer optimization module is in communication connection with the answer generation module; the result output module is in communication connection with the answer optimization module and receives the optimized answers sent by the answer optimization module; and the optimized answers are sent to the user side, so that the real-time questions of the students can be answered, the students can be continuously perfected through continuous learning, and the method adapts to dynamic changes in the education field.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of intelligent education and artificial intelligence technology, and more specifically, to an intelligent education Q&A information processing system based on an AI teaching assistant. Background Art

[0002] With the rapid development of educational informatization, intelligent education Q&A systems, as an important auxiliary teaching tool, are receiving increasing attention. Such systems aim to simulate human teachers and provide students with real-time question answers and learning guidance. However, despite certain progress in this field, there are still many limitations and deficiencies in the existing technologies.

[0003] Currently, the mainstream intelligent education Q&A systems are mainly divided into two categories: retrieval-based systems and generative model-based systems. Retrieval-based systems usually rely on a pre-built Q&A pair database and find the most similar questions through keyword matching or shallow semantic analysis, and return the corresponding answers. The advantages of such systems are fast response speed and high answer controllability, but their limitations are also very obvious. They are difficult to handle complex and reasoning-required questions, are also sensitive to changes in question expressions, and often cannot understand the true intentions of students. In addition, the knowledge coverage of such systems is limited and it is difficult to cope with the ever-changing knowledge updates in the education field.

[0004] On the other hand, generative model-based systems, especially those using large-scale pre-trained language models, perform well in understanding complex questions and generating fluent answers. These systems can process various forms of natural language inputs and generate seemingly reasonable answers. However, such systems also face serious problems. First, they often lack in-depth understanding of knowledge in specific education fields and are prone to generating answers that seem correct on the surface but actually contain incorrect information. Second, the black-box characteristics of these models make it difficult to explain and control their outputs, which may lead to serious consequences in an educational scenario. Moreover, such systems usually require huge computing resources and have long response times, making it difficult to meet the needs of real-time interaction. Finally, they lack effective learning and updating mechanisms and are difficult to adapt to the rapidly changing educational content and personalized learning needs.

[0005] These deficiencies in the existing technologies have led to a series of problems: students cannot obtain accurate, comprehensive, and personalized answers; the system is difficult to provide a coherent knowledge context, which is not conducive to students establishing a systematic knowledge system; educators are difficult to trust and effectively utilize these systems as auxiliary teaching tools. These problems have severely restricted the practical application and popularization of intelligent education Q&A systems. Summary of the Invention

[0006] The present invention aims to solve the deficiencies of existing intelligent education Q&A systems in terms of accuracy, integrity, knowledge association, adaptability, and real-time performance. Through innovative system architecture and algorithm design, the present invention provides an intelligent education Q&A information processing system that can deeply understand educational questions, accurately extract relevant knowledge, generate high-quality answers, and continuously improve itself.

[0007] The present invention provides an intelligent education Q&A information processing system based on an AI teaching assistant, including:

[0008] An education Q&A collection module, used for:

[0009] Obtain the educational questions input by the user;

[0010] Preprocess the obtained educational questions to generate preprocessing information;

[0011] A question division module, communicatively connected to the education Q&A collection module, used for:

[0012] Receive the preprocessing information sent by the education Q&A collection module;

[0013] Based on the preprocessing information, perform question semantic analysis and classification;

[0014] A text vectorization module, communicatively connected to the question division module, used for:

[0015] Receive the classification result sent by the question division module;

[0016] Based on the classification result, convert the question into a special random matrix representation;

[0017] A semantic space mapping module, communicatively connected to the text vectorization module, used for:

[0018] Receive the random matrix representation sent by the text vectorization module;

[0019] Based on the random matrix representation, perform non-linear mapping to a high-dimensional semantic space;

[0020] A knowledge graph matching module, communicatively connected to the semantic space mapping module, used for:

[0021] Receive the high-dimensional semantic representation sent by the semantic space mapping module;

[0022] Based on the high-dimensional semantic representation, perform matching and embedding in the knowledge graph;

[0023] An answer generation module, communicatively connected to the knowledge graph matching module, used for:

[0024] Receive the matching result sent by the knowledge graph matching module;

[0025] Based on the matching result, construct an answer generation probability field;

[0026] An answer optimization module, communicatively connected to the answer generation module, for:

[0027] Receive the probability field information sent by the answer generation module;

[0028] Based on the probability field information, extract the optimal answer and refine it;

[0029] A result output module, communicatively connected to the answer optimization module, for:

[0030] Receive the optimized answer sent by the answer optimization module;

[0031] Send the optimized answer to the user side.

[0032] Preferably, the educational Q&A collection module includes:

[0033] A voice input unit for receiving the voice question input from the user;

[0034] A text input unit for receiving the text question input from the user;

[0035] A preprocessing unit, connected to the voice input unit and the text input unit, for performing preprocessing operations such as word segmentation and stop word removal on the input question.

[0036] Preferably, the question partitioning module includes:

[0037] A semantic analysis unit for performing semantic parsing on the preprocessing information;

[0038] A classification unit, connected to the semantic analysis unit, for partitioning the question into predefined categories based on the semantic parsing result;

[0039] A label generation unit, connected to the classification unit, for generating knowledge point labels and domain labels for the partitioned question.

[0040] Preferably, the text vectorization module uses the following formula to transform the question into a special random matrix representation:

[0041]

[0042] where ω i is the word weight, |w i > is the quantum state representation of the word vector, n is the number of words in the question, H is a random Hermitian matrix, and ∈ is a small perturbation parameter.

[0043] Preferably, the semantic space mapping module performs a non-linear mapping using the following formula:

[0044] S = Φ(Q) = tr(e iQT )P + λL(Q),

[0045] where Φ is a non - linear mapping function, T is a time - evolution operator, P is a projection matrix, L is a Laplace operator, and λ is a regularization parameter.

[0046] Preferably, the knowledge graph matching module performs matching and embedding using the following formula:

[0047]

[0048] where is a knowledge graph manifold, G(x) is the local structure of the knowledge graph at point x, f is a matching function, μ is a measure, and D is a diffusion tensor.

[0049] Preferably, the answer generation module constructs an answer generation probability field using the following formula:

[0050]

[0051] where A is a potential answer, Z is a partition function, β is an inverse temperature parameter, H is an energy function, and φ i is a local potential function.

[0052] Preferably, the answer optimization module extracts the optimal answer using the following formula:

[0053]

[0054] where A * is the optimal answer, γ is a balance parameter, and Ric is the Ricci curvature.

[0055] Preferably, it further includes a learning and updating module, which is communicatively connected to the result output module and is used for:

[0056] receiving feedback information from the user on the output result;

[0057] updating the knowledge graph and model parameters based on the feedback information.

[0058] Preferably, the system further includes a cloud server, which is used for:

[0059] storing the knowledge graph and model parameters;

[0060] providing distributed computing resources to support large - scale parallel processing.

[0061] The beneficial effects of the present invention are mainly reflected in the following aspects:

[0062] The intelligent education Q&A information processing system based on AI teaching assistants of the present invention has achieved breakthrough progress in many aspects, bringing significant technological improvements and application values to the field of intelligent education.

[0063] From a macroscopic perspective, the present invention constructs a complete and adaptive intelligent education ecosystem. It can not only answer students' immediate questions, but also continuously improve itself through continuous learning to adapt to the dynamic changes in the education field. This self-improving ability enables the system to maintain high efficiency and relevance in the long term, greatly extending the system's life cycle and reducing the costs of maintenance and update.

[0064] At the system architecture level, the present invention realizes high flexibility and scalability through modular design. The close cooperation among core components such as the education Q&A collection module, question division module, text vectorization module, semantic space mapping module, knowledge graph matching module, answer generation module, answer optimization module, and result output module forms a complete information processing chain. This design not only improves the overall performance of the system, but also enables each module to be independently optimized and upgraded, leaving room for future technological innovation.

[0065] In terms of knowledge representation and processing, the innovation of the present invention lies in introducing quantum concepts into the text vectorization process and introducing new mathematical tools for natural language processing through a special random matrix representation method. This method not only retains the semantic information of the question, but also introduces the concept of quantum superposition, which can better capture the potential relationships between words. This innovation provides a powerful mathematical basis for dealing with complex education problems.

[0066] In terms of semantic understanding and knowledge matching, the present invention adopts a semantic space conversion method based on non-linear mapping and a knowledge graph matching algorithm based on continuous manifolds. The combination of these methods enables the system to deeply understand the semantic connotation of the question and accurately locate relevant knowledge points in the huge knowledge graph. This not only improves the accuracy of the answer, but also discovers potential knowledge associations, helping students build a more systematic and comprehensive knowledge system.

[0067] In terms of answer generation and optimization, the present invention introduces a probability field model based on statistical physics and an answer optimization method based on geometric concepts. This innovative combination not only ensures the accuracy and completeness of the generated answer, but also takes into account the internal consistency and readability of the answer. This is crucial for generating high-quality education answers, because a good answer should not only be accurate, but also easy to understand and learn.

[0068] In terms of the adaptability of the system, the present invention designs an update mechanism based on incremental learning and a distributed storage and computing architecture. This enables the system to efficiently process user feedback, update the knowledge base and model parameters in real time, while maintaining the stability of the system. This design greatly enhances the learning ability and scalability of the system, enabling it to continuously adapt to new knowledge and teaching requirements.

[0069] From a microscopic perspective, the present invention has achieved significant improvements in multiple key performance indicators. In terms of the accuracy and completeness of answers, the present system has reached 95.8% and 92.5% respectively, far exceeding traditional retrieval-based systems and systems based on pre-trained language models. This means that students can obtain more accurate and comprehensive knowledge support. In terms of knowledge relevance, the present system has reached a high level of 90.3%, which helps students establish a more systematic and in-depth understanding of knowledge. In terms of learning ability, the system has shown a 15.2% performance improvement, and this ability to continuously progress ensures that the system can maintain high efficiency in the long term.

[0070] Generally speaking, the intelligent education Q&A information processing system based on AI teaching assistant of the present invention represents a major breakthrough in educational technology. It not only solves the problems of insufficient accuracy, lack of knowledge association, and low adaptability in the prior art, but also realizes innovations in multiple aspects such as system architecture, algorithm innovation, and knowledge representation. This system opens up new possibilities for intelligent education, has the potential to completely change the education mode, and promotes the development of personalized learning and adaptive education. It can not only significantly improve the quality and efficiency of online education, but also provide students with more comprehensive and in-depth learning support, bringing a revolutionary change to the future education cause. BRIEF DESCRIPTION OF THE DRAWINGS

[0071] Figure 1 It is the overall logic block diagram of the system of the present invention.

[0072] Figure 2 It is the logic block diagram of the education Q&A collection module of the present invention.

[0073] Figure 3 It is the logic block diagram of the problem division module of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0074] Please refer to Figures 1-3 , the present invention provides an intelligent education Q&A information processing system based on an AI teaching assistant. This system can intelligently process Q&A information in the field of education and provide students with efficient and accurate question-answering services. The following will describe the detailed implementation manners of the present invention.

[0075] The system of the present invention includes multiple functional modules that work together to achieve the full - process intelligent processing from problem input to answer output. Specifically, the system of the present invention includes an educational Q&A collection module 1, a problem division module 2, a text vectorization module 3, a semantic space mapping module 4, a knowledge graph matching module 5, an answer generation module 6, an answer optimization module 7, and a result output module 8.

[0076] The educational Q&A collection module 1 is used to obtain the educational questions input by the user and pre - process the obtained educational questions to generate pre - processed information. In an embodiment of the present invention, the user can input questions in various ways, such as text input or voice input. The pre - processing process may include operations such as removing stop words, correcting spelling mistakes, and identifying key words. For example, when the user inputs "What is Newton's second law?", the pre - processed information may be "Newton's second law".

[0077] The problem division module 2 is communicatively connected to the educational Q&A collection module 1 and is used to receive the pre - processed information sent by the educational Q&A collection module 1 and perform problem semantic analysis and classification based on this pre - processed information. The present invention preferably uses a deep - learning model for semantic analysis, such as the BERT (Bidirectional Encoder Representations from Transformers) model. The classification may include subject classification (such as physics, chemistry, mathematics, etc.) and difficulty classification (such as primary, intermediate, advanced). For example, the above - mentioned question may be classified as a physics subject and of intermediate difficulty.

[0078] The text vectorization module 3 is communicatively connected to the problem division module 2 and is used to receive the classification result sent by the problem division module 2 and convert the problem into a special random matrix representation based on this classification result. This step is an innovation point of the present invention, which converts natural - language questions into a mathematical representation that can be processed by machines.

[0079] The semantic space mapping module 4 is communicatively connected to the text vectorization module 3 and is used to receive the random matrix representation sent by the text vectorization module 3 and perform a non - linear mapping to a high - dimensional semantic space based on this random matrix representation. This step further extracts the deep semantic features of the problem and prepares for the subsequent knowledge graph matching.

[0080] The knowledge graph matching module 5 is communicatively connected to the semantic space mapping module 4 and is used to receive the high - dimensional semantic representation sent by the semantic space mapping module 4 and perform matching and embedding in the knowledge graph based on this high - dimensional semantic representation. The knowledge graph used in the present invention is a structured database containing a large amount of knowledge in the education field. By performing matching in this graph, the system can find the knowledge points most relevant to the problem.

[0081] The answer generation module 6 is communicatively connected to the knowledge graph matching module 5, and is used to receive the matching result sent by the knowledge graph matching module 5, and construct an answer generation probability field based on the matching result. This step uses the concept of statistical physics to transform the answer generation problem into a probability optimization problem.

[0082] The answer optimization module 7 is communicatively connected to the answer generation module 6, and is used to receive the probability field information sent by the answer generation module 6, and extract and refine the optimal answer based on the probability field information. This step ensures that the generated answer is not only accurate but also easy to understand.

[0083] Finally, the result output module 8 is communicatively connected to the answer optimization module 7, and is used to receive the optimized answer sent by the answer optimization module 7, and send the optimized answer to the user side. Users can receive the answers generated by the system through various devices (such as mobile phones, tablets, personal computers, etc.).

[0084] The educational Q&A collection module 1 of the present invention includes a voice input unit 11, a text input unit 12, and a preprocessing unit 13. This design enables the system to adapt to the usage habits and scenario requirements of different users.

[0085] The voice input unit 11 is used to receive the voice question input of the user. The present invention preferably adopts a high-precision speech recognition technology, such as an end-to-end speech recognition model based on deep learning, to ensure accurate capture of the user's voice input. For example, when the user says "Please explain the process of photosynthesis", the voice input unit 11 can accurately convert this piece of speech into text.

[0086] The text input unit 12 is used to receive the text question input of the user. This can be achieved through keyboard input, handwriting recognition, or other text input methods. The text input unit 12 supports multiple languages and input methods to meet the needs of users from different regions and cultural backgrounds.

[0087] The preprocessing unit 13 is connected to the voice input unit 11 and the text input unit 12, and is used to perform preprocessing operations such as word segmentation and stop word removal on the input questions. Preprocessing is a key step to improve the system performance. For example, for the input question "What is the process of photosynthesis?", the preprocessing unit 13 may perform the following operations:

[0088] 1. Word segmentation: Split the sentence into "photosynthesis / of / process / is / how / of / ?";

[0089] 2. Stop word removal: Remove words such as "of", "is", "how" that have little impact on the main idea of the question, and get "photosynthesis process";

[0090] 3. Lemmatization: Restore verbs, adjectives, etc. to their basic forms;

[0091] 4. Standardization: Unify case, remove redundant punctuation marks, etc.;

[0092] These preprocessing operations can significantly improve the efficiency and accuracy of subsequent processing steps.

[0093] The problem division module 2 of the present invention includes a semantic analysis unit 21, a classification unit 22, and a tag generation unit 23. This module design can deeply understand the semantic content of the problem and provide important structured information for subsequent processing.

[0094] The semantic analysis unit 21 is used to perform semantic parsing on the preprocessed information. The present invention preferably adopts advanced natural language processing technologies, such as pre-trained language models like BERT (Bidirectional Encoder Representations from Transformers) or GPT (Generative Pre-trained Transformer), to understand the deep semantics of the problem. For example, for the question "What is the process of photosynthesis?", the semantic analysis unit 21 can not only identify the key concepts "photosynthesis" and "process", but also understand that this is a question asking about the process, rather than a question asking about the definition or application.

[0095] The classification unit 22 is connected to the semantic analysis unit 21 and is used to divide the problem into predefined categories based on the semantic parsing result. These categories may include:

[0096] 1. Subject categories: such as biology, physics, chemistry, mathematics, etc.;

[0097] 2. Problem types: such as definition explanation, process description, principle elaboration, application examples, etc.;

[0098] 3. Difficulty levels: such as beginner, intermediate, advanced;

[0099] 4. Knowledge point scope: such as Chapter 3 of high school textbooks, undergraduate courses, etc.;

[0100] For the above photosynthesis problem, the classification unit 22 may classify it as a biology subject, a process description type, intermediate difficulty, and within the knowledge point scope of Chapter 4 of high school biology textbooks.

[0101] The tag generation unit 23 is connected to the classification unit 22 and is used to generate knowledge point tags and domain tags for the divided problem. These tags further refine the attributes of the problem and help the system more accurately match relevant knowledge and generate answers in subsequent steps. For example, for the photosynthesis problem, the possible generated tags may include:

[0102] Knowledge point tags: #Photosynthesis#Chloroplast#Light reaction#Dark reaction;

[0103] Field labels: #Plant Physiology#Energy Conversion#Carbon Cycle;

[0104] These labels not only help the system better understand and locate problems, but can also be used for advanced functions such as related problem recommendation and learning path planning.

[0105] From the above detailed description, it can be seen that the system of the present invention has significant advantages in problem collection, preprocessing, and semantic understanding. It can accurately capture the problems input by users, deeply understand the semantic content of the problems, and provide rich structured information for subsequent knowledge matching and answer generation. This design not only improves the accuracy and efficiency of the question-answering system, but also provides a solid foundation for personalized learning and intelligent education. The system of the present invention further includes a series of innovative algorithm modules, which work together to achieve the full-process intelligent processing from text vectorization to the generation of the optimal answer. The working principles and specific implementation methods of these modules will be described in detail below.

[0106] The text vectorization module 3 of the present invention adopts a special random matrix representation method. Specifically, this module uses the following formula to convert the problem into a special random matrix representation:

[0107]

[0108] In this formula, ω i represents the word weight, |w i > represents the quantum state representation of the word vector, n is the number of words in the problem, H is a random Hermitian matrix, and ∈ is a small perturbation parameter. This representation method integrates the concepts of quantum mechanics and introduces new mathematical tools for natural language processing.

[0109] In the preferred embodiment of the present invention, the word weight ω i can be calculated by the TF-IDF (Term Frequency-Inverse Document Frequency) method. For example, for the question "What is the process of photosynthesis?", the weight of the word "photosynthesis" may be much higher than the weight of the word "what". The word vector |w i > can be obtained using a pre-trained word embedding model (such as Word2Vec or GloVe).

[0110] In an intelligent education Q&A system based on an AI teaching assistant, the questions raised by students are usually in natural language form. To effectively process and understand these questions, they need to be transformed into a mathematical model. Here, the concept of quantum state superposition is used to represent the questions; data on students' questions are collected from an education platform, including text content, context information, etc. Preprocessing steps such as word segmentation, part-of-speech tagging, and syntactic analysis are performed using NLP techniques to extract word weights and word vectors.

[0111] For example, when answering "What is photosynthesis?", the system can identify "photosynthesis" as the core vocabulary and other words as auxiliary information, thus more accurately matching relevant knowledge points.

[0112] The introduction of the small perturbation parameter ∈ is to increase the robustness of the system. In practice, the value of ∈ is usually set between 0.01 and 0.1. The existence of this parameter enables the system to better handle unknown words or rare words and improves the generalization ability of the model.

[0113] The introduction of the random Hermitian matrix H is inspired by the density matrix in quantum mechanics. It introduces a certain degree of randomness to the question representation and helps to capture the potential relationships between words. In practical applications, H can be constructed by generating a random complex matrix and adding it to its conjugate transpose.

[0114] This special method of representing random matrices not only preserves the semantic information of the questions but also introduces quantum concepts, providing a rich mathematical structure for subsequent semantic space mapping.

[0115] Transform the question representation into a high-dimensional semantic space to better capture the deep meaning of the questions.

[0116] Next, the semantic space mapping module 4 of the present invention adopts a non-linear mapping method. This module uses the following formula for non-linear mapping:

[0117] S = Φ(Q) = tr(e iQY )P + λL(Q),

[0118] In this formula, Φ represents the non-linear mapping function, T is the time evolution operator, P represents the projection matrix, L is the Laplace operator, and λ is the regularization parameter. This mapping method combines the concepts of dynamic system theory and differential geometry and can effectively map the question representation to a high-dimensional semantic space.

[0119] In an embodiment of the present invention, the time evolution operator T can be set as the identity matrix multiplied by a small time step, such as 0.1 or 0.01. This setting can simulate the short-time evolution of the quantum system. The projection matrix P can be designed according to specific educational domain knowledge to highlight certain important semantic dimensions. The introduction of the Laplace operator L is to capture the "curvature" information of the problem representation, which can help the system understand the relative relationships between words. The regularization parameter λ is usually set between 0.1 and 1, and its role is to balance the relationship between structure preservation and information extraction.

[0120] The advantage of this non-linear mapping method is that it can effectively extract the deep semantic features of the problem while maintaining the key structure of the original representation. This lays a solid foundation for subsequent knowledge graph matching. Extract high-frequency questions and their answers from historical Q&A records to construct a knowledge graph. Perform graph embedding on the knowledge graph to obtain vector representations of nodes and edges.

[0121] For example, when answering the question "Why do plants need sunlight?", the system can find relevant biological knowledge points through semantic space mapping and further explain the process of photosynthesis.

[0122] The knowledge graph matching module 5 of the present invention adopts an innovative matching and embedding method. This module uses the following formula for matching and embedding:

[0123]

[0124] In this formula, represents the knowledge graph manifold, G(x) represents the local structure of the knowledge graph at point x, f is the matching function, μ is the measure, and D is the diffusion tensor. This method transforms the knowledge graph matching problem into an integral problem on a continuous manifold and introduces a diffusion process to enhance knowledge association.

[0125] In a preferred embodiment of the present invention, the knowledge graph manifold can be obtained through manifold learning of large-scale educational knowledge data. The local structure G(x) can be represented by a graph neural network, which can effectively capture the complex relationships between knowledge points. The matching function f can be designed as a similarity metric such as cosine similarity or dot product. The measure μ can be defined according to the importance or usage frequency of knowledge points. The introduction of the diffusion tensor D enables the matching process to consider not only directly relevant knowledge points but also explore potential associated knowledge.

[0126] The advantage of this matching and embedding method is that it can comprehensively consider the relationship between the question and the knowledge graph, not only finding directly relevant knowledge points but also discovering potential valuable associated knowledge. This has significant advantages for answering complex educational questions, especially those that require integrating multiple knowledge points.

[0127] Obtain various knowledge points from the educational resource library to construct a knowledge graph. Use a graph neural network (GNN) for node classification and link prediction to optimize the structure of the knowledge graph. For example, when answering the question "How to conduct a photosynthesis experiment?", the system can provide detailed experimental guidance based on the experimental steps and required materials in the knowledge graph.

[0128] Based on the knowledge graph embedding results, construct a probability distribution for answer generation and select the most appropriate answer.

[0129] The answer generation module 6 of the present invention adopts a method based on the concepts of statistical physics to construct an answer generation probability field. This module uses the following formula to construct the answer generation probability field:

[0130]

[0131] In this formula, A represents the potential answer, Z is the partition function, β is the inverse temperature parameter, H is the energy function, and φ i represents the local potential function. This method transforms the answer generation problem into an energy minimization problem similar to a physical system.

[0132] In an embodiment of the present invention, the energy function H(A, K) can be designed as a measure of the degree of inconsistency between the answer A and the knowledge graph matching result K. The inverse temperature parameter β controls the "sharpness" of the distribution, and usually, the optimal value can be selected through cross-validation.

[0133] The local potential function φ i can be used to capture the local consistency between various parts of the answer. For example, for the question of explaining the photosynthesis process, φ i can ensure that the logical order between the various steps mentioned in the answer is correct. The advantage of this answer generation method based on the probability field is that it can comprehensively consider global consistency and local rationality, and the generated answer not only matches the question and the knowledge graph but also ensures internal logical coherence.

[0134] Extract high-quality answer samples from the past answer database. Train through a machine learning model to optimize the probability distribution of answer generation. For example, when answering the question "What is the chemical equation of photosynthesis?", the system can generate multiple possible answers and select the most accurate one according to the energy function.

[0135] Extract the optimal answer from the generated probability field and perform post - processing to improve accuracy.

[0136] Finally, the answer optimization module 7 of the present invention adopts an innovative optimal answer extraction method. This module extracts the optimal answer using the following formula:

[0137]

[0138] In this formula, A * represents the optimal answer, γ is the balance parameter, and Ric represents the Ricci curvature. This method not only considers the probability of the answer but also introduces geometric concepts to evaluate the "smoothness" of the answer.

[0139] In the preferred embodiment of the present invention, the balance parameter γ can be determined by experiments on the validation set to obtain the optimal value, usually between 0.1 and 1. The introduction of the Ricci curvature term is to ensure that the generated answer is "smooth" in the semantic space, that is, the transition between various parts of the answer is natural without abrupt jumps.

[0140] Obtain the verified high - quality answers from the final answer library. Post - processing techniques (such as grammar checking, logical verification) further improve the accuracy and readability of the answers.

[0141] For example, when answering "What is the significance of photosynthesis?", the system not only gives the standard answer but also adds practical application scenarios and examples through post - processing to make the answer more rich and practical.

[0142] These algorithms together constitute a complete intelligent education Q&A system, which can efficiently process and answer various questions raised by students, significantly improving the quality and efficiency of education.

[0143] The advantage of this optimal answer extraction method is that it can not only select the answer with the highest probability but also ensure the coherence and readability of the answer. This is particularly important for generating high - quality education Q&A, because a good answer should not only be accurate but also easy to understand and learn.

[0144] As can be seen from the above detailed description, the system of the present invention adopts innovative algorithms in aspects such as text vectorization, semantic space mapping, knowledge graph matching, answer generation and optimization. These algorithms not only integrate advanced concepts from multiple disciplines such as quantum mechanics, dynamical systems, differential geometry, and statistical physics, but also are specifically designed and optimized for the characteristics of educational Q&A. This enables the system of the present invention to handle complex educational problems, generate high-quality answers, and provide powerful technical support for intelligent education. The system of the present invention not only includes the aforementioned innovative algorithm modules, but also designs a learning update mechanism and a cloud support architecture to achieve continuous optimization and efficient operation of the system. The following will detail these additional key components.

[0145] The system of the present invention further includes a learning update module 9, which is communicatively connected to the result output module 8. The main function of the learning update module 9 is to receive feedback information from the user on the output result and update the knowledge graph and model parameters based on this feedback information. This design enables the system to continuously improve itself and adapt to changes in user needs and the development of knowledge domains.

[0146] In a preferred embodiment of the present invention, the learning update module 9 adopts an incremental learning algorithm. When the system receives feedback from the user, for example, the user points out an error or inaccuracy in the answer, the learning update module 9 will first analyze this feedback and determine the knowledge points or model parameters that need to be updated. Then, it will perform a local update on the relevant part without affecting the overall performance of the system.

[0147] Specifically, for the update of the knowledge graph, the present invention adopts a dynamic graph embedding technique. Suppose the user points out a new knowledge point association. The system will add this new edge to the original knowledge graph and update the embedding representations of the relevant nodes simultaneously. This process can be expressed as:

[0148]

[0149] where v i represents the embedding representation of the i-th node in the knowledge graph, α is the learning rate, GNN represents the graph neural network operation, and E new is the set of newly added edges.

[0150] When the user points out a new knowledge point association, the system will add this new edge to the original knowledge graph and update the embedding representations of the relevant nodes. This method can dynamically expand the knowledge graph and make it more comprehensive and accurate.

[0151] When the user points out that the relationship between "photosynthesis" and "plant growth" is not fully covered, the system will, based on this feedback, add corresponding edges in the knowledge graph and update the embedding representations of relevant nodes to better capture this relationship.

[0152] For the update of model parameters, the present invention adopts the idea of online learning. Whenever the system receives a feedback, it will perform a gradient update using this sample. This process can be expressed as:

[0153]

[0154] where θ represents the model parameters, η is the learning rate, L is the loss function, and f θ represents the model with parameters θ, x is the input, and y is the correct answer feedback by the user.

[0155] Whenever the system receives a feedback, it will perform a gradient update using this sample. This enables the model to gradually improve without interrupting the service and adapt to new knowledge and user needs.

[0156] If the user points out that the answer to a certain question is not accurate enough, the system will fine-tune the model parameters according to the user's feedback, thereby improving the quality of answers to future similar questions.

[0157] The advantage of this learning update mechanism is that it can enable the system to continuously adapt to new knowledge and user needs while maintaining the stability of the system. This is particularly important in the field of education because educational knowledge and teaching methods are constantly evolving.

[0158] Next, the invented system further includes a cloud server 10. The main function of the cloud server 10 is to store the knowledge graph and model parameters and provide distributed computing resources to support large-scale parallel processing. This design greatly improves the storage capacity and computing power of the system, enabling the system to handle more complex problems and larger-scale data.

[0159] In a preferred embodiment of the present invention, the cloud server 10 adopts a distributed storage and computing architecture. The knowledge graph is divided into multiple subgraphs and stored on different server nodes. When the system needs to perform knowledge graph matching, it will perform parallel processing on multiple nodes simultaneously and then merge the results. This process can be expressed as:

[0160] K = Merge({K i | i = 1, 2,..., N}),

[0161] where K is the final matching result, K i is the matching result of the i-th subgraph, N is the number of subgraphs, and Merge is the result merging function.

[0162] The knowledge graph is partitioned into multiple sub - graphs and stored on different server nodes. When knowledge graph matching is required, the system processes in parallel on multiple nodes and then merges the results. This method can significantly improve the processing speed and efficiency.

[0163] When answering complex biological questions, the system can process different sub - graphs simultaneously on multiple nodes and merge the results, thus quickly generating accurate answers.

[0164] For large - scale parallel processing, the present invention adopts a combination of model parallelism and data parallelism. In model parallelism, a large - scale neural network model is divided into multiple parts and distributed on different computing nodes. In data parallelism, the same model is replicated to multiple nodes, and each node processes different data batches. This way can make full use of the computing resources of the distributed system and greatly improve the processing speed.

[0165] This way makes full use of the computing resources of the distributed system and greatly improves the processing speed. For example, when processing a large number of students' questions, the system can process multiple requests simultaneously through data parallelism technology, thus realizing an efficient question - answering service.

[0166] During peak periods, the system can process the questions of thousands of students simultaneously and quickly generate high - quality answers through distributed computing resources.

[0167] Another important function of the cloud server 10 is to support the online update of the model. When the learning update module 9 generates new model parameters, these updates are pushed to the cloud server 10 and then broadcast to all computing nodes. This process ensures that the system uses the latest model for question - answering processing at any time.

[0168] In addition, the cloud server 10 also provides data backup and fault - tolerance mechanisms. Through data redundant storage and fault detection and recovery technologies, the high availability and data security of the system are ensured. Even if a certain node fails, the system can still run normally.

[0169] This design ensures that the system always uses the latest model for question - answering processing and that even if a certain node fails, the system can still run normally.

[0170] If a certain node fails, the system can automatically switch to the backup node to ensure uninterrupted service.

[0171] The advantage of this cloud - based distributed architecture is that it can handle large - scale educational question - answering requests, support real - time model updates, and provide highly reliable services. This enables the system of the present invention to meet the needs of large - scale online education platforms and provide high - quality intelligent question - answering services for millions of users.

[0172] Through the learning update module 9, the system can continuously adapt to new knowledge and user needs, maintaining the latest and most accurate information. The distributed computing resources provided by the cloud server 10 enable the system to handle more complex problems and larger-scale data, significantly improving the processing speed and efficiency.

[0173] The fault tolerance mechanism ensures the high availability and data security of the system, allowing it to continue running even in the event of partial node failures. The system can provide personalized learning suggestions and Q&A services according to the needs of different students, enhancing the learning effect. It supports real-time model updates and large-scale parallel processing, capable of meeting the simultaneous access needs of millions of users and providing high-quality intelligent Q&A services. The dynamically updated knowledge graph can help teachers and educational institutions better manage and organize educational resources, improving teaching quality and efficiency.

[0174] In summary, the intelligent education Q&A information processing system based on AI teaching assistant of the present invention realizes efficient, accurate, and scalable education Q&A services through innovative algorithm design, adaptive learning mechanism, and cloud distributed architecture. This system can not only answer complex education questions but also continuously improve itself to adapt to the development and changes in the education field. Its application will greatly improve the quality and efficiency of online education, opening up new possibilities for personalized learning and intelligent education.

[0175] To verify the superiority of the intelligent education Q&A information processing system based on AI teaching assistant of the present invention, a set of examples and comparative examples were designed and detailed performance tests were conducted. The following is a detailed description of the test plan and results.

[0176] Example 1: Fully implement the system of the present invention

[0177] In this example, all the modules described in the present invention are fully implemented, including innovative text vectorization, semantic space mapping, knowledge graph matching, answer generation and optimization algorithms, as well as the learning update module and cloud distributed architecture. The system uses a large-scale knowledge graph in the education field, containing more than 1 million knowledge points and 10 million relationships.

[0178] Comparative Example 1: Traditional retrieval-based Q&A system

[0179] This system uses traditional information retrieval techniques to retrieve answers from a pre-prepared Q&A database through keyword matching. It does not use deep semantic understanding or knowledge graph technology.

[0180] Comparative Example 2: Q&A system based on pre-trained language model

[0181] This system uses the latest pre-trained language models (such as GPT-3) to generate answers, but does not have a dedicated knowledge graph matching or answer optimization module.

[0182] The following test metrics were designed to evaluate the system performance:

[0183] 1. Answer accuracy: The correctness of the answers was evaluated by domain experts.

[0184] 2. Answer completeness: Evaluate whether the answers comprehensively cover all aspects of the questions.

[0185] 3. Response time: The time from receiving the question to generating the answer.

[0186] 4. Knowledge relevance: Evaluate whether the system can associate relevant knowledge points.

[0187] 5. Learning ability: The ability of the system to improve performance based on feedback.

[0188] 1000 educational questions from different disciplines and difficulty levels were randomly selected for testing. Each question was evaluated by 3 education domain experts. The following are the test results:

[0189] Test indicators Example 1 Comparative Example 1 Comparative Example 2 Answer accuracy rate 95.8 78.3 89.2 Answer integrity 92.5 70.1 85.7 Average response time 1.2 0.8 2.5 Knowledge correlation degree 90.3 60.5 75.8 Improvement of learning ability 15.2 2.1 8.7

[0190] From the test results, it can be seen that the system of the present invention (Example 1) is significantly superior to the comparative system in most metrics. The specific analysis is as follows:

[0191] Answer accuracy and completeness: The system of the present invention performs best in these two core metrics. This is mainly due to the innovative knowledge graph matching algorithm and answer generation optimization method. The system can not only accurately understand the questions, but also extract comprehensive relevant information from the knowledge graph to generate high-quality answers.

[0192] Response time: Although the response time of the system of the present invention is slightly longer than that of traditional retrieval-based systems, considering the quality and completeness of the answers, this response time is very ideal. Compared with systems based on large language models, the system of the present invention has a faster response speed, which benefits from the efficient knowledge graph matching algorithm and distributed computing architecture.

[0193] Knowledge relevance: The advantage of the system of the present invention in this metric is particularly obvious. This reflects that the system can not only answer specific questions, but also provide relevant background knowledge and associated concepts, which helps students build a more comprehensive knowledge system.

[0194] Learning ability: After a week of continuous use and feedback, the system of the present invention has demonstrated remarkable learning ability. The system performance has been improved by 15.2%, which is much higher than the other two systems. This proves the effectiveness of the learning and updating module, enabling the system to continuously adapt to new knowledge points and user needs.

[0195] Based on the test results, it is considered that Example 1 represents the best implementation mode of the present invention. In this example, it is found that the scale and quality of the knowledge graph have an important impact on the system performance. When the knowledge graph contains more than 1 million knowledge points, the answer accuracy and integrity of the system are significantly improved. In addition, setting the learning rate to 0.01 and updating the model parameters every 1000 Q&A sessions seems to be an ideal balance point, which can not only adapt to new knowledge in a timely manner but also prevent the system from being unstable.

[0196] These test results fully prove the superiority of the system of the present invention. It not only performs excellently in answer accuracy and integrity but also demonstrates strong knowledge association ability and learning ability. Such a system can provide students with higher-quality learning support, not only answering questions but also stimulating thinking and helping to establish knowledge connections. At the same time, the adaptive ability of the system ensures that it can continuously improve and keep up with the latest developments in the education field.

[0197] Generally speaking, the intelligent education Q&A information processing system based on AI teaching assistant of the present invention represents an important breakthrough in educational technology. It combines advanced natural language processing technology, knowledge graph, and machine learning methods, opening up new possibilities for intelligent education. This system can not only improve the quality and efficiency of online education but also has the potential to promote the development of personalized learning and adaptive education, bringing a revolutionary change to the future education model.

[0198] It should be noted that the above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. An intelligent education question-answering information processing system based on AI teaching assistants, characterized in that: include: Educational question and answer collection module, used for: Get user input for educational questions; Preprocessing the acquired educational questions to generate preprocessing information; The question division module is in communication with the educational question and answer collection module and is used to: Receiving preprocessing information sent by the educational question and answer collection module; Based on the preprocessed information, performing question semantic analysis and classification; A text vectorization module is connected to the problem partitioning module and is used to: Receiving the classification result sent by the problem division module; Based on the classification results, the problem is transformed into a special random matrix representation; A semantic space mapping module is connected to the text vectorization module for: Receiving the random matrix representation sent by the text vectorization module; Based on the random matrix representation, performing nonlinear mapping to a high-dimensional semantic space; A knowledge graph matching module, which is in communication with the semantic space mapping module and is used to: receive the high-dimensional semantic representation sent by the semantic space mapping module; Based on the high-dimensional semantic representation, matching and embedding are performed in the knowledge graph; The answer generation module is connected to the knowledge graph matching module and is used to: Receiving the matching result sent by the knowledge graph matching module; Based on the matching results, construct an answer generation probability field; The answer optimization module is in communication with the answer generation module and is used to: Receiving probability field information sent by the answer generation module; Based on the probability field information, extract the best answer and refine it; A result output module is connected to the answer optimization module for: Receiving the optimized answer sent by the answer optimization module; The optimized answer is sent to the user end.

2. The system according to claim 1, characterized in that The educational question and answer collection module includes: A voice input unit, used to receive voice question input from a user; A text input unit, used to receive text question input from a user; The preprocessing unit is connected to the speech input unit and the text input unit, and is used for performing preprocessing operations such as word segmentation and stop word removal on the input question.

3. The system according to claim 1, characterized in that The problem division module includes: A semantic analysis unit, used for semantically analyzing the preprocessed information; A classification unit, connected to the semantic analysis unit, for classifying questions into predefined categories based on the semantic analysis results; The label generation unit is connected to the classification unit and is used to generate knowledge point labels and domain labels for the divided questions.

4. The system according to claim 1, characterized in that The text vectorization module uses the following formula to convert the problem into a special random matrix representation: Among them, ω i is the word weight, |w i > is the quantum state representation of the word vector, n is the number of words in the question, H is a random Hermitian matrix, and ∈ is a small perturbation parameter.

5. The system according to claim 1, characterized in that The semantic space mapping module uses the following formula for nonlinear mapping: S=Φ ( Q ) =tr ( from iQT ) P+λL ( Q ) , Among them, Φ is the nonlinear mapping function, T is the time evolution operator, P is the projection matrix, L is the Laplacian operator, and λ is the regularization parameter.

6. The system according to claim 1, characterized in that The knowledge graph matching module uses the following formula for matching and embedding: in, is the knowledge graph manifold, G ( x ) is the local structure of the knowledge graph at point x, f is the matching function, μ is the measure, and D is the diffusion tensor.

7. The system according to claim 1, characterized in that The answer generation module uses the following formula to construct the answer generation probability field: Among them, A is the potential answer, Z is the partition function, β is the inverse temperature parameter, H is the energy function, φ i is the local potential function.

8. The system according to claim 1, characterized in that The answer optimization module uses the following formula to extract the best answer: Among them, A * is the optimal answer, γ is the balance parameter, and Ric is the Ricci curvature.

9. The system according to claim 1, characterized in that It also includes a learning update module, which is in communication with the result output module and is used to: Receive user feedback on output results; Based on the feedback information, the knowledge graph and model parameters are updated.

10. The system according to claim 1, characterized in that The system also includes a cloud server for: Store knowledge graphs and model parameters; Provides distributed computing resources and supports large-scale parallel processing.

Citation Information

Cited By

  • Knowledge graph-based low-altitude economic domain question and answer teaching interaction method and system

    CN120561254A

  • Distributed intelligent smoke monitoring and alarm management system

    CN120766425A

  • A distributed intelligent smoke monitoring and alarm management system

    CN120766425B

  • Question and answer information processing method for AI online education based on big data

    CN121581215A