An artificial intelligence-based knowledge question and answer rapid processing method and system
By constructing a multimodal knowledge graph and a hybrid retrieval strategy, combined with the RAG enhancement framework and cognitive reinforcement learning, the problem of response speed and accuracy of existing knowledge question answering systems in complex scenarios is solved, and personalized and efficient knowledge question answering services are achieved.
Patent Information
- Application Number
- CN202511092863.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-06
- Publication Date
- 2025-12-05
- Estimated Expiration
- 2045-08-06
AI Technical Summary
Existing knowledge-based question-answering systems are slow to respond to complex and ever-changing question scenarios, have low accuracy rates, struggle to meet personalized needs, and lack a deep understanding of individual users' learning history and interests.
By constructing a multimodal knowledge graph, employing a hybrid retrieval strategy and the RAG enhancement framework, and combining graph neural networks, federated learning, quantum optimization, and cognitive reinforcement learning, cross-modal data fusion and efficient question answering are achieved.
It improves the accuracy and efficiency of the knowledge question-answering system, enhances the system's adaptability and cross-domain adaptability, provides personalized learning resources and immersive learning experiences, and improves response speed and flexibility.
Smart Images

Figure CN120596639B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of artificial intelligence, in particular to the field of intelligent question answering systems, and specifically to a knowledge question answering rapid processing method and system based on artificial intelligence. BACKGROUND
[0002] With the rapid development of artificial intelligence technology, especially the continuous progress of large language models (LLM) and multi-modal technology, various intelligent question answering systems have been widely applied in education, scientific research and learning fields. However, existing knowledge question answering systems face many challenges in practical applications.
[0003] Existing knowledge question answering systems often face slow response speed, low accuracy of answers and difficulty in meeting individual needs when dealing with complex and variable problem scenarios. When dealing with complex problems, the system often needs a long time to search, analyze and generate answers, resulting in poor user experience. Existing systems generate answers that are not accurate and comprehensive when dealing with polysemy, ambiguity or problems that require comprehensive knowledge from multiple aspects. Moreover, the system usually lacks a deep understanding of individual learning history and interests of users, making it difficult to recommend personalized learning resources and paths. With the development of artificial intelligence technology, especially the application of large language models and multi-modal technology, it is necessary to design a more efficient, intelligent and flexible knowledge question answering system to improve the efficiency of user learning and scientific research. SUMMARY
[0004] Therefore, the purpose of the present application is to provide a knowledge question answering rapid processing method and system based on artificial intelligence, which can improve the professional skill level of teachers and students in the field of artificial intelligence and large model application, and meet their individual needs in teaching, scientific research and innovation courses.
[0005] To achieve the above-mentioned purpose, the present application provides the following technical solutions:
[0006] Based on the above purpose, in the first aspect, the present application provides a knowledge question answering rapid processing method based on artificial intelligence, comprising the following steps:
[0007] Collecting multi-source teaching data including text, image and voice, defining graph node features according to the data modalities of the multi-source teaching data, and extracting inter-node relationships from the multi-source teaching data;
[0008] Training node relationship weights and similarity through graph neural network, periodically updating nodes using time decay mechanism, aligning and integrating image, text and voice features into unified semantic space using cross-modal, and constructing multi-modal knowledge graph;
[0009] Fusing the collected multi-source teaching data through a hybrid retrieval strategy, weighting and sorting the retrieval results to generate a comprehensive retrieval list;
[0010] Based on the RAG enhanced framework, multi-level question and answer processing is performed, multi-modal input intentions of multi-source teaching data are parsed, high-frequency questions are cross-library combined retrieval, retrieval results of the integrated comprehensive retrieval list are integrated, and optimized answers are generated in combination with teaching scenes;
[0011] Through federated learning to collaboratively update the global model with multi-node model parameters, the global model is distilled to the lightweight TinyBERT architecture, the quality of the question and answer is dynamically optimized through the cognitive reinforcement learning framework, the key document is located from the comprehensive retrieval list, the optimized answer is evaluated and verified with the key document, and the answer is reconstructed.
[0012] As a further scheme of the present application, when the collected multi-source teaching data is fused through the hybrid retrieval strategy, the hybrid retrieval strategy includes semantic retrieval, vector retrieval and metadata retrieval, and three modalities of retrieval are performed in parallel, including:
[0013] Semantic retrieval: calculating the semantic similarity of queries and documents based on the BERT-Whitening model;
[0014] Vector retrieval: achieving millisecond-level neighbor search of hundreds of millions of vectors through FAISS index;
[0015] Metadata retrieval: using ElasticSearch to match structured information such as title, author and keyword.
[0016] As a further scheme of the present application, by fusing the collected multi-source teaching data through the hybrid retrieval strategy, the fusion result obtained is weighted and sorted according to the preset weight ratio 6:3:1 for the three types of retrieval results of semantic retrieval, vector retrieval and metadata retrieval, and a comprehensive retrieval list is generated.
[0017] As a further scheme of the present application, a multi-modal knowledge graph is constructed, and the collected multi-source teaching data is fused through a hybrid retrieval strategy, including the following steps:
[0018] Collecting teaching data, defining the characteristics of graph nodes according to data modalities, and extracting the relationships between nodes from multi-source data;
[0019] Training the multi-modal knowledge graph through a graph neural network, learning the relationship weights and similarities between nodes, and periodically updating the nodes in the graph using a time decay mechanism;
[0020] Cross-modal alignment is adopted to integrate the features of images, texts and voices, and semantic retrieval, vector retrieval and metadata retrieval are performed in parallel, and for the three types of retrieval results, weighted sorting is performed according to the preset weight, and a comprehensive retrieval list is generated;
[0021] Fuse the documents or data obtained from different retrieval strategies, remove duplicate results, and form a comprehensive retrieval result.
[0022] As a further scheme of the present application, when performing multi-level question answering processing based on the RAG enhanced framework, the multi-modal input intention includes image input, voice input and text input, wherein the visual feature vector is extracted through the CLIP model, and is mapped to the semantic space for image input; the Whisper model is used for voice-to-text conversion, and the intention keywords are extracted for voice input; the BERT model is used to generate an intention vector, and the classifier outputs the question type for text input.
[0023] As a further scheme of the present application, when performing cross-database joint retrieval, quantum optimization cross-database retrieval is adopted, for problems with daily call quantity greater than a preset number of times, a quantum optimization cross-database retrieval model is constructed, and annealing is performed on a set quantum computer to output an optimal document set; wherein the quantum optimization cross-database retrieval is designed for high-frequency problems, and when performing cross-database joint retrieval, for high-frequency problems with daily call quantity greater than 1000 times, a quantum optimization model is constructed, and the quantum computer is used to simulate the annealing process to output the optimal document set; the quantum annealing algorithm is used to filter the optimal set from the billion-level documents, and the speed is improved by more than 40% compared with the traditional brute force retrieval, and the core difference lies in that the document association strength and independent weight are mapped to the spin interaction of the quantum bit, and the quantum tunneling effect is used to accelerate the search for the optimal solution.
[0024] As a further scheme of the present application, a quantum optimization cross-database retrieval model is constructed, which converts the document selection problem into a quadratic unconstrained binary optimization problem, and the objective function of the quantum optimization cross-database retrieval model is:
[0025]
[0026] wherein, is a binary variable, and indicates whether the i-th document is selected, wherein 1=selected and 0=not selected; is a binary variable, and indicates whether the i-th document is selected, wherein 1=selected and 0=not selected; represents the association strength of the document and the document , which is a positive value indicating synergistic effect or a negative value indicating repulsion; represents the independent weight of the document , which is calculated by weighting the click rate, authority and other metadata.
[0027] As a further scheme of the present application, the global model is distilled to the lightweight TinyBERT architecture, including:
[0028] Based on the GPT-4 architecture model trained locally by each university, the global model is updated by federated averaging;
[0029] The global model is distilled to a TinyBERT architecture student model, and Gaussian noise is added during parameter transmission for differential protection data.
[0030] As a further scheme of the application, the quality of question and answer is dynamically optimized through a cognitive reinforcement learning framework, and the learning framework includes a retrieval agent, a verification agent and a generation agent, key documents are located from mixed retrieval results, scores of answer fact consistency are calculated, and answers are reconstructed according to verification results.
[0031] In a second aspect, the application provides a knowledge question and answer rapid processing system based on artificial intelligence, comprising:
[0032] A data acquisition module is configured to collect multi-source teaching data including text, image and voice, define graph node features according to data modalities of the multi-source teaching data, and extract inter-node relationships from the multi-source teaching data;
[0033] A multi-modal knowledge graph construction module is configured to train node relationship weights and similarity through a graph neural network, periodically update nodes using a time decay mechanism, integrate image, text and voice features into a unified semantic space using cross-modal alignment, and construct a multi-modal knowledge graph;
[0034] A hybrid retrieval module is configured to fuse the collected multi-source teaching data through a hybrid retrieval strategy, weight and sort retrieval results to generate a comprehensive retrieval list;
[0035] A RAG enhanced framework module is configured to perform multi-level question and answer processing based on a RAG enhanced framework, analyze multi-modal input intentions of the multi-source teaching data, perform cross-database joint retrieval on high-frequency questions, integrate retrieval results of the comprehensive retrieval list and generate optimized answers in combination with teaching scenarios;
[0036] A dynamic knowledge distillation module is configured to update a global model through federated learning and collaborative multi-node model parameter updating, and distill the global model to a lightweight TinyBERT architecture;
[0037] A question and answer optimization verification module is configured to dynamically optimize the quality of question and answer through a cognitive reinforcement learning framework, locate key documents from the comprehensive retrieval list, evaluate optimized answers and key document verification scores, and reconstruct answers.
[0038] In still another aspect of the application, a computer device is also provided, comprising a memory and a processor, the memory storing a computer program, and the computer program being executed by the processor to perform any one of the above knowledge question and answer rapid processing methods based on artificial intelligence according to the application.
[0039] Still another aspect of the present application also provides a computer readable storage medium storing computer program instructions, which, when executed, implement any of the above-mentioned artificial intelligence-based knowledge question and answer rapid processing methods according to the present application.
[0040] Compared with the prior art, the artificial intelligence-based knowledge question and answer rapid processing method and system proposed by the present application has the following beneficial effects:
[0041] 1. Through the construction of multi-modal knowledge graph, mixed retrieval strategy and RAG enhancement framework, the accuracy and efficiency of knowledge question and answer are improved. The present application constructs a comprehensive knowledge graph through the fusion of multi-modal data, so that the system can understand and process user input from different modalities, greatly improving the accuracy and adaptability of the system in complex scenarios. The mixed retrieval strategy combining semantic retrieval, vector retrieval and metadata retrieval ensures efficient response to user queries. Through the way of weighted sorting, the system can fully utilize different types of data sources, and comprehensively retrieve the results to improve the accuracy of question and answer. Multi-level question and answer processing based on RAG framework can optimize the query processing process at different levels, ensure accurate analysis and generation of cross-modal input intent, and provide more accurate answers.
[0042] 2. Through dynamic knowledge distillation, federated learning, quantum heuristic optimization and cognitive reinforcement learning, the adaptive ability and dynamic optimization of the system are strengthened. The present application realizes the collaborative optimization between multiple nodes through federated learning, so that the distributed model can improve the generalization ability of the model while protecting the data. The dynamic knowledge distillation technology enables the global model to be converted into a lightweight model suitable for edge devices, improving the response speed and flexibility of the system. For high-frequency retrieval tasks, the quantum heuristic optimization method significantly improves the task processing efficiency and retrieval quality, especially in the case of large-scale data sets and complex queries, which can accelerate the optimization process through quantum computing, greatly reducing the response time. Through the cognitive reinforcement learning framework, the system's answer quality and retrieval strategy are continuously optimized, enhancing the adaptability and intelligence level of the system when facing different user demands. The agent model improves the recognition and optimization ability of the answer quality through continuous interactive learning.
[0043] 3. Through the virtual-real fusion training environment, the depth fusion of virtual and reality is realized, and the interactivity and immersion of the training environment are improved. The present application constructs a virtual-real fusion training environment through real-time interaction between physical devices and virtual models, greatly improving the interactivity and immersion in the teaching and training process. Users can not only obtain knowledge through the question and answer system, but also dynamically learn and operate through the combination of virtual and reality, which provides more vivid and practical support for actual teaching and training scenarios.
[0044] 4. The system's cross-domain and cross-platform adaptability is enhanced through cross-database joint retrieval and the fusion of multi-source teaching data. The system can quickly retrieve relevant knowledge from multiple different databases through cross-database joint retrieval, greatly improving the system's knowledge coverage and flexibility. This allows the system to handle data from different domains or platforms, meeting the knowledge question and answer needs in different domains and scenarios. The fusion of various teaching data sources (such as courseware, cases, images, videos, etc.) into the knowledge graph ensures that the system can handle various forms of knowledge and provide more comprehensive answers. This multi-source data fusion capability enhances the system's versatility in various application scenarios.
[0045] 5. The real-time and scalability of the knowledge question and answer system is improved through fast response and efficient processing. The invention uses quantum heuristic optimization technology to accelerate high-frequency retrieval tasks, greatly improving the system's real-time response capability. Especially in scenarios with large-scale concurrent requests, the system can still maintain high processing speed and low latency, adapting to rapidly changing question and answer needs. Through modular design and distributed optimization based on federated learning, the system can support simultaneous access by a large number of users and achieve efficient resource sharing and processing capacity under multi-node collaboration. The scalability of the system architecture allows it to adapt to the growing needs of future teaching and question and answer tasks, accommodating more data sources and technology upgrades.
[0046] In summary, the invention greatly improves the accuracy, efficiency, flexibility, and adaptability of the knowledge question and answer system, meeting complex and diverse educational and practical training needs. Through the fusion of virtual and real training environments, it provides a more immersive learning experience, while enhancing the system's cross-platform adaptability and efficiency, allowing it to flexibly handle different scenarios and task requirements in practical applications. In addition, the innovative introduction of differential data protection mechanisms and personalized learning functions provides users with safer, smarter, and more personalized knowledge question and answer services.
[0047] These aspects or other aspects of the present application will be more apparent in the following description of the embodiments. It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present application. BRIEF DESCRIPTION OF DRAWINGS
[0048] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the related art, the following will briefly introduce the drawings needed to be used in the exemplary embodiments or the related art description. The drawings are used to provide further understanding of the present application, and constitute a part of the specification. Together with the embodiments of the present application, they are used to explain the present application, and do not constitute a limitation to the present application. In the drawings:
[0049] Figure 1A flow chart of a knowledge question and answer rapid processing method based on artificial intelligence according to an embodiment of the present application.
[0050] Figure 2 A flow chart of a knowledge question and answer rapid processing method based on artificial intelligence according to an embodiment of the present application.
[0051] Figure 3 A structural block diagram of a knowledge question and answer rapid processing system based on artificial intelligence according to an embodiment of the present application. DETAILED DESCRIPTION
[0052] Hereinafter, the present application will be further described in conjunction with the accompanying drawings and specific embodiments. It should be noted that the following described embodiments or technical features can be combined with each other to form new embodiments without conflict.
[0053] In order to make the objectives, technical solutions and advantages of the present application clearer, the embodiments of the present application are further described in detail below with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are only used to explain the present application and should not be used to limit the present application.
[0054] It should be noted that all the expressions of "first" and "second" in the embodiments of the present application are used to distinguish two non-identical entities or non-identical parameters with the same name. It can be seen that "first" and "second" are only used for the convenience of description and should not be understood as a limitation of the embodiments of the present application. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, the process, method, system, product or device inherently includes other steps or units.
[0055] The technical solutions in the embodiments of the present application will be described clearly and completely below in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present application, not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.
[0056] The flow chart shown in the accompanying drawings is only an example and does not necessarily include all the contents and operations / steps, nor does it necessarily be executed in the described order. For example, some operations / steps can be further divided, combined or partially merged, so the actual execution order may be changed according to the actual situation.
[0057] Some embodiments of the present application will be described in detail below in conjunction with the accompanying drawings. The following embodiments and features in the embodiments can be combined with each other without conflict.
[0058] Since the system often needs a long time to search, analyze and generate answers when dealing with complex problems, resulting in poor user experience, the existing system generates answers that are not accurate and comprehensive enough when dealing with problems with ambiguity, fuzziness or the need for comprehensive knowledge in multiple aspects, and the system usually lacks deep understanding of the individual learning history and interests of the user, making it difficult to recommend personalized learning resources and paths. The present application proposes a knowledge question and answer rapid processing method and system based on artificial intelligence, which can improve the professional skill level of teachers and students in the field of application of artificial intelligence and large models, and meet their individual needs in teaching, scientific research and innovation courses.
[0059] Referring to Figure 1 The embodiments of the present application provide a knowledge question and answer rapid processing method based on artificial intelligence, which comprises the following steps:
[0060] Step S10, collect multi-source teaching data including text, image and voice, define graph node features according to the data modalities of the multi-source teaching data, and extract the relationship between nodes from the multi-source teaching data.
[0061] Step S20, train node relationship weight and similarity through graph neural network, update nodes periodically using time decay mechanism, integrate image, text and voice features into unified semantic space using cross-modal alignment, and construct multi-modal knowledge graph.
[0062] Step S30, fuse the collected multi-source teaching data through a hybrid retrieval strategy, and generate a comprehensive retrieval list by weighting and sorting the retrieval results.
[0063] Step S40, perform multi-level question and answer processing based on the RAG enhanced framework, analyze the multi-modal input intent of the multi-source teaching data, perform cross-database joint retrieval on high-frequency questions, integrate the retrieval results of the comprehensive retrieval list and generate optimized answers combined with the teaching scene.
[0064] Step S50, update the global model by federated learning and collaborative multi-node model parameter, distill the global model to the lightweight TinyBERT architecture, dynamically optimize the question and answer quality through the cognitive reinforcement learning framework, locate the key documents from the comprehensive retrieval list, evaluate the optimized answers and key document verification scores to reconstruct the answers.
[0065] In the knowledge question and answer rapid processing method of the present application, the node embedding can be learned through the graph neural network (GNN), and the message passing mechanism is used to update the node features to capture the dependence strength of “chloroplast-photosynthesis”, for example: in the multi-source teaching data collection and node relationship extraction stage of step S10, the following botanical course data can be collected:
[0066] Text: textbook paragraph (such as “photosynthesis depends on chloroplast”);
[0067] Image: Chloroplast structure diagram, photosynthesis process schematic diagram;
[0068] Voice: Teacher's explanation recording ("ATP is produced in the light reaction stage");
[0069] Then when defining the graph node features according to the data modalities of the multi-source teaching data, the defined graph node features are as follows:
[0070] Text node: Key words (chloroplast, light reaction), semantic vector (BERT encoding);
[0071] Image node: Visual features (chloroplast structure vector extracted by CLIP);
[0072] Voice node: Intention encoding after text conversion (Whisper output).
[0073] When extracting the relationship between nodes, the node association is established:
[0074] "Photosynthesis" → composition relationship → "Chloroplast";
[0075] "Light reaction" → time sequence relationship → "Dark reaction".
[0076] In the multi-modal knowledge graph construction phase of step S20, first, cross-modal alignment is performed, and multi-modal Transformer fuses image, text and voice features to output unified semantic vectors; then dynamic updating is performed, and a time decay factor λ=0.8 is used to eliminate old nodes (such as the outdated "C3 plant model") with a weight <0.2 every quarter, for example: a new research is published "C4 plant efficient photosynthesis", a new node is added and associated with "light reaction optimization path", which can solve the data timeliness problem and improve the subsequent retrieval accuracy in the unified semantic space.
[0077] In step S30, when the collected multi-source teaching data is fused through the mixed retrieval strategy, parallel three-modal retrieval can be used, in this embodiment, the mixed retrieval strategy includes semantic retrieval, vector retrieval and metadata retrieval, and the three-modal retrieval is performed in parallel, including:
[0078] Semantic retrieval: Calculate the semantic similarity between the query and the document based on the BERT-Whitening model;
[0079] Vector retrieval: Achieve millisecond-level neighbor search of hundreds of millions of vectors through FAISS index;
[0080] Metadata retrieval: Use ElasticSearch to match the structured information of title, author and keywords.
[0081] The multi-source teaching data collected by the mixed retrieval strategy is fused, and the fusion result is obtained. The three types of retrieval results of semantic retrieval, vector retrieval and metadata retrieval are weighted and sorted according to the preset weight ratio 6:3:1, and a comprehensive retrieval list is generated.
[0082] For example, if the user needs to query: "the key organ of photosynthesis", the following three-mode retrieval is performed:
[0083] Semantic retrieval: BERT-Whitening matches the "chloroplast function" document (similarity 0.92);
[0084] Vector retrieval: FAISS index 1 billion vectors, return the nearest neighbor of chloroplast electron microscope image features;
[0085] Metadata retrieval: ElasticSearch hits the paper with "chloroplast" in the title.
[0086] Then, the three types of retrieval results of semantic retrieval, vector retrieval and metadata retrieval are weighted and sorted according to the preset weight ratio 6:3:1, and the weighted sorting is as follows:
[0087] Weight distribution: semantic result x 0.6 + vector result x 0.3 + metadata x 0.1.
[0088] In step S40, when performing multi-level question answering processing based on the RAG enhanced framework, the RAG multi-level question answering processing and high-frequency optimization are performed. First, the intention analysis is performed, for example: the user uploads the chloroplast picture and asks "What is the function of this structure?" The CLIP output image semantics is "chloroplast" -> Whisper text -> BERT classified as "functional question". The high-frequency question is retrieved across the library, the quantum optimization cross-library retrieval is adopted, and for the question with daily call quantity greater than the preset number of times, the quantum optimization cross-library retrieval model is constructed, and the annealing of the set quantum computer is performed. Output the optimal document set.
[0089] Wherein, the quantum optimization cross-library retrieval model is constructed as follows:
[0090]
[0091] Wherein, is a binary variable, indicating whether the i-th document is selected, wherein 1=selected, 0=not selected; represents the independent weight of the i-th document. represents the association strength of the i-th document with the j-th document. is a positive value indicating synergistic effect, and a negative value indicating repulsion.
[0092] For example, if a user uploads an image describing the process of "photosynthesis" and also asks a voice question: "How does photosynthesis affect plant growth?". First, the CLIP model analyzes the image content and associates it with "photosynthesis"; then, the Whisper model converts the voice into text and extracts the intent keywords; next, the system combines the image and text, and through quantum optimization, it searches for relevant literature and tutorials across databases, and fuses the information of different modalities to generate an optimized answer.
[0093] In step S50, when updating the global model through federated learning and collaborative multi-node model parameters, federated distillation can be used to distill the global model to the TinyBERT architecture by updating the global model through local training and federated averaging as in 10: add δ = 0.01 Gaussian noise to satisfy ϵ = 0.5 differential data protection; when dynamically optimizing the quality of the question and answer through the cognitive reinforcement learning framework, the verification agent BLEU score is calculated, for example: the generated answer vs the key sentence of the authoritative document (score 0.82 > threshold 0.7); if the score < 0.7 (e.g. "chloroplast produces oxygen" is incorrect), the agent reconstructs the answer.
[0094] By federated learning, the global model is updated through collaborative multi-node model parameters, and the global model is distilled to the lightweight TinyBERT architecture. Through the cognitive reinforcement learning framework, the quality of the question and answer is dynamically optimized, the key document is located from the comprehensive retrieval list, and the optimized answer is reconstructed based on the verification score of the key document.
[0095] In this embodiment, the global model is distilled to the lightweight TinyBERT architecture, which includes:
[0096] Based on the local training of GPT-4 architecture model in each university, the global model is updated through federated averaging;
[0097] The global model is distilled to the TinyBERT architecture student model, and Gaussian noise is added for differential data protection during parameter transmission.
[0098] In this embodiment, when training the model through federated learning, each university (or educational institution) trains a GPT-4 architecture model locally. These models are trained based on local teaching data, and the parameters of each node model are combined and updated to the global model through federated averaging algorithm. Each node uploads the trained parameters to the server, and the server performs average update of the global model. When knowledge distillation, the updated global large model (such as GPT-4) is distilled to a lightweight model (such as TinyBERT) to run on resource-limited devices. When transmitting parameters, Gaussian noise is added for differential data protection to ensure user data.
[0099] For example, a teaching platform has multiple universities participating, and each university's teaching data is slightly different. Through federated learning, each university trains a GPT-4 model based on its own data, uploads the parameters of the model, and updates the global model through federated averaging. Subsequently, the global model is converted into a lightweight TinyBERT model through knowledge distillation, which can efficiently run on students' mobile phones, tablets, and other edge devices, while protecting data.
[0100] In this embodiment, referring to Figure 2 A multi-modal knowledge graph is constructed, and the collected multi-source teaching data is fused through a hybrid retrieval strategy, including the following steps:
[0101] Step S101, collect teaching data, define the features of the graph nodes according to the data modalities, and extract the relationships between the nodes from the multi-source data.
[0102] In this embodiment, the graph neural network (GNN) is used to train the knowledge graph, learn the relationship weights and similarities between the nodes in the graph, and the graph neural network can identify complex dependency relationships between nodes and periodically update the knowledge in the nodes according to the time decay mechanism. First, the nodes are trained, the graph is trained through the graph neural network (such as GCN or GAT), and the features of each node and the relationship with other nodes are learned. In the education scenario, this may mean learning which knowledge points are associated with each other (for example, "photosynthesis" has a high degree of association with "plant growth"); then, according to the time decay mechanism, for education-related knowledge, some new teaching resources and theories will be added to the knowledge graph over time. To ensure the timeliness of the knowledge, the graph neural network periodically updates the graph nodes, removes outdated knowledge, and adds new information.
[0103] Step S102, train the multi-modal knowledge graph through the graph neural network, learn the relationship weights and similarities between the nodes, and periodically update the nodes in the graph using the time decay mechanism.
[0104] In this embodiment, information of different modalities (such as text, pictures, and voice) is integrated together and aligned across modalities. Then, through a hybrid retrieval strategy, semantic retrieval, vector retrieval, and metadata retrieval are performed in parallel, and different modalities of information are aligned into a unified representation through image description generation and voice recognition. For example, a picture describing the process of "photosynthesis" may be aligned with a voice explaining photosynthesis and related text materials (such as "photosynthesis is the process by which plants synthesize organic matter through light energy") in the knowledge graph, forming a unified node representation. When a user asks a question, the system will perform three retrievals simultaneously:
[0105] Semantic Retrieval: Calculate semantic similarity between query and documents based on BERT-Whitening model. For example, if a user asks "What is the definition of photosynthesis?", the system will find relevant documents (e.g., "Photosynthesis is...") through semantic retrieval.
[0106] Vector Retrieval: Perform large-scale vector space retrieval using FAISS index to quickly find the most similar vectors. Suppose there are a large number of pictures and videos in the system, these vectors can be image features or keyframe features in videos.
[0107] Metadata Retrieval: Perform retrieval based on title, author, keywords, etc. metadata through ElasticSearch.
[0108] Step S103: Integrate image, text and speech features using cross-modal alignment, and perform semantic retrieval, vector retrieval and metadata retrieval in parallel. For the three retrieval results, weighted ranking is performed according to the preset weight, and a comprehensive retrieval list is generated.
[0109] Weighted ranking of results from different retrieval strategies to generate a comprehensive retrieval list, and removing duplicate content, where each retrieval result (semantic, vector, metadata) is ranked according to the preset weight during weighted ranking. In this example, the semantic retrieval weight is 60%, the vector retrieval is 30%, and the metadata retrieval is 10%. This weighting strategy ensures that results with strong semantic relevance are displayed first. To avoid duplicate retrieval results, the results from multiple sources are de-duplicated to ensure that users do not receive multiple documents or resources with the same content.
[0110] Step S104: Fuse the documents or data obtained from different retrieval strategies, remove duplicate results, and form a comprehensive retrieval result.
[0111] Among them, when constructing a multi-modal knowledge graph and fusing the collected multi-source teaching data through mixed retrieval strategy, first collect teaching data, which can collect teaching resources in various forms such as teaching materials, academic articles, video explanations, pictures and voice. These resources are obtained from different education platforms and databases through API or crawler to establish a preliminary data source. Then define the node features in the graph, where each node represents a knowledge point, concept, term, picture or video, etc. Node features can include text description, image content, voice information, etc. For example, node A may be the knowledge point "photosynthesis", and its features include content descriptions related to "botany", "chemical reactions", etc., pictures (such as process diagrams of photosynthesis), voice explanations, etc.
[0112] Then the relationship between the nodes of the graph is extracted, which can be "belongs to" (such as "cell respiration" belongs to the category of "biology"), "causal relationship" (such as "temperature rise leads to acceleration of photosynthesis") or "similarity" (such as "photosynthesis" and "respiration" are similar concepts). These relationships are extracted from multi-source data and modeled through NLP techniques such as word embedding, ultimately forming a multi-modal knowledge graph.
[0113] In this embodiment, when performing multi-level question answering processing based on the RAG enhanced framework, the multi-modal input intent includes image input, voice input and text input. The visual feature vector is extracted through the CLIP model, mapped to the semantic space for image input; the Whisper model is used for voice-to-text conversion, and the intent keywords are extracted for voice input; the BERT model is used to generate an intent vector, and the classifier outputs the question type for text input.
[0114] In the multi-modal input intent analysis, for image input, the CLIP model is used to extract the visual feature vector of the image and map it to the semantic space. This step can effectively associate image content with text content. For example, if a user uploads a picture of the process of "photosynthesis", the CLIP model will extract the visual features of the picture and associate them with the concepts related to "photosynthesis". For voice input, the Whisper model is used to transcribe the voice input into text. Then, through keyword extraction, the user's intent is identified, such as the voice input of "What is photosynthesis?" After being converted into text, the system can identify the keywords "photosynthesis" and "definition". For text input, the BERT model is used to generate an intent vector from the input text, and then classify the question type, such as factual questions ("What is the definition of photosynthesis?") or application questions ("How to improve the efficiency of photosynthesis?").
[0115] In this embodiment, the quality of the question and answer is dynamically optimized through the cognitive reinforcement learning framework, and the learning framework includes a retrieval agent, a verification agent and a generation agent. The key documents are located from the mixed retrieval results, the consistency of the answer facts is scored, and the answer is reconstructed according to the verification result.
[0116] For example, a user asks "What is the relationship between photosynthesis and plant growth?", the system first searches for relevant literature and materials through the retrieval agent, then checks the consistency of the answer through the verification agent, such as checking whether there are different descriptions or data. Finally, the generation agent generates a comprehensive answer, and adjusts it according to the feedback of the verification agent, so that the final answer is more accurate and comprehensive.
[0117] In this embodiment, the knowledge question and answer rapid processing method of the present application further includes constructing a virtual-real fusion training environment for real-time interaction between physical equipment and virtual models.
[0118] The artificial intelligence-based knowledge question and answer rapid processing method of the present application reproduces the operation and use context of the physical device in the virtual model when interacting with the virtual model and the physical device, such as simulating the growth process of a plant through a virtual simulation platform. The user interacts with the physical device through the virtual environment, simulates the experiment, and obtains real-time feedback. During real-time interaction, the user can interact with the virtual plant model in the virtual environment, adjust different environmental parameters (such as temperature, light, etc.), and see the changes in the plant growth process. By combining these virtual operations with the data feedback of the actual device, an immersive learning experience is provided.
[0119] For example, the user is learning botany and using a virtual training platform to conduct a "photosynthesis experiment". The system displays the photosynthesis process of the plant through the virtual model and allows the user to adjust the temperature, light, and other conditions in the virtual environment to observe the changes in plant growth. At the same time, the physical devices operated by the user (such as temperature sensors, light intensity adjusters, etc.) interact with the virtual environment in real time, providing a more immersive learning experience.
[0120] The present application improves the accuracy and efficiency of knowledge question and answer through the construction of multi-modal knowledge graph, mixed retrieval strategy and RAG enhancement framework. The present application constructs a comprehensive knowledge graph through the fusion of multi-modal data, so that the system can understand and process user input from different modalities, greatly improving the accuracy and adaptability of the system in complex scenarios. The mixed retrieval strategy combining semantic retrieval, vector retrieval and metadata retrieval ensures efficient response to user queries. Through weighted sorting, the system can fully utilize different types of data sources to integrate retrieval results and improve question and answer accuracy. Multi-level question and answer processing based on the RAG framework can optimize query processing at different levels, ensure accurate analysis and generation of cross-modal input intent, and provide more accurate answers.
[0121] It should be noted that the above figures are only schematic illustrations of the processes included in the method according to the exemplary embodiments of the present application, and are not for limiting purposes. It is easy to understand that the processes shown in the above figures do not indicate or limit the time sequence of these processes. In addition, it is also easy to understand that these processes can be executed synchronously or asynchronously, for example, in multiple modules.
[0122] It should be understood that although the above steps are described in a certain order, these steps are not necessarily executed in the above order. Unless explicitly stated herein, the execution of these steps is not strictly limited in order, and these steps can be executed in other orders. Moreover, part of the steps of the present embodiment can include multiple steps or multiple stages, which are not necessarily executed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily sequential, but can be executed alternately or alternately with at least part of other steps or steps or stages in other steps.
[0123] In a second aspect of the embodiments of the present application, referring to Figure 3 The present application also provides an artificial intelligence-based knowledge question and answer rapid processing system, which comprises:
[0124] The data acquisition module 100 is configured to collect multi-source teaching data including text, images and voice, define graph node features according to the data modalities of the multi-source teaching data, and extract inter-node relationships from the multi-source teaching data.
[0125] The multi-modal knowledge graph construction module 200 is configured to train node relationship weights and similarities through a graph neural network, update nodes periodically using a time decay mechanism, integrate image, text and voice features into a unified semantic space using cross-modal alignment, and construct a multi-modal knowledge graph.
[0126] The hybrid retrieval module 300 is configured to fuse the collected multi-source teaching data through a hybrid retrieval strategy, weight and sort the retrieval results to generate a comprehensive retrieval list.
[0127] The RAG enhanced framework module 400 is configured to perform multi-level question and answer processing based on a RAG enhanced framework, analyze multi-modal input intentions of the multi-source teaching data, perform cross-database joint retrieval on high-frequency questions, integrate retrieval results of the comprehensive retrieval list and generate optimized answers in combination with teaching scenarios.
[0128] The dynamic knowledge distillation module 500 is configured to update a global model through federated learning and collaborative multi-node model parameter updating, and distill the global model to a lightweight TinyBERT architecture.
[0129] The question and answer optimization verification module 600 is configured to dynamically optimize question and answer quality through a cognitive reinforcement learning framework, locate key documents from the comprehensive retrieval list, evaluate optimized answers and key document verification scores to reconstruct answers.
[0130] Through the above detailed steps, the artificial intelligence-based knowledge question and answer rapid processing system of the present application is used to execute the steps of the artificial intelligence-based knowledge question and answer rapid processing method in the above embodiments, which will not be repeated here.
[0131] The present application employs quantum heuristic optimization technology to accelerate high-frequency retrieval tasks, greatly improving the real-time response capability of the system. Especially in the scene of large-scale concurrent requests, the system can still maintain high processing speed and low delay, adapting to the rapidly changing question and answer demand. Through modular design and distributed optimization based on federated learning, the system can support simultaneous access of large-scale users and realize efficient resource sharing and processing capacity under multi-node cooperation. The scalability of the system architecture enables it to adapt to the growing demand of future teaching and question and answer tasks, and adapt to more data sources and technology upgrades.
[0132] The artificial intelligence-based knowledge question and answer rapid processing system of the present application greatly improves the accuracy, efficiency, flexibility and adaptability of the knowledge question and answer system, and can meet the complex and diverse education and training needs. Through the virtual-real integrated training environment, a more immersive learning experience is provided, while the cross-platform adaptability and efficiency of the system are enhanced, so that it can flexibly cope with different scenes and task requirements in actual application. In addition, the differential protection mechanism and personalized learning function are innovatively introduced, providing users with safer, smarter and personalized knowledge question and answer services.
[0133] The third aspect of the embodiment of the present application also provides a computer device comprising a memory and a processor, the memory storing a computer program, and the computer program is executed by the processor to realize the method of any one of the above embodiments.
[0134] The computer device includes a processor and a memory, and can also include an input system and an output system. The processor, memory, input system and output system can be connected through a bus or other means. The input system can receive input digital or character information and generate signal input related to the migration of the artificial intelligence-based knowledge question and answer rapid processing. The output system can include a display device such as a display screen.
[0135] The memory, as a non-volatile computer readable storage medium, can be used to store non-volatile software programs, non-volatile computer executable programs and modules, such as program instructions / modules corresponding to the knowledge question and answer rapid processing method based on artificial intelligence in the embodiments of the present application. The memory can include a program storage area and a data storage area, wherein the program storage area can store an operating system and application programs required by at least one function; the data storage area can store data created by use of the knowledge question and answer rapid processing method based on artificial intelligence, etc. In addition, the memory can include a high-speed random access memory, and can also include a non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage device. In some embodiments, the memory can optionally include a memory remotely arranged with respect to the processor, and these remote memories can be connected to the local module through a network. Examples of the above network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network, and a combination thereof.
[0136] The processor in some embodiments can be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chip. The processor is generally used to control the overall operation of the computer device. In the present embodiment, the processor is used to run the program code stored in the memory or process data. The processors of the plurality of computer devices of the computer device in the present embodiment execute various function applications and data processing of the server by running the non-volatile software programs, instructions and modules stored in the memory, that is, implement the steps of the knowledge question and answer rapid processing method based on artificial intelligence of the above-mentioned method embodiments.
[0137] It should be understood that all the embodiments, features and advantages described above for the knowledge question and answer rapid processing method based on artificial intelligence according to the present application are equally applicable to the knowledge question and answer rapid processing method based on artificial intelligence and the storage medium according to the present application without mutual conflict.
[0138] Those skilled in the art will also appreciate that the various illustrative logical blocks, modules, circuits, and algorithm steps described in connection with the disclosure herein can be implemented as electronic hardware, computer software, or combinations of both. To clearly illustrate this interchangeability of hardware and software, various illustrative components, blocks, modules, circuits, and steps have been described above generally in terms of their functionality. Whether such functionality is implemented as hardware or software depends on the particular application and design constraints imposed on the overall system. Skilled persons can implement the functions described in various ways for each particular application, but such implementation decisions should not be interpreted as causing a departure from the scope of the present embodiments.
[0139] Finally, it is noted that the computer-readable storage media of the present disclosure (e.g., the memory) can be volatile or nonvolatile storage. By way of example, and not limitation, nonvolatile memory can include read only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM), which can act as external cache memory. By way of example, and not limitation, RAM is available in many forms such as Static RAM (SRAM), Dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), Synchlink DRAM (SLDRAM), and direct Rambus RAM (DRRAM). The memory of the disclosed aspects is intended to include, without being limited to, these and any other suitable types of memory.
[0140] The various illustrative logical blocks, modules, and circuits described in connection with the disclosure herein can be implemented or performed with a general purpose processor, a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field Programmable Gate Array (FPGA) or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or any combination thereof designed to perform the functions described herein. A general purpose processor can be a microprocessor, but, in the alternative, the processor can be any conventional processor, controller, microcontroller, or state machine. A processor can also be implemented as a combination of computing devices, e.g., a combination of a DSP and a microprocessor, a plurality of microprocessors, one or more microprocessors in conjunction with a DSP core, or any other such configuration.
[0141] The above are exemplary embodiments of the present disclosure, but it should be noted that various changes and modifications can be made without departing from the scope of the embodiments disclosed by the present disclosure. The functions, steps and / or actions of the method steps described in the embodiments disclosed herein need not be performed in any particular order. Furthermore, although elements of the embodiments disclosed by the present disclosure can be described or claimed in individual forms, they can also be understood as plural unless explicitly limited to a single.
[0142] It should be understood that, as used herein, the singular forms "a", "an" and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that the terms "comprises", "comprising", "includes" and / or "including", as used herein, specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.
[0143] Those skilled in the art should understand that the above discussion of any embodiment is only exemplary, and is not intended to mean that the scope of the embodiments disclosed by the present application is limited to these examples; the technical features in the above embodiments or different embodiments can also be combined, and there are many other changes of different aspects of the embodiments of the present application as above, which are not provided in details for the sake of brevity. Therefore, any omission, modification, equivalent replacement, improvement, etc. made within the spirit and principle of the embodiments of the present application shall be included in the protection scope of the embodiments of the present application.
Claims
1. A rapid knowledge question-answering processing method based on artificial intelligence, characterized in that, The method includes the following steps: Collect multi-source teaching data including text, images and voice, define graph node features according to the data modalities of the multi-source teaching data, and extract the relationships between nodes from the multi-source teaching data; By training node relationship weights and similarities through graph neural networks, updating nodes periodically using a time decay mechanism, and integrating image, text, and speech features into a unified semantic space through cross-modal alignment, a multimodal knowledge graph is constructed. By integrating multi-source teaching data collected through a hybrid retrieval strategy, a comprehensive retrieval list is generated by weighting and sorting the retrieval results. Based on the RAG enhancement framework, multi-level question answering processing is performed, multi-modal input intent of multi-source teaching data is parsed, cross-database joint retrieval is carried out for high-frequency questions, and the retrieval results of the comprehensive retrieval list are integrated and optimized answers are generated in combination with teaching scenarios. By using federated learning to collaboratively update the parameters of multi-node models, the global model is distilled into a lightweight TinyBERT architecture. The question-answering quality is dynamically optimized through a cognitive reinforcement learning framework. Key documents are located from the comprehensive search list, and the optimized answer is evaluated and the key document validation score is used to reconstruct the answer.
2. The rapid knowledge question answering method based on artificial intelligence as described in claim 1, characterized in that, When integrating multi-source teaching data collected through a hybrid retrieval strategy, the hybrid retrieval strategy includes semantic retrieval, vector retrieval, and metadata retrieval, and executes the three-modal retrieval in parallel, including: Semantic retrieval: Calculates the semantic similarity between the query and the document based on the BERT-Whitening model; Vector retrieval: Achieve millisecond-level nearest neighbor search for hundreds of millions of vectors using the FAISS index; Metadata retrieval: Use ElasticSearch to match structured information such as title, author, and keywords.
3. The rapid knowledge question answering method based on artificial intelligence as described in claim 2, characterized in that, By integrating multi-source teaching data collected through a hybrid retrieval strategy, the fusion results are weighted and sorted according to a preset weight ratio of 6:3:1 for semantic retrieval, vector retrieval, and metadata retrieval, generating a comprehensive retrieval list.
4. The rapid knowledge question answering method based on artificial intelligence as described in claim 3, characterized in that, Constructing a multimodal knowledge graph and fusing collected multi-source teaching data through a hybrid retrieval strategy includes the following steps: Collect teaching data, define the characteristics of graph nodes according to data modalities, and extract the relationships between nodes from multi-source data; The multimodal knowledge graph is trained using a graph neural network to learn the relationship weights and similarities between nodes, and the nodes in the graph are updated periodically using a time decay mechanism. Cross-modal alignment is used to integrate features of images, text and speech, and semantic retrieval, vector retrieval and metadata retrieval are performed in parallel. The three retrieval results are weighted and sorted according to preset weights to generate a comprehensive retrieval list. Documents or data obtained from different search strategies are merged, duplicate results are removed, and a comprehensive search result is formed.
5. The rapid knowledge question answering method based on artificial intelligence as described in claim 1, characterized in that, When performing multi-level question answering based on the RAG enhancement framework, the multimodal input intent includes image input, voice input, and text input. Specifically, the CLIP model extracts visual feature vectors and maps them to the semantic space for image input; the Whisper model is used for speech-to-text conversion and extracts intent keywords for voice input; and the BERT model is used to generate intent vectors, and the classifier outputs the question type for text input.
6. The rapid knowledge question answering method based on artificial intelligence as described in claim 5, characterized in that, During cross-database joint retrieval, quantum-optimized cross-database retrieval is employed. For problems with daily call volumes exceeding a preset number, a quantum-optimized cross-database retrieval model is constructed, and annealing is performed on a designated quantum computer to output the optimal document set. Specifically, quantum-optimized cross-database retrieval is designed for high-frequency problems. During cross-database joint retrieval, for high-frequency problems with daily call volumes exceeding 1000, a quantum-optimized model is constructed, and the optimal document set is output through quantum computer simulation of the annealing process. The quantum annealing algorithm filters the optimal set from hundreds of millions of documents, improving speed by more than 40% compared to traditional brute-force retrieval. The core difference lies in the increased document association strength. and independent weights The spin interaction of qubits is mapped to accelerate the search for the optimal solution through the quantum tunneling effect.
7. The rapid knowledge question answering method based on artificial intelligence as described in claim 6, characterized in that, A quantum-optimized cross-database retrieval model is constructed, which transforms the document selection problem into a quadratic unconstrained binary optimization problem. The objective function of the quantum-optimized cross-database retrieval model is: ; in, , represents a binary variable, indicating whether to select the first . This document contains 1 selected documents and 0 unselected documents. Document With Documents The correlation strength is positive when it indicates synergy and negative when it indicates repulsion. Document The independent weight is calculated by weighting metadata such as click-through rate and authority.
8. The rapid knowledge question answering method based on artificial intelligence as described in claim 1, characterized in that, Distilling the global model into the lightweight TinyBERT architecture includes: The global model is updated by federated averaging, based on the local training of GPT-4 architecture models in various universities. The global model is distilled into the TinyBERT architecture student model, and Gaussian noise is added during parameter transmission to protect the differential data.
9. The rapid knowledge question answering method based on artificial intelligence as described in claim 8, characterized in that, The cognitive reinforcement learning framework dynamically optimizes the quality of question answering. The learning framework includes a retrieval agent, a verification agent, and a generation agent. It locates key documents from mixed retrieval results, scores the factual consistency of answers, and reconstructs answers based on verification results.
10. A knowledge question-answering rapid processing system based on artificial intelligence, characterized in that, The system is used to perform the AI-based rapid knowledge question answering method as described in any one of claims 1-9, and includes: The data acquisition module is used to collect multi-source teaching data including text, images and voice, define graph node features according to the data modalities of the multi-source teaching data, and extract the relationships between nodes from the multi-source teaching data. The multimodal knowledge graph construction module is used to train node relationship weights and similarities through graph neural networks, update nodes periodically using a time decay mechanism, and integrate image, text, and speech features into a unified semantic space using cross-modal alignment to construct a multimodal knowledge graph. The hybrid retrieval module is used to integrate multi-source teaching data collected through a hybrid retrieval strategy, and generate a comprehensive retrieval list by weighting and sorting the retrieval results. The RAG Enhancement Framework module is used to perform multi-level question-answering processing based on the RAG Enhancement Framework, parse the multimodal input intent of multi-source teaching data, perform cross-database joint retrieval of high-frequency questions, integrate the retrieval results of the comprehensive retrieval list and generate optimized answers in combination with teaching scenarios; The Dynamic Knowledge Distillation module is used to update the global model through federated learning in collaboration with multi-node model parameters, and distill the global model into the lightweight TinyBERT architecture. The question-answering optimization and verification module is used to dynamically optimize the quality of question-answering through a cognitive reinforcement learning framework. It locates key documents from the comprehensive search list, evaluates the optimized answer and the key document verification score, and reconstructs the answer.
Citation Information
Patent Citations
Large language model knowledge question-answering method and system fused with multi-modal knowledge graph
CN118627628A
Intelligent question answering method, device and equipment based on combination of large model and knowledge graph
CN119807386A