Database generation method, question and answer method and system and computer equipment
By extracting and parsing data from multimodal information and determining contextual relationships, a variety of database structures are generated, which solves the problem of low accuracy of the generated model and improves the accuracy of the question-answering model and the flexibility of data storage.
Patent Information
- Application Number
- CN202510681419.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-24
- Publication Date
- 2025-09-05
AI Technical Summary
The accuracy of generated content by existing generative models is low, which limits their application.
By extracting multimodal parsing data from multimodal information, determining the contextual relationship between different parsing data, generating multiple databases with different knowledge data structures, and utilizing these databases in the question-answering model to improve content accuracy and flexibility.
It enhances the accuracy of answers in question-answering models and the flexibility of data storage, supports efficient management and analysis of large-scale data, and improves the accuracy of generated content and the scalability of the database.
Smart Images

Figure CN120596578A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and in particular to a database generation method, a question-answering method and system, and a computer device. Background Art
[0002] Currently, knowledge bases can be used to improve the content generation capabilities of large models. However, the accuracy of the generated content by the generative models is relatively low, which limits the application of the generative models. Summary of the Invention
[0003] The purpose of the embodiments of the present application is to provide a database generation method, a question-answering method and system, and a computer device to improve the accuracy of content generated by the retrieval enhancement generation method.
[0004] In a first aspect, an embodiment of the present application provides a method for generating a database, comprising:
[0005] In response to the received multimodal information, extracting multimodal parsed data from the multimodal information, the multimodal parsed data including a plurality of parsed data;
[0006] Based on multimodal parsed data, determine the contextual relationship between different parsed data;
[0007] Based on the multimodal parsing data and the contextual relationships between different parsing data, multiple databases are generated, and the data structures of the knowledge data included in each database are different.
[0008] In the technical solution of the embodiment of the present application, multimodal parsing data can be extracted from multimodal information, and the contextual relationship between different parsing data can be determined based on the multimodal parsing data. At this time, the knowledge data included in the generated multiple databases based on the contextual relationship between the multimodal parsing data and the different parsing data are relatively complete and accurate. In addition, the data structures of the knowledge data included in each database are different. Therefore, when implementing a question-and-answer solution based on multiple databases, even for the same user question, different retrieval results can be recalled from different databases. In this case, the retrieval structures recalled from different databases can provide more sufficient reference information for the question-and-answer model, thereby ensuring the accuracy of the content generated by the retrieval enhancement generation method. Moreover, the embodiment of the present application can store knowledge data with different data structures in different databases, which can enhance the flexibility and scalability of data storage, better support the efficient management and analysis of large-scale data, and enhance the reliability and performance of the question-and-answer model.
[0009] In addition, since the multimodal analysis data extracted from the multimodal information makes the data extracted from the multimodal information more comprehensive, the embodiment of the present application is based on the contextual relationship between the multimodal analysis data and different analysis data. The knowledge data in the multiple databases generated is cross-modal knowledge data, and its content is relatively complete, which can fully improve the answer accuracy of the question-answering model.
[0010] In one possible implementation, the multimodal information includes multiple target files. The multiple target files include a first target file. The first target file may refer to any one of the multiple target files. In this case, extracting multimodal parsed data from the multimodal information includes:
[0011] When it is detected that the type of the first target file is a single-modal file, single-modal data parsing is performed on the first target file to obtain parsed data of the single-modal file; when it is detected that the type of the first target file is a multi-modal file, multi-modal data detection is performed on the first target file to obtain parsed data of at least one modality.
[0012] When extracting multimodal parsed data, the embodiment of the present application can select different data parsing methods according to the type of the target file to completely extract the data in the target file, thereby reducing the possibility of data omission in the target file.
[0013] In one possible implementation, performing multimodal data detection on the first target file to obtain parsed data of at least one modality includes:
[0014] Modality recognition is performed on the file data contained in the first target file to obtain at least one data modality, and data detection is performed on the file data contained in the first target file based on the at least one data modality to obtain parsed data of the at least one modality.
[0015] The embodiment of the present application performs data detection on the file data contained in the first target file through the modality of the file data of the first target file, so that the parsed data of each modality contained in the first target file can be extracted, thereby ensuring the integrity of the knowledge data contained in the various generated databases.
[0016] In a possible implementation, when the modalities of at least two parsed data in a plurality of parsed data are different, the contextual relationship between the different parsed data includes both the contextual relationship between the different parsed data included in the same target file and the contextual relationship between the parsed data included in different target files. It can be seen that the embodiment of the present application determines the contextual relationship between different parsed data based on multimodal parsed data, and can not only establish the contextual relationship between different parsed data of the same file, but also establish the contextual relationship between different parsed data across files. Therefore, the embodiment of the present application uses multiple target files as data sources, and can fully explore the relationship between the parsed data of the same target file and different target files, and multiply more contextual relationships between different parsed data, thereby improving the data diversity of the knowledge data contained in different types of databases, so as to further improve the answer accuracy of the question-answering model.
[0017] In one possible implementation, multiple databases are generated based on multimodal parsed data and contextual relationships between different parsed data, including:
[0018] The multimodal parsing data is segmented into multiple parsed data to obtain multimodal data segments. The multimodal data segments include multiple data segments, which is equivalent to segmenting the longer parsed data so that each longer parsed data becomes multiple shorter data segments. Therefore, based on the contextual relationship between the multimodal data segments and different parsed data, multiple databases are generated. This not only ensures the comprehensiveness of the knowledge data included in each database, but also improves the detailedness of the content of each database, so that the accuracy of the question-answering model can be enhanced based on multiple databases.
[0019] In a second aspect, the embodiments of the present application further provide a question-and-answer method, including:
[0020] In response to a user inquiry request from a user terminal, determining a user question corresponding to the user inquiry request;
[0021] The user question is input into the target question-answering model to obtain the answer to the user question. The target question-answering model is used to obtain the retrieval results of the user question in at least one database, and based on the retrieval results of the user question in at least one database, the answer to the user question is determined. The data structures of the knowledge data included in the various databases are different.
[0022] When the technical solution of the embodiment of the present application is adopted, the data structures of the knowledge data included in various databases are different. Therefore, when implementing a question-and-answer solution based on multiple databases, even for the same user question, different retrieval results can be recalled from different databases. In this case, the retrieval structures recalled by different databases can provide more sufficient reference information for the question-and-answer model, thereby ensuring the accuracy of the content generated by the retrieval enhancement generation method. Moreover, the embodiment of the present application can store knowledge data of different data structures in different databases, which can enhance the flexibility and scalability of data storage, better support the efficient management and analysis of large-scale data, and enhance the reliability and performance of the question-and-answer model. In addition, since the knowledge data included in the same database is cross-modal knowledge data, the knowledge data of the database is relatively complete, thereby fully improving the accuracy of the answers of the question-and-answer model.
[0023] In one possible implementation, the method further includes: sending the answer to the question to the user terminal, with the display interface of the user terminal being used to display the answer; responding to an evaluation result of the answer from the user terminal, determining the satisfaction level of the answer based on the evaluation result; and if the satisfaction level of the answer is lower than a preset satisfaction level, updating the accuracy parameters and retrieval parameters of the target question-answering model. When the question-answering model is used to re-answer the user's question, the accuracy of the obtained answer can be improved.
[0024] In one possible implementation, the method further includes: if the answer satisfaction level is lower than a preset satisfaction level, sending a message to the user terminal suggesting a knowledge update; and in response to the user terminal confirming the suggestion message, updating the knowledge data in the multiple databases so that each database contains knowledge data relevant to the user's question. This approach ensures the real-time availability of knowledge data in the databases, meeting user needs.
[0025] In a possible implementation, the method further includes:
[0026] In response to a question-answering model selection instruction from a user terminal, a target question-answering model is selected from multiple candidate question-answering models. In this way, the user can select a target question-answering model to answer the user's question according to actual needs to improve the accuracy of the answer to the question obtained.
[0027] In a third aspect, an embodiment of the present application provides an electronic device, comprising: a processor and a memory storing a program. The program includes instructions that, when executed by the processor, cause the processor to perform the method according to the first aspect of the embodiment of the present application or any possible implementation of the first aspect.
[0028] In a fourth aspect, an embodiment of the present application further provides a computer storage medium, which stores computer instructions. When the program instructions are run on an electronic device, the processor of the electronic device executes the method described in the first aspect of the embodiment of the present application or any possible implementation method of the first aspect.
[0029] In a fifth aspect, an embodiment of the present application further provides a computer program product, comprising a computer program, wherein the computer program, when executed by a processor, implements the method according to the first aspect or any possible implementation manner of the first aspect.
[0030] In a sixth aspect, an embodiment of the present application provides an electronic device, comprising: a processor and a memory storing a program. The program includes instructions that, when executed by the processor, cause the processor to perform the method according to the second aspect of the embodiment of the present application or any possible implementation of the second aspect.
[0031] In the seventh aspect, an embodiment of the present application also provides a computer storage medium, which stores computer instructions. When the program instructions are run on an electronic device, the processor of the electronic device executes the method described in the second aspect of the embodiment of the present application or any possible implementation method of the second aspect.
[0032] In an eighth aspect, an embodiment of the present application further provides a computer program product, comprising a computer program, wherein the computer program, when executed by a processor, implements the method described in the second aspect or any possible implementation manner of the second aspect.
[0033] In a ninth aspect, an embodiment of the present application further provides a question-answering system, comprising: a server and a storage device communicatively connected to the server; wherein,
[0034] The storage device is used to store at least multiple databases generated by the method described in the first aspect of the embodiment of the present application or any possible implementation manner of the first aspect;
[0035] The server is used to execute the method described in the second aspect of the embodiment of the present application or any possible implementation of the second aspect.
[0036] The beneficial effects brought about by the third to ninth aspects of the embodiments of the present application can refer to the beneficial effects of the method described in the first aspect of the embodiments of the present application or any possible implementation method of the first aspect, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] Further details, features and advantages of the present application are claimed in the following description of exemplary embodiments in conjunction with the accompanying drawings, in which:
[0038] Figure 1A schematic diagram illustrating an example system in which the various methods described herein may be implemented according to an exemplary embodiment of the present application;
[0039] Figure 2 A schematic diagram showing a framework of a RAG service in which various methods described herein may be implemented according to an exemplary embodiment of the present application;
[0040] Figure 3 A schematic diagram showing a flow chart of a method for generating a database according to an embodiment of the present application is shown;
[0041] Figure 4 A schematic diagram showing a flow chart of a question-answering method according to an embodiment of the present application is shown;
[0042] Figure 5A A schematic diagram showing an interface display interface of a user terminal according to an embodiment of the present application is shown;
[0043] Figure 5B Another schematic diagram of the display interface of the user terminal according to an embodiment of the present application is shown;
[0044] Figure 6 A schematic block diagram of a device for generating a database according to an exemplary embodiment of the present application is shown;
[0045] Figure 7 shows a schematic block diagram of a question-answering device according to an exemplary embodiment of the present application;
[0046] Figure 8 shows a schematic block diagram of a chip according to an exemplary embodiment of the present application;
[0047] Figure 9 A structural block diagram of an exemplary electronic device that can be used to implement the embodiments of the present application is shown. DETAILED DESCRIPTION
[0048] The following describes embodiments of the present application in more detail with reference to the accompanying drawings. Although certain embodiments of the present application are shown in the accompanying drawings, it should be understood that the present application can be implemented in various forms and should not be construed as limited to the embodiments described herein. Instead, these embodiments are provided to provide a more thorough and complete understanding of the present application. It should be understood that the drawings and embodiments of the present application are for illustrative purposes only and are not intended to limit the scope of protection of the present application.
[0049] It should be understood that the various steps described in the method embodiments of the present application can be performed in different orders and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present application is not limited in this respect.
[0050] The term "including" and its variations used in this document are open inclusions, that is, "including but not limited to". The term "based on" means "based at least in part on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one other embodiment"; the term "some embodiments" means "at least some embodiments". The relevant definitions of other terms will be given in the description below. It should be noted that the concepts of "first", "second", etc. mentioned in this application are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.
[0051] It should be noted that the modifications of "one" and "multiple" mentioned in this application are illustrative rather than restrictive. Those skilled in the art should understand that unless otherwise clearly indicated in the context, they should be understood as "one or more".
[0052] Before introducing the embodiments of the present application, the following definitions are given for the relevant terms involved in the embodiments of the present application:
[0053] Multimodality refers to the multiple forms of data or information. The modality of data refers to the source, form, and mode of expression or perception of information. For example, human senses such as touch, hearing, vision, and smell, as well as different information media (voice, video, text) and sensors (radar, infrared, accelerometers) can all be referred to as modalities.
[0054] Retrieval-Augmented Generation (RAG) is a natural language processing technology that combines retrieval and generation. It enhances the ability of the generation model by introducing external knowledge, making the generated content more accurate, rich, and factual.
[0055] Knowledge-Augmented Generation (KAG) is a method that enhances generative models (such as the Generative Pre-trained Transformer (GPT) and the Text-to-Text Transfer Transformer (T5) model) by combining them with external knowledge bases (such as knowledge graphs and databases). Unlike traditional generative models, KAG incorporates external knowledge into the generation process, making the generated content more accurate and reliable, and better able to answer complex questions or generate fact-based content.
[0056] Contextual Retrieval-Augmented Generation (CRAG) is an extension of RAG. It further incorporates contextual information based on retrieval-augmented generation. CRAG not only relies on an external knowledge base but also dynamically adjusts its retrieval strategy and generated content based on the current context (such as conversation history and user input context).
[0057] Self-Refine Assisted Generation (Self-RAG) is a framework designed to improve the quality and accuracy of large language models by using on-demand retrieval and self-reflection mechanisms. In contrast to traditional retrieval-augmented generation methods, Self-RAG retrieves information from the knowledge base on demand, meaning it can retrieve multiple times or even not at all, depending on the query it encounters.
[0058] The Contrastive Language-Image Pre-Training (CLIP) model is a multimodal pre-training neural network. It performs image-text joint learning through contrastive learning, mapping images and text into a shared vector space. This enables unsupervised joint learning and is applicable to a variety of vision and language tasks.
[0059] MinIO Object Storage is a high-performance object storage service suitable for storing large amounts of unstructured data, such as images, videos, log files, etc. MinIO can be deployed in a variety of environments, including containers, physical servers, and virtual machines.
[0060] The MinerU multi-model framework is a powerful open source PDF, Word, and PPT data extraction tool. It is particularly capable of converting complex multimodal PDF / PPT documents into Markdown / JSON structured data formats. When complex content such as photocopied text, mixed text and images, mathematical formulas, tables, footnotes, etc. appears in the document, the MinerU multi-model framework can accurately identify them, and the extracted content retains the original text hierarchy, ensuring content coherence, and greatly improving the efficiency of AI corpus collection.
[0061] The embodiments of the present application provide a database generation method, a question-answering method and system, and a computer device to improve the versatility of the intelligent agent in actively triggering tasks, thereby increasing the possibility of the intelligent agent being applied in products that actively trigger tasks.
[0062] Figure 1 Schematic diagram of an example system in which the various methods described herein may be implemented according to an exemplary embodiment of the present application. Figure 1As shown, the system of the embodiment of the present application may include a user terminal 101 and a question-answering system 102. The user terminal 101 and the question-answering system 102 may be connected via a network.
[0063] like Figure 1 As shown, the user terminal 101 in the embodiment of the present application can be a terminal with a display function. The terminal can be a mobile phone, a tablet computer, a wearable device, an in-vehicle device, a laptop computer, an ultra-mobile personal computer (UMPC), a netbook, a personal digital assistant (PDA), and a wearable device based on augmented reality (AR) and / or virtual reality (VR) technology.
[0064] For example, when the terminal is a wearable device, the wearable device can also be a general term for wearable devices that are intelligently designed and developed by applying wearable technology to daily wear, such as glasses, gloves, watches, clothing and shoes. A wearable device is a portable device that is worn directly on the body or integrated into the user's clothes or accessories. Wearable devices are not only hardware devices, but also achieve powerful functions through software support, data interaction, and cloud interaction. Broadly speaking, wearable smart devices include those that are full-featured, large in size, and can achieve complete or partial functions without relying on smartphones, such as smart watches or smart glasses, as well as those that only focus on a certain type of application function and need to be used in conjunction with other devices such as smartphones, such as various smart bracelets and smart jewelry for vital sign monitoring.
[0065] like Figure 1 As shown, in an embodiment of the present application, the user terminal 101 can send a user inquiry request to the question-answering system 102 in response to the user's question input operation. The question-answering system 102 can parse the user inquiry request, obtain the user question, and return the answer to the user question to the user terminal 101 based on the target question-answering model and in combination with at least one database, so that the user terminal 101 displays the answer to the question. The data structures of the knowledge data of various databases are different here, and the knowledge data included in the same database is cross-modal knowledge data, that is, the database includes both homomodal knowledge data and cross-modal knowledge data, thereby ensuring the comprehensiveness and integrity of the knowledge base.
[0066] Optional, such as Figure 1As shown, the question-answering system 102 is deployed with multiple candidate question-answering models, which can be retrieval-enhanced generation models such as RAG, CRAG, KAG, and Self-RAG. In response to a question-answering model selection instruction from the management terminal, the question-answering system 102 can select a question-answering model from the multiple candidate question-answering models as the target question-answering model.
[0067] For example, when the management terminal and the user terminal are the same terminal, the user terminal has the authority to manage the question-and-answer system. When the management terminal and the user terminal are different terminals, the user terminal may not have the authority to manage the question-and-answer system, but only has the authority to use the management system.
[0068] Optionally, the answer to the question may include, but is not limited to, at least one of text, images, and audio. Accordingly, the answer to the question may be presented in a variety of ways, including, but not limited to, visual or auditory presentation. For example, if the user terminal has a display screen, the user terminal may use the display screen as an interactive medium to receive a question input by the user and also display the answer to the question on the display screen.
[0069] In an alternative approach, Figure 1 As shown, after displaying the answer to the question, the user terminal 101 can also determine the evaluation result of the answer to the question in response to the user's evaluation operation on the answer to the question, and send the evaluation result of the answer to the question to the question and answer system 102.
[0070] Optionally, after receiving the evaluation result of the question answer, the question system can update the accuracy parameter and retrieval relevance of the target question answering model when the evaluation result is not good. For example, the accuracy parameter can be top_k and the retrieval parameter can be minimum relevance. When the evaluation result is not good, the minimum relevance can be lowered and top_k can be increased. In this way, the accuracy of the question answer obtained by the question answering model based on the user's question will be higher, and the question answer will be sent to Figure 1 The user terminal 101 shown can effectively improve user satisfaction.
[0071] Optional, such as Figure 1 As shown, after receiving the evaluation result of the answer to the question, the question system can send a knowledge update initiation suggestion message to the user terminal 101 when the evaluation result is not good, so that the user terminal 101 can display the knowledge update initiation suggestion message, and the knowledge update initiation suggestion message can be used to prompt the user to initiate a knowledge update request.
[0072] like Figure 1As shown, in response to the user's confirmation operation for initiating a suggestion message, user terminal 101 sends a confirmation indication of the suggestion message initiation to question-and-answer system 102. In response to the confirmation indication of the suggestion message initiation from user terminal 101, question-and-answer system 102 can update the knowledge data in multiple databases so that each database contains knowledge data related to the user's question. This ensures the integrity and real-time nature of the knowledge data in the databases. As a result, the accuracy of the answers obtained by the question-and-answer model based on the user's question is relatively high. Sending these answers to user terminal 101 can effectively improve user satisfaction.
[0073] In an alternative approach, Figure 1 As shown, the question-answering system 102 of the embodiment of the present application includes: a server 1021 and a storage device 1022 communicatively connected to the server 1021. For example, the server 1021 can be connected to the storage device 1022 via a network.
[0074] like Figure 1 As shown, the storage device 1022 can store multiple databases. Optionally, the knowledge data contained in the multiple databases stored in the storage device 1022 all come from multimodal information. Therefore, the storage device 1022 can also store multimodal information to facilitate query and download. The multimodal information here can include multiple target files of various types.
[0075] Optionally, the target file format can be used to preliminarily determine whether the target file is a unimodal file or a multimodal file. The target file format may include, but is not limited to, doc, pdf, table, txt, and image formats (e.g., jpg, png, etc.). In terms of the source of the target file, the target file may be, but is not limited to, an invoice, a contract, a report, a bill, etc.
[0076] like Figure 1 As shown, the server 1021 can be used to execute the question-answering method. Optionally, the server 1021 can deploy multiple question-answering models, receive a user query request from a user device, determine the user question corresponding to the user query request, and then use one or more question-answering models and multiple databases to determine the answer to the user question, and finally send the answer to the user device to display the answer on a display interface of the user device.
[0077] In an alternative approach, Figure 1As shown, the server 1021 of the embodiment of the present application can also execute a database generation method. Optionally, the server 1021 can perform multimodal analysis on multimodal information and perform data fusion in the form of different data structures, thereby obtaining knowledge data of multiple data structures, and storing the knowledge data of each data structure in the corresponding database in the storage device 1022.
[0078] In an alternative approach, Figure 1 As shown, the server 1021 can be a cloud server, and its deployment mode can be stand-alone deployment, cluster deployment, or distributed deployment. The storage device 1022 can be deployed as a closed system (such as mainframe storage) or an open system (such as a server based on a related operating system). The storage of the open system can be built-in storage and / or external storage.
[0079] For example, Figure 1 As shown, external storage can be divided into direct-attached storage (DAS) and fabric-attached storage (FAS) based on the connection method between storage device 1022 and server 1021. FAS can be further divided into network-attached storage (NAS) and storage area network (SAN) based on the transmission protocol.
[0080] In an optional manner, the network of the exemplary embodiments of the present disclosure may include one or more networks, and any suitable network may be considered. By way of example and not limitation, one or more portions of the network may include an ad hoc network, an intranet, an extranet, a virtual private network (VPN), a local area network (LAN), a wireless local area network (WLAN), a wide area network (WAN), a wireless wide area network (WWAN), a metropolitan area network (MAN), a portion of the Internet, a portion of a public switched telephone network (PSTN), a cellular telephone network, or a combination of two or more of these.
[0081] Figure 2FIG2 shows a schematic diagram of a framework of a RAG service in which various methods described herein may be implemented according to an exemplary embodiment of the present application. Figure 1 The question answering system 102 in the example may support a RAG service 200. The RAG service 200 may include a data processing service 201, a data storage service 202, and a reasoning search service 203.
[0082] like Figure 2 As shown, the data processing service 201 of the embodiment of the present application can obtain various raw data through the first interface A. The modalities of these raw data are diverse. Therefore, the various raw data obtained by the first interface A can be defined as multimodal information. After obtaining the raw data, the first interface A can, on the one hand, store the raw data through MinIO object storage, and on the other hand, construct multiple databases with the raw data as knowledge data through the data processing service 201 and the data storage service 202.
[0083] like Figure 2 As shown, the data processing service 201 can perform data analysis on the original data (i.e., multimodal information) input by the first interface A to obtain multimodal analysis data. The various engines included in the data processing engine can process the multimodal analysis data into knowledge data of various data structures under the KAG framework to obtain knowledge data of different data structures.
[0084] like Figure 2 As shown, the knowledge data for each data structure includes both homomodal knowledge data and cross-modal knowledge data. The data storage service 202 can support knowledge data storage in various types of databases within the knowledge base. As can be seen, the data processing service 201 can parse various modal parsed data from the original data, preventing the omission of original data due to incomplete or unparseable data, thereby improving the comprehensiveness and accuracy of the knowledge data in each database, avoiding the incompleteness of single modal information, and improving the data parsing quality of the original data.
[0085] Optional, such as Figure 2 As shown, databases may include vector databases (such as Milvus), relational databases, graph databases, and Elasticsearch databases. To optimize query performance, data storage service 202 also creates corresponding indexes for each database, such as inverted indexes in Elasticsearch, vector indexes in Milvus, attribute indexes in GraphDB, and B-tree indexes in relational databases. Each database can be considered a knowledge base for subsequent querying by question-answering system 102.
[0086] like Figure 2As shown, the data processing engine may include a general engine, a knowledge graph engine, and a special engine. Under the RAG framework, the multimodal parsed data may be fully input into the general engine, the knowledge graph engine, and the special engine, etc., so that the general engine, the knowledge graph engine, and the special engine, etc. may be used to process the multimodal parsed data and generate knowledge data of the corresponding data structure. Of course, the data processing service 201 may also divide the multimodal parsed data into multiple parts and process them through the general engine, the knowledge graph engine, and the special engine, etc. to generate knowledge data of the corresponding data structure.
[0087] In the embodiments of the present application, the data characteristics and usage scenarios of multimodal parsed data can be taken into consideration. For example, professional data in the multimodal parsed data can be processed through a knowledge graph engine to obtain knowledge data of a graph database. Non-professional data in the multimodal parsed data can be processed through general engines and special engines into knowledge data of a vector database, knowledge data of a relational database, and knowledge data of an Elasticsearch database.
[0088] Optional, such as Figure 2 As shown, when the original data is received through the first interface A in the form of various original documents, the efficient scheduling capability of the MinerU multi-model framework can be used to dynamically select and combine parsing models suitable for the original documents, or directly use the multimodal model (Contrastive Language-Image Pre-Training, CLIP) to parse the original data, thereby optimizing resource utilization and improving parsing efficiency, so that the data storage service 202 supports efficient storage and retrieval of large-scale data. Therefore, the embodiment of the present application can parse the original data by selecting different types of parsing models through the MinerU multi-model framework, which can prevent the limitation of a single file processing engine that cannot process different modal data or different types of files, improve the accuracy and efficiency of data parsing, reduce resource waste, and adapt to diverse file processing needs.
[0089] Optionally, a general engine can convert raw data into knowledge data with common data structures. A special engine can convert raw data into special data structures, and a knowledge graph engine can convert raw data into graph data. For example, Figure 2As shown, the raw data can be converted into knowledge data for a relational database and stored in the relational database through a general engine, or converted into knowledge data for a vector database and stored in the vector database through a general engine. The raw data can be converted into knowledge data for an Elasticsearch database and stored in the Elasticsearch database through a special engine, or converted into knowledge data for a graph database and stored in the graph database through a knowledge graph engine.
[0090] like Figure 2 As shown, the second interface B can receive a user question in the form of a user query request. The question-answering system 102 can receive the user question and the reasoning search service 203 can obtain the answer to the user question through one of the multiple RAG schemes (defined as the target question-answering model). When obtaining the answer to the user question, the reasoning search service 203 can use the retrieval function of the data storage service 202 to recall the retrieval results from multiple databases (such as vector databases, Elasticsearch databases, relational databases, and graph databases), and obtain the answer to the user question through the target question-answering model. It can be seen that by performing multimodal analysis on the original data and converting it into knowledge data with different data structures, and then storing it in different databases, the flexibility and scalability of data storage are improved, and the efficient management and analysis of large-scale data are supported, thereby enhancing the reliability and performance of the question-answering system 102. At the same time, since the data structures of the knowledge data in different databases are different, the corresponding retrieval methods will also be different. Therefore, by recalling the retrieval results of the user question from multiple databases in the data storage service 202 in multiple ways, the accuracy of the answer to the question generated by the question-answering system 102 can be guaranteed.
[0091] Optionally, multiple RAG solutions can be implemented through multiple question-answering models. Figure 2 As shown, the reasoning search service 203 can select a question-answering model from multiple candidate question-answering models in response to the question-answering model selection instruction from the user terminal 101. For example, for questions with a relatively high degree of answer certainty, the Naive-RAG model can be selected; for questions with a relatively high degree of professionalism, the KAG model can be selected; for questions where the timeliness of the answer is relatively high or the knowledge data stored in various databases may not meet the requirements for answering the question, the CRAG model can be selected; for user questions with multiple answers but requiring the best response, the Self-RAG model can be selected.
[0092] Optional, permissions for the question-answering model, such as Figure 2As shown, the reasoning search service 203 can configure the usage permissions of the CRAG model, Self-RAG model, KAG model and Naive-RAG model in response to the first permission configuration indication of the management terminal, so that the CRAG model, Self-RAG model, KAG model and Naive-RAG model can be optionally opened to the user terminal 101, so that the user terminal 101 can select different question-answering models to answer user questions according to actual needs.
[0093] Optional, permissions for using the database, such as Figure 2 As shown, the inference search service 203 can configure the usage permissions of the vector database, relational database, graph database and Elasticsearch database in response to the second permission configuration indication of the management terminal, so that the query functions of the vector database, relational database, graph database and Elasticsearch database can be optionally opened to the target question and answer model to improve the search accuracy of the target question and answer model.
[0094] like Figure 1 and Figure 2 As shown, the data processing service 201 of the embodiment of the present application can obtain various original documents through the first interface A. The data processing service 201 can store various original documents in the MinIO object storage to facilitate user viewing and downloading, and then detect various document types to obtain single-modal files and multi-modal files.
[0095] Optional original documents may include, but are not limited to, doc files, pdf files, ppt files, spreadsheet files (such as excel files), txt files, image files (such as jpg and png formats), and even multimedia files. For example, original documents may include invoices, contracts, reports, and bills. For example, contracts may be in doc files, reports may be in ppt files, and invoices and bills may be in jpg files.
[0096] like Figure 1 and Figure 2As shown, the data processing service 201 can perform data parsing on the original documents by scheduling and selecting the required models through the MinerU multi-model architecture according to the type of each original document. On the one hand, this parsing solution can enhance the flexibility and adaptability of the parsing solution, improve the parsing efficiency, and ensure the completeness and comprehensiveness of the parsing of each original document, so that the data processing service 201 has high scalability and high accuracy; on the other hand, it can provide strong data support for the subsequent database construction and the combination with the question-answering model, ensuring the accuracy and completeness of the answers to the questions output by the question-answering model, thereby meeting the diverse needs in complex business scenarios. The following examples are given with doc files, pdf files, excel files, txt files, jpg files and even multimedia.
[0097] like Figure 1 and Figure 2 As shown, when the data processing service 201 detects that the original document is a txt file, it indicates that the original document is a unimodal file, and the unimodal file is a text modal file. The text modal file can be processed by a natural language model such as the Bidirectional Encoder Representation from Transformers (BERT) series model to obtain text parsing data of the original document.
[0098] When the data processing service 201 detects that the original document is a jpg file, it indicates that the original document is a unimodal file, and the unimodal file is a picture modal file. The original document can be processed by an image recognition model such as YOLOv8 to obtain image parsing data of the original document.
[0099] For PDF files, PPT files, doc files and spreadsheet files, there may be data of various possible modalities such as text, pictures, tables, etc. Therefore, a multimodal model such as the Contrastive Language-Image Pre-Training (CLIP) model can be used to perform multimodal analysis on PDF files, PPT files, doc files and spreadsheet files to completely parse out the data of various modalities contained in the PDF files, PPT files, doc files and spreadsheet files.
[0100] In the embodiment of the present application, the integration of different original documents and various data parsed from the same original document can be defined as multimodal parsing data. Figure 2As shown, the data processing service 201 can call multiple data processing engines to process multimodal analysis data to generate knowledge data with different data structures, and store them in different databases according to the data type of the knowledge data, so that the knowledge data in each database includes both homomodal knowledge data and cross-modal knowledge data.
[0101] The embodiment of the present application can extract entities from multiple parsed data included in the multimodal parsed data to obtain multimodal entity data, which includes entity data of multiple modalities, and entity data of the same modality includes multiple entity data. At the same time, the contextual relationship between different parsed data can be extracted, and then based on the entity data of multiple modalities and the contextual relationship between different parsed data, knowledge data of multiple data structures can be generated, and then the knowledge data can be imported into the data structure according to the data structure. Figure 2 The data storage service 202 shows different databases. The corresponding relationship between each type of knowledge data and the database can be referred to the relevant description above and will not be repeated here.
[0102] Optionally, from the perspective of the source of the parsed data, the different parsed data may be contextual relationships between different parsed data of the same original document, or between different parsed data of different original documents. From the perspective of the modality of the parsed data, the different parsed data may be contextual relationships between different parsed data of the same modality, or between different parsed data of different modalities.
[0103] Optionally, different databases may store the same knowledge data but different knowledge data structures. This allows the target question-answering model to retrieve results for a user's question from different databases. Each database can return results based on different query strategies, and the results returned by each database may differ. This allows the target question-answering model to more accurately generate answers to user questions based on the query results returned by multiple databases, ensuring the accuracy of the answers.
[0104] Optionally, the multimodal entity data can be divided into multiple groups of entity data, and the same entity data can exist between different groups of entity data. Then, based on the contextual relationship between the different parsed data of different original documents and each group of entity data, knowledge data of the corresponding data structure is generated and imported into the corresponding database. In this way, in the subsequent retrieval process, the required database can be selected according to the type of user question to be retrieved, thereby improving the search efficiency of the search results of the user question. Figure 2The data storage service 202 shown supports efficient storage and retrieval of large-scale data. Furthermore, after storing knowledge data of various data structures in different databases within the knowledge base, a database index can be generated to facilitate preliminary retrieval of the target question-answering model, avoiding the inefficient retrieval caused by searching all databases.
[0105] like Figure 1 and Figure 2 As shown, the user terminal 101 of the embodiment of the present application can install the client of the question-answering system 102. The user can enter the user question in the input box of the client interaction interface and select the answer generation scheme to be used through the model selection control. The client can generate a selection indication of the question-answering model in response to the selection operation of the answer generation scheme, and then encapsulate the user question and the selection indication of the generated question-answering model into data, and send it to the server 1021 through the user terminal 101 in the form of a user query request.
[0106] like Figure 1 and Figure 2 As shown, the reasoning search service 203 of the question-answering system 102 in the server 1021 can receive a user inquiry request through the second interface B. After receiving the user inquiry request, the user question and the target question-answering model can be determined based on the user inquiry request. The reasoning search service 203 can input the user question into the selected target question-answering model.
[0107] Among them, such as Figure 1 and Figure 2 As shown, the target question-answering model can extract entities and relationships from user questions, and based on the extracted entities and relationships, send a query request to the data storage service 202. The data storage service 202 can call the retrieval function to retrieve user questions from various databases, and return the retrieval results of the user questions in at least one database to the reasoning search service 203. The reasoning search service 203 can determine the answer to the user question based on the retrieval results in at least one database.
[0108] Optionally, when the knowledge data in different databases are partially the same, such as Figure 1 and Figure 2As shown, the data storage service 202 can perform a preliminary search on the knowledge base index in response to the query request. After retrieving the target knowledge base index that matches the user question, data retrieval can be performed in the database corresponding to the target knowledge base index to obtain the retrieval result of the user question in the database, and then the retrieval result of the user question in the database is returned to the reasoning search service 203. The target question model in the reasoning search service 203 can accurately output the answer to the user question based on the retrieval result of the user question in the database and the user question. It can be seen that this method can narrow the search scope of knowledge data and improve the search efficiency and the efficiency of generating answers to questions.
[0109] Optionally, when the knowledge data in different databases are exactly the same, such as Figure 1 and Figure 2 As shown, data storage service 202 can directly search from various databases in response to query requests, allowing each database to obtain retrieval results for the user's question using different retrieval strategies. Due to the different retrieval strategies of each database, the retrieval results retrieved from each database may also vary. Therefore, after data storage service 202 returns the retrieval results from each database to reasoning search service 203, the target question model in reasoning search service 203 can generate a more accurate answer to the user's question based on the retrieval results from multiple databases and the user's question.
[0110] Optionally, after the target question answering model generates the answer to the user's question, such as Figure 1 and Figure 2 As shown, the inference search service 203 can return the answer to the user question to the user terminal 101 through the server 1021, and the user terminal 101 displays the answer to the user question in the message display area of the client interaction interface.
[0111] like Figure 1 and Figure 2 As shown, the user can also evaluate the answer to the user's question. The client can generate an evaluation result of the question answer in response to the question answer evaluation operation, and then send the evaluation result of the question answer to the server 1021 through the user terminal 101. After receiving the evaluation result of the question answer through the second interface B, the question-answering system 102 of the server 1021 can determine the satisfaction level of the answer based on the evaluation result.
[0112] In one example, when the answer satisfaction is lower than the preset satisfaction, it indicates that the answer to the question is not good. The accuracy parameters and retrieval parameters of the target question-answering model can be updated to improve the accuracy of answer generation. For example, the accuracy parameter is top_k. Before the update, top_k = 3, and after the update, top_k = 2. By increasing top_k, the possibility of the user's answer is increased. For another example, the retrieval parameter is minimum relevance. Before the update, the minimum relevance is 7, and after the update, the minimum relevance is 6. This can increase the search scope to obtain more search results, thereby providing more references for the target question and answer, thereby improving the accuracy of the user's answer.
[0113] In another example, Figure 1 and Figure 2 As shown, if the answer satisfaction is lower than the preset satisfaction, a knowledge update initiation suggestion message can be sent to the user terminal 101, and the user terminal 101 can display the knowledge update initiation suggestion message in the message display area of the client. After seeing the initiation suggestion message, the user can choose to accept or cancel it according to actual needs.
[0114] When the user chooses to accept, Figure 1 and Figure 2 As shown, in response to the acceptance operation, the client can send a confirmation indication regarding the initiation of the suggestion message to the server 1021 via the user terminal 101. The question-answering system 102 of the server 1021 receives the confirmation indication regarding the initiation of the suggestion message from the user terminal 101 via the second interface B and can update the knowledge data of the various databases with reference to the generation process of each database described above, so that at least one database contains the knowledge data associated with the user's question. This ensures the real-time nature of the knowledge data in the databases.
[0115] It can be seen that the question-answering system 102 of the embodiment of the present application can obtain the evaluation results of the answer to the question by exchanging information with the user terminal 101 through the question-answering system 102, and thus use the evaluation results of the answer to the question to optimize the parameters (accuracy parameters and retrieval parameters) of the target question-answering model and dynamically update the knowledge data of the database. Therefore, the question-answering method of the embodiment of the present application introduces a dynamic knowledge update and evaluation feedback mechanism to ensure the continuous optimization and adaptability of the question-answering system 102 and provide users with higher quality information extraction services.
[0116] The embodiment of the present application provides a method for generating a database, which can be executed by an electronic device or a chip applied to the electronic device. The electronic device can be a server or a user terminal.
[0117] Figure 3 FIG. 1 is a flow chart showing a method for generating a database according to an embodiment of the present application. Figure 3As shown, the database generation method 300 of the embodiment of the present application includes steps 301 to 303.
[0118] In step 301, in response to received multimodal information, multimodal parsed data is extracted from the multimodal information. The multimodal parsed data includes multiple parsed data. By extracting the multiple modal data contained in the multimodal information, the integrity of the extracted multimodal information can be ensured, and omissions in the extraction of the multimodal information can be reduced.
[0119] Optionally, the multimodal information of the embodiment of the present application may include multiple target files. In terms of the source of the target file, the target file may be an invoice, a contract, a report, a bill, etc., but is not limited thereto. The file format of the target file may include, but is not limited to, doc format, pdf format, table format, txt format, and image format (such as jpg format, png format, etc.).
[0120] For any one of the multiple target files, it can be defined as the first target file. In order to improve the extraction integrity and extraction efficiency of multimodal parsing data, extracting multimodal parsing data from multimodal information includes:
[0121] The file type of the first target file is detected. When it is detected that the type of the first target file is a unimodal file, the unimodal data of the first target file is parsed to obtain parsed data of the unimodal file; when it is detected that the type of the first target file is a multimodal file, the multimodal data detection is performed on the first target file to obtain parsed data of at least one modality.
[0122] In one example, when the first target file is directly determined to contain only unimodal data based on its file type, the first target file can be considered a unimodal file. For example, when the first target file is in txt format, the first target file is a text modal file, and the BERT series model can be used to extract the text data contained in the first target file; when the first target file is in jpg format, png format, or other image format, the first target file is an image modal file, and the YOLOv model can be used to extract the image data contained in the first target file.
[0123] In one example, when the file type of the first target file cannot determine that the first target file contains only unimodal data, the first target file can be considered a multimodal file. In this case, modality identification can be performed on the file data contained in the first target file to obtain at least one data modality. Based on the at least one data modality, data detection can be performed on the file data contained in the first target file to obtain parsed data of the at least one modality.
[0124] When the first target file is in the doc format, the first target file may contain text, pictures, or tables, or may contain only one of them. The following example uses the first target file containing text, pictures, and tables as an example for explanation.
[0125] Optionally, the text data contained in the first target file can be extracted using the BERT series model, the image data contained in the first target file can be extracted using the YOLOv model, and the table data contained in the first target file can be extracted using TableNet. Optionally, the efficient scheduling capabilities of the MinerU multi-model architecture can be used to dynamically select and combine models suitable for extracting the multimodal data required for the first target file. This can optimize resource utilization and improve the parsing efficiency of the target file, adapting to diverse file processing needs.
[0126] Alternatively, a multimodal model, such as the CLIP model, can be used to extract data from the first target file. The CLIP model can simultaneously extract data from multiple modalities, such as text, images, and tables, and leverage structured knowledge to enhance the logic and accuracy of the extracted parsed data. This improves the comprehensiveness and accuracy of information extraction, avoids the incompleteness of single-modal information, and enhances the quality of file parsing.
[0127] It can be seen that when extracting multimodal analysis data, the embodiment of the present application can select different data analysis methods according to the type of the target file to completely extract the data in the target file, thereby reducing the possibility of data omission in the target file, so that the multimodal analysis data includes both multiple analysis data of the same modality and multiple analysis data of different modalities.
[0128] In step 302, contextual relationships between different parsed data are determined based on the multimodal parsed data. A natural model can be used to analyze the multimodal parsed data to fully explore the relationships between different parsed data, thereby obtaining the contextual relationships between different parsed data.
[0129] Optionally, because the multimodal parsed data includes multiple parsed data of both the same modality and different modalities, the contextual relationships between different parsed data obtained based on the multimodal parsed data include both contextual relationships between different parsed data included in the same target file and contextual relationships between parsed data included in different target files. This allows for the establishment of contextual relationships not only between different parsed data within the same file, but also between different parsed data across different files.
[0130] It can be seen that the embodiment of the present application uses multiple target files as data sources, which can fully explore the relationship between the parsed data of the same target file and different target files, and generate more contextual relationships between different parsed data, thereby improving the data diversity of the knowledge data contained in different types of databases, so as to further improve the answer accuracy of the question and answer model.
[0131] In step 303, based on the multimodal parsing data and the contextual relationships between different parsing data, multiple databases are generated, and each database includes knowledge data having a different data structure.
[0132] In step 301, multimodal parsing data can be fully and completely extracted from the multimodal information, so that the knowledge data included in the multiple databases generated based on the multimodal parsing data and the contextual relationships between different parsing data are relatively complete and accurate. In addition, the data structure of the knowledge data included in each database is different. Therefore, when implementing a question-and-answer solution based on multiple databases, even for the same user question, different search results can be recalled from different databases. In this case, the search structures recalled from different databases can provide more sufficient reference information for the question-and-answer model, thereby ensuring the accuracy of the content generated by the search-enhanced generation method.
[0133] Moreover, the embodiments of the present application can store knowledge data of different data structures in different databases, which can enhance the flexibility and scalability of data storage to better support the efficient management and analysis of large-scale data and enhance the reliability and performance of the question-answering model.
[0134] In addition, since the multimodal analysis data extracted from the multimodal information makes the data extracted from the multimodal information more comprehensive, the embodiment of the present application is based on the contextual relationship between the multimodal analysis data and different analysis data. The knowledge data in the multiple databases generated is cross-modal knowledge data, and its content is relatively complete, which can fully improve the answer accuracy of the question-answering model.
[0135] Optionally, when generating multiple databases, the multiple parsed data included in the multimodal parsed data can be segmented to obtain multimodal data segments, which include multiple data segments. Based on the contextual relationship between the multimodal data segments and different parsed data, multiple databases are generated.
[0136] For example, a Natural Language Processing (NLP) model can be used to segment the multiple parsed data included in the multimodal parsed data to extract the entity data contained in each parsed data. This is equivalent to breaking each longer parsed data into multiple shorter data segments, and then organizing the multiple data segments based on the contextual relationships between different parsed data to increase the comprehensiveness and accuracy of the knowledge in the knowledge base.
[0137] Since the multiple parsed data included in the multimodal parsed data include parsed data of various modalities, after data segmentation is performed on the multiple parsed data included in the multimodal parsed data, a collection of data segments including various modalities can be obtained. Therefore, when organizing the multiple data segments based on the contextual relationships between the different parsed data, data segments of the same modality can be grouped together, and data segments of different modalities can also be grouped together. Furthermore, the different data segments grouped together can be data segments from the same target file or from different target files, and the different target files can be target files of the same type or different types.
[0138] Considering the different data structures of different databases, when organizing multiple data fragments using the contextual relationships between different parsed data, we can organize the different data fragments into knowledge data that meets the requirements of each database based on its data structure. The knowledge data of each database can then be stored in the corresponding database. For example, databases can include relational databases, Elasticsearch databases, vector databases, Elasticsearch databases, relational databases, and graph databases. This supports efficient data management and retrieval. This improves the flexibility and scalability of data storage, supports efficient management and analysis of large-scale data, and enhances the reliability and performance of the system.
[0139] The embodiment of the present application provides a question-answering method, which can be executed by an electronic device or a chip applied to the electronic device. The electronic device can be a server or a user terminal.
[0140] Figure 4 FIG. 1 shows a flow chart of a question-answering method according to an embodiment of the present application. Figure 4 As shown, the question-answering method 400 of the embodiment of the present application includes step 401 and step 402.
[0141] In step 401, in response to a user query request from a user terminal, a user question corresponding to the user query request is determined.
[0142] In step 402, the user question is input into the target question-answering model to obtain the answer to the user question. The gauge target question-answering model is used to obtain the retrieval results of the user question in at least one database, and determine the answer to the user question based on the retrieval results of the user question in at least one database. The data structures of the knowledge data included in various databases are different, and the knowledge data included in the same database is cross-modal knowledge data.
[0143] The various databases in the embodiments of the present application include knowledge data with different data structures. Therefore, when implementing a question-and-answer solution based on multiple databases, even for the same user question, different retrieval results can be recalled from different databases. In this case, the retrieval structures recalled by different databases can provide more sufficient reference information for the question-and-answer model, thereby ensuring the accuracy of the content generated by the retrieval enhancement generation method. Moreover, the embodiments of the present application can store knowledge data with different data structures in different databases, which can enhance the flexibility and scalability of data storage, better support the efficient management and analysis of large-scale data, and enhance the reliability and performance of the question-and-answer model. In addition, since the knowledge data included in the same database is cross-modal knowledge data, the knowledge data of the database is relatively complete, thereby fully improving the accuracy of the answers of the question-and-answer model.
[0144] In one possible implementation, the target question-answering model of the embodiment of the present application may be a pre-specified target question-answering model or a customized target question-answering model. For example, in response to a question-answering model selection instruction from a user terminal, the target question-answering model is selected from multiple candidate question-answering models.
[0145] The above-mentioned multiple candidate question-answering models may include the CRAG model, the Self-RAG model, the KAG model and the Naive-RAG model. According to the actual needs of the user's question, one of the candidate question-answering models can be selected to answer the user's question, thereby improving the accuracy of the answer to the question.
[0146] For example, for user questions with fixed answers, the Naive-RAG model can be used to recall retrieval structures from various databases and combined with the target question-answering model to determine the answer to the question. For user questions that require professional responses, the KAG model can be used and combined with various databases to ensure the professionalism of the answer.
[0147] It can be seen that the embodiment of the present application can selectively obtain the target question-answering model from different candidate question-answering models according to actual needs, so as to use the target question-answering model in combination with multiple databases to generate answers to questions, thereby preventing the problem of lack of innovation in question answers caused by using simple retrieval.
[0148] In one possible implementation, in order to improve the accuracy of answers to user questions and enhance system performance, the embodiment of the present application may also add an evaluation and feedback link. In this case, the method of the embodiment of the present application may further include: sending the answer to the question to the user terminal, and the display interface of the user terminal is used to display the answer to the question; responding to the evaluation result of the answer to the question from the user terminal, and determining the satisfaction of the answer based on the evaluation result; if the satisfaction of the answer is lower than the preset satisfaction, updating the accuracy parameters and retrieval parameters of the target question-answering model. When the question-answering model is used to re-answer the user's question, the accuracy of the answer to the question can be improved.
[0149] For example, the accuracy parameter is top_k. Before the update, top_k = 3, and after the update, top_k = 2. Increasing top_k increases the likelihood of the user's answer. For another example, the search parameter is minimum relevance. Before the update, the minimum relevance was 7, and after the update, the minimum relevance was 6. This increases the search scope and obtains more search results, providing more references for the target question and answer, thereby improving the accuracy of the user's answer.
[0150] In order to ensure the timeliness of the knowledge data in the database, a dynamic knowledge update mechanism can be configured. At this time, the method of the embodiment of the present application may also include sending the answer to the question to the user terminal, and the display interface of the user terminal is used to display the answer to the question; in response to the evaluation result of the answer to the question from the user terminal, the answer satisfaction is determined based on the evaluation result; if the answer satisfaction is lower than the preset satisfaction, a knowledge update prompt message can be automatically triggered to remind the administrator to update the knowledge data in the database according to the user question, so that at least one database contains knowledge data associated with the user question. Of course, a knowledge update initiation suggestion message can also be sent to the user terminal 101, and the user terminal 101 can display the knowledge update initiation suggestion message in the message display area of the client. After seeing the initiation suggestion message, the user can choose to accept or cancel according to actual needs.
[0151] When the user chooses to accept, Figure 1 and Figure 2 As shown, in response to the acceptance operation, the client can send a confirmation indication regarding the initiation of the suggestion message to server 1021 via the user terminal. Upon receiving the confirmation indication regarding the initiation of the suggestion message from user terminal 101, server 1021 can update the knowledge data in the various databases, referring to the generation process of each database described above, so that each database contains knowledge data related to the user's question. This ensures the real-time nature of the knowledge data in the databases.
[0152] Figure 5A FIG1 shows an interface diagram of a display interface of a user terminal according to an embodiment of the present application. Figure 5AAs shown, the display interface 500 is divided into a message display area 501 and a message input area 502. A user question can be entered in the message input area 502, and the user question is displayed in the message display area 501. After the user terminal receives the answer to the user question, the answer to the user question is also displayed in the message display area 501.
[0153] like Figure 5A As shown, after the answer to the user's question is displayed in the message display area 501, an answer evaluation pop-up box 503 may also pop up in the message display area 501. The user can rate the answer to the user's question based on their satisfaction. For example, the highest score in the evaluation pop-up box is five points, and the lowest score is 0 points. The lower the score, the lower the user's satisfaction with the answer to the question.
[0154] The server is the execution entity of the question-answering method of the embodiment of the present application. The user can select a score from 0 to 5 based on the satisfaction of the user's question and click Submit. In response to the submission operation, the user terminal can generate an evaluation result of the answer to the question and send the evaluation result of the answer to the question to the server. The server can determine the satisfaction of the answer based on the evaluation result of the answer to the question.
[0155] For example, the user-selected rating is used as the answer satisfaction, and the preset satisfaction is set to 3. If the user selects a rating of 4, it can be considered that the answer satisfaction is higher than the preset satisfaction, and the server does not need to take any action; if the user selects a rating of 2, it can be considered that the answer satisfaction is lower than the preset satisfaction. On the one hand, the server updates the accuracy parameters and retrieval parameters of the target question-answering model, and on the other hand, it can send a knowledge update initiation suggestion message to the user terminal.
[0156] Figure 5B Another interface diagram of the display interface of the user terminal of the embodiment of the present application is shown. Figure 5B As shown, after the user terminal receives the suggestion message, the user terminal can display the suggestion message in the message display area 501 as a system suggestion 504. The user can choose to "accept" or "reject" according to actual needs. For example, when the user selects "accept", the user can send a confirmation instruction of the suggestion message to the server in response to the user's operation of selecting "accept". After receiving the confirmation instruction of the suggestion message, the server can refer to Figure 2 Update various databases of the embodiments of the present application.
[0157] It can be seen that the technical solution of the embodiment of the present application can continuously optimize the function and suitability of the question-answering system with the help of dynamic knowledge update and feedback loop mechanism, and provide users with high-quality information services.
[0158] The above mainly introduces the solution provided by the embodiment of the present application from the perspective of an electronic device. It is understandable that, in order to realize the above functions, the electronic device includes a hardware structure and / or software module corresponding to the execution of each function. Those skilled in the art should easily realize that, in combination with the units and algorithm steps of each example described in the embodiment applied for herein, the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a function is executed in a hardware or computer software driven hardware manner depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
[0159] The embodiment of the present application can divide the functional units of the electronic device according to the above method example. For example, each functional module can be divided according to each function, or two or more functions can be integrated into one processing module. The above integrated modules can be implemented in the form of hardware or in the form of software functional modules. It should be noted that the division of modules in the embodiment of the present application is schematic and is only a logical function division. There may be other division methods in actual implementation.
[0160] In the case of dividing each functional module according to each function, an exemplary embodiment of the present application provides a device for generating a database, which may be an electronic device or a chip applied to the electronic device. Figure 6 FIG. 1 shows a schematic block diagram of functional modules of a database generation device according to an exemplary embodiment of the present application. Figure 6 As shown, the database generation device 600 includes:
[0161] An extraction module 601 is configured to extract multimodal parsed data from the received multimodal information in response to the received multimodal information, wherein the multimodal parsed data includes a plurality of parsed data;
[0162] A determination module 602 is configured to determine, based on the multimodal parsed data, contextual relationships between different parsed data;
[0163] The generating module 603 is configured to generate multiple databases based on the multimodal parsed data and the contextual relationships between different parsed data, wherein each database includes knowledge data having a different data structure.
[0164] In a possible implementation, the multimodal information includes multiple target files, and the multiple target files include a first target file.
[0165] The extraction module 601 is used to perform unimodal data analysis on the first target file when it is detected that the type of the first target file is a unimodal file, and obtain parsed data of the unimodal file; and to perform multimodal data detection on the first target file when it is detected that the type of the first target file is a multimodal file, and obtain parsed data of at least one modality.
[0166] In one possible implementation, the extraction module 601 is used to perform modality recognition on the file data contained in the first target file to obtain at least one data modality; and perform data detection on the file data contained in the first target file based on the at least one data modality to obtain parsed data of at least one modality.
[0167] In a possible implementation, at least two of the plurality of parsed data have different modalities;
[0168] The contextual relationship between the different parsed data includes the contextual relationship between different parsed data included in the same target file, and the contextual relationship between the parsed data included in different target files.
[0169] In one possible implementation, the generation module 603 is used to perform data segmentation on the multiple parsed data included in the multimodal parsed data to obtain multimodal data segments, where the multimodal data segments include multiple data segments; and generate multiple databases based on the contextual relationship between the multimodal data segments and different parsed data.
[0170] In the case of dividing each functional module according to each function, an exemplary embodiment of the present application provides a question-answering device, which may be an electronic device or a chip applied to an electronic device. Figure 7 FIG1 shows a schematic block diagram of the functional modules of the question-answering device according to an exemplary embodiment of the present application. Figure 7 As shown, the question-answering device 700 includes:
[0171] A determination module 701 is configured to determine, in response to a user query request from a user terminal, a user question corresponding to the user query request;
[0172] The reasoning module 702 is used to input the user question into the target question-answering model to obtain the answer to the user question. The target question-answering model is used to obtain the retrieval results of the user question in at least one database, and determine the answer to the user question based on the retrieval results of the user question in at least one database. The data structures of the knowledge data included in various databases are different, and the knowledge data included in the same database is cross-modal knowledge data.
[0173] In a possible implementation, the apparatus further includes a sending module 703, configured to send the answer to the question to the user terminal, and the display interface of the user terminal is configured to display the answer to the question.
[0174] The reasoning module 702 is also used to respond to the evaluation results of the answer to the question from the user terminal, and determine the answer satisfaction based on the evaluation results; if the answer satisfaction is lower than the preset satisfaction, update the accuracy parameters and retrieval parameters of the target question-answering model.
[0175] In one possible implementation, the reasoning module 702 is also used to send a knowledge update initiation suggestion message to the user terminal if the answer satisfaction is lower than a preset satisfaction level; in response to the user terminal's confirmation indication of the initiation suggestion message, update the knowledge data of multiple databases so that each database contains knowledge data associated with the user question.
[0176] In a possible implementation, the determination module 701 is further configured to select the target question-answering model from a plurality of candidate question-answering models in response to a question-answering model selection instruction from the user terminal.
[0177] Figure 8 FIG. 1 shows a schematic block diagram of a chip according to an exemplary embodiment of the present application. Figure 8 As shown, the chip 800 includes one or more (including two) processors 801 and a communication interface 802. The communication interface 802 can support the server to execute the data sending and receiving steps in the above method, and the processor 801 can support the server to execute the data processing steps in the above method.
[0178] Optional, such as Figure 8 As shown, the chip 800 also includes a memory 803, which may include a read-only memory and a random access memory, and provides operation instructions and data to the processor. Part of the memory may also include a non-volatile random access memory (NVRAM).
[0179] In some embodiments, as Figure 8As shown, the processor 801 performs corresponding operations by calling the operation instructions stored in the memory (the operation instructions may be stored in the operating system). The processor 801 controls the processing operations of any one of the terminal devices, and the processor may also be called a central processing unit (CPU). The memory 803 may include a read-only memory and a random access memory, and provides instructions and data to the processor 801. A portion of the memory 803 may also include NVRAM. For example, in an application, the memory, the communication interface, and the memory are coupled together through a bus system, wherein the bus system may include a power bus, a control bus, and a status signal bus in addition to a data bus. However, for the sake of clarity, in Figure 8 Various buses are labeled as bus system 804 .
[0180] The methods disclosed in the above embodiments of the present application can be applied to or implemented by a processor. The processor may be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method can be completed by hardware integrated logic circuits in the processor or by software instructions. The above processor may be a general-purpose processor, a digital signal processor (DSP), an ASIC, a field-programmable gate array (FPGA), or other programmable logic device, a discrete gate or transistor logic device, or a discrete hardware component. The various methods, steps, and logic block diagrams in the embodiments of the present application can be implemented or executed. The general-purpose processor may be a microprocessor or any conventional processor. The steps of the methods applied in conjunction with the embodiments of the present application can be directly implemented and executed by a hardware decoding processor, or by a combination of hardware and software modules in the decoding processor. The software module can be located in a storage medium well-known in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, etc. The storage medium is located in a memory, and the processor reads the information in the memory and, in conjunction with its hardware, completes the steps of the above method.
[0181] The exemplary embodiments of the present application further provide an electronic device, comprising: at least one processor; and a memory communicatively connected to the at least one processor. The memory stores a computer program executable by the at least one processor, wherein the computer program, when executed by the at least one processor, causes the electronic device to perform a method according to an embodiment of the present application.
[0182] An exemplary embodiment of the present application further provides a non-transitory computer-readable storage medium storing a computer program, wherein the computer program, when executed by a processor of a computer, is used to cause the computer to perform a method according to an embodiment of the present application.
[0183] An exemplary embodiment of the present application further provides a computer program product, including a computer program, wherein when the computer program is executed by a processor of a computer, it is used to cause the computer to perform the method according to the embodiment of the present application.
[0184] refer to Figure 9 , a block diagram of an electronic device 900 that can serve as a server or client of the present application will now be described, which is an example of a hardware device that can be applied to various aspects of the present application. The electronic device is intended to represent various forms of digital electronic computer equipment, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present application described and / or required herein.
[0185] like Figure 9 As shown, the electronic device 900 includes a computing unit 901, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 902 or a computer program loaded from a storage unit 908 into a random access memory (RAM) 903. Various programs and data required for the operation of the device 900 can also be stored in the RAM 903. The computing unit 901, the ROM 902, and the RAM 903 are connected to each other via a bus 904. An input / output (I / O) interface 905 is also connected to the bus 904.
[0186] like Figure 9As shown, multiple components within electronic device 900 are connected to I / O interface 905, including an input unit 906, an output unit 907, a storage unit 908, and a communication unit 909. Input unit 906 can be any type of device capable of inputting information into electronic device 900. Input unit 906 can receive input digital or character information and generate key input signals related to user settings and / or function control of the electronic device. Output unit 907 can be any type of device capable of presenting information and may include, but is not limited to, a display, a speaker, a video / audio output terminal, a vibrator, and / or a printer. Storage unit 908 may include, but is not limited to, a magnetic disk or an optical disk. Communication unit 909 allows electronic device 900 to exchange information / data with other devices via computer networks such as the Internet and / or various telecommunication networks, and may include, but is not limited to, a modem, a network card, an infrared communication device, a wireless communication transceiver and / or a chipset, such as a Bluetooth™ device, a WiFi device, a WiMax device, a cellular communication device, and / or the like.
[0187] like Figure 9 As shown, the computing unit 901 can be various general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 901 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units that run machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 901 performs the various methods and processes described above. For example, in some embodiments, the method of the embodiment of the present application can be implemented as a computer software program, which is tangibly included in a machine-readable medium, such as a storage unit 908. In some embodiments, part or all of the computer program can be loaded and / or installed on the electronic device 900 via the ROM 902 and / or the communication unit 909. In some embodiments, the computing unit 901 can be configured to perform the method of the embodiment of the present application in any other appropriate manner (e.g., by means of firmware).
[0188] The program code for implementing the methods of the present application can be written in any combination of one or more programming languages. Such program code can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that when the program code is executed by the processor or controller, the functions / operations specified in the flow charts and / or block diagrams are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0189] In the context of the present application, a machine-readable medium can be a tangible medium that can contain or store a program for use by an instruction execution system, device or equipment or used in combination with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0190] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0191] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.
[0192] Computer systems may include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The client and server relationship arises through computer programs running on the respective computers and having a client-server relationship to each other.
[0193] In the above embodiments, they can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using software, they can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer programs or instructions. When the computer program or instructions are loaded and executed on a computer, the process or function described in the embodiment of the present application is performed in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, a terminal, a user device or other programmable device. The computer program or instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer program or instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center via wired or wireless means. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium, such as a floppy disk, a hard disk, or a tape; it can also be an optical medium, such as a digital video disc (DVD); it can also be a semiconductor medium, such as a solid state drive (SSD).
[0194] Although the present application has been described with reference to specific features and embodiments thereof, it is apparent that various modifications and combinations may be made thereto without departing from the spirit and scope of the present application. Accordingly, this specification and the drawings are merely illustrative of the present application as defined by the appended claims and are deemed to cover any and all modifications, variations, combinations or equivalents within the scope of the present application. Obviously, those skilled in the art may make various modifications and variations to the present application without departing from the spirit and scope of the present application. Thus, the present application is intended to include such modifications and variations as fall within the scope of the claims of the present application and their equivalents.
Claims
1. A method for generating a database, characterized in that: include: In response to the received multimodal information, extracting multimodal parsed data from the multimodal information, the multimodal parsed data comprising a plurality of parsed data; Determining, based on the multimodal parsed data, contextual relationships between different parsed data; Based on the multimodal analysis data and the contextual relationships between different analysis data, multiple databases are generated, and the data structures of the knowledge data included in each database are different.
2. The method according to claim 1, characterized in that The multimodal information includes a plurality of target files, the plurality of target files include a first target file, and extracting multimodal parsed data from the multimodal information includes: When detecting that the type of the first target file is a single-modal file, performing single-modal data parsing on the first target file to obtain parsed data of the single-modal file; When it is detected that the type of the first target file is a multimodal file, multimodal data detection is performed on the first target file to obtain parsed data of at least one modality.
3. The method according to claim 2, characterized in that The performing multimodal data detection on the first target file to obtain parsed data of at least one modality includes: Performing modality recognition on file data contained in the first target file to obtain at least one data modality; Based on the at least one data modality, data detection is performed on the file data contained in the first target file to obtain parsed data of the at least one modality.
4. The method according to claim 2, characterized in that At least two of the plurality of parsed data have different modalities; The contextual relationship between the different parsed data includes the contextual relationship between different parsed data included in the same target file, and the contextual relationship between the parsed data included in different target files.
5. The method according to any one of claims 1 to 4, characterized in that The generating of multiple databases based on the multimodal parsed data and the contextual relationships between different parsed data includes: Performing data segmentation on a plurality of parsed data included in the multimodal parsed data to obtain multimodal data segments, wherein the multimodal data segments include a plurality of data segments; Based on the contextual relationships between the multimodal data segments and different parsed data, a plurality of databases are generated.
6. A question-answering method, characterized in that: include: In response to a user inquiry request from a user terminal, determining a user question corresponding to the user inquiry request; The user question is input into a target question-answering model to obtain an answer to the user question. The target question-answering model is used to obtain a search result of the user question in at least one database, and based on the search result of the user question in at least one database, determine the answer to the user question. The data structures of the knowledge data included in various databases are different, and the knowledge data included in the same database is cross-modal knowledge data.
7. The method according to claim 6, characterized in that The method further comprises: Sending the answer to the question to the user terminal, where the display interface of the user terminal is used to display the answer to the question; In response to an evaluation result of the answer to the question from the user terminal, determining answer satisfaction based on the evaluation result; If the answer satisfaction is lower than the preset satisfaction, the accuracy parameters and retrieval parameters of the target question-answering model are updated.
8. The method according to claim 6, characterized in that The method further comprises: If the satisfaction level of the answer is lower than a preset satisfaction level, sending a knowledge update initiation suggestion message to the user terminal; In response to a confirmation indication from the user terminal regarding the initiating suggestion message, the knowledge data in the plurality of databases are updated so that each of the databases contains knowledge data associated with the user problem.
9. The method according to claim 6, characterized in that The method further comprises: In response to a question-answering model selection instruction from the user terminal, the target question-answering model is selected from a plurality of candidate question-answering models.
10. A computer device, characterized in that: include: processor; as well as, Memory for storing programs; The program includes instructions, which, when executed by the processor, cause the processor to perform the method according to any one of claims 1 to 5.
11. A computer device, characterized in that: include: processor; as well as, Memory for storing programs; The program includes instructions, which, when executed by the processor, cause the processor to perform the method according to any one of claims 6 to 9.
12. A question-answering system, characterized in that: include: A server and a storage device connected to the server in communication; wherein, The storage device is used to store at least a plurality of databases generated by the method according to any one of claims 1 to 5; The server is used to execute the method according to any one of claims 6 to 9.