Knowledge question and answer method and system, storage medium and program product
By performing physical extraction and multi-level database search on welding workers' questions, the problem of poor Q&A quality caused by user non-standard questions is solved, and high-quality knowledge Q&A results are achieved.
Patent Information
- Application Number
- CN202510088222.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-20
- Publication Date
- 2025-05-06
AI Technical Summary
When welding workers use the intelligent knowledge question and answer system, the question statements have the characteristics of "colloquialization", "less information", and "irregularity", which leads to poor matching of the search content, affecting the quality of the question and answer results.
By analyzing user problems, extracting entities and entity relationships, and performing multi-level searches in the knowledge graph database, text database, picture vector database and semantic vector database, obtaining relevant content, and sorting and integrating these contents to generate high-quality question and answer results.
Even if users ask questions casually or have insufficient information, the quality and accuracy of question-and-answer results can be significantly improved through multi-level information expansion and content integration methods.
Smart Images

Figure CN119938940A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technology, and in particular to a knowledge question answering method, system, storage medium and program product. Background Art
[0002] With the rapid development of large language model technology, enterprises are increasingly in need of efficient and intelligent knowledge question answering solutions. In key links such as welding process development and on-site problem solving, welding practitioners are in urgent need of obtaining accurate and detailed technical knowledge and information, so intelligent knowledge question answering systems have become an ideal choice to meet this demand. However, despite the continuous advancement of intelligent knowledge question answering technology, it still faces many challenges in practical applications.
[0003] For welding technicians, when using knowledge question-and-answer tools, they can often ask detailed questions based on specific situations and professional information. After obtaining sufficient information, the large model can retrieve the reference content needed to answer the questions and answer them accordingly. However, for workers on site, due to their limited cultural level, they do not pay enough attention to prompt words. When asking questions, they use sentences with characteristics such as "colloquial", "little information" and "irregularity". These sentences are difficult to form a clear orientation in the semantic vector space, resulting in poor matching of the search content, which directly affects the question-and-answer results. Summary of the invention
[0004] A technical problem to be solved by the present disclosure is that the present disclosure provides a knowledge question and answer method, system, storage medium and program product, which can obtain high-quality question and answer results.
[0005] According to one aspect of the present disclosure, a knowledge question and answer method is proposed, including: analyzing a first question, extracting entities and entity relationships in the first question; retrieving entities and entity relationships in a knowledge graph database to obtain multiple key node entities and a first content; based on the multiple key node entities, searching in a text database to obtain a second content; based on the multiple key node entities, searching in a picture vector database to obtain a third content; based on the multiple key node entities, searching in a semantic vector database to obtain a fourth content; sorting knowledge blocks in the first content, the second content, the third content and the fourth content to output content related to the first question, wherein the knowledge graph database, the text database, the picture vector database and the semantic vector database are generated based on knowledge files.
[0006] In some embodiments, retrieving entities and entity relationships in a knowledge graph database to obtain multiple key node entities and a first content includes: retrieving entities and entity relationships in the knowledge graph database to obtain an initial retrieval node entity, and along the nodes and edges of the knowledge graph, obtaining related node entities within a predetermined layer in the knowledge graph database, the initial retrieval node entity and the related node entities constitute multiple key node entities; obtaining content corresponding to each of the multiple key node entities in the knowledge graph database, the content corresponding to each of the multiple key node entities constitutes the first content.
[0007] In some embodiments, searching in a semantic vector database based on multiple key node entities to obtain the fourth content includes: integrating the multiple key node entities and the first question to obtain the second question, and searching in the semantic vector database based on the second question to obtain the fourth content.
[0008] In some embodiments, there are multiple second questions, and based on the second questions, searching in the semantic vector database to obtain the fourth content includes: determining a target question among the multiple second questions; searching the target question in the semantic vector database to obtain the fourth content.
[0009] In some embodiments, determining a target question among multiple second questions includes: determining the target question in response to a user selecting from multiple second questions; or sorting the multiple second questions using a deep learning model to obtain the target question.
[0010] In some embodiments, sorting the knowledge blocks in the first content, the second content, the third content, and the fourth content so as to output the knowledge blocks related to the first question includes: searching for knowledge blocks with overlapping contents in the first content, the second content, the third content, and the fourth content; sorting the knowledge blocks according to the content to which each knowledge block belongs and the weight of each content; selecting a predetermined number of knowledge blocks for feature extraction; and sorting the predetermined number of knowledge blocks according to the extracted features to obtain knowledge blocks related to the first question.
[0011] In some embodiments, searching in an image vector database based on multiple key node entities to obtain the third content includes: searching in an image vector database based on multiple key node entities to obtain a reference image; and converting the content in the reference image into the third content.
[0012] In some embodiments, the content of the knowledge document is textualized; the textualized content is matched with the corresponding outline according to paragraphs to obtain split pairs; and a text database is constructed based on the split pairs.
[0013] In some embodiments, matching the textual content with the corresponding outline according to paragraphs to obtain split pairs includes: merging multiple paragraphs in the textual content around a fixed theme to form new paragraphs; matching the new paragraphs with the corresponding outline to obtain split pairs.
[0014] In some embodiments, the split pairs are vectorized; and a semantic vector database is constructed based on the vectorized split pairs.
[0015] In some embodiments, based on the entities in the split pairs and the relationships between the entities, the entity names, entity descriptions, relationship descriptions, and related content are constructed into a knowledge graph database, and the nodes and edges of the knowledge graph are established.
[0016] In some embodiments, the knowledge documents are visualized; the visualized knowledge documents are vectorized using a multimodal vector model to construct a picture vector library.
[0017] According to another aspect of the present disclosure, a knowledge question and answer system is also proposed, including: an extraction module, configured to analyze a first question, and extract entities and entity relationships in the first question; a first retrieval module, configured to retrieve entities and entity relationships in a knowledge graph database, and obtain multiple key node entities and a first content; a second retrieval module, configured to search in a text database based on multiple key node entities, and obtain a second content; a third retrieval module, configured to search in an image vector database based on multiple key node entities, and obtain a third content; a fourth retrieval module, configured to search in a semantic vector database based on multiple key node entities, and obtain a fourth content; a knowledge output module, configured to sort knowledge blocks in the first content, the second content, the third content, and the fourth content, so as to output content related to the first question, wherein the knowledge graph database, the text database, the image vector database, and the semantic vector database are generated based on knowledge files.
[0018] According to another aspect of the present disclosure, a knowledge question answering system is also proposed, including: a memory; and a processor coupled to the memory, the processor being configured to execute the knowledge question answering method as described above based on instructions stored in the memory.
[0019] According to another aspect of the present disclosure, a computer-readable storage medium is also provided, on which computer program instructions are stored, and when the instructions are executed by a processor, the knowledge question and answer method as described above is implemented.
[0020] According to another aspect of the present disclosure, a computer program product is also proposed, including a computer program or instructions, and the computer program or instructions implement the above-mentioned knowledge question and answer method when executed by a processor.
[0021] Other features and advantages of the present disclosure will become apparent from the following detailed description of exemplary embodiments of the present disclosure with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] The accompanying drawings, which constitute a part of the specification, illustrate embodiments of the present disclosure and, together with the description, serve to explain the principles of the present disclosure.
[0023] The present disclosure may be more clearly understood from the following detailed description with reference to the accompanying drawings, in which:
[0024] Figure 1 Schematic diagram of some embodiments of the knowledge question answering method disclosed in the present invention;
[0025] Figure 2 This is a schematic diagram of keyword expansion based on knowledge graph disclosed in the present invention;
[0026] Figure 3 Schematic diagram of the flow chart of other embodiments of the knowledge question answering method disclosed in the present invention;
[0027] Figure 4 Schematic diagram of the flow chart of some further embodiments of the knowledge question answering method disclosed in the present invention;
[0028] Figure 5 It is a schematic diagram of the structure of some embodiments of the knowledge question answering system disclosed in the present invention;
[0029] Figure 6 Schematic diagrams of structures of other embodiments of the knowledge question answering system disclosed in the present invention;
[0030] Figure 7 Schematic diagrams of structures of other embodiments of the knowledge question answering system disclosed in the present invention. DETAILED DESCRIPTION
[0031] Various exemplary embodiments of the present disclosure will now be described in detail with reference to the accompanying drawings. It should be noted that the relative arrangement of components and steps, numerical expressions and numerical values set forth in these embodiments do not limit the scope of the present disclosure unless otherwise specifically stated.
[0032] At the same time, it should be understood that for the convenience of description, the sizes of the various parts shown in the drawings are not drawn according to the actual proportional relationship.
[0033] The following description of at least one exemplary embodiment is merely illustrative in nature and is in no way intended to limit the present disclosure, its application, or uses.
[0034] Technologies, methods, and equipment known to ordinary technicians in the relevant art may not be discussed in detail, but where appropriate, the technologies, methods, and equipment should be considered as part of the specification.
[0035] In all examples shown and discussed herein, any specific values should be interpreted as merely exemplary and not as limiting. Therefore, other examples of the exemplary embodiments may have different values.
[0036] It should be noted that like reference numerals and letters refer to similar items in the following figures, and therefore, once an item is defined in one figure, it need not be further discussed in subsequent figures.
[0037] In order to make the objectives, technical solutions and advantages of the present disclosure more clearly understood, the present disclosure is further described in detail below in combination with specific embodiments and with reference to the accompanying drawings.
[0038] In the related art, the intelligent knowledge question-and-answer system can be used normally when the user's questions contain as much key information as possible. However, in actual applications, for users who have not received special training, due to insufficient attention to the prompt words, the sentences used when asking questions have the characteristics of "colloquial", "little information", and "irregularity". These sentences are difficult to form a clear orientation in the semantic vector space, resulting in poor matching of the retrieval content, which directly affects the question-and-answer results.
[0039] The present invention provides a knowledge question-answering solution, which can still provide high-quality answers even when user questions are too casual, not specific, and contain little information.
[0040] Figure 1 1 is a flowchart of some embodiments of the knowledge question answering method disclosed in the present invention, and the embodiment includes steps S11 to S16.
[0041] In step S11, the first question is analyzed to extract entities and entity relationships in the first question.
[0042] For example, after receiving a question from a user, the system uses a large language model to analyze the user's question and extract the entities in the question and the relationships between the entities. Taking a welding question as an example, the user's question is "Why did the welded part I welded yesterday crack today?" The extracted entity is "weld", the relationship is "crack", and the relationship description is "weld, from yesterday to today, crack".
[0043] In step S12, entities and entity relationships are retrieved in the knowledge graph database to obtain multiple key node entities and the first content.
[0044] The knowledge graph database is generated based on knowledge files, such as welding process knowledge files.
[0045] In some embodiments, entities and entity relationships are retrieved in a knowledge graph database to obtain an initial retrieval node entity, and along the nodes and edges of the knowledge graph, related node entities within a predetermined layer are obtained in the knowledge graph database, and the initial retrieval node entity and the related node entities constitute a plurality of key node entities; content corresponding to each of the plurality of key node entities is obtained in the knowledge graph database, and content corresponding to each of the plurality of key node entities constitutes a first content.
[0046] For example, in the knowledge graph database, entities and entity relationships contained in the question are searched. After obtaining the search content, nodes are searched outward layer by layer along the nodes and edges of the knowledge graph to obtain all relevant key node entities K within m layers and the first content A. The first content A includes the original text corresponding to the initial search entity and the original text corresponding to the surrounding related node entities. m is a positive integer greater than or equal to 1.
[0047] Taking welding as an example, the user's question is "Why did the welded part yesterday crack today?" The extracted entity is "weld", the relationship is "crack", and the relationship description is "weld, yesterday to today, crack". The search results in the knowledge graph database are Figure 2 The "welding crack" shown, that is, the central word is "welding crack", the first layer C1 outward is the constraint node, joint form node, metallurgical factor node, constraint node, etc., the second layer C2 is the cold crack node, hydrogen diffusion aggregation node, high-strength steel cold crack node, welding heat input node, lower critical stress node, etc., the third layer C3 is the welding process node, small heat input node, post-heat input node, etc. At this time, if m is 3, the entity K at least includes the keywords: "welding crack", "constraint", "cold crack" and "welding process", etc., and the original texts attached to these nodes constitute document A. Those skilled in the art should understand that this is only used as an example, and the number of expansion layers can be set according to actual conditions.
[0048] In this embodiment, full use is made of the close connections between entities in the knowledge graph to expand the information of user questions outward layer by layer, effectively solving the problem of poor answer quality caused by user questions being too casual, non-specific, and lacking in information.
[0049] In step S13, based on the multiple key node entities, a search is performed in the text database to obtain the second content.
[0050] The text database is generated based on knowledge documents.
[0051] In some embodiments, the key node entity K is used as a keyword to perform a keyword search in a common text database to obtain the second content.
[0052] In step S14, based on the multiple key node entities, a search is performed in the image vector database to obtain the third content.
[0053] The image vector database is generated based on the knowledge file.
[0054] In some embodiments, based on multiple key node entities, a search is performed in an image vector database to obtain a reference image; and the content in the reference image is converted into a third content.
[0055] For example, the key node entity K is used as a keyword to search in the image vector database to obtain a reference image, and the content in the reference image is converted into the third content C using the OCR (Optical Character Recognition) model.
[0056] In step S15, based on the multiple key node entities, a search is performed in the semantic vector database to obtain the fourth content.
[0057] The semantic vector database is generated based on the knowledge documents.
[0058] In some embodiments, the key node entity K is used as a keyword to search in the semantic vector database to obtain the fourth content.
[0059] In some embodiments, multiple key node entities and the first question are integrated to obtain a second question, and based on the second question, a search is performed in a semantic vector database to obtain fourth content.
[0060] For example, multiple key node entities K and the original question are combined to form professional and detailed questions, so that more accurate answers can be retrieved from the semantic vector database.
[0061] In some embodiments, there are multiple second questions, and based on the second questions, searching in the semantic vector database to obtain the fourth content includes: determining a target question among the multiple second questions; searching the target question in the semantic vector database to obtain the fourth content.
[0062] For example, multiple key node entities K and the original question are reintegrated multiple times, and the target question is the best question. By determining the best question, the retrieval efficiency and accuracy can be improved. The target question can be determined in two modes.
[0063] In some embodiments, the target question is determined in response to a user selecting among a plurality of second questions.
[0064] For example, multiple second questions are fed back to the user, and the user selects a target question, which is vectorized and searched in a semantic vector database to obtain the fourth content.
[0065] In some embodiments, a deep learning model is used to sort the second questions to obtain a target question. The deep learning model is, for example, a large language model.
[0066] For example, a plurality of second questions are analyzed and sorted through a large model to obtain a target question, the target question is vectorized, and then searched in a semantic vector database to obtain the fourth content.
[0067] In step S16, the knowledge blocks in the first content, the second content, the third content, and the fourth content are sorted so as to output the content related to the first question.
[0068] For example, the large model answers questions based on the acquired knowledge blocks and outputs content.
[0069] In this embodiment, knowledge questions and answers are conducted by taking the knowledge graph as the basis, expanding through multi-level information, and integrating the four-dimensional overlapping information of "graph, text, semantic vector, and picture", and sorting the knowledge blocks, thereby ensuring that the corpus provided to the large model has "high relevance" and "strong concentration". Even if the questions asked by users are relatively casual, high-quality answers can still be obtained when conducting intelligent knowledge questions and answers.
[0070] In some embodiments, sorting the knowledge blocks in the first content, the second content, the third content, and the fourth content so as to output the knowledge blocks related to the first question includes: searching for knowledge blocks with overlapping contents in the first content, the second content, the third content, and the fourth content; sorting the knowledge blocks according to the content to which each knowledge block belongs and the weight of each content; selecting a predetermined number of knowledge blocks for feature extraction; and sorting the predetermined number of knowledge blocks according to the extracted features to obtain knowledge blocks related to the first question.
[0071] For example, a comprehensive comparison is performed on the first content A, the second content B, the third content C, and the fourth content D to determine the overlapping paragraphs, and all the paragraphs are scored. For example, if a paragraph belongs to the first content A, the score is 0.2; if it belongs to the second content B, the score is 0.2; if it belongs to the third content C, the score is 0.2; if it belongs to the fourth content D, the score is 0.4. Those skilled in the art should understand that the scores here are only for example purposes, and the scores can be set according to actual conditions. All the scored paragraphs are initially sorted, and then the sorted paragraphs are feature extracted. For example, features such as semantic relevance, word frequency, TF-IDF (Term Frequency–Inverse Document Frequency) values in the paragraphs are extracted, and based on the extracted features, the documents are secondary sorted to obtain the reference paragraphs with the highest relationship and overlap with the user's questions and surrounding keywords.
[0072] Those skilled in the art should understand that when extracting features from sorted paragraphs, feature extraction may be performed on all sorted paragraphs, or a predetermined number of paragraphs ranked at the top may be selected for feature extraction.
[0073] In the above embodiment, scoring and two-round sorting are performed on the knowledge blocks in the content obtained through four dimensions, which can achieve efficient screening of highly matching reference content and effectively improve the effect of question and answer.
[0074] Figure 3 1 is a flowchart of some other embodiments of the knowledge question answering method disclosed in the present invention. In addition to steps S11 to S16, this embodiment also includes step S31.
[0075] In step S31, a knowledge graph database, a text database, an image vector database and a semantic vector database are generated based on the knowledge file.
[0076] The knowledge file is, for example, a welding process knowledge file.
[0077] In some embodiments, the content of the knowledge document is textualized; the textualized content is matched with the corresponding outline according to paragraphs to obtain split pairs; and a text database is constructed based on the split pairs.
[0078] For example, after a user uploads a knowledge file, the file is converted to plain text, that is, the pictures, tables, and text in the file are described in text using a large language model. Then, the paragraphs are matched with the corresponding outlines to obtain split pairs, which are stored in a normal text database.
[0079] In some embodiments, matching the textual content with the corresponding outline according to paragraphs to obtain split pairs includes: merging multiple paragraphs in the textual content around a fixed theme to form new paragraphs; matching the new paragraphs with the corresponding outline to obtain split pairs.
[0080] For example, use the large model to review each paragraph one by one, merge the scattered paragraphs around a fixed theme to form a new paragraph, and use this fixed theme as the title of the paragraph. Then, together with the outlines of all levels of the original document, split the entire document into paragraphs to form split pairs.
[0081] In this embodiment, the file is converted into plain text and scattered paragraphs are merged into a paragraph with a fixed topic as the title, which can ensure the integrity of the information, thereby increasing the probability of matching information being retrieved, and also ensuring that once retrieved, the reference information obtained by the large model is complete.
[0082] In some embodiments, the split pairs are vectorized; and a semantic vector database is constructed based on the vectorized split pairs.
[0083] For example, the split pairs are vectorized and stored in a semantic vector database, making it easier to parse the problem through semantics and obtain the corresponding answer.
[0084] In some embodiments, based on the entities in the split pairs and the relationships between the entities, the entity names, entity descriptions, relationship descriptions, and related content are constructed into a knowledge graph database, and the nodes and edges of the knowledge graph are established.
[0085] For example, a large language model is used to extract the entities in each part of the split pair and the relationship between each entity, and the entity name, entity description, relationship description, and related original text are constructed into a knowledge graph database, and the nodes and edges of the knowledge graph are established, and the complexity of the node relationship is evaluated. After the user asks a question, the large language model is used to extract the entities in the question and the relationship between the entities, and then search in the knowledge graph database. After obtaining the search content, the nodes are searched layer by layer along the nodes and edges of the knowledge graph to obtain all relevant key node entities and original texts within m layers. The multi-level expansion of user question information is realized to improve the search efficiency and accuracy.
[0086] In some embodiments, the knowledge documents are visualized; the visualized knowledge documents are vectorized using a multimodal vector model to construct a picture vector library.
[0087] For example, each page of content in a file is converted into an image, and the image is vectorized using a multimodal vector model to establish an image vector library, thereby achieving comprehensive perception of information.
[0088] The following will introduce the knowledge question answering method by taking a specific example. Figure 4 As shown, Figure 4 Schematic diagram of the flow of some further embodiments of the knowledge question answering method disclosed herein, including steps S401-S420.
[0089] In step S401, the local file is analyzed and processed by LLM (Large Language Model) to achieve plain text.
[0090] The local file is, for example, a file related to welding.
[0091] In step S402, the text is split into paragraphs level by level to construct split pairs.
[0092] In step S403, a common text knowledge base is constructed based on the split pairs.
[0093] In step S404, LLM is used to extract entities and relationships in the split pairs.
[0094] In step S405, a knowledge graph database is constructed.
[0095] In step S406, the split pairs are vectorized to construct a semantic vector database.
[0096] In step S407, the local file is converted into a picture by processing the script.
[0097] In step S408, the image is vectorized using a multimodal vector model to construct an image vector database.
[0098] The process of constructing a general text knowledge base, a knowledge graph database, a semantic vector database, and an image vector database may be performed simultaneously or in no particular order.
[0099] In step S409, the user asks a question.
[0100] In step S410, the problem is analyzed using LLM to extract entities and relationships.
[0101] In step S411, a search is performed in the knowledge graph database to obtain the multi-layer key node entity K and the original text A corresponding to the multi-layer node entity.
[0102] In step S412, a search is performed in a common text database using the multi-layer key node entity K as a keyword to obtain text B.
[0103] In step S413, the multi-layer key node entity K is used as a keyword to search in the image vector library to obtain relevant image content C.
[0104] In step S414, the multi-layer key node entity K is used as a keyword to search in the semantic vector database to obtain the text D.
[0105] In step S415, the original text A, text B, image content C and text D are comprehensively compared to obtain multiple knowledge blocks.
[0106] In step S416, the knowledge blocks are preliminarily sorted.
[0107] In step S417, feature extraction is performed on the sorted knowledge blocks.
[0108] In step S418, the knowledge blocks are sorted a second time.
[0109] In step S419, the prompt words are combined to generate standardized document-related content.
[0110] In step S420, the answers are integrated and output through LLM.
[0111] In the above embodiment, knowledge questions and answers are performed by taking the knowledge graph as the base, expanding the information at multiple levels, and integrating the four-dimensional overlapping information of "graph, text, semantic vector, and picture". When constructing "reference content", the system uses the knowledge graph for the multi-level expansion of user question information. After searching the information sources at four levels, the obtained knowledge blocks are scored and sorted in two rounds, so as to ensure that the corpus provided to the large model has "high relevance" and "strong concentration". Based on this, even if the user asks a very casual question, when calling the large language model for intelligent knowledge questions and answers, high-quality answers can still be obtained.
[0112] Figure 5 This is a structural diagram of some embodiments of the knowledge question answering system disclosed in the present invention. The knowledge question answering system 5 includes an extraction module 51, a first retrieval module 52, a second retrieval module 53, a third retrieval module 54, a fourth retrieval module 55 and a knowledge output module 56.
[0113] The extraction module 51 is configured to analyze the first question and extract entities and entity relationships in the first question.
[0114] The first retrieval module 52 is configured to retrieve entities and entity relationships in the knowledge graph database to obtain multiple key node entities and the first content.
[0115] In some embodiments, entities and entity relationships are retrieved in a knowledge graph database to obtain an initial retrieval node entity, and along the nodes and edges of the knowledge graph, related node entities within a predetermined layer are obtained in the knowledge graph database, and the initial retrieval node entity and the related node entities constitute a plurality of key node entities; content corresponding to each of the plurality of key node entities is obtained in the knowledge graph database, and content corresponding to each of the plurality of key node entities constitutes a first content.
[0116] The second retrieval module 53 is configured to search in the text database based on a plurality of key node entities to obtain the second content.
[0117] The third retrieval module 54 is configured to search the image vector database based on a plurality of key node entities to obtain the third content.
[0118] In some embodiments, based on multiple key node entities, a search is performed in an image vector database to obtain a reference image; and the content in the reference image is converted into a third content.
[0119] The fourth retrieval module 55 is configured to search in the semantic vector database based on multiple key node entities to obtain fourth content.
[0120] In some embodiments, multiple key node entities and the first question are integrated to obtain a second question, and based on the second question, a search is performed in a semantic vector database to obtain fourth content.
[0121] In some embodiments, there are multiple second questions, and based on the second questions, searching in the semantic vector database to obtain the fourth content includes: determining a target question among the multiple second questions; searching the target question in the semantic vector database to obtain the fourth content.
[0122] In some embodiments, determining a target question among multiple second questions includes: determining the target question in response to a user selecting from multiple second questions; or sorting the multiple second questions using a deep learning model to obtain the target question.
[0123] The knowledge output module 56 is configured to sort the knowledge blocks in the first content, the second content, the third content and the fourth content so as to output content related to the first question, wherein the knowledge graph database, the text database, the image vector database and the semantic vector database are generated based on the knowledge files.
[0124] In some embodiments, knowledge blocks with overlapping contents in the first content, the second content, the third content, and the fourth content are searched; the knowledge blocks are sorted according to the contents to which each knowledge block belongs and the weight of each content; a predetermined number of knowledge blocks are selected for feature extraction; and a predetermined number of knowledge blocks are sorted according to the extracted features to obtain knowledge blocks related to the first question.
[0125] In this embodiment, knowledge questions and answers are conducted by taking the knowledge graph as the basis, expanding through multi-level information, and integrating the four-dimensional overlapping information of "graph, text, semantic vector, and picture", and sorting the knowledge blocks, thereby ensuring that the corpus provided to the large model has "high relevance" and "strong concentration". Even if the questions asked by users are relatively casual, high-quality answers can still be obtained when conducting intelligent knowledge questions and answers.
[0126] Figure 6 It is a structural schematic diagram of some other embodiments of the knowledge question and answer system disclosed in the present invention. The knowledge question and answer system also includes a database construction module 61, which is configured to construct a knowledge graph database, a text database, an image vector database and a semantic vector database based on knowledge files.
[0127] The knowledge file is, for example, a welding process knowledge file.
[0128] In some embodiments, the database construction module 61 is configured to textualize the content of the knowledge document; match the textualized content with the corresponding outline according to paragraphs to obtain split pairs; and construct a text database based on the split pairs.
[0129] Matching the textual content with the corresponding outline according to the paragraphs to obtain the split pairs includes: merging multiple paragraphs in the textual content with a fixed theme as the center to form new paragraphs; matching the new paragraphs with the corresponding outline to obtain the split pairs.
[0130] In some embodiments, the database construction module 61 is further configured to vectorize the split pairs; and construct a semantic vector database based on the vectorized split pairs.
[0131] In some embodiments, the database construction module 61 is also configured to construct the entity names, entity descriptions, relationship descriptions, and related content into a knowledge graph database based on the entities in the split pairs and the relationships between the entities, and to establish nodes and edges of the knowledge graph.
[0132] In some embodiments, the database construction module 61 is further configured to image the knowledge files; use a multimodal vector model to image vectorize the imaged knowledge files, and construct an image vector library.
[0133] It should be noted that the above modules are only logical modules divided according to the specific functions they implement, and are not used to limit the specific implementation methods. For example, they can be implemented in software, hardware, or a combination of software and hardware. In actual implementation, the above modules can be implemented as independent physical entities, or can also be implemented by a single entity (for example, a processor (CPU or DSP, etc.), an integrated circuit, etc.). In addition, the above modules are shown with dotted lines in the accompanying drawings to indicate that these modules may not actually exist, and the operations / functions they implement can be implemented by the processing circuit itself.
[0134] Figure 7 Schematic diagram of the structure of some other embodiments of the knowledge question and answer system of the present disclosure. The knowledge question and answer system 7 includes a memory 710 and a processor 720. Among them: the memory 710 can be a disk, a flash memory or any other non-volatile storage medium. The memory is used to store the instructions in the above embodiments. The processor 720 is coupled to the memory 710 and can be implemented as one or more integrated circuits, such as a microprocessor or a microcontroller. The processor 720 is used to execute the instructions stored in the memory.
[0135] In some embodiments, the processor 720 is coupled to the memory 710 via a BUS 730. The target tracking device 700 can also be connected to an external storage system 770 via a storage interface 740 to call external data, and can also be connected to a network or another computer system (not shown) via a network interface 760. No further detailed description will be given here.
[0136] In this embodiment, by storing data instructions in a memory and then processing the instructions through a processor, the problem of poor answer quality caused by user questions being too casual, non-specific, and lacking in information can be solved.
[0137] In other embodiments, a computer-readable storage medium stores computer program instructions, which, when executed by a processor, implement the steps of the method in the above-mentioned embodiment. It should be understood by those skilled in the art that the embodiments of the present disclosure may be provided as methods, devices, or computer program products. Therefore, the present disclosure may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present disclosure may take the form of a computer program product implemented on one or more computer-usable non-transient storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0138] In some embodiments, a computer program product is protected, including a computer program or an instruction, which implements the above method when the computer program or instruction is executed by a processor. The computer program product includes a computer program carried on a computer-readable medium, and the computer program contains program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network through a communication device, or installed from a storage device, or installed from a ROM. When the computer program is executed by the CPU, the above functions defined in the method of the embodiment of the present disclosure are executed.
[0139] The present disclosure is described with reference to flowcharts and / or block diagrams of methods, devices (systems) and computer program products according to embodiments of the present disclosure. It should be understood that each process and / or block in the flowchart and / or block diagram and the combination of processes and / or blocks in the flowchart and / or block diagram can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0140] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0141] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0142] So far, the present disclosure has been described in detail. In order to avoid obscuring the concept of the present disclosure, some details known in the art are not described. Based on the above description, those skilled in the art can fully understand how to implement the technical solution disclosed here.
[0143] The method and apparatus of the present disclosure may be implemented in many ways. For example, the method and apparatus of the present disclosure may be implemented by software, hardware, firmware, or any combination of software, hardware, and firmware. The above order of steps for the method is for illustration only, and the steps of the method of the present disclosure are not limited to the order specifically described above, unless otherwise specifically stated. In addition, in some embodiments, the present disclosure may also be implemented as a program recorded in a recording medium, which includes machine-readable instructions for implementing the method according to the present disclosure. Therefore, the present disclosure also covers a recording medium storing a program for executing the method according to the present disclosure.
[0144] Although some specific embodiments of the present disclosure have been described in detail by way of example, it should be understood by those skilled in the art that the above examples are for illustration only and are not intended to limit the scope of the present disclosure. It should be understood by those skilled in the art that the above embodiments may be modified without departing from the scope and spirit of the present disclosure. The scope of the present disclosure is defined by the appended claims.
Claims
1. A knowledge question answering method, comprising: Analyze the first question and extract entities and entity relationships in the first question; Retrieving the entity and entity relationship in the knowledge graph database to obtain multiple key node entities and the first content; Based on the multiple key node entities, searching in a text database to obtain second content; Based on the multiple key node entities, searching in the image vector database to obtain third content; Based on the multiple key node entities, searching in a semantic vector database to obtain fourth content; The knowledge blocks in the first content, the second content, the third content and the fourth content are sorted so as to output content related to the first question, wherein the knowledge graph database, the text database, the image vector database and the semantic vector database are generated based on knowledge files.
2. The knowledge question answering method according to claim 1, wherein: The retrieving the entities and entity relationships in the knowledge graph database to obtain multiple key node entities and the first content includes: The entity and entity relationship are retrieved in the knowledge graph database to obtain an initial retrieval node entity, and along the nodes and edges of the knowledge graph, relevant node entities within a predetermined layer are obtained in the knowledge graph database, wherein the initial retrieval node entity and the relevant node entities constitute the multiple key node entities; The content corresponding to each of the multiple key node entities is obtained in the knowledge graph database, and the content corresponding to each of the multiple key node entities constitutes the first content.
3. The knowledge question answering method according to claim 1, wherein: The searching in the semantic vector database based on the multiple key node entities to obtain the fourth content includes: The multiple key node entities and the first question are integrated to obtain a second question, and based on the second question, a search is performed in a semantic vector database to obtain fourth content.
4. The knowledge question answering method according to claim 3, wherein: The second question is multiple, and the searching in the semantic vector database based on the second question to obtain the fourth content includes: determining a target problem among a plurality of the second problems; The target question is retrieved in the semantic vector database to obtain the fourth content.
5. The knowledge question answering method according to claim 4, wherein: Determining a target question among the plurality of second questions comprises: In response to a user selecting a plurality of the second questions, determining the target question; or The plurality of the second questions are sorted by using a deep learning model to obtain the target question.
6. The knowledge question answering method according to claim 1, wherein: The step of sorting the knowledge blocks in the first content, the second content, the third content, and the fourth content so as to output the knowledge blocks related to the first question comprises: Finding knowledge blocks with overlapping contents among the first content, the second content, the third content, and the fourth content; Sorting the knowledge blocks according to the content to which each knowledge block belongs and the weight of each content; Selecting a predetermined number of knowledge blocks for feature extraction; The predetermined number of knowledge blocks are sorted according to the extracted features to obtain knowledge blocks related to the first question.
7. The knowledge question answering method according to claim 1, wherein: The searching in the picture vector database based on the multiple key node entities to obtain the third content includes: Based on the multiple key node entities, searching in a picture vector database to obtain a reference picture; The content in the reference picture is converted into the third content.
8. The knowledge question answering method according to any one of claims 1 to 7, further comprising: Textualizing the content of the knowledge document; Match the textual content with the corresponding outline according to the paragraphs to obtain split pairs; Based on the split pairs, the text database is constructed.
9. The knowledge question answering method according to claim 8, wherein: The textualized content is matched with the corresponding outline according to paragraphs, and the split pairs obtained include: Merging multiple paragraphs in the textual content around a fixed topic to form a new paragraph; The new paragraph is matched with the corresponding outline to obtain the split pair.
10. The knowledge question answering method according to claim 8, further comprising: vectorizing the split pairs; Based on the vectorized split pairs, the semantic vector database is constructed.
11. The knowledge question answering method according to claim 8, further comprising: According to the entities in the split pairs and the relationships between the entities, the entity names, entity descriptions, relationship descriptions, and related contents are constructed into the knowledge graph database, and the nodes and edges of the knowledge graph are established.
12. The knowledge question answering method according to any one of claims 1 to 7, further comprising: Picturing the knowledge document; The image-based knowledge files are vectorized using a multimodal vector model to construct the image vector library.
13. A knowledge question answering system, comprising: An extraction module is configured to analyze the first question and extract entities and entity relationships in the first question; A first retrieval module is configured to retrieve the entity and entity relationship in the knowledge graph database to obtain a plurality of key node entities and a first content; A second retrieval module is configured to search in a text database based on the multiple key node entities to obtain second content; A third retrieval module is configured to search in the image vector database based on the multiple key node entities to obtain third content; A fourth retrieval module is configured to search in a semantic vector database based on the multiple key node entities to obtain fourth content; A knowledge output module is configured to sort the knowledge blocks in the first content, the second content, the third content and the fourth content so as to output content related to the first question, wherein the knowledge graph database, the text database, the image vector database and the semantic vector database are generated based on knowledge files.
14. A knowledge question answering system, comprising: Memory; as well as A processor coupled to the memory, the processor being configured to execute the knowledge question answering method according to any one of claims 1 to 12 based on instructions stored in the memory.
15. A computer-readable storage medium having computer program instructions stored thereon, which, when executed by a processor, implement the knowledge question and answer method as claimed in any one of claims 1 to 12.
16. A computer program product, comprising a computer program or instructions, wherein when the computer program or instructions are executed by a processor, the knowledge question and answer method according to any one of claims 1 to 12 is implemented.
Citation Information
Cited By
Intelligent Knowledge Debugging and Management System Using LLM based Semantic Block Architecture and Method Thereof
KR102993692B1