Private domain data question and answer retrieval enhancement method and device, equipment, medium and product
By performing entity recognition, relationship extraction and data vectorization on private domain data, combined with semantic reordering model and large model, the efficiency and accuracy of graphic and text retrieval in private domain knowledge base are solved, and user experience and satisfaction are improved.
Patent Information
- Application Number
- CN202510402634.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-01
- Publication Date
- 2025-07-18
AI Technical Summary
When existing graphic and text search technology is oriented towards private domain knowledge bases, the search efficiency and accuracy are limited, which affects user experience and satisfaction, and ignores the inherent connection and context between graphics and text.
The private domain data is subject to entity recognition and relationship extraction, and then data vectorization is carried out to the vector database and relational database. The entity, mapped pictures, relationships and original text candidate sets are obtained through vector search, and the semantic reordering model and large model are used to construct prompt words to answer questions.
It has achieved higher search efficiency, deeper search depth and wider search breadth, improving user experience and satisfaction.
Smart Images

Figure CN120336469A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of data retrieval, and particularly relates to a method, device, equipment, medium and product for enhancing private domain data question-answering retrieval. Background Art
[0002] With the rapid progress of information technology, network media and digital content have shown explosive growth. As an important platform for enterprises and individuals to accumulate, manage and share knowledge, the importance of private domain knowledge bases has become increasingly prominent. Private domain knowledge bases not only contain a large amount of text information, but also integrate rich multimedia content such as pictures and charts, forming a knowledge system with both pictures and texts. However, how to efficiently and accurately retrieve the information required by users from such complex and diverse private domain data has become a major challenge faced by current data retrieval technologies.
[0003] Currently, most traditional picture-text retrieval methods rely on keyword matching or text tokenization techniques. These methods perform well when dealing with structured text data, but when faced with unstructured or semi-structured picture-text mixed data in private domain knowledge bases, their retrieval efficiency and accuracy often decrease significantly. Especially when it comes to professional domain knowledge or emerging vocabulary, the limitations of traditional methods are particularly obvious. In addition, existing picture-text retrieval technologies often ignore the internal connection and context between pictures and texts, resulting in retrieval results that only focus on surface text matching and ignore the important information and deep meaning contained in images. This not only limits the depth and breadth of retrieval, but also affects user experience and satisfaction. Summary of the Invention
[0004] The purpose of the present invention is to provide a method, device, computer equipment, computer-readable storage medium and computer program product for enhancing private domain data question-answering retrieval, so as to solve the problems of limited retrieval efficiency and accuracy and affecting user experience and satisfaction existing in existing picture-text retrieval technologies when facing private domain knowledge bases.
[0005] To achieve the above purpose, the present invention adopts the following technical solutions:
[0006] In a first aspect, a method for enhancing private domain data question-answering retrieval is provided, including:
[0007] Performing entity recognition processing and relationship extraction processing on private domain data respectively to obtain a first entity, a first relationship, and a first original text on which entity recognition and relationship extraction are performed;
[0008] Perform data vectorization processing on the first entity, the first relationship, and the first original text respectively to obtain first entity vector data, first relationship vector data, and first original text vector data, and store them in different vector databases. In addition, store the mapping relationship between the first entity and the picture in the private domain data in a relational database;
[0009] Perform entity extraction processing on the input question to obtain a second entity, then perform the data vectorization processing on the second entity to obtain second entity vector data, and retrieve an entity candidate set from the entity vector database according to the second entity vector data, and query a mapping picture candidate set from the relational database according to the entity candidate set;
[0010] Perform the data vectorization processing on the input question to obtain question vector data, and retrieve a relationship candidate set from the relationship vector database and retrieve an original text candidate set from the original text vector database according to the question vector data;
[0011] Input the input question, the original text to which the entity candidate set and the relationship candidate set belong, and the original text candidate set into a semantic re-ranking model to obtain the top N candidate original texts sorted in descending order of question relevance, where N represents a positive integer greater than 2, and the top N candidate original texts are located in the original text to which they belong and the original text candidate set;
[0012] Construct a prompt according to the input question, the top N candidate original texts, and the mapping picture candidate set, and input it into a large model with semantic understanding ability to output a rich text for answering the input question.
[0013] Based on the above invention content, a new graphic retrieval scheme for a private domain knowledge base is provided. First, perform entity recognition processing and relationship extraction processing on the private domain data respectively, then perform data vectorization processing on the processing results to enrich the vector database, and store the mapping relationship between the recognized entity and the picture in a relational database. Then, based on the input question, obtain an entity candidate set, a mapping picture candidate set, a relationship candidate set, and an original text candidate set through vector retrieval, and input the input question, the original text to which the entity and relationship candidate sets belong, and the original text candidate set into a semantic re-ranking model to obtain the top N candidate original texts sorted in descending order of question relevance. Finally, construct a prompt according to the input question, the top N candidate original texts, and the mapping picture candidate set, and input it into a large model to output a rich text for answering the input question. In this way, compared with the existing graphic retrieval technology, a graphic answer result with higher retrieval efficiency, deeper retrieval depth, wider retrieval breadth, higher accuracy, and improved user experience and satisfaction can be obtained, which is convenient for practical application and promotion.
[0014] In a possible design, data vectorization processing is respectively performed on the first entity, the first relationship, and the first original text to obtain first entity vector data, first relationship vector data, and first original text vector data, including:
[0015] The M3E Chinese embedding model is used to perform data vectorization processing on the first entity, the first relationship, and the first original text respectively, to obtain first entity vector data corresponding to the first entity, first relationship vector data corresponding to the first relationship, and first original text vector data corresponding to the first original text.
[0016] In a possible design, the relational database uses a mysql database.
[0017] In a possible design, entity extraction processing is performed on the input question to obtain a second entity, including:
[0018] Construct a dictionary library based on all the entities obtained through the entity recognition processing;
[0019] Based on the forward maximum matching word segmentation algorithm, keyword matching is performed on the input question and the dictionary library, and the matching result is used as the second entity extracted from the input question.
[0020] In a possible design, the semantic re-ranking model uses the bce-reranker-base model.
[0021] In a second aspect, a private domain data question-answering retrieval enhancement device is provided, including a data preprocessing unit, a processed data storage unit, a candidate entity and picture retrieval unit, a candidate relationship and original text retrieval unit, a candidate original text ranking unit, and a large model application unit;
[0022] The data preprocessing unit is used to perform entity recognition processing and relationship extraction processing on the private domain data respectively, to obtain a first entity, a first relationship, and a first original text that has undergone entity recognition and relationship extraction;
[0023] The processed data storage unit is communicatively connected to the data preprocessing unit, and is used to perform data vectorization processing on the first entity, the first relationship, and the first original text respectively, to obtain first entity vector data, first relationship vector data, and first original text vector data and store them in different vector databases, and also store the mapping relationship between the first entity and the picture in the private domain data in a relational database;
[0024] The candidate entity and picture retrieval unit are communicatively connected to the processed data storage unit, and are used for performing entity extraction processing on the input question to obtain a second entity, then performing data vectorization processing on the second entity to obtain second entity vector data, retrieving an entity candidate set from the entity vector database according to the second entity vector data, and querying a mapped picture candidate set from the relational database according to the entity candidate set;
[0025] The candidate relationship and original text retrieval unit are communicatively connected to the processed data storage unit, and are used for performing data vectorization processing on the input question to obtain question vector data, retrieving a relationship candidate set from the relationship vector database and retrieving an original text candidate set from the original text vector database according to the question vector data;
[0026] The candidate original text sorting unit is communicatively connected to the candidate entity and picture retrieval unit and the candidate relationship and original text retrieval unit respectively, and is used for inputting the input question, the original text to which the entity candidate set and the relationship candidate set belong, and the original text candidate set into a semantic re-ranking model together, to obtain the top N candidate original texts sorted in descending order of question relevance, where N represents a positive integer greater than 2, and the top N candidate original texts are located in the original text to which they belong and the original text candidate set;
[0027] The large model application unit is communicatively connected to the candidate entity and picture retrieval unit and the candidate original text sorting unit respectively, and is used for constructing a prompt word according to the input question, the top N candidate original texts, and the mapped picture candidate set, and inputting the prompt word into a large model with semantic understanding ability, and outputting a rich text for answering the input question.
[0028] In a possible design, the candidate entity and picture retrieval unit includes a dictionary library construction subunit and a keyword matching subunit that are communicatively connected;
[0029] The dictionary library construction subunit is used for constructing a dictionary library according to all entities obtained through the entity recognition processing;
[0030] The keyword matching subunit is used for performing keyword matching on the input question and the dictionary library based on the forward maximum matching word segmentation algorithm, and using the matching result as the second entity extracted from the input question.
[0031] In a third aspect, the present invention provides a computer device, including a memory, a processor, and a transceiver that are communicatively connected in sequence, where the memory is used for storing a computer program, the transceiver is used for sending and receiving messages, and the processor is used for reading the computer program and executing the private domain data question and answer retrieval enhancement method as described in the first aspect or any possible design in the first aspect.
[0032] In a fourth aspect, the present invention provides a computer-readable storage medium, on which instructions are stored. When the instructions run on a computer, they execute the private domain data Q&A retrieval enhancement method as described in the first aspect or any possible design in the first aspect.
[0033] In a fifth aspect, the present invention provides a computer program product, including a computer program or instructions. When the computer program or the instructions are executed by a computer, they implement the private domain data Q&A retrieval enhancement method as described in the first aspect or any possible design in the first aspect.
[0034] Beneficial effects of the above solutions:
[0035] (1) The present invention creatively provides a new graphic retrieval solution for a private domain knowledge base, that is, first perform entity recognition processing and relationship extraction processing on private domain data respectively, then perform data vectorization processing on the processing results to enrich the vector database, and store the mapping relationship between the recognized entities and pictures in a relational database. Then, based on the input question, obtain an entity candidate set, a mapped picture candidate set, a relationship candidate set, and an original text candidate set through vector retrieval, and input the input question, the original text to which the entity and relationship candidate sets belong, and the original text candidate set into a semantic re-ranking model to obtain the top N candidate original texts sorted in descending order of question relevance. Finally, according to the input question, the top N candidate original texts, and the mapped picture candidate set, construct a prompt word and input it into a large model, and output a rich text for answering the input question. In this way, compared with the existing graphic retrieval technology, graphic-based answer results with higher retrieval efficiency, deeper retrieval depth, wider retrieval breadth, higher accuracy, and improved user experience and satisfaction can be obtained, which is convenient for practical application and promotion;
[0036] (2) By presenting entities and their relationships in a structured manner, more accurate retrieval information can be provided to help the private domain knowledge base better handle complex question answering tasks;
[0037] (3) By directly storing the vectors of entities and relationships in the vector library, retrieval based on entities and relationships can better answer relational questions (such as "Who is Su Shi's wife? Who is his father?", etc.) compared to directly retrieving the original text, and retains the feature of directly retrieving the original text to assist other types of retrievals. Description of the Drawings
[0038] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0039] Figure 1 It is a schematic flowchart of the method for enhancing private domain data Q&A retrieval provided by the embodiments of this application.
[0040] Figure 2 It is an example diagram of the constructed prompt words and output results provided by the embodiments of this application.
[0041] Figure 3 It is a schematic structural diagram of the device for enhancing private domain data Q&A retrieval provided by the embodiments of this application.
[0042] Figure 4 It is a schematic structural diagram of the computer device provided by the embodiments of this application. Detailed implementation manners
[0043] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the present invention in combination with the drawings and the description of the embodiments or the prior art. Obviously, the following description of the drawing structure is only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other embodiments can be obtained based on these embodiments. It should be noted here that the description of these embodiments is used to help understand the present invention, but does not constitute a limitation to the present invention.
[0044] It should be understood that although terms such as first and second may be used herein to describe various objects, these objects should not be limited by these terms. These terms are only used to distinguish one object from another. For example, the first object can be called the second object, and similarly, the second object can be called the first object, without departing from the scope of the exemplary embodiments of the present invention.
[0045] It should be understood that for the term "and / or" that may appear in this text, it is merely a relational description of associated objects, indicating that there can be three relationships. For example, A and / or B can represent three situations: A exists alone, B exists alone, or both A and B exist simultaneously. Another example, A, B, and / or C can represent any one of A, B, and C or any combination of them. For the term " / and" that may appear in this text, it is a description of another relational object, indicating that there can be two relationships. For example, A / and B can represent two situations: A exists alone or both A and B exist simultaneously. Additionally, for the character " / " that may appear in this text, it generally indicates that the associated objects before and after are in an "or" relationship.
[0046] Embodiment
[0047] As Figure 1 and Figure 2 shown, the private domain data question-answering retrieval enhancement method provided in the first aspect of this embodiment can, but is not limited to, be executed by a computer device with certain computing resources, such as a private domain knowledge base management server, a cloud server, a personal computer (Personal Computer, PC, referring to a multi-purpose computer suitable for personal use in terms of size, price, and performance; desktop computers, laptops, small laptops, tablets, and ultrabooks, etc. all belong to personal computers), a smart phone, a personal digital assistant (Personal Digital Assistant, PDA), or a wearable device and other electronic devices. As Figure 1 shown, the private domain data question-answering retrieval enhancement method can, but is not limited to, include the following steps S1 to S6.
[0048] S1. Perform entity recognition processing and relationship extraction processing on the private domain data respectively to obtain a first entity, a first relationship, and a first original text on which entity recognition and relationship extraction are performed.
[0049] In step S1, the private domain data is the unstructured or semi-structured graphic and text mixed data in the private domain knowledge base. The specific methods of the entity recognition processing and the relationship extraction processing are existing conventional methods. For example, if the first original text is "Modern Yuan Kangsheng's Calligraphy of Chen Fan's Poem, Dynasty: Modern, Dimensions: 179 cm in length and 96 cm in width; Mounted Piece: 228 cm in length and 112 cm in width, Texture: Paper, Collection Unit: Meishan Three Su Ancestral Hall Museum. It belongs to movable cultural relics, and the cultural relic classification is cultural relic calligraphy and painting", then the following first entity can be conventionally recognized: "Modern Yuan Kangsheng's Calligraphy of Chen Fan's Poem", and the following first relationship can be conventionally extracted: "The dimensions of Modern Yuan Kangsheng's Calligraphy of Chen Fan's Poem are 179 cm in length and 96 cm in width; the mounted piece is 228 cm in length and 112 cm in width".
[0050] S2. Perform data vectorization processing on the first entity, the first relationship, and the first original text respectively to obtain first entity vector data, first relationship vector data, and first original text vector data, and store them in different vector databases. In addition, store the mapping relationship between the first entity and the picture in the private domain data in a relational database.
[0051] In step S2, specifically, the M3E Chinese embedding model (abbreviation of Moka Massive Mixed Embedding, an open-source Chinese embedding model that uses advanced training methods and large-scale corpora to generate high-quality word vectors) or other vector models can be used to perform the data vectorization processing on the first entity, the first relationship, and the first original text respectively, so as to obtain first entity vector data corresponding to the first entity, first relationship vector data corresponding to the first relationship, and first original text vector data corresponding to the first original text. Then, the first entity vector data can be stored in the entity vector database, the first relationship vector data can be stored in the relationship vector database, and the first original text vector data can be stored in the original text vector database. Since there is a certain mapping relationship between the entity itself in the private domain data (i.e., unstructured or semi-structured text and image mixed data in the private knowledge base) and the picture, that is, for the entity "Modern Yuan Kangsheng's Book of Chen Fan's Poems", its corresponding picture is the photo of the cultural relic, it is also necessary to store the mapping relationship between the first entity and the picture in the private domain data in a relational database. In addition, specifically, the relational database can be but is not limited to using a mysql database.
[0052] S3. Perform entity extraction processing on the input question to obtain a second entity, then perform the data vectorization processing on the second entity to obtain second entity vector data, and retrieve an entity candidate set from the entity vector database according to the second entity vector data, and query a mapping picture candidate set from the relational database according to the entity candidate set.
[0053] In the step S3, the input question is obtained by the user's regular input. For example, an example is "What is the size of the modern calligraphy of Chen Fan's poem by Yuan Kangsheng in cultural relics?". Specifically, entity extraction processing is performed on the input question to obtain a second entity, including but not limited to the following steps: First, construct a dictionary library based on all entities obtained through the entity recognition processing; then, based on the forward maximum matching word segmentation algorithm, perform keyword matching on the input question and the dictionary library, and use the matching result as the second entity extracted from the input question. The aforementioned forward maximum matching word segmentation algorithm is a word segmentation algorithm based on a word list. The main algorithm idea is: scan the string in the sentence (in this embodiment, scan the input question) from front to back, and try to find the longer words in the "dictionary" as the result of word segmentation. Therefore, when the input question is "What is the size of the modern calligraphy of Chen Fan's poem by Yuan Kangsheng in cultural relics?", the second entity that can be extracted is "modern calligraphy of Chen Fan's poem by Yuan Kangsheng". The specific process of entity vector retrieval in the entity vector database is a conventional method, such as using the ANN vector retrieval method (Approximate Nearest Neighbor, which is a method of retrieving K vectors similar to the query vector in a given vector dataset through a certain metric, but usually only focuses on approximate nearest neighbors rather than exact nearest neighbors). The entity candidate set contains at least one third entity to be selected, such as the entity "modern calligraphy of Chen Fan's poem by Yuan Kangsheng", etc. The mapped picture candidate set contains at least one picture to be selected and having a mapping relationship with the aforementioned third entity, which can be obtained by conventional query in the relational database. In addition, specifically, the M3E Chinese embedding model or other vector models can also be used to perform the data vectorization processing on the second entity, so as to obtain the second entity vector data corresponding to the second entity.
[0054] S4. Perform the data vectorization processing on the input question to obtain question vector data, and based on the question vector data, retrieve a relationship candidate set in the relational vector database and a source text candidate set in the source text vector database.
[0055] In the step S4, specifically, the problem vector data is used as the second relationship vector data, and the relationship candidate set is retrieved from the relationship vector database according to the second relationship vector data (for example, by using the ANN vector retrieval method); and the problem vector data is used as the second original text vector data, and the original text candidate set is retrieved from the original text vector database according to the second original text vector data (for example, by using the ANN vector retrieval method). The relationship candidate set contains at least one second relationship to be selected, and the original text candidate set contains at least one second original text to be selected. In addition, the M3E Chinese embedding model or other vector models can also be specifically used to perform the data vectorization processing on the input problem to obtain the problem vector data.
[0056] S5. Input the input problem, the entity candidate set, the original text to which the relationship candidate set belongs, and the original text candidate set into the semantic re-ranking model to obtain the top N candidate original texts sorted in descending order of problem relevance, where N represents a positive integer greater than 2, and the top N candidate original texts are located in the original text to which the relationship candidate set belongs and the original text candidate set.
[0057] In the step S5, the original text to which the entity candidate set and the relationship candidate set belong refers to the original text that can identify the third entity in the entity candidate set and extract the second relationship in the relationship candidate set, and it can be obtained by conventional query. The semantic re-ranking model (Semantic Reranking Model) is an existing model that utilizes the speed and efficiency of the fast retrieval method and performs hierarchical semantic search on it. It can re-rank a set of retrieved documents to improve the relevance of the search; specifically, the semantic re-ranking model can, but is not limited to, adopt the bce-reranker-base model (i.e., a problem and recall relevance secondary ranking model specifically used in the fields of information retrieval and natural language processing. It is developed by the Beijing Academy of Artificial Intelligence and is part of the BGE series; the BGE series of models focuses on providing general embedding representations, and the bce-reranker-base model re-ranks the results of the preliminary retrieval on this basis to improve the relevance and quality of the final retrieval results). In addition, N can be exemplified as 10.
[0058] S6. Construct a prompt word according to the input problem, the top N candidate original texts, and the mapping picture candidate set, and input it into a large model with semantic understanding ability to output a rich text for answering the input problem.
[0059] In step S6, examples of the large model with semantic understanding ability include but are not limited to Baidu Wenxin ERNIE (Wenxin Yiyan) model or ChatGLM series (such as ChatGLM-6B) model, etc. The Rich Text refers to a text format that embeds multimedia elements such as formats, styles, images, and / or links in the text content. When the input question is, for example, "What is the size of the modern calligraphy of the poem by Chen Fan written by Yuan Kangsheng?", the specific constructed prompt words and output results are as Figure 2 shown, and a graphic answer result with higher retrieval efficiency, deeper retrieval depth, wider retrieval breadth, greater accuracy, and capable of improving user experience and satisfaction can be obtained.
[0060] Based on the private domain data Q&A retrieval enhancement method described in the foregoing steps S1 to S6, a new graphic retrieval scheme for the private domain knowledge base is provided. That is, the private domain data is first subjected to entity recognition processing and relationship extraction processing respectively, and then the processing results are vectorized to enrich the vector database, and the mapping relationship between the recognized entities and pictures is stored in the relational database. Then, based on the input question, entity candidate sets, mapped picture candidate sets, relationship candidate sets, and original text candidate sets are obtained through vector retrieval. The input question, the original text to which the entity and relationship candidate sets belong, and the original text candidate sets are input into the semantic re-ranking model together to obtain the top N candidate original texts sorted in descending order according to the question relevance. Finally, according to the input question, the top N candidate original texts, and the mapped picture candidate sets, prompt words are constructed and input into the large model, and a rich text for answering the input question is output. In this way, compared with the existing graphic retrieval technology, a graphic answer result with higher retrieval efficiency, deeper retrieval depth, wider retrieval breadth, greater accuracy, and capable of improving user experience and satisfaction can be obtained, which is convenient for practical application and promotion.
[0061] As Figure 3 shown, in the second aspect of this embodiment, a virtual device for implementing the private domain data Q&A retrieval enhancement method described in the first aspect is provided, including a data preprocessing unit, a processed data storage unit, a candidate entity and picture retrieval unit, a candidate relationship and original text retrieval unit, a candidate original text ranking unit, and a large model application unit;
[0062] The data preprocessing unit is used to perform entity recognition processing and relationship extraction processing on the private domain data respectively to obtain the first entity, the first relationship, and the first original text subjected to entity recognition and relationship extraction;
[0063] The processed data storage unit is communicatively connected to the data preprocessing unit, and is used to perform data vectorization processing on the first entity, the first relationship, and the first original text respectively, obtain the first entity vector data, the first relationship vector data, and the first original text vector data, and store them in different vector databases. In addition, it also stores the mapping relationship between the first entity and the picture in the private domain data in a relational database;
[0064] The candidate entity and picture retrieval unit is communicatively connected to the processed data storage unit, and is used to perform entity extraction processing on the input question to obtain a second entity, then perform the data vectorization processing on the second entity to obtain the second entity vector data, and retrieve an entity candidate set from the entity vector database according to the second entity vector data, and query a mapping picture candidate set from the relational database according to the entity candidate set;
[0065] The candidate relationship and original text retrieval unit is communicatively connected to the processed data storage unit, and is used to perform the data vectorization processing on the input question to obtain question vector data, and retrieve a relationship candidate set from the relationship vector database and a original text candidate set from the original text vector database according to the question vector data;
[0066] The candidate original text sorting unit is communicatively connected to the candidate entity and picture retrieval unit and the candidate relationship and original text retrieval unit respectively, and is used to input the input question, the original text to which the entity candidate set and the relationship candidate set belong, and the original text candidate set into a semantic re-ranking model together, and obtain the top N candidate original texts sorted in descending order of question relevance, where N represents a positive integer greater than 2, and the top N candidate original texts are located in the original text to which they belong and the original text candidate set;
[0067] The large model application unit is communicatively connected to the candidate entity and picture retrieval unit and the candidate original text sorting unit respectively, and is used to construct a prompt according to the input question, the top N candidate original texts, and the mapping picture candidate set, and input it into a large model with semantic understanding ability, and output a rich text for answering the input question.
[0068] In a possible design, the candidate entity and picture retrieval unit includes a dictionary library construction subunit and a keyword matching subunit that are communicatively connected;
[0069] The dictionary library construction subunit is used to construct a dictionary library according to all the entities obtained through the entity recognition processing;
[0070] The keyword matching subunit is configured to perform keyword matching on the input question and the dictionary library based on the forward maximum matching word segmentation algorithm, and use the matching result as the second entity extracted from the input question.
[0071] For the working process, working details and technical effects of the foregoing apparatus provided in the second aspect of this embodiment, reference may be made to the private domain data question-answering retrieval enhancement method described in the first aspect, which will not be elaborated herein.
[0072] As Figure 4 shown, a computer device for executing the private domain data question-answering retrieval enhancement method described in the first aspect is provided in the third aspect of this embodiment, including a memory, a processor, and a transceiver that are communicatively connected in sequence. Among them, the memory is used to store a computer program, the transceiver is used to send and receive messages, and the processor is used to read the computer program and execute the private domain data question-answering retrieval enhancement method described in the first aspect. Specifically, by way of example, the memory may include, but is not limited to, a random access memory (RAM), a read-only memory (ROM), a flash memory, a first-in-first-out memory (FIFO), and / or a first-in-last-out memory (FILO), etc.; the processor may include, but is not limited to, a microprocessor of the STM32F105 series. In addition, the computer device may further include, but is not limited to, a power module, a display screen, and other necessary components.
[0073] For the working process, working details and technical effects of the foregoing computer device provided in the third aspect of this embodiment, reference may be made to the private domain data question-answering retrieval enhancement method described in the first aspect, which will not be elaborated herein.
[0074] A computer-readable storage medium storing instructions including the private domain data question-answering retrieval enhancement method described in the first aspect is provided in the fourth aspect of this embodiment, that is, instructions are stored on the computer-readable storage medium, and when the instructions run on a computer, the private domain data question-answering retrieval enhancement method described in the first aspect is executed. Among them, the computer-readable storage medium refers to a carrier for storing data, and may include, but is not limited to, a floppy disk, an optical disk, a hard disk, a flash memory, a USB flash drive, and / or a memory stick, etc. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices.
[0075] For the working process, working details and technical effects of the aforementioned computer-readable storage medium provided in the fourth aspect of this embodiment, reference may be made to the private domain data Q&A retrieval enhancement method described in the first aspect, which will not be elaborated here.
[0076] In the fifth aspect of this embodiment, a computer program product is provided, including a computer program or instruction, and the computer program or the instruction, when executed by a computer, implements the private domain data Q&A retrieval enhancement method described in the first aspect. Among them, the computer may be a general-purpose computer, a special-purpose computer, a computer network or other programmable devices.
[0077] Finally, it should be noted that the above are only the preferred embodiments of the present invention and are not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A method for enhancing private domain data Q&A retrieval, characterized in that Including: Performing entity recognition processing and relationship extraction processing on private domain data respectively to obtain a first entity, a first relationship, and a first original text subjected to entity recognition and relationship extraction; Performing data vectorization processing on the first entity, the first relationship, and the first original text respectively to obtain first entity vector data, first relationship vector data, and first original text vector data, and storing them in different vector databases, and also storing the mapping relationship between the first entity and the picture in the private domain data in a relational database; Performing entity extraction processing on the input question to obtain a second entity, then performing the data vectorization processing on the second entity to obtain second entity vector data, retrieving an entity candidate set from the entity vector database according to the second entity vector data, and querying a mapping picture candidate set from the relational database according to the entity candidate set; Performing the data vectorization processing on the input question to obtain question vector data, and retrieving a relationship candidate set from the relationship vector database and an original text candidate set from the original text vector database according to the question vector data; Inputting the input question, the original text to which the entity candidate set and the relationship candidate set belong, and the original text candidate set into a semantic re-ranking model to obtain the top N candidate original texts sorted in descending order of question relevance, where N represents a positive integer greater than 2, and the top N candidate original texts are located in the original text to which they belong and the original text candidate set; Constructing a prompt word according to the input question, the top N candidate original texts, and the mapping picture candidate set, and inputting it into a large model with semantic understanding ability to output a rich text for answering the input question; 2. The enhanced method for private domain data question-answer retrieval according to claim 1, wherein Performing data vectorization processing on the first entity, the first relationship, and the first original text respectively to obtain first entity vector data, first relationship vector data, and first original text vector data, including: Using the M3E Chinese embedding model to perform data vectorization processing on the first entity, the first relationship, and the first original text respectively to obtain first entity vector data corresponding to the first entity, first relationship vector data corresponding to the first relationship, and first original text vector data corresponding to the first original text; 3. The enhanced method for private domain data Q&A retrieval according to claim 1, wherein The relational database uses a mysql database.
4. The enhanced method for private domain data Q&A retrieval according to claim 1, wherein Performing entity extraction processing on the input question to obtain a second entity, including: Constructing a dictionary library according to all entities obtained through the entity recognition processing; Based on the forward maximum matching word segmentation algorithm, performing keyword matching on the input question and the dictionary library, and taking the matching result as the second entity extracted from the input question; 5. The method for enhancing private domain data Q&A retrieval according to claim 1, characterized in that The semantic re-ranking model uses the bce-reranker-base model.
6. A private domain data Q&A retrieval enhancement device, characterized in that, Including a data preprocessing unit, a processed data storage unit, a candidate entity and picture retrieval unit, a candidate relationship and original text retrieval unit, a candidate original text sorting unit, and a large model application unit; The data preprocessing unit is used to perform entity recognition processing and relationship extraction processing on private domain data respectively to obtain a first entity, a first relationship, and a first original text subjected to entity recognition and relationship extraction; The processed data storage unit is communicatively connected to the data preprocessing unit, and is used to perform data vectorization processing on the first entity, the first relationship, and the first original text respectively, obtain the first entity vector data, the first relationship vector data, and the first original text vector data and store them in different vector databases, and also store the mapping relationship between the first entity and the picture in the private domain data in a relational database; The candidate entity and picture retrieval unit is communicatively connected to the processed data storage unit, and is used to perform entity extraction processing on the input question to obtain a second entity, then perform the data vectorization processing on the second entity to obtain the second entity vector data, and retrieve an entity candidate set from the entity vector database according to the second entity vector data, and query a mapping picture candidate set from the relational database according to the entity candidate set; The candidate relationship and original text retrieval unit is communicatively connected to the processed data storage unit, and is used to perform the data vectorization processing on the input question to obtain question vector data, and retrieve a relationship candidate set from the relationship vector database and retrieve an original text candidate set from the original text vector database according to the question vector data; The candidate original text sorting unit is communicatively connected to the candidate entity and picture retrieval unit and the candidate relationship and original text retrieval unit respectively, and is used to input the input question, the original text to which the entity candidate set and the relationship candidate set belong, and the original text candidate set into a semantic re-ranking model together, and obtain the top N candidate original texts sorted in descending order of question relevance, where N represents a positive integer greater than 2, and the top N candidate original texts are located in the original text to which they belong and the original text candidate set; The large model application unit is communicatively connected to the candidate entity and picture retrieval unit and the candidate original text sorting unit respectively, and is used to construct a prompt according to the input question, the top N candidate original texts, and the mapping picture candidate set and input it into a large model with semantic understanding ability, and output a rich text for answering the input question.
7. The private domain data Q&A retrieval enhancement device according to claim 6, wherein The candidate entity and picture retrieval unit includes a dictionary library construction subunit and a keyword matching subunit that are communicatively connected; The dictionary library construction subunit is used to construct a dictionary library according to all the entities obtained through the entity recognition processing; The keyword matching subunit is used to perform keyword matching on the input question and the dictionary library based on the forward maximum matching word segmentation algorithm, and use the matching result as the second entity extracted from the input question.
8. A computer device, characterized in that, It includes a memory, a processor, and a transceiver that are communicatively connected in sequence. Among them, the memory is used to store a computer program, the transceiver is used to send and receive messages, and the processor is used to read the computer program and execute the private domain data question-answering retrieval enhancement method according to any one of claims 1 to 5.
9. A computer-readable storage medium, characterized in that , An instruction is stored on the computer-readable storage medium, and when the instruction runs on the computer, it executes the private domain data question-answering retrieval enhancement method according to any one of claims 1 to 5.
10. A computer program product, comprising a computer program or instructions, characterized in that, The computer program or the instructions, when executed by a computer, implement the method for enhancing private domain data Q&A retrieval as described in any one of claims 1 to 5.