Archive information question and answer method and system
By combining the mixed method of graph database and keyword retrieval, high-dimensional vector matching and comprehensive scoring, the problems of inefficient archival information retrieval and inaccurate results are solved, and efficient and accurate archival information acquisition and answer generation are achieved.
Patent Information
- Application Number
- CN202510296035.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-13
- Publication Date
- 2025-06-27
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
In the prior art, the search efficiency of archive information is inefficient and the results are inaccurate, making it difficult to obtain archive information efficiently and accurately.
Using a mixed method of graph database and keyword retrieval, a problem answer is generated by converting archive query questions into high-dimensional vectors, vector matching and keyword retrieval, comprehensive scoring and filtering.
It improves the efficiency and accuracy of obtaining archive information and enhances the accuracy and efficiency of answer generation.
Smart Images

Figure CN120216728A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and particularly to a method and system for answering questions about archival information. Background Art
[0002] With the continuous increase in the number of archives, how to efficiently and accurately retrieve and utilize this information and intelligently generate answers to users' questions about archives has become an urgent problem to be solved. Traditional information retrieval methods, such as keyword-based retrieval, have problems such as low retrieval efficiency and inaccurate results. How to improve the efficiency and accuracy of obtaining archival information has become an urgent problem in answering questions about archival information. Summary of the Invention
[0003] The present invention provides a method and system for answering questions about archival information to solve the defect in the prior art that the efficiency and accuracy of obtaining archival information cannot be improved, and to achieve the improvement of the efficiency and accuracy of obtaining archival information.
[0004] The present invention provides a method for answering questions about archival information, including the following steps: in response to receiving an archival query question, obtaining an archival information retrieval result corresponding to the archival query question through graph database retrieval and keyword retrieval; inputting the archival query question and the archival information retrieval result into an answer generation model to generate an answer to the archival query question.
[0005] According to the method for answering questions about archival information provided by the present invention, the obtaining of the archival information retrieval result corresponding to the archival query question through graph database retrieval and keyword retrieval includes: converting the archival query question into a high-dimensional vector in a preset dimension; performing vector matching in a knowledge graph index structure based on the high-dimensional vector to obtain a first retrieval result with a semantic similarity greater than a preset threshold to the archival query question; filtering the first retrieval result to obtain a second retrieval result; performing keyword retrieval on the second retrieval result based on the archival query question to obtain a third retrieval result; comprehensively scoring the third retrieval result based on the vector similarity and term frequency similarity between the archival query question and the third retrieval result, and sorting the third retrieval result in descending order according to the result of the comprehensive scoring to obtain the archival information retrieval result.
[0006] A method for answering questions about archival information provided by the present invention. Before performing vector matching in the knowledge graph index structure based on the high-dimensional vector, the method further includes: preprocessing the pre-stored archival files; wherein, the preprocessing includes at least one of text cleaning, word segmentation, and stop word removal; segmenting the preprocessed archival files into text blocks, extracting entity nodes and entity node description information from the text blocks by using a large model, and constructing an archival knowledge graph according to the relationship information between the entity nodes; wherein, the entity nodes include at least one of name, position, work experience, and educational background; vectorizing the entity node description information to obtain entity node description vectors; performing community detection and clustering on the entity nodes based on the archival knowledge graph to obtain a community report; wherein, the community report includes information such as community description, community importance, community associated entities, and internal relationships within the community; constructing the knowledge graph index structure based on the entity nodes, the entity node description vectors, and the community report.
[0007] A method for answering questions about archival information provided by the present invention. The method further includes: using an intelligent adaptive learning engine to optimize and adjust the retrieval parameters of the graph database retrieval and the keyword retrieval and the model weights of the answer generation model according to user satisfaction and the relevance between the question answer and the archival query question; wherein, the intelligent adaptive learning engine adopts a deep reinforcement learning framework.
[0008] A method for answering questions about archival information provided by the present invention. The reward function of the intelligent adaptive learning engine is expressed as: R(s,a)=r(positive)+r(ctr)+r(dwell)+r(negative) Wherein, R(s,a) represents the comprehensive reward value, r(positive) represents the positive reward value; r(ctr) represents the reward value of the query result click-through rate, r(dwell) represents the reward value of the dwell time, and r(negative) represents the negative reward value.
[0009] A method for answering questions about archival information provided by the present invention. The method further includes: using the intelligent adaptive learning engine to dynamically adjust the retrieval parameters of the graph database retrieval and the keyword retrieval and the model weights of the answer generation model based on a pre-constructed user profile; wherein, the user profile is constructed by the intelligent adaptive learning engine based on at least one type of data among user historical query records, click behaviors, and preference settings.
[0010] A method for answering questions about archival information provided by the present invention. The method further includes: using the answer generation model to output personalized archival recommendations and query suggestions based on the user profile.
[0011] A method for answering questions about archival information provided by the present invention, generating the question answers for the archival query questions includes: generating the question answers in a preset style based on the user's questioning manner and preferences; wherein, the user's questioning manner and preferences are learned by the intelligent adaptive learning engine based on the user profile and fed back to the answer generation model.
[0012] A method for answering questions about archival information provided by the present invention, the method further includes: controlling the user's access to and operation on archival information according to the authority scope of users at different levels; and / or, adopting encryption technology and data desensitization processing during the transmission and storage of archival data.
[0013] The present invention also provides an archival information question-answering system, including the following modules: a retrieval module, configured to: in response to receiving an archival query question, obtain an archival information retrieval result corresponding to the archival query question through graph database retrieval and keyword retrieval; an answer generation module, configured to: input the archival query question and the archival information retrieval result into an answer generation model to generate the question answers for the archival query question.
[0014] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein when the processor executes the computer program, the method for answering questions about archival information as described in any one of the above is implemented.
[0015] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the method for answering questions about archival information as described in any one of the above is implemented.
[0016] The present invention also provides a computer program product, including a computer program, and when the computer program is executed by a processor, the method for answering questions about archival information as described in any one of the above is implemented.
[0017] The method and system for answering questions about archival information provided by the present invention, by in response to receiving an archival query question, obtaining an archival information retrieval result corresponding to the archival query question through graph database retrieval and keyword retrieval, and inputting the archival query question and the archival information retrieval result into an answer generation model to generate the question answers for the archival query question, improve the efficiency and accuracy of answer generation in archival information query. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] To more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0019] Figure 1 It is a schematic flowchart of the file information question-answering method provided by the present invention.
[0020] Figure 2 It is a schematic structural diagram of the file information question-answering system provided by the present invention.
[0021] Figure 3 It is a schematic structural diagram of the electronic device provided by the present invention. Specific Embodiments
[0022] To make the objectives, technical solutions, and advantages of the present invention clearer, the following will clearly and completely describe the technical solutions in the present invention with reference to the accompanying drawings in the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art without creative efforts based on the embodiments in the present invention fall within the protection scope of the present invention.
[0023] Figure 1 It is a schematic flowchart of the file information question-answering method provided by the present invention. As Figure 1 shown, the method includes: Step S1: In response to receiving a file query question, obtain a file information retrieval result corresponding to the file query question through graph database retrieval and keyword retrieval.
[0024] In the description of the embodiments of the present invention, the cadre file information is taken as an example for introduction, and it is not used to limit the type of the file.
[0025] Receive a file query question, which is used to obtain file-related information. Obtain a file information retrieval result corresponding to the file query question through a hybrid retrieval method of graph database retrieval and keyword retrieval. Among them, the graph database can be a database constructed based on a file knowledge graph.
[0026] Among them, when constructing the file knowledge graph, using graph database technology, key information such as basic information, job changes, and reward and punishment records in the file is structurally stored in the form of entity relationships to form an easily retrievable and understandable file knowledge graph.
[0027] Step S2: Input the file query question and the file information retrieval result into an answer generation model to generate an answer to the file query question.
[0028] The file information retrieval result is information related to the file query question. Inputting the file query question and the file information retrieval result into the answer generation model together can improve the accuracy and efficiency of question answer generation. The file information question-answering method provided by the present invention belongs to answer generation based on RAG (Retrieval-Augmented Generation) technology.
[0029] Among them, the answer generation model can adopt an existing large model and can be fine-tuned before use. For example, file query question samples and file information retrieval result samples can be used as the input of the answer generation model, and question answer samples can be used as output labels to further train the large model to obtain the answer generation model.
[0030] The file information question-answering method provided by the present invention, by responding to receiving a file query question, obtains a file information retrieval result corresponding to the file query question through graph database retrieval and keyword retrieval, inputs the file query question and the file information retrieval result into an answer generation model, and generates an answer to the file query question, improving the efficiency and accuracy of answer generation in file information query.
[0031] According to a file information question-answering method provided by the present invention, the obtaining of the file information retrieval result corresponding to the file query question through graph database retrieval and keyword retrieval includes: converting the file query question into a high-dimensional vector of a preset dimension; performing vector matching in a knowledge graph index structure based on the high-dimensional vector to obtain a first retrieval result whose semantic similarity to the file query question is greater than a preset threshold; filtering the first retrieval result to obtain a second retrieval result; performing keyword retrieval on the second retrieval result based on the file query question to obtain a third retrieval result; comprehensively scoring the third retrieval result based on the vector similarity and word frequency similarity between the file query question and the third retrieval result, and sorting the third retrieval result in descending order according to the result of the comprehensive scoring to obtain the file information retrieval result.
[0032] When retrieving the file information retrieval result corresponding to the file query problem through graph database retrieval and keyword retrieval, first convert the file query problem into a high-dimensional vector, and then use the vector retrieval algorithm to quickly match the documents or answers semantically similar to the file query problem in the knowledge graph index structure to obtain the first retrieval result; filter the first retrieval result, such as filtering out the results with low relevance, to obtain a set of highly relevant documents, that is, the second retrieval result; based on the file query problem, perform keyword retrieval on the second retrieval result, and obtain the refined result set by precisely matching keywords, that is, the third retrieval result; apply a hybrid scoring mechanism to score the refined result set, comprehensively consider the vector similarity and word frequency similarity, and sort them in descending order according to the comprehensive score to obtain the final file information retrieval result. Among them, when comprehensively scoring the third retrieval result based on the vector similarity and word frequency similarity between the file query problem and the third retrieval result, the first score can be obtained according to the vector similarity between the file query problem and the third retrieval result, the second score can be obtained according to the number of keywords matched between the file query problem and the third retrieval result, and the first score and the second score are weighted and summed to obtain the result of the comprehensive score.
[0033] The present invention designs and implements a hybrid retrieval mechanism, combines the efficiency of vector retrieval with the accuracy of graph database retrieval, and designs a hybrid retrieval algorithm based on context understanding. This algorithm can intelligently identify the intention of the user's question and accurately locate relevant information in the knowledge graph. At the same time, by introducing semantic similarity calculation, the accuracy and relevance of the retrieval are further improved.
[0034] The file information answering method provided by the present invention converts the file query problem into a high-dimensional vector of a preset dimension, performs vector matching in the knowledge graph index structure based on the high-dimensional vector to obtain the first retrieval result whose semantic similarity to the file query problem is greater than the preset threshold, filters the first retrieval result to obtain the second retrieval result, performs keyword retrieval on the second retrieval result based on the file query problem to obtain the third retrieval result, comprehensively scores the third retrieval result based on the vector similarity and word frequency similarity between the file query problem and the third retrieval result, sorts the third retrieval result in descending order according to the result of the comprehensive score to obtain the file information retrieval result, improves the accuracy and acquisition efficiency of the file information retrieval result, and thus further improves the efficiency and accuracy of answer generation in file information query.
[0035] A method for answering questions about archival information provided by the present invention. Before performing vector matching in the knowledge graph index structure based on the high-dimensional vector, the method further includes: preprocessing the pre-stored archival files; wherein, the preprocessing includes at least one of text cleaning, word segmentation, and stop word removal; segmenting the preprocessed archival files into text blocks, using a large model to extract entity nodes and entity node description information from the text blocks, and constructing an archival knowledge graph according to the relationship information between the entity nodes; wherein, the entity nodes include at least one of name, position, work experience, and educational background; performing vectorization processing on the entity node description information to obtain entity node description vectors; performing community detection and clustering on the entity nodes based on the archival knowledge graph to obtain a community report; wherein, the community report includes information such as community description, community importance, community associated entities, and internal relationships within the community; constructing the knowledge graph index structure based on the entity nodes, the entity node description vectors, and the community report.
[0036] Before performing vector matching in the knowledge graph index structure based on the high-dimensional vector, it is necessary to pre-construct the knowledge graph index structure. The construction process of the knowledge graph index structure includes: Preprocessing the pre-stored archival files; wherein, the preprocessing includes at least one of text cleaning, word segmentation, and stop word removal.
[0037] Exemplarily, Python regular expressions, NLTK (Natural Language Toolkit), or SpaCy natural language processing libraries can be used to clean the text, specifically including: removing special characters and punctuation marks, removing HTML tags, handling case, removing extra whitespace, removing or replacing numbers, etc.
[0038] Exemplarily, jieba word segmentation, Stanford NLP, or SpaCy can be used for word segmentation. The specific methods can be rule-based word segmentation (such as maximum forward matching, minimum segmentation, etc. algorithms), statistics-based word segmentation (such as hidden Markov models, conditional random fields, etc.), or deep learning word segmentation (such as LSTM, Transformer, etc.).
[0039] Exemplarily, stop words can be directly removed through natural language processing libraries such as NLTK and SpaCy, or a custom stop word list can be defined according to specific situations.
[0040] The preprocessed archival files are segmented into text chunks and stored in a database. An entity node and entity node description information are extracted from the text chunks by using a large model, and an archival knowledge graph is constructed according to the relationship information between the entity nodes; wherein, the entity nodes include at least one of name, position, work experience, and educational background. The entity node description information is the specific content information of the entity node.
[0041] Exemplarily, large models such as BERT, GPT series or ERNIE can be fine-tuned to achieve the extraction of entity nodes and entity node description information.
[0042] The entity node description information is vectorized to obtain an entity node description vector.
[0043] Community detection and clustering are performed on all entity nodes to obtain a community report; the community report includes information such as community description, community importance, community-related entities, and internal relationships within the community.
[0044] Based on the entity nodes, entity node description vectors, and community reports, a knowledge graph index structure is constructed. The knowledge graph index structure includes the corresponding relationships between entity nodes, entity node description vectors, and community reports.
[0045] The archival information question-answering method provided by the present invention preprocesses the pre-stored archival files, segments the preprocessed archival files into text chunks, extracts entity nodes and entity node description information from the text chunks by using a large model, constructs an archival knowledge graph according to the relationship information between the entity nodes, vectorizes the entity node description information to obtain an entity node description vector, performs community detection and clustering on the entity nodes based on the archival knowledge graph to obtain a community report, and constructs a knowledge graph index structure based on the entity nodes, entity node description vectors, and community reports, improving the accuracy and comprehensiveness of the knowledge graph index structure and facilitating further improvement of the accuracy of answer generation in archival information queries.
[0046] According to an archival information question-answering method provided by the present invention, the method further includes: using an adaptive learning engine to optimize and adjust the retrieval parameters of the graph database retrieval and keyword retrieval and the model weights of the answer generation model according to user satisfaction and the relevance between the question answer and the archival query question; wherein, the adaptive learning engine adopts a deep reinforcement learning framework.
[0047] Introduce an intelligent adaptive learning engine, which continuously optimizes the retrieval strategy and answer generation model based on user feedback and historical query data. The specific implementation includes: using reinforcement learning technology to optimize and adjust the retrieval parameters of graph database retrieval and keyword retrieval, as well as the model weights of the answer generation model, according to user satisfaction and the relevance between the question answer and the archive query question. When optimizing and adjusting the retrieval parameters of graph database retrieval and keyword retrieval, as well as the model weights of the answer generation model, according to user satisfaction and the relevance between the question answer and the archive query question, it can be executed according to a preset time period, so as to continuously adjust the retrieval parameters and model weights.
[0048] The intelligent adaptive learning engine adopts a deep reinforcement learning (DRL) framework. Among them, the archive information Q&A system for implementing the archive information Q&A method serves as an agent, user satisfaction and query result accuracy serve as reward signals, the retrieval parameters and the model weights serve as an action space, and historical query data and user feedback constitute the environmental state.
[0049] An instant response mechanism for user feedback can be established. Users can evaluate the query results or provide correction information, and these feedbacks are directly used for the next iteration optimization of the model, forming a closed-loop feedback system. In the model performance monitoring and evaluation, A / B testing and multi-dimensional performance metric monitoring (such as accuracy, recall rate, F1 score, user satisfaction, etc.) can be implemented to regularly evaluate the model performance and timely discover and solve problems.
[0050] At the same time, by continuously learning new archive information and user query patterns, the archive knowledge graph and Q&A library are continuously updated to ensure the continuous progress and adaptability of the system. Combining the characteristics of natural language processing technology and graph database, entities, relationships, and attributes in the knowledge graph are updated regularly or in real time to ensure the timeliness and accuracy of the knowledge base.
[0051] Incremental learning of archives can be achieved through an incremental learning mechanism. Whenever new archives are added or existing archives are updated, the system can automatically identify and learn new information without retraining the model, improving the learning efficiency.
[0052] The archive information Q&A method provided by the present invention further improves the accuracy of answer generation in archive information query by using an intelligent adaptive learning engine to optimize and adjust the retrieval parameters of graph database retrieval and keyword retrieval, as well as the model weights of the answer generation model, and the intelligent adaptive learning engine adopts a deep reinforcement learning framework.
[0053] A method for answering questions about file information provided by the present invention, the reward function of the intelligent adaptive learning engine is expressed as: R(s,a)=r(positive)+r(ctr)+r(dwell)+r(negative) Among them, R(s,a) represents the comprehensive reward value, r(positive) represents the positive reward value; r(ctr) represents the reward value of the click-through rate of the query result, r(dwell) represents the reward value of the dwell time, and r(negative) represents the negative reward value.
[0054] Design a multi-level reward mechanism, including direct rewards (such as explicit positive reviews from users) and indirect rewards (such as implicit feedback such as click-through rate of query results and dwell time). Give positive rewards for highly accurate answers and answers with user positive reviews, and give negative rewards for answers with low satisfaction or incorrect answers, so as to encourage the model to continuously optimize.
[0055] In the expression of the comprehensive reward function, the positive reward value r(positive) and the negative reward value r(negative) belong to direct rewards, and the query result click-through rate reward value r(ctr) and the dwell time reward value r(dwell) belong to indirect rewards.
[0056] Among them, the positive reward r(positive) can take a value of 1 or a higher value, which is adjusted according to the specific application scenario. The query result click-through rate reward value , where CTR represents the click-through rate of the query result, is the adjustment coefficient. The dwell time reward value , where DwellTime represents the dwell time, is the adjustment coefficient. Using the penalty mechanism, for answers with low satisfaction or incorrect answers, give negative rewards (or adjusted according to the degree of error).
[0057] Strategic gradient methods or algorithms such as Q-learning can be used. Through continuous trial and error and learning, iteratively update the retrieval parameters and the model weights of the answer generation model to maximize the long-term cumulative reward.
[0058] The method for answering questions about file information provided by the present invention improves the accuracy of the comprehensive reward value by setting the comprehensive reward value of the intelligent adaptive learning engine as the sum of the positive reward value, the query result click-through rate reward value, the dwell time reward value and the negative reward value, which is beneficial to further improving the accuracy of answer generation in file information query.
[0059] A method for answering questions about archival information provided by the present invention, the method further includes: using the intelligent adaptive learning engine to dynamically adjust the retrieval parameters of the graph database retrieval and the keyword retrieval and the model weights of the answer generation model based on a pre-constructed user profile; wherein, the user profile is constructed by the intelligent adaptive learning engine based on at least one type of data among user historical query records, click behaviors, and preference settings.
[0060] The intelligent adaptive learning engine dynamically adjusts the retrieval parameters of the graph database retrieval and the keyword retrieval and the model weights of the answer generation model based on the user profile.
[0061] For example, according to the user profile, dynamically adjust the weight distribution during the graph database retrieval and the keyword retrieval. For users who frequently query a certain type of information, increase the weight of this type of information in the retrieval results to improve the personalization degree of the retrieval. The optimized parameters output by the intelligent adaptive learning engine are directly applied to the hybrid retrieval algorithm to improve the retrieval efficiency and accuracy. At the same time, the quality of the retrieval results is used as feedback and input into the intelligent adaptive learning engine to form a virtuous cycle.
[0062] For another example, by analyzing a large number of user query logs, use sequence models (such as LSTM, Transformer) to capture the changing trends of user query patterns, predict future query needs, and adjust the model weights of the answer generation model according to future query needs to adapt to future queries.
[0063] Among them, the intelligent adaptive learning engine constructs a user profile based on multi-dimensional data such as user historical query records, click behaviors, and preference settings. The user profile includes information such as interest preferences, query habits, and job relevance.
[0064] The method for answering questions about archival information provided by the present invention further improves the accuracy of answer generation in archival information queries by using the intelligent adaptive learning engine to dynamically adjust the retrieval parameters of the graph database retrieval and the keyword retrieval and the model weights of the answer generation model based on the user profile.
[0065] A method for answering questions about archival information provided by the present invention, the method further includes: using the answer generation model to output personalized archival recommendations and query suggestions based on the user profile.
[0066] The answer generation model can not only generate answers to questions corresponding to archival query questions, but also output personalized archival recommendations and query suggestions based on a pre-constructed user profile. The intelligent adaptive learning engine constructs a user profile based on multi-dimensional data such as user historical query records, click behaviors, and preference settings. The user profile includes information such as interest preferences, query habits, and job relevance, providing a basis for personalized recommendations and query optimization.
[0067] The file information Q&A method provided by the present invention uses an answer generation model to output personalized file recommendations and query suggestions based on a pre-constructed user profile, obtaining richer personalized query results and improving the user experience.
[0068] According to a file information Q&A method provided by the present invention, generating the question answer of the file query question includes: generating the question answer in a preset style based on the user's question-asking manner and preferences; wherein, the user's question-asking manner and preferences are learned by the intelligent adaptive learning engine based on the user profile and fed back to the answer generation model.
[0069] The intelligent adaptive learning engine learns the user's question-asking manner and preferences based on the user profile and feeds them back to the answer generation model. The answer generation model generates the question answer in a style that meets the user's expectations based on the user's question-asking manner and preferences. By learning the user's question-asking manner and preferences, the intelligent adaptive learning engine can guide the answer generation model to generate a response style that better meets the user's expectations and enhance the interaction experience.
[0070] The file information Q&A method provided by the present invention further improves the user experience by generating the question answer in a style that meets the user's expectations based on the user's question-asking manner and preferences.
[0071] According to a file information Q&A method provided by the present invention, the method further includes: controlling the user's access to and operation on file information according to the permission scope of users at different levels; and / or, adopting encryption technology and data desensitization processing during the transmission and storage of file data.
[0072] Strengthening permission management and data security to achieve a perfect permission management function. Among them, according to the permission scope of users at different levels, strictly controlling the user's access to and operation on file information; in addition, the security and privacy protection of file data during transmission and storage can be ensured by adopting advanced encryption technology and data desensitization processing.
[0073] The file information Q&A method provided by the present invention improves the system security by controlling the user's access to and operation on file information according to the permission scope of users at different levels and / or adopting encryption technology and data desensitization processing during the transmission and storage of file data.
[0074] The file information Q&A method provided by the present invention is based on the RAG technology and deeply integrates deep learning, natural language processing, graph database, and adaptive learning technologies, aiming to solve the complexity and diversity problems in file queries such as cadre files and improve the intelligent level of file management and query. The present invention combines the adaptive learning technology with the RAG technology to achieve the intelligence, personalization, and security of cadre file management and query. This innovation not only improves the efficiency and accuracy of information query but also provides more convenient and efficient support for work such as cadre selection, appointment, and assessment.
[0075] The file information Q&A system provided by the present invention will be described below. The file information Q&A system described below can be mutually corresponding and referred to the file information Q&A method described above.
[0076] Figure 2 It is a schematic structural diagram of the file information Q&A system provided by the present invention. As Figure 2 shown, the system includes a retrieval module 10 and an answer generation module 20, where: The retrieval module 10 is used for: in response to receiving a file query question, retrieving file information retrieval results corresponding to the file query question through graph database retrieval and keyword retrieval; The answer generation module 20 is used for: inputting the file query question and the file information retrieval results into an answer generation model to generate an answer to the file query question.
[0077] The file information Q&A system provided by the present invention, by responding to receiving a file query question, retrieving file information retrieval results corresponding to the file query question through graph database retrieval and keyword retrieval, and inputting the file query question and the file information retrieval results into an answer generation model to generate an answer to the file query question, improves the efficiency and accuracy of answer generation in file information query.
[0078] Figure 3 Illustrates a schematic structural diagram of an electronic device. As Figure 3 shown, the electronic device may include: a processor 310, a communication interface 320, a memory 330, and a communication bus 340. Among them, the processor 310, the communication interface 320, and the memory 330 communicate with each other through the communication bus 340. The processor 310 can call the logical instructions in the memory 330 to execute the file information Q&A method, which includes: in response to receiving a file query question, retrieving file information retrieval results corresponding to the file query question through graph database retrieval and keyword retrieval; inputting the file query question and the file information retrieval results into an answer generation model to generate an answer to the file query question.
[0079] In addition, when the logical instructions in the above-mentioned memory 330 are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.
[0080] On the other hand, the present invention also provides a computer program product. The computer program product includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the file information Q&A method provided by the above-mentioned various methods. The method includes: in response to receiving a file query question, obtaining a file information retrieval result corresponding to the file query question through graph database retrieval and keyword retrieval; inputting the file query question and the file information retrieval result into an answer generation model to generate an answer to the file query question.
[0081] In yet another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it is implemented to execute the file information Q&A method provided by the above-mentioned various methods. The method includes: in response to receiving a file query question, obtaining a file information retrieval result corresponding to the file query question through graph database retrieval and keyword retrieval; inputting the file query question and the file information retrieval result into an answer generation model to generate an answer to the file query question.
[0082] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without creative labor.
[0083] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0084] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for questioning and answering archival information, characterized in that: include: In response to receiving the archive query question, obtaining archive information retrieval results corresponding to the archive query question through graph database retrieval and keyword retrieval; The archive query question and the archive information retrieval result are input into the answer generation model to generate the answer to the archive query question.
2. The archival information question-and-answer method according to claim 1, characterized in that: The obtaining of archive information retrieval results corresponding to the archive query question through graph database retrieval and keyword retrieval includes: Converting the archive query question into a high-dimensional vector of a preset dimension; Performing vector matching in the knowledge graph index structure based on the high-dimensional vector, obtaining a first search result whose semantic similarity with the archive query question is greater than a preset threshold; Filtering the first search result to obtain a second search result; Based on the archive query question, perform a keyword search on the second search result to obtain a third search result; The third search result is comprehensively scored based on the vector similarity and word frequency similarity between the archive query question and the third search result, and the third search result is arranged in descending order according to the result of the comprehensive score to obtain the archive information retrieval result.
3. The archival information question-and-answer method according to claim 2, characterized in that: Before performing vector matching in the knowledge graph index structure based on the high-dimensional vector, the method further includes: Preprocessing the pre-stored archive files; wherein the preprocessing includes at least one of text cleaning, word segmentation, and stop word removal; The preprocessed archive files are divided into text blocks, entity nodes and entity node description information are extracted from the text blocks using a large model, and an archive knowledge graph is constructed according to the relationship information between the entity nodes; wherein the entity nodes include at least one of name, position, work experience, and educational background; Vectorizing the entity node description information to obtain an entity node description vector; Performing community detection and clustering on the entity nodes based on the archive knowledge graph to obtain a community report; wherein the community report includes information on community description, community importance, community-related entities, and community internal relationships; The knowledge graph index structure is constructed based on the entity nodes, the entity node description vectors and the community reports.
4. The archival information question-and-answer method according to claim 2, characterized in that: The method further comprises: By using an intelligent adaptive learning engine, the search parameters of the graph database search and the keyword search and the model weights of the answer generation model are optimized and adjusted according to user satisfaction and the relevance of the answer to the archive query question; wherein the intelligent adaptive learning engine adopts a deep reinforcement learning framework.
5. The archival information question-and-answer method according to claim 4, characterized in that: The reward function of the intelligent adaptive learning engine is expressed as: R(s,a)=r(positive)+r(ctr)+r(dwell)+r(negative) Among them, R(s,a) represents the comprehensive reward value, r (positive) represents the positive reward value; r (ctr) represents the query result click rate reward value, r (dwell) represents the dwell time reward value, and r (negative) represents the negative reward value.
6. The archival information question-and-answer method according to claim 4, characterized in that: The method further comprises: Utilizing the intelligent adaptive learning engine, based on a pre-built user portrait, the search parameters of the graph database search and the keyword search and the model weights of the answer generation model are dynamically adjusted; wherein the user portrait is constructed by the intelligent adaptive learning engine based on at least one of the user's historical query records, click behaviors, and preference settings.
7. The archival information question-and-answer method according to claim 4, characterized in that: The method further comprises: The answer generation model is used to output personalized profile recommendations and query suggestions based on the user portrait.
8. The archival information question-and-answer method according to claim 4, characterized in that: The step of generating an answer to the archive query question includes: Based on the user's questioning style and preferences, generate answers to the questions in a preset style; wherein the user's questioning style and preferences are learned by the intelligent adaptive learning engine based on the user portrait and fed back to the answer generation model.
9. The archival information question-and-answer method according to claim 1, characterized in that: The method further comprises: Control users' access to and operations on archive information according to the authority scope of users at different levels; and / or, Encryption technology and data desensitization are used in the transmission and storage of archival data.
10. A file information question and answer system, characterized in that: include: A retrieval module is used to: in response to receiving an archive query question, obtain an archive information retrieval result corresponding to the archive query question through a graph database search and a keyword search; The answer generation module is used to: input the archive query question and the archive information retrieval result into the answer generation model to generate the answer to the archive query question.
Citation Information
Patent Citations
Searching system and searching method based on industry characteristics
CN110427547A
Method and device for determining standard problem and related equipment
CN116414940A
Parameter searching method and device in recommended scene, equipment and medium
CN116821513A
Personalized intention dynamic identification method and device and related equipment
CN118733727A
Method and system for generating enhanced knowledge questions and answers for mixed retrieval of heterogeneous database
CN119311831A
Cited By
Retrieval enhancement method for threat intelligence mapping knowledge domain
CN120832420A
Retrieval method and device for automobile standard document
CN121188169A