Question answering method and device based on knowledge base, equipment and medium

By preprocessing, vectorizing, and performing high-dimensional semantic analysis on knowledge base document data, combined with an initial performance scoring optimization strategy, the problem of insufficient accuracy of answers in financial and medical scenarios of question-answering systems was solved, enabling efficient and reliable deployment of question-answering systems and ensuring the accuracy and interpretability of output results.

CN121166752APending Publication Date: 2025-12-19PING AN TECH (SHENZHEN) CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511285499.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-09
Publication Date
2025-12-19

AI Technical Summary

Technical Problem

Existing knowledge base-based question answering systems lack effective context modeling when processing long documents and cross-paragraph information, resulting in insufficient answer accuracy. Furthermore, the lack of iterative optimization mechanisms means that the generated answers may contain logical errors or factual biases, making it difficult to promote and apply them in financial and medical scenarios.

Method used

By preprocessing, vectorizing, and performing high-dimensional semantic analysis on knowledge base document data, a high-dimensional context representation is constructed. An initial performance scoring optimization strategy is used to iteratively optimize the question-answering pipeline, generating a target question-answering pipeline to improve answer accuracy and efficiency.

Benefits of technology

It has improved the accuracy, stability and interpretability of question-answering systems in the financial and medical fields, reduced optimization costs and the need for manual intervention, and enhanced the adaptability and maintainability of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121166752A_ABST
    Figure CN121166752A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence, and discloses a knowledge base-based question and answer method, device and equipment and a medium, and the method comprises the following steps: obtaining document data and input query in a knowledge base, and carrying out preprocessing operation on the document data to obtain a preprocessed document set; constructing a text vectorization representation based on the preprocessed document set, generating a vectorization feature matrix, and analyzing the vectorization feature matrix through a language model to obtain a high-dimensional context representation; performing question answering on the high-dimensional context representation through an initial question answering assembly line based on the input query, outputting initial answer information, and determining an initial performance score through a test query set; and determining an optimization strategy according to the initial performance score, and optimizing the initial question and answer assembly line based on the optimization strategy to obtain a target question and answer assembly line so as to output and obtain target answer information. The method can be applied to financial science and technology or medical care service program systems, and the accuracy and efficiency of question and answer results can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of data processing, and in particular to a knowledge base-based question answering method, device, equipment and medium. BACKGROUND

[0002] In high-risk fields such as finance and medicine, knowledge base question answering technology has gradually become an important tool for improving the quality of information services. The financial field has characteristics such as complex policies and regulations, numerous professional terms, and frequent real-time updates. Users often have difficulty quickly obtaining accurate answers in insurance clause interpretation, financial product descriptions, risk assessment, and compliance review. The medical field also faces challenges such as the complexity of medical literature, the rapid updating of clinical knowledge, and the diversity of personalized patient problems. Traditional retrieval methods cannot meet the needs of doctors and patients for accurate and efficient question answering.

[0003] Existing knowledge base-based question answering systems mostly use vector retrieval combined with generative models, but still have two shortcomings: first, the system lacks effective context modeling when processing long documents and cross-paragraph information, resulting in insufficient answer accuracy; second, existing models lack iterative optimization mechanisms for output results, and the generated answers may have logical errors, factual deviations, or non-standard expressions, making it difficult to apply in key scenarios such as financial audit, risk control, or medical diagnosis assistance. Therefore, there is an urgent need for a knowledge base-based question answering method that can improve the accuracy and efficiency of question answering results to ensure that the output results have accuracy, stability, and explainability in financial and medical scenarios. SUMMARY

[0004] The present application provides a knowledge base-based question answering method, device, equipment and medium to solve the technical problem that the accuracy and efficiency of the question answering results of the question answering system in related technologies are poor, and thus the output results cannot have accuracy, stability and explainability in financial and medical scenarios.

[0005] In a first aspect, a knowledge base-based question answering method is provided, the method comprising: Obtaining document data in a knowledge base and an input query, and performing a preprocessing operation on the document data to obtain a preprocessed document set; Based on the preprocessed document set, a text vectorization representation is constructed to generate a vectorization feature matrix, and a language model is used to analyze the vectorization feature matrix to obtain a high-dimensional context representation; Based on the input query, the high-dimensional context representation is processed by an initial question answering pipeline to output an initial answer information corresponding to the input query, and a test query set is used to determine an initial performance score corresponding to the initial answer information; According to the initial performance score, a corresponding optimization strategy is determined, and the initial question and answer pipeline is optimized based on the optimization strategy to obtain a target question and answer pipeline, so that the target question and answer pipeline outputs target answer information corresponding to the input query based on the high-dimensional context representation.

[0006] In a second aspect, a knowledge base-based question and answer device is provided, comprising: An acquisition module is configured to acquire document data in a knowledge base and an input query, and perform a preprocessing operation on the document data to obtain a preprocessed document set. An analysis module is configured to construct a text vectorization representation based on the preprocessed document set, generate a vectorization feature matrix, and analyze the vectorization feature matrix through a language model to obtain a high-dimensional context representation. A determination module is configured to perform question and answer processing on the high-dimensional context representation through an initial question and answer pipeline based on the input query, output initial answer information corresponding to the input query, and determine an initial performance score corresponding to the initial answer information through a test query set. An output module is configured to determine a corresponding optimization strategy according to the initial performance score, and optimize the initial question and answer pipeline based on the optimization strategy to obtain a target question and answer pipeline, so that the target question and answer pipeline outputs target answer information corresponding to the input query based on the high-dimensional context representation.

[0007] In a third aspect, a computer device is provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the above-described knowledge base-based question and answer method.

[0008] In a fourth aspect, a computer-readable storage medium is provided, which stores a computer program, and the computer program is executed by a processor to implement the steps of the above-described knowledge base-based question and answer method.

[0009] The scheme realized by the knowledge base-based question answering method, device, computer device and storage medium includes: obtaining document data in a knowledge base and an input query, and performing a preprocessing operation on the document data to obtain a preprocessed document set; constructing a text vectorization representation based on the preprocessed document set, generating a vectorization feature matrix, and analyzing the vectorization feature matrix through a language model to obtain a high-dimensional context representation; performing question answering processing on the high-dimensional context representation through an initial question answering pipeline based on the input query, outputting an initial answer information corresponding to the input query, and determining an initial performance score corresponding to the initial answer information through a test query set; determining a corresponding optimization strategy according to the initial performance score, and optimizing the initial question answering pipeline based on the optimization strategy to obtain a target question answering pipeline, so that the target question answering pipeline outputs target answer information corresponding to the input query based on the high-dimensional context representation. In the present application, by preprocessing, vectorization representation and high-dimensional semantic analysis of the document data in the knowledge base, accurate understanding and question answering generation of the input query are realized, and the accuracy and reliability of the question answering system in high-risk fields such as finance and medicine are effectively improved. Through the optimization strategy selection and pipeline iterative update of the initial performance score, the performance bottleneck of the retrieval or generation stage can be quickly located and optimized, and the tuning cost and manual intervention demand are reduced. At the same time, the complex optimization experience is solidified into a standardized process and a strategy library, thereby the high-performance question answering system can be efficiently deployed, the system adaptability, maintainability and long-term scalability are enhanced, the closed-loop optimization from knowledge processing, semantic understanding to question answering output is realized, and the output result has accuracy, stability and explainability in the financial and medical scenarios. BRIEF DESCRIPTION OF DRAWINGS

[0010] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings needed to be used in the description of the embodiments of the present application. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.

[0011] Figure 1 is an application environment schematic diagram of the knowledge base-based question answering method in an embodiment of the present application; Figure 2 is a flow schematic diagram of the knowledge base-based question answering method in an embodiment of the present application; Figure 3 is Figure 2 is a specific implementation flow schematic diagram of step S10 in Figure 4 is a structure schematic diagram of the knowledge base-based question answering device in an embodiment of the present application; Figure 5is a structural schematic diagram of a computer device in an embodiment of the present application; Figure 6 is another structural schematic diagram of a computer device in an embodiment of the present application. DETAILED DESCRIPTION

[0012] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are some but not all of the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative work fall within the protection scope of the present application.

[0013] The knowledge base-based question answering method provided in the embodiments of the present application can be applied in an application environment such as Figure 1 , wherein a client communicates with a server through a network. The server can obtain document data in a knowledge base and an input query through the client, and perform a preprocessing operation on the document data to obtain a preprocessed document set. A text vectorization representation is constructed based on the preprocessed document set, a vectorization feature matrix is generated, and the vectorization feature matrix is analyzed through a language model to obtain a high-dimensional context representation. The high-dimensional context representation is processed through an initial question answering pipeline based on the input query, and an initial answer information corresponding to the input query is output. An initial performance score corresponding to the initial answer information is determined through a test query set. An optimization strategy is determined according to the initial performance score, and the initial question answering pipeline is optimized based on the optimization strategy to obtain a target question answering pipeline, so that the target question answering pipeline outputs target answer information corresponding to the input query based on the high-dimensional context representation. In the present application, through preprocessing, vectorization representation and high-dimensional semantic analysis of the document data in the knowledge base, accurate understanding and question answering generation of the input query are realized, and the accuracy and reliability of the question answering system in high-risk fields such as finance and medicine are effectively improved. Through optimization strategy selection and pipeline iterative updating of the initial performance score, the performance bottleneck in the retrieval or generation stage can be quickly located and optimized, and the tuning cost and manual intervention demand are reduced. At the same time, the present application solidifies complex optimization experience into a standardized process and a strategy library, thereby efficiently deploying a high-performance question answering system and enhancing the adaptability, maintainability and long-term scalability of the system, realizing closed-loop optimization from knowledge processing, semantic understanding to question answering output, and ensuring that the output result has accuracy, stability and explainability in finance and medical scenarios. The client can be, but is not limited to, various personal computers, notebook computers, smart phones, tablet computers and portable wearable devices. The server can be implemented by an independent server or a server cluster composed of multiple servers. The present application will be described in detail through specific embodiments.

[0014] Referring to Figure 2 as shown, Figure 2 A flowchart of a knowledge base-based question answering method provided by an embodiment of the present application is shown in FIG. 1. The method comprises the following steps: S10: obtaining document data in a knowledge base and an input query, and performing a preprocessing operation on the document data to obtain a preprocessed document set.

[0015] For example, document data containing financial, medical or health care fields can be first extracted from the knowledge base, and at the same time, query information input by a user is obtained; then, the extracted document data is preprocessed, including format analysis, text cleaning, word segmentation, noise information removal and standardization processing, etc., which are not limited in the present application, so as to convert the document data into a structured and consistent text set. Thus, the preprocessed document set is obtained, which provides reliable and clean basic data for subsequent vectorization representation and high-dimensional semantic analysis, ensuring that the question answering system can efficiently process document content in the retrieval and generation stages and improve the accuracy and consistency of answers.

[0016] As shown in Figure 3 S10, that is, the preprocessing operation on the document data to obtain the preprocessed document set comprises the following steps: S11: sequentially performing format analysis and text cleaning on the document data to obtain a basic cleaning text.

[0017] S12: performing word segmentation on the basic cleaning text, and performing part-of-speech tagging on the word-segmented basic cleaning text to obtain word sequence information.

[0018] S13: performing structured processing on the word sequence information to obtain the preprocessed document set.

[0019] For example, in step S11, the document data can be first format analyzed to uniformly convert financial reports, insurance clauses, medical case records or research literature of different sources into a standardized text format; then, text cleaning is performed, including removing redundant symbols, noise characters, invalid line breaks and special encodings, so as to obtain a basic cleaning text with standardized content and easy for subsequent analysis. This process ensures that the document data has a unified structure and high data purity without losing semantic information.

[0020] Further, in step S12 and step S13, the basic cleaned text can be further processed at the linguistic level. First, in S12, the basic cleaned text is split into word units by a word segmentation tool, and combined with a professional dictionary in the financial and medical fields to identify insurance contract terms, financial risk indicators or medical terms; at the same time, the word segmentation result is tagged with part of speech, to obtain word sequence information containing syntax and semantic attributes. Then in S13, based on the word sequence information, structured processing is performed, such as establishing dependency relationships at the sentence level, extracting entities and relationships, and organizing them into structured text representations, to finally form a pre-processed document set. This set not only retains the semantic context, but also has machine readability, laying a solid foundation for subsequent vectorization representation and semantic modeling.

[0021] S20: constructing a text vectorization representation based on the pre-processed document set, generating a vectorization feature matrix, and analyzing the vectorization feature matrix through a language model to obtain a high-dimensional context representation.

[0022] For example, first, based on the pre-processed document set, a text vectorization representation is constructed, the text is encoded and processed, a text vectorization representation that can reflect semantic and contextual relationships is constructed, and a vectorization feature matrix is generated, which can comprehensively represent key information such as financial clauses, risk indicators, medical terms or clinical descriptions; then a large language model is used to deeply analyze the vectorization feature matrix, and through a multi-layer attention mechanism, semantic dependencies and implicit logical relationships across sentences and paragraphs are captured, thereby obtaining a high-dimensional context representation with rich semantic expression ability, providing a precise and contextually closely related semantic basis for subsequent question and answer processing and optimization.

[0023] In some embodiments, constructing a text vectorization representation based on the pre-processed document set, generating a vectorization feature matrix, includes: performing block processing on the pre-processed document set to obtain a document block set; analyzing the document block set through a word vector encoding model to obtain an initial semantic vector set; performing context modeling on the initial semantic vector set through a deep neural network to obtain a target semantic vector set; and concatenating the target semantic vector set according to the order of text segments to obtain the vectorization feature matrix.

[0024] For example, the pre-processed document set can be first divided into document blocks, such as long financial reports, insurance clauses or medical cases, which are suitable for computing processing; then the document blocks are encoded by using a word vector encoding model to generate an initial semantic vector set that can represent basic word meanings and syntactic features; on this basis, a deep neural network is introduced to model the context of the initial semantic vector set, capture cross-block semantic dependencies and logical connections, and form a target semantic vector set with richer semantic expression; finally, the target semantic vector set is spliced according to the order of the original document fragments to construct a complete vectorized feature matrix, which provides a unified and structured input representation for subsequent high-dimensional context analysis of the language model.

[0025] In some embodiments, the analysis of the vectorized feature matrix by the language model to obtain a high-dimensional context representation includes: inputting the vectorized feature matrix into the language model for analysis to obtain an encoded representation; fusing the encoded representation with preset domain prior knowledge to obtain a semantically enhanced feature representation; hierarchically modeling the semantically enhanced feature representation to obtain a hierarchical representation result; and vector splicing the hierarchical representation result to obtain the high-dimensional context representation.

[0026] For example, the vectorized feature matrix can be first input into the language model for deep semantic analysis to obtain an encoded representation that can represent syntactic and semantic relationships; then the encoded representation is fused with preset domain prior knowledge such as financial regulatory clauses, insurance claim rules or medical diagnosis and treatment specifications to form a semantically enhanced feature representation, so as to improve the understanding ability of the model for domain-specific concepts and logic; then the semantically enhanced feature representation is hierarchically modeled to capture multi-level semantic associations from words to sentences and from paragraphs to entire documents, to obtain a hierarchical representation result; finally, the representation results of each level are vector spliced according to the semantic logic order to generate a complete high-dimensional context representation, which provides deep support for the accurate reasoning and answering of the question-answering system in complex financial and medical scenarios.

[0027] S30: performing question-answering processing on the high-dimensional context representation based on the input query by an initial question-answering pipeline, outputting an initial answer information corresponding to the input query, and determining an initial performance score corresponding to the initial answer information by using a test query set.

[0028] For example, first, a user input query can be received and input to an initial question and answer pipeline together with the high-dimensional context representation obtained in the previous step; the initial question and answer pipeline can sequentially perform retrieval matching, context fusion, and inference generation, and output initial answer information corresponding to the input query. Subsequently, the initial answer information can be verified and quantitatively evaluated using a preset test query set, and an initial performance score can be calculated from the accuracy, error rate, logical reasonableness, and domain compliance, thereby providing a quantitative basis for subsequent optimization strategy selection and iterative improvement.

[0029] In some embodiments, the initial question and answer processing of the high-dimensional context representation based on the input query includes: retrieving a candidate result set from the preprocessed document set based on the high-dimensional context representation; wherein the semantic similarity between the candidate result set and the high-dimensional context representation is higher than a preset threshold; performing relevance sorting on the candidate result set to filter the most relevant answer information to the input query; and determining the most relevant answer information as the initial answer information.

[0030] For example, first, a user input query can be received and input to an initial question and answer pipeline together with the high-dimensional context representation obtained in the previous step; the initial question and answer pipeline can sequentially perform retrieval matching, context fusion, and inference generation, and output initial answer information corresponding to the input query. Subsequently, the initial answer information can be verified and quantitatively evaluated using a preset test query set, and an initial performance score can be calculated from the accuracy, error rate, logical reasonableness, and domain compliance, thereby providing a quantitative basis for subsequent optimization strategy selection and iterative improvement.

[0031] Further, after obtaining the initial answer information, a preset test query set can be called to verify and quantitatively evaluate the result. The evaluation process is carried out from multiple dimensions, including the accuracy, error rate, logical reasonableness, compliance, and applicability in the professional field of the answer. Thus, an initial performance score can be calculated according to these evaluation indexes, and used as the core basis for subsequent optimization. Through the performance score, the system can determine whether the current pipeline meets the expectations, thereby providing reliable data support for subsequent strategy selection and iterative improvement.

[0032] In some embodiments, the determining the initial performance score corresponding to the initial answer information by testing the query set comprises: inputting the test query set into the initial question and answer pipeline to obtain a predicted answer set; comparing the predicted answer set with the initial answer information to obtain an answer matching result; determining performance evaluation data corresponding to the answer matching result based on the answer matching result, and weighting the performance evaluation data to obtain the initial performance score; wherein the performance evaluation data comprises accuracy, recall rate and semantic similarity.

[0033] For example, first, a preset test query set can be input into the initial question and answer pipeline to obtain a corresponding predicted answer set; then the predicted answer set is compared with the initial answer information generated by the system one by one to form an answer matching result; based on the matching result, performance evaluation data is calculated, which includes accuracy, recall rate, semantic similarity and other indicators of the correctness, coverage and semantic consistency of the answer; finally, the above multi-dimensional performance evaluation data is weighted and processed to obtain an initial performance score, thereby objectively reflecting the actual performance of the question and answer pipeline in the financial and medical fields.

[0034] S40: determining a corresponding optimization strategy according to the initial performance score, and optimizing the initial question and answer pipeline based on the optimization strategy to obtain a target question and answer pipeline, so that the target question and answer pipeline outputs target answer information corresponding to the input query based on the high-dimensional context representation.

[0035] In some embodiments, the determining a corresponding optimization strategy according to the initial performance score, and optimizing the initial question and answer pipeline based on the optimization strategy to obtain a target question and answer pipeline comprises: determining error source information corresponding to the initial question and answer pipeline according to the initial performance score, and labeling the error source information to obtain a labeling result; wherein the labeling result comprises an error type; determining an optimization strategy matching the error type from a preset optimization strategy library based on the labeling result; optimizing the initial question and answer pipeline based on the optimization strategy to obtain an optimized question and answer pipeline; functionally verifying the optimized question and answer pipeline, and determining the question and answer pipeline that passes the verification as the target question and answer pipeline.

[0036] For example, initially, the error sources of the initial question and answer pipeline can be determined according to the initial performance score, and the errors are labeled to form specific labeling results, wherein the labeling results include the types of errors, such as insufficient retrieval, inaccurate context matching, hallucination in the generation stage, or non-standard output format, etc. Through error labeling, the possible clause understanding errors, risk indicator omissions in the financial scenario, or the possible clinical description deviations, medical term misjudgments in the medical scenario can be explicitly identified, thereby providing a basis for subsequent strategy selection.

[0037] Subsequently, the optimization strategy corresponding to the error type can be found in the preset optimization strategy library based on the error labeling result, such as adopting sparse and dense mixed retrieval, strengthening context prompt, or introducing structured output template, etc. Thus, the selected optimization strategy can be applied to the initial question and answer pipeline to adjust and optimize the related modules, generate an optimized question and answer pipeline, and verify its function to ensure that the optimized pipeline can correctly process the query and maintain the stability and standardization of the output. The verified question and answer pipeline is determined as the target question and answer pipeline, which can output more accurate target answer information for the user input query based on high-dimensional context representation.

[0038] As can be seen, in the above scheme, by preprocessing, vectorizing representation, and high-dimensional semantic analysis of the document data in the knowledge base, accurate understanding and question and answer generation for the input query are realized, effectively improving the accuracy and reliability of the question and answer system in high-risk fields such as finance and medicine. Through optimization strategy selection and pipeline iteration update based on initial performance score, the performance bottleneck in the retrieval or generation stage can be quickly located and optimized accordingly, reducing the tuning cost and the need for manual intervention. At the same time, the complex optimization experience is solidified into a standardized process and strategy library, thereby enabling efficient deployment of high-performance question and answer systems, enhancing the adaptability, maintainability, and long-term scalability of the system, realizing closed-loop optimization from knowledge processing, semantic understanding to question and answer output, and ensuring that the output results have accuracy, stability, and explainability in financial and medical scenarios.

[0039] It should be understood that the size of the serial number of each step in the above embodiments does not mean the order of execution, and the execution order of each process should be determined according to its function and inherent logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.

[0040] In an embodiment, a knowledge base-based question and answer device is provided, which corresponds one-to-one to the knowledge base-based question and answer method in the above embodiments. As shown in the figure, the knowledge base-based question and answer device includes an acquisition module 101, an analysis module 102, a determination module 103, and an output module 104. The detailed description of each functional module is as follows: Figure 4 ​The acquisition module 101 is configured to acquire document data and an input query in a knowledge base, and perform a preprocessing operation on the document data to obtain a preprocessed document set. The analysis module 102 is configured to construct a text vectorization representation based on the preprocessed document set, generate a vectorization feature matrix, and analyze the vectorization feature matrix through a language model to obtain a high-dimensional context representation. The determination module 103 is configured to perform question and answer processing on the high-dimensional context representation through an initial question and answer pipeline based on the input query, output initial answer information corresponding to the input query, and determine an initial performance score corresponding to the initial answer information through a test query set. The output module 104 is configured to determine a corresponding optimization strategy according to the initial performance score, optimize the initial question and answer pipeline based on the optimization strategy to obtain a target question and answer pipeline, so that the target question and answer pipeline outputs target answer information corresponding to the input query based on the high-dimensional context representation.

[0041] The acquisition module 101 is configured to sequentially perform format analysis and text cleaning on the document data to obtain a basic cleaned text, perform word segmentation on the basic cleaned text, and perform part-of-speech tagging on the segmented basic cleaned text to obtain word sequence information, and perform structured processing on the word sequence information to obtain the preprocessed document set.

[0042] The analysis module 102 is configured to perform block processing on the preprocessed document set to obtain a document block set, analyze the document block set through a word vector encoding model to obtain an initial semantic vector set, perform context modeling on the initial semantic vector set through a deep neural network to obtain a target semantic vector set, and splice the target semantic vector set according to a text segment order to obtain the vectorization feature matrix.

[0043] The analysis module 102 is configured to input the vectorization feature matrix into the language model for analysis to obtain an encoding representation, fuse the encoding representation with preset domain priori knowledge to obtain a semantic enhanced feature representation, perform hierarchical modeling on the semantic enhanced feature representation to obtain a hierarchical representation result, and perform vector splicing on the hierarchical representation result to obtain the high-dimensional context representation.

[0044] The determining module 103 is configured to search from the preprocessed document set based on the high-dimensional context representation to obtain a candidate result set, wherein a semantic similarity between the candidate result set and the high-dimensional context representation is higher than a preset threshold value, perform relevance sorting on the candidate result set to filter out the most relevant answer information to the input query, and determine the most relevant answer information as the initial answer information.

[0045] The determining module 103 is configured to input the test query set into the initial question and answer pipeline to obtain a predicted answer set, compare the predicted answer set with the initial answer information to obtain an answer matching result, determine performance evaluation data corresponding to the answer matching result based on the answer matching result, and weight the performance evaluation data to obtain the initial performance score, wherein the performance evaluation data includes accuracy, recall rate and semantic similarity.

[0046] The output module 104 is configured to determine error source information corresponding to the initial question and answer pipeline according to the initial performance score, label the error source information to obtain a labeling result, wherein the labeling result includes an error type, determine an optimization strategy matched with the error type from a preset optimization strategy library based on the labeling result, optimize the initial question and answer pipeline based on the optimization strategy to obtain an optimized question and answer pipeline, perform function verification on the optimized question and answer pipeline, and determine the question and answer pipeline that passes the verification as the target question and answer pipeline.

[0047] The application provides a knowledge base-based question and answer device, which realizes accurate understanding of input queries and question and answer generation through preprocessing, vectorization representation and high-dimensional semantic analysis of document data in the knowledge base, effectively improves the accuracy and reliability of the question and answer system in high-risk fields such as finance and medicine. Through optimization strategy selection and pipeline iterative updating of the initial performance score, the performance bottleneck in the retrieval or generation stage can be quickly located and optimized, reducing the tuning cost and the need for manual intervention. At the same time, the application solidifies complex optimization experience into a standardized process and strategy library, thereby efficiently deploying a high-performance question and answer system, enhancing the system adaptability, maintainability and long-term scalability, realizing closed-loop optimization from knowledge processing, semantic understanding to question and answer output, and ensuring that the output result has accuracy, stability and explainability in the financial and medical scenarios.

[0048] The specific definition of the knowledge base based question answering device can refer to the definition of the knowledge base based question answering method in the foregoing, and will not be described here. Each module in the knowledge base based question answering device described above can be realized by software, hardware and a combination thereof in whole or in part. The above-mentioned modules can be embedded in or independent of the processor in the computer device in hardware form, or can be stored in the memory in the computer device in software form, so that the processor calls and executes the operations corresponding to each of the above-mentioned modules.

[0049] In an embodiment, a computer device is provided, which can be a server, and an internal structure diagram thereof can be as shown in Figure 5 The computer device includes a processor, a memory, a network interface and a database connected through a system bus. The processor of the computer device is configured to provide computing and control capabilities. The memory of the computer device includes a non-volatile and / or volatile storage medium, an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operating system and the computer program in the non-volatile storage medium to run. The network interface of the computer device is configured to communicate with an external client through a network connection. The computer program is executed by the processor to implement the functions or steps of the server side of the knowledge base based question answering method.

[0050] In an embodiment, a computer device is provided, which can be a client, and an internal structure diagram thereof can be as shown in Figure 6 The computer device includes a processor, a memory, a network interface, a display screen and an input device connected through a system bus. The processor of the computer device is configured to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operating system and the computer program in the non-volatile storage medium to run. The network interface of the computer device is configured to communicate with an external server through a network connection. The computer program is executed by the processor to implement the functions or steps of the client side of the knowledge base based question answering method.

[0051] In an embodiment, a computer device is provided, which includes a memory, a processor and a computer program stored in the memory and executable on the processor, and the processor implements the following steps when executing the computer program: obtaining document data in a knowledge base and an input query, and performing a preprocessing operation on the document data to obtain a preprocessed document set; constructing a text vectorization representation based on the preprocessed document set, generating a vectorization feature matrix, and analyzing the vectorization feature matrix through a language model to obtain a high-dimensional context representation; perform question and answer processing on the high-dimensional context representation based on the input query through an initial question and answer pipeline, output initial answer information corresponding to the input query, and determine an initial performance score corresponding to the initial answer information through a test query set; determine an optimization strategy corresponding to the initial performance score, and optimize the initial question and answer pipeline based on the optimization strategy to obtain a target question and answer pipeline, so that the target question and answer pipeline outputs target answer information corresponding to the input query based on the high-dimensional context representation.

[0052] In one embodiment, a computer readable storage medium is provided, which stores a computer program. When the computer program is executed by a processor, the following steps are implemented: obtain document data in a knowledge base and an input query, and perform preprocessing operations on the document data to obtain a preprocessed document set; construct a text vectorization representation based on the preprocessed document set, generate a vectorization feature matrix, and analyze the vectorization feature matrix through a language model to obtain a high-dimensional context representation; perform question and answer processing on the high-dimensional context representation based on the input query through an initial question and answer pipeline, output initial answer information corresponding to the input query, and determine an initial performance score corresponding to the initial answer information through a test query set; determine an optimization strategy corresponding to the initial performance score, and optimize the initial question and answer pipeline based on the optimization strategy to obtain a target question and answer pipeline, so that the target question and answer pipeline outputs target answer information corresponding to the input query based on the high-dimensional context representation.

[0053] It should be noted that the functions or steps that the above computer readable storage medium or computer device can implement can correspond to the related descriptions of the server side and the client side in the foregoing method embodiments. To avoid repetition, they will not be described one by one here.

[0054] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer readable storage medium, and when executed, can include the processes of the above-mentioned embodiment methods. Any reference to memory, storage, database or other medium used in the embodiments provided by the present application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration but not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0055] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the division of the above-mentioned functional units and modules is exemplified, and in actual application, the above-mentioned functions can be completed by different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above.

[0056] The above-mentioned embodiments are only used to illustrate the technical solutions of the present application, but not limit it. Although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that they can modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement for part of the technical features. The modification or replacement does not make the essence of the corresponding technical solution deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should be included in the protection scope of the present application.

Claims

1. A knowledge base based question answering method, characterized by, The method comprises: acquiring document data in a knowledge base and an input query, and performing a preprocessing operation on the document data to obtain a preprocessed document set; constructing a text vectorization representation based on the preprocessed document set, generating a vectorization feature matrix, and analyzing the vectorization feature matrix through a language model to obtain a high-dimensional context representation; performing question and answer processing on the high-dimensional context representation through an initial question and answer pipeline based on the input query, outputting an initial answer information corresponding to the input query, and determining an initial performance score corresponding to the initial answer information through a test query set; determining a corresponding optimization strategy according to the initial performance score, and optimizing the initial question and answer pipeline based on the optimization strategy to obtain a target question and answer pipeline, so that the target question and answer pipeline outputs target answer information corresponding to the input query based on the high-dimensional context representation.

2. The method of claim 1, wherein, The preprocessing operation on the document data to obtain a preprocessed document set comprises: performing format analysis and text cleaning on the document data in sequence to obtain a basic cleaned text; performing word segmentation on the basic cleaned text, and performing part-of-speech tagging on the word-segmented basic cleaned text to obtain word sequence information; performing structured processing on the word sequence information to obtain the preprocessed document set.

3. The method of claim 1, wherein, The construction of a text vectorization representation based on the preprocessed document set to generate a vectorization feature matrix comprises: performing block processing on the preprocessed document set to obtain a document block set; analyzing the document block set through a word vector encoding model to obtain an initial semantic vector set; performing context modeling on the initial semantic vector set through a deep neural network to obtain a target semantic vector set; splicing the target semantic vector set according to the order of text segments to obtain the vectorization feature matrix.

4. The method of claim 1, wherein, The analysis of the vectorization feature matrix through a language model to obtain a high-dimensional context representation comprises: inputting the vectorization feature matrix into the language model for analysis to obtain an encoding representation; fusing the encoding representation with preset domain prior knowledge to obtain a semantically enhanced feature representation; performing hierarchical modeling on the semantically enhanced feature representation to obtain a hierarchical representation result; performing vector splicing on the hierarchical representation result to obtain the high-dimensional context representation.

5. The method of claim 1, wherein, The question and answer processing on the high-dimensional context representation through an initial question and answer pipeline based on the input query to output an initial answer information corresponding to the input query comprises: performing retrieval from the preprocessed document set based on the high-dimensional context representation to obtain a candidate result set; wherein the semantic similarity between the candidate result set and the high-dimensional context representation is higher than a preset threshold; performing relevance sorting on the candidate result set to filter out the most relevant answer information to the input query; determining the most relevant answer information as the initial answer information.

6. The method of claim 1, wherein, The determination of an initial performance score corresponding to the initial answer information through a test query set comprises: inputting the test query set into the initial question and answer pipeline to obtain a predicted answer set; comparing the predicted answer set with initial answer information to obtain an answer matching result; determining performance evaluation data corresponding to the answer matching result based on the answer matching result, and weighting the performance evaluation data to obtain an initial performance score; wherein the performance evaluation data includes accuracy, recall rate, and semantic similarity.

7. The method of claim 1, wherein, determining a corresponding optimization strategy according to the initial performance score, and optimizing the initial question and answer pipeline based on the optimization strategy to obtain a target question and answer pipeline, including: determining error source information corresponding to the initial question and answer pipeline according to the initial performance score, and labeling the error source information to obtain a labeling result; wherein the labeling result includes an error type; determining an optimization strategy matching the error type from a pre-set optimization strategy library based on the labeling result; optimizing the initial question and answer pipeline based on the optimization strategy to obtain an optimized question and answer pipeline; functionally verifying the optimized question and answer pipeline, and determining the question and answer pipeline that passes the verification as the target question and answer pipeline.

8. A knowledge base based question answering apparatus characterized by comprising: including: an acquisition module configured to acquire document data in a knowledge base and an input query, and perform a preprocessing operation on the document data to obtain a preprocessed document set; an analysis module configured to construct a text vectorization representation based on the preprocessed document set, generate a vectorization feature matrix, and analyze the vectorization feature matrix through a language model to obtain a high-dimensional context representation; a determination module configured to perform question and answer processing on the high-dimensional context representation through an initial question and answer pipeline based on the input query, output initial answer information corresponding to the input query, and determine an initial performance score corresponding to the initial answer information through a test query set; an output module configured to determine a corresponding optimization strategy according to the initial performance score, and optimize the initial question and answer pipeline based on the optimization strategy to obtain a target question and answer pipeline, so that the target question and answer pipeline outputs target answer information corresponding to the input query based on the high-dimensional context representation.

9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, The processor executes the computer program to implement the steps of the knowledge base-based question and answer method according to any one of claims 1 to 7.

10. A computer-readable storage medium storing a computer program, the computer program comprising instructions that, when executed by a computer, cause the computer to perform the method of any one of claims 1 to 9. The computer program is executed by the processor to implement the steps of the knowledge base-based question and answer method according to any one of claims 1 to 7.

Citation Information

Cited By

  • Metallographic structure data-oriented AI high-availability evaluation method and related system

    CN121765324A