Knowledge question and answer method and device, computer equipment and storage medium

By building a knowledge question and answer database and using pre-trained knowledge large models, the accuracy problems of the existing technology when dealing with complex documents and professional fields are solved, and efficient and accurate knowledge question and answer effects are achieved.

CN120104760APending Publication Date: 2025-06-06CHINA PING AN PROPERTY INSURANCE CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510532584.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-25
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

Existing knowledge Q&A technology is difficult to effectively understand and generate accurate answers when dealing with complex documents and expertise, especially in the fields of insurance, healthcare and fintech.

Method used

By constructing a knowledge question and answer database, the original prompt text of the question to be query is obtained and word segmentation and vectorization are performed, the similarity between the word vector representation and the statement in the database is calculated, the relevant content is filtered to generate the target prompt text, and the answer is generated through the pre-trained knowledge model.

Benefits of technology

It realizes effective processing of query questions and accurate understanding of key information, generates high-quality answers, and meets users' needs for efficient and accurate knowledge questions and answers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120104760A_ABST
    Figure CN120104760A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence, can be applied to service system platforms of medical health, financial science and technology and the like, and discloses a knowledge question and answer method and device, computer equipment and a storage medium. Obtaining an original prompt text corresponding to a to-be-queried question, and performing word segmentation and vectorization processing on the original prompt text to obtain word vector representation of the original prompt text; calculating the similarity between the word vector representation and each statement in the knowledge question and answer database, screening out the content related to the word vector representation in the knowledge question and answer database according to a calculation result, and generating a target prompt text; taking the target prompt text as input, and generating a target answer of the to-be-queried question through a pre-trained knowledge large model; therefore, the original prompt text of the to-be-queried question can be effectively processed, the key information can be accurately understood, the high-quality answer is generated, and the requirement of a user for efficient and accurate knowledge questions and answers can be met.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a knowledge question answering method, device, computer equipment and computer-readable storage medium. Background Art

[0002] At present, in today's era of information explosion, knowledge question answering systems play a vital role in various fields. Whether in the fields of finance, law, medicine or education, users hope to obtain the required information quickly and accurately. However, existing knowledge question answering technologies still face many challenges when processing complex documents and professional domain knowledge. Take the insurance industry as an example. Its documents and clauses are usually very long and contain a large number of legal terms and professional terms, which makes it difficult for traditional knowledge question answering systems to process effectively and the generated answers are not accurate enough. However, these problems are not limited to the insurance industry, but are widely present in various fields that need to process complex knowledge documents.

[0003] In the healthcare field, medical texts usually contain a large number of professional terms and complex medical concepts, such as disease names, drug names, treatment plans, etc. These terms are not only numerous, but also semantically complex, making it difficult for traditional knowledge question-answering systems to accurately understand them and generate inaccurate answers. For example, a query about a specific disease may involve multiple medical terms, and the knowledge question-answering system needs to be able to accurately identify and understand the relationship between these terms in order to provide accurate answers.

[0004] In the field of FinTech, financial documents often involve complex laws, regulations and compliance requirements, such as financial regulatory policies, contract terms, etc. Knowledge question answering systems need to be able to accurately understand and interpret these regulations and terms to ensure that the answers provided are accurate and meet legal requirements. For example, for a query about the compliance of a financial product, the knowledge question answering system needs to be able to accurately quote relevant regulatory terms and provide compliance advice.

[0005] That is, the existing knowledge question answering technology has difficulty in effectively understanding the relevant information when processing the original prompt text of the question to be queried, resulting in inaccurate answers. Therefore, how to provide a knowledge question answering method, device, computer equipment and computer-readable storage medium that can effectively process the original prompt text of the question to be queried and accurately understand the key information, and generate high-quality answers to meet the user's needs for efficient and accurate knowledge question answering is a problem that needs to be solved urgently by those skilled in the art. Summary of the invention

[0006] In view of the above-mentioned shortcomings of the prior art, the purpose of the present invention is to provide a knowledge question and answer method, device, computer equipment and computer-readable storage medium, aiming to solve the problem of how to effectively process the original prompt text of the question to be queried and accurately understand the key information, and generate high-quality answers to meet the user's needs for efficient and accurate knowledge question and answering.

[0007] In order to achieve the above object, the present invention adopts the following technical solutions:

[0008] In a first aspect, the present invention provides a knowledge question answering method, which includes:

[0009] Build a knowledge question-answering database;

[0010] Obtaining the original prompt text corresponding to the question to be queried, performing word segmentation and vectorization processing on the original prompt text, and obtaining a word vector representation of the original prompt text;

[0011] Calculate the similarity between the word vector representation and each sentence in the knowledge question and answer database, filter out the content related to the word vector representation in the knowledge question and answer database according to the calculation result, and generate a target prompt text;

[0012] The target prompt text is used as input, and the target answer to the question to be queried is generated through a pre-trained knowledge model.

[0013] In a second aspect, the present invention provides a knowledge question answering device, which includes:

[0014] A construction module, used to construct a knowledge question-answering database;

[0015] An acquisition module is used to acquire the original prompt text corresponding to the question to be queried, perform word segmentation and vectorization processing on the original prompt text, and obtain a word vector representation of the original prompt text;

[0016] A screening module, used to calculate the similarity between the word vector representation and each sentence in the knowledge question and answer database, and screen out the content related to the word vector representation in the knowledge question and answer database according to the calculation result to generate a target prompt text;

[0017] The answer generation module is used to take the target prompt text as input and generate the target answer to the question to be queried through a pre-trained knowledge model.

[0018] In a third aspect, the present invention provides a computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the knowledge question and answer method as described above when executing the computer program.

[0019] In a fourth aspect, the present invention provides a computer-readable storage medium storing a computer program, wherein the computer program, when executed by a processor, implements the knowledge question and answer method as described above.

[0020] Compared with the prior art, the present invention provides a knowledge question and answer method, apparatus, computer equipment and computer-readable storage medium, wherein a knowledge question and answer database is constructed; the original prompt text corresponding to the question to be queried is obtained, and the original prompt text is segmented and vectorized to obtain a word vector representation of the original prompt text; the similarity between the word vector representation and each sentence in the knowledge question and answer database is calculated, and the content related to the word vector representation in the knowledge question and answer database is screened out according to the calculation result to generate a target prompt text; the target prompt text is used as input to generate a target answer to the question to be queried through a pre-trained knowledge model; thus, the present invention can effectively process the original prompt text of the question to be queried and accurately understand the key information, and generate high-quality answers, which can meet the user's needs for efficient and accurate knowledge question and answer. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0022] Figure 1 A schematic diagram of an application environment of a knowledge question answering method provided by an embodiment of the present invention.

[0023] Figure 2 A flowchart of a knowledge question-answering method provided in one embodiment of the present invention.

[0024] Figure 3 A schematic diagram of a program module of a knowledge question and answer device provided in one embodiment of the present invention.

[0025] Figure 4 A schematic diagram of the structure of a computer device provided by an embodiment of the present invention.

[0026] Figure 5 Another structural schematic diagram of a computer device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0027] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0028] It should be understood that when used in the present specification and the appended claims, the term "comprising" indicates the presence of described features, integers, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or combinations thereof.

[0029] It should also be understood that the term "and / or" used in the present description and the appended claims refers to and includes any and all possible combinations of one or more of the associated listed items.

[0030] As used in the present specification and the appended claims, the term "if" can be interpreted as "when" or "uponce" or "in response to determining" or "in response to detecting", depending on the context. Similarly, the phrase "if it is determined" or "if [described condition or event] is detected" can be interpreted as meaning "uponce it is determined" or "in response to determining" or "uponce [described condition or event] is detected" or "in response to detecting [described condition or event]", depending on the context.

[0031] In addition, in the description of the present specification and the appended claims, the terms "first", "second", "third", etc. are only used to distinguish the descriptions and cannot be understood as indicating or implying relative importance.

[0032] References to "one embodiment" or "some embodiments" etc. described in the present specification mean that one or more embodiments of the present invention include specific features, structures or characteristics described in conjunction with the embodiment. Therefore, the statements "in one embodiment", "in some embodiments", "in some other embodiments", "in some other embodiments", etc. that appear in different places in this specification do not necessarily refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in other ways. The terms "including", "comprising", "having" and their variations all mean "including but not limited to", unless otherwise specifically emphasized in other ways.

[0033] It should be understood that the order of execution of the steps in the following embodiments does not imply a precedence of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0034] In order to illustrate the technical solution of the present invention, specific embodiments are provided below for illustration.

[0035] A knowledge question answering method provided by an embodiment of the present invention can be applied in Figure 1 In the application environment shown, the client and the server communicate through the network. The client includes but is not limited to PDAs, desktop computers, laptops, ultra-mobile personal computers (UMPCs), netbooks, cloud computing devices, personal digital assistants (PDAs) and other computer devices. The server can be an independent server or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms.

[0036] See also Figure 2 An embodiment of the present invention provides a knowledge question answering method, wherein the method comprises the following steps:

[0037] S100, building a knowledge question and answer database;

[0038] S200, obtaining an original prompt text corresponding to the question to be queried, performing word segmentation and vectorization processing on the original prompt text, and obtaining a word vector representation of the original prompt text;

[0039] S300, calculating the similarity between the word vector representation and each sentence in the knowledge question and answer database, filtering out the content related to the word vector representation in the knowledge question and answer database according to the calculation result, and generating a target prompt text;

[0040] S400: taking the target prompt text as input, and generating a target answer to the question to be queried through a pre-trained knowledge model.

[0041] In specific implementation, the knowledge question answering method of this embodiment realizes the effective processing of the original prompt text of the query question, the accurate understanding of key information and the generation of high-quality answers through a series of carefully designed steps, thereby meeting the user's needs for efficient and accurate knowledge question answering. Specifically:

[0042] 1. Build a knowledge question and answer database: A structured knowledge question and answer database can be built by collecting, organizing and annotating a large amount of document data related to the target knowledge field. This step provides a rich and accurate knowledge foundation for subsequent information screening and answer generation, ensuring that the system can quickly locate the part related to the query question from the massive data, thereby improving the system's response speed and the accuracy of the answer.

[0043] 2. Get the original prompt text and perform preprocessing: After getting the original prompt text corresponding to the query question, perform word segmentation and vectorization to obtain a word vector representation. This process not only converts the text into a numerical form that can be processed by the machine, but also removes redundant information by removing stop words and other operations, retaining the key content. The vectorized word vector representation can more effectively capture the semantic features of the original prompt text, providing accurate input for subsequent information screening.

[0044] 3. Filter relevant content to generate target prompt text: Based on the knowledge question and answer database, we can use algorithms such as cosine similarity to calculate the similarity between the word vector representation and each sentence in the database, and filter out target sentences with similarity scores higher than the similarity threshold, and combine them into target prompt text. This process accurately filters out the sentences most relevant to the query question in a quantitative way, which can avoid the subjectivity and inaccuracy caused by manual judgment in traditional methods. At the same time, by combining target sentences to generate coherent target prompt text, the quality of input information is further optimized, laying the foundation for generating high-quality answers.

[0045] 4. Generate target answers through pre-trained knowledge big models: Input the target prompt text into the pre-trained knowledge big model to generate the target answer to the question to be queried. Among them, the pre-trained knowledge big model has strong language understanding and generation capabilities, and can generate accurate, fluent and natural language answers based on the target prompt text. This process fully utilizes the advantages of the pre-trained knowledge big model and combines the key information in the target prompt text to generate high-quality answers, which can meet users' needs for efficient and accurate knowledge questions and answers.

[0046] The knowledge question answering method of this embodiment achieves effective processing of the original prompt text and accurate understanding of key information by building a professional knowledge base, accurate text preprocessing, efficient sentence screening and professional large model generation, and can generate high-quality answers. This method not only improves the efficiency and accuracy of the knowledge question answering system, but also improves the user experience, and has significant technical effects and application value.

[0047] It can be understood that the knowledge question answering method provided in the embodiment of the present invention can also be applied to knowledge question answering scenarios related to the medical and health field. The following are some specific examples:

[0048] Example 1: Disease diagnosis assistance

[0049] Scenario description:

[0050] When a patient visits a hospital, the doctor needs to quickly make an accurate diagnosis based on the patient's symptoms, medical history, and other information. However, the symptoms of some diseases may be similar to other diseases, and may involve complex medical knowledge and a large amount of medical literature. In this case, the knowledge question-answering system can provide auxiliary diagnosis support for doctors.

[0051] Specific process:

[0052] 1. Build a knowledge question and answer database: Collect and organize a large amount of medical literature, clinical guidelines, disease diagnosis standards and other materials in advance to build a medical knowledge question and answer database covering a variety of disease information. Label and classify the content in the database, for example, by disease type, symptoms, treatment methods, etc.

[0053] 2. Obtain the original prompt text and perform preprocessing: The doctor enters the question to be queried consisting of information such as the patient's symptom description and medical history as the original prompt text. The system performs word segmentation and vectorization on these texts and extracts the corresponding word vector representation.

[0054] 3. Filter relevant content to generate target prompt text: Based on the knowledge question and answer database, the system calculates the similarity between the word vector representation of the original prompt text and each sentence in the database. It selects the medical literature fragments, diagnostic criteria and other sentences that are most relevant to the patient's symptoms and medical history, and combines them into the target prompt text.

[0055] 4. Generate target answers through pre-trained knowledge big model: Input the target prompt text into the pre-trained medical knowledge big model. The model generates target answers such as possible disease diagnosis results and suggestions for further examinations based on the input target prompt text.

[0056] Technical effects:

[0057] Through the knowledge question-and-answer method of the present invention, doctors can quickly obtain accurate medical information related to the patient's symptoms, assisting them in making a more accurate diagnosis. This not only improves the efficiency of diagnosis, but also reduces the risk of misdiagnosis due to insufficient information or misjudgment, thereby improving the quality of medical services.

[0058] Example 2: Drug Information Query

[0059] Scenario description:

[0060] When prescribing medicines for patients, doctors need to accurately understand the indications, contraindications, side effects, and other information of the medicines. In addition, patients may also have questions about the use and precautions of the medicines. Traditional drug information query methods may require consulting a large number of drug instructions or medical literature, which is time-consuming and prone to missing key information. The knowledge question-answering system can quickly provide accurate drug information to meet the needs of doctors and patients.

[0061] Specific process:

[0062] 1. Build a knowledge question and answer database: Collect and organize drug instructions, drug research literature, clinical medication guidelines and other materials in advance to build a medical knowledge question and answer database containing various drug information. Label and classify the content in the database, for example, by drug type, mechanism of action, adverse reactions, etc.

[0063] 2. Get the original prompt text and preprocess it: The doctor or patient enters the query question about the drug as the original prompt text. The system performs word segmentation and vectorization on these texts and extracts the corresponding word vector representation.

[0064] 3. Filter relevant content to generate target prompt text: Based on the knowledge question and answer database, the system calculates the similarity between the word vector representation of the original prompt text and each sentence in the database. It selects the instructions fragments most relevant to the queried drug, key information in the research literature, and other sentences to combine into the target prompt text.

[0065] 4. Generate target answers through pre-trained knowledge big model: Input the target prompt text into the pre-trained medical knowledge big model. The model generates target answers such as drug side effects, precautions, and usage methods based on the input target prompt text.

[0066] Technical effects:

[0067] The knowledge question-answering method of the present invention can quickly and accurately provide drug-related information, help doctors prescribe more reasonably, and also answer patients' questions about drugs. This not only improves the efficiency and quality of medical services, but also enhances patients' satisfaction and trust in medical services.

[0068] It can be understood that the knowledge question answering method provided in the embodiment of the present invention can also be applied to knowledge question answering scenarios related to the field of financial technology. The following are some specific examples:

[0069] Example 1: Financial product consultation

[0070] Scenario description:

[0071] When customers consult banks or financial institutions about financial products such as wealth management products, funds, and insurance, they usually need to know key information such as product returns, risk levels, and investment periods. However, the terms and instructions of financial products are often complex and lengthy, making it difficult for customers to quickly obtain the information they need. The knowledge question and answer system can help customers quickly obtain accurate product information and improve consultation efficiency.

[0072] Specific process:

[0073] 1. Build a knowledge question and answer database: Collect and organize financial product instructions, contract terms, investment strategies and other information from financial institutions in advance to build a knowledge question and answer database covering a variety of financial product information. Label and classify the content in the database, for example, by product type (such as financial products, funds, insurance, etc.), risk level, investment period, etc.

[0074] 2. Get the original prompt text and pre-process it: The customer enters the consulting question about the financial product (i.e. the question to be queried) as the original prompt text. The system performs word segmentation and vectorization on these texts and extracts the corresponding word vector representation.

[0075] 3. Filter relevant content to generate target prompt text: Based on the knowledge question and answer database, the system calculates the similarity between the word vector representation of the original prompt text and each sentence in the database. It selects the financial product manual fragments, income data and other sentences that are most relevant to the customer's consultation questions and combines them into the target prompt text.

[0076] 4. Generate target answers through pre-trained knowledge big model: Input the target prompt text into the pre-trained financial knowledge big model. The model generates target answers such as the income, risk level, and investment period of the financial product based on the input target prompt text.

[0077] Technical effects:

[0078] Through the knowledge question and answer method of the present invention, customers can quickly obtain accurate financial product information without spending a lot of time reading complicated instructions. This not only improves the efficiency of consultation, but also enhances customers' trust and satisfaction with financial institutions.

[0079] Example 2: Financial regulations compliance query

[0080] Scenario description:

[0081] When handling business (such as insurance business), employees of financial institutions need to ensure that all operations comply with relevant laws, regulations and regulatory requirements. Financial regulations and compliance clauses are usually complex and frequently updated, and employees may find it difficult to quickly find accurate regulatory basis in actual operations. The knowledge question and answer system can help employees quickly query and understand relevant regulations to ensure compliance of business operations.

[0082] Specific process:

[0083] 1. Build a knowledge question and answer database: Collect and organize laws, regulations, regulatory policies, compliance guidelines and other information in the financial field in advance, and build a knowledge question and answer database covering a variety of financial regulatory information. Label and classify the content in the database, for example, by regulatory type (such as anti-money laundering regulations, securities regulations, etc.), scope of application, etc.

[0084] 2. Obtain the original prompt text and preprocess it: Financial institution employees input consulting questions about financial regulations (i.e., questions to be queried) as the original prompt text. The system performs word segmentation and vectorization on these texts and extracts the corresponding word vector representations.

[0085] 3. Filter relevant content to generate target prompt text: Based on the knowledge question and answer database, the system calculates the similarity between the word vector representation of the original prompt text and each sentence in the database. It filters out the regulatory clauses, compliance requirements and other sentences that are most relevant to the query question and combines them into the target prompt text.

[0086] 4. Generate target answers through pre-trained knowledge big model: Input the target prompt text into the pre-trained financial knowledge big model. The model generates target answers such as specific requirements for customer identification in anti-money laundering regulations based on the input target prompt text.

[0087] Technical effects:

[0088] The knowledge question-and-answer method of the present invention can quickly and accurately provide financial regulatory information, helping financial institution employees to quickly find compliance basis when handling business. This not only improves work efficiency, but also reduces compliance risks caused by unfamiliarity with regulations, ensuring the legality and compliance of business operations.

[0089] Furthermore, in one embodiment, the knowledge question answering method, wherein the step S100, constructing a knowledge question answering database, specifically comprises the steps of:

[0090] Collecting several document data related to knowledge question answering from multiple data sources;

[0091] Cleaning and normalizing the document data, and organizing it into question-and-answer data including questions and corresponding answers;

[0092] Each of the question and answer data is categorized, and the knowledge question and answer database is constructed based on all the question and answer data after categorization.

[0093] In specific implementation, the knowledge question answering method of this embodiment collects several document data related to knowledge question answering from multiple data sources in advance, cleans and normalizes them, organizes them into question answering data containing questions and corresponding answers, and then performs category labeling to finally construct a knowledge question answering database. This process can achieve the following technical effects:

[0094] 1. Improve data quality: Cleaning and normalizing document data can remove irrelevant information and formatted text content to ensure the neatness and consistency of data. This helps improve the efficiency and accuracy of subsequent data processing and reduce errors or deviations caused by data quality issues. Document data is organized into question-and-answer data containing questions and corresponding answers, thus achieving data structuring.

[0095] 2. Classification management and retrieval optimization: The question and answer data is categorized and labeled, and a knowledge question and answer database is constructed based on the categorized data, thus realizing the classification management of the data. This not only facilitates the targeted retrieval and processing of knowledge in different fields, but also improves the retrieval efficiency and accuracy of the system. For example, in the field of medical health, it can be labeled and retrieved according to categories such as disease type and treatment method; in the field of financial technology, it can be labeled and retrieved according to financial product type, regulatory category, etc.

[0096] 3. Knowledge base construction and expansion: The knowledge question-answering database constructed through the above steps provides a rich knowledge foundation for the knowledge question-answering system. This database can be continuously expanded and updated to cover more knowledge fields and question types. This enables the knowledge question-answering system to adapt to the needs of different fields and scenarios and provide a wider range of knowledge support.

[0097] The specific implementation process of the steps in this embodiment is roughly as follows:

[0098] 1. Data Collection

[0099] Collect document data: Collect several document data related to knowledge question answering from multiple data sources. This ensures that the collected data covers all aspects of the target knowledge domain to improve the comprehensiveness and accuracy of the knowledge question answering system.

[0100] 2. Data Preprocessing

[0101] Data cleaning: Clean the collected document data, remove irrelevant information, format the text content, and ensure the neatness and consistency of the data. The cleaning process includes removing HTML tags, extra spaces, special characters, etc., and correcting obvious spelling errors.

[0102] Data normalization: Normalize the cleaned data to unify the text format and encoding method. For example, convert all texts to a unified encoding format (such as UTF8), unify uppercase and lowercase, etc. Segment the text to ensure that the structure of each paragraph or sentence is clear for subsequent processing.

[0103] 3. Data collation

[0104] Organize into question-answer pairs: Organize the normalized document data into question-answer data containing questions and corresponding answers. This step requires analyzing the document content to extract clear question and answer pairs. For some documents, manual annotation or natural language processing tools may be required to identify questions and answers. For example, extracting common questions and their answers from user manuals, and extracting research questions and their conclusions from academic papers.

[0105] 4. Data Annotation

[0106] Category labeling: Label the sorted question and answer data into categories to clarify the category to which each question and answer pair belongs. Categories can be divided according to knowledge fields, such as "health care", "financial technology", "legal consulting", etc. Category labeling can be completed through manual labeling or semi-automatic labeling. Manual labeling requires the participation of domain experts to ensure the accuracy and professionalism of the labeling; semi-automatic labeling can use machine learning models to assist labeling and improve labeling efficiency.

[0107] 5. Database construction

[0108] Construct a knowledge question and answer database: Construct a knowledge question and answer database based on the question and answer data after category annotation. The database can be a relational database or a non-relational database. In the database, each question and answer pair can be used as a record, including fields such as question, answer, and category. At the same time, an index is established for each category to improve retrieval efficiency.

[0109] Data optimization and maintenance: Optimize the knowledge question and answer database, including establishing efficient indexes, optimizing query statements, configuring cache mechanisms, etc., to improve retrieval speed and response time. Regularly maintain and update the database to ensure the timeliness and accuracy of the data. For example, regularly add new question and answer pairs and update outdated information.

[0110] Through the above process, this embodiment can build a structured, clearly classified, and efficient knowledge question and answer database, providing a solid foundation for the subsequent knowledge question and answer system.

[0111] Furthermore, in one embodiment, the knowledge question answering method, wherein the step S200, obtaining the original prompt text corresponding to the question to be queried, performing word segmentation and vectorization processing on the original prompt text, and obtaining the word vector representation of the original prompt text, specifically comprises the steps of:

[0112] Acquire the original prompt text corresponding to the question to be queried input by the target user through the user interface;

[0113] Using a natural language processing tool, the original prompt text is segmented, and based on a predefined stop word list, stop words in the segmentation result are removed to obtain a segmented text;

[0114] The vectorization model is used to obtain the word vector of each word in the segmented text, and all the word vectors are aggregated to obtain the word vector representation of the original prompt text.

[0115] In specific implementation, the knowledge question answering method of this embodiment obtains the original prompt text corresponding to the query question input by the target user, and performs word segmentation, removes stop words and vectorization on it, and finally obtains the word vector representation of the original prompt text. This process can achieve the following technical effects:

[0116] 1. Accurate information extraction: By using natural language processing tools to segment the original prompt text and generate segmentation results, the original prompt text is decomposed into independent vocabulary units, which can more accurately extract the key information in the original prompt text. Based on the predefined stop word list, the stop words in the segmentation results are removed, which further removes the redundant information in the original prompt text and retains the words that contribute significantly to semantic understanding, thereby improving the efficiency and accuracy of subsequent data processing.

[0117] 2. Semantic vectorization representation: Using the vectorization model to convert the segmented text into a word vector representation can map the text information into a high-dimensional space. This vectorization representation not only retains the semantic features of the vocabulary, but also measures the semantic relevance between texts by calculating the similarity between word vectors, providing a basis for subsequent similarity calculations and information screening.

[0118] 3. Efficient information processing: Word vector representation enables text information to be processed by machine learning models in numerical form, greatly improving the efficiency of information processing. By aggregating all word vectors, the overall word vector representation of the original prompt text is obtained, and the text information is further compressed into a compact vector form, which is convenient for subsequent large model reasoning and answer generation, thereby improving the response speed and performance of the entire knowledge question answering system.

[0119] 4. Enhanced semantic understanding capability: The word vector representation obtained through the above steps can better capture the semantic features of the original prompt text, allowing the knowledge question answering system to more accurately understand the target user's question intent. This provides strong support for the subsequent generation of high-quality and accurate answers, thereby improving the overall performance and user experience of the knowledge question answering system.

[0120] Among them, the specific implementation process of the steps in this embodiment is roughly as follows:

[0121] 1. Obtain the original prompt text

[0122] User input: Obtain the original prompt text corresponding to the query problem input by the target user. This can be achieved through a user interface (such as a web form, a mobile application input box, etc.) or an API interface. Ensure that the text input by the user is captured completely and accurately for subsequent processing.

[0123] 2. Word segmentation processing

[0124] Word segmentation operation: Use natural language processing tools to perform word segmentation on the original prompt text, splitting the text into independent lexical units.

[0125] Remove stop words: Based on a predefined stop word list, remove the stop words in the word segmentation result. Stop words are usually common words that contribute less to semantics, such as "de" "shi" "he" (Chinese) or "the" "is" "and" (English). The stop word list can be customized according to the specific language and application scenario to improve the word segmentation effect.

[0126] 3. Vectorization processing

[0127] Select a vectorization model: Select a suitable pre-trained vectorization model, which can convert words into high-dimensional vectors and capture the semantic features of the words. For example, a pre-trained Word2Vec model can be selected because it performs well in text vectorization and has high computational efficiency.

[0128] Obtain word vectors: For each word in the word-segmented text, use the selected vectorization model to obtain its word vector. If a word is not in the model's vocabulary, a predefined default vector (such as a zero vector) can be used or the word can be skipped. For example, when using the Word2Vec model, the API of the model can be called to obtain the word vector of each word.

[0129] 4. Aggregate word vectors

[0130] Word vector aggregation: Aggregate the word vectors of all words in the word-segmented text to obtain an overall word vector representation of the original prompt text. Aggregation methods can include taking the average value, weighted summation, etc.

[0131] 5. Output the word vector representation

[0132] Generate word vector representation: Use the aggregated word vector as the word vector representation of the original prompt text for subsequent similarity calculation and information retrieval. Ensure that the generated word vector representation can accurately reflect the semantic features of the original prompt text and provide high-quality input for subsequent steps.

[0133] Through the above process, this embodiment can efficiently convert the original prompt text input by the user into a word vector representation, provide accurate and efficient input for the knowledge question answering system, thereby improving the overall performance of the system and user experience.

[0134] Further, in one embodiment, the knowledge question answering method, wherein the step S300, calculating the similarity between the word vector representation and each sentence in the knowledge question answering database, filtering out the content related to the word vector representation in the knowledge question answering database according to the calculation result, and generating the target prompt text, specifically comprises the steps of:

[0135] Performing vectorization processing on each of the statements in the knowledge question and answer database;

[0136] Calculating the similarity between the word vector representation and the vector representation of each of the sentences, and filtering out the target sentence in the knowledge question and answer database according to the calculation result;

[0137] All of the target sentences are combined to obtain the target prompt text.

[0138] During specific implementation, the knowledge question and answer method of this embodiment realizes efficient screening of sentences highly relevant to the user's question to be queried from the knowledge question and answer database through a series of fine operations, and generates a target prompt text. First, each sentence in the knowledge question and answer database is vectorized. This step converts the text information into a numerical vector form, so that the computer can efficiently process and compare the text content. Then, by calculating the similarity between the word vector representation of the user's question to be queried and the vector representation of each sentence in the database, the system can quantitatively evaluate the relevance of each sentence to the question to be queried. The target sentence can be screened out based on the set similarity threshold, ensuring that only sentences highly relevant to the question to be queried are selected, thereby improving the accuracy and efficiency of information retrieval. Finally, the screened target sentences are reasonably combined to generate a coherent and complete target prompt text, which provides high-quality input for subsequent answer generation, and further improves the performance and user experience of the knowledge question and answer system. This process not only optimizes the accuracy of information retrieval, but also enhances the response speed and answer quality of the system, so that the knowledge question and answer system can meet the needs of users more efficiently and accurately.

[0139] The specific implementation process of the steps in this embodiment is roughly as follows:

[0140] 1. Vectorized processing of knowledge question and answer database statements

[0141] Sentence extraction: Extract each sentence from the constructed knowledge question and answer database. These sentences can be questions, answers, or background information related to the questions.

[0142] Sentence segmentation: Segment each sentence into independent vocabulary units. The same natural language processing tool as in step S200 can be used for segmentation to ensure consistency.

[0143] Vectorization: Each sentence can be vectorized using the same vectorization model (such as Word2Vec) as in step S200 to obtain a vector representation of each sentence. Ensure that each sentence in the database is converted into a vector of a fixed dimension.

[0144] 2. Similarity calculation

[0145] Select similarity calculation method: Select an appropriate similarity calculation method, such as cosine similarity, Euclidean distance, etc. Cosine similarity is a commonly used similarity calculation method that can effectively measure the cosine value of the angle between two vectors, thereby reflecting their similarity.

[0146] Calculate similarity: For each sentence vector representation in the knowledge question answering database, calculate the similarity between it and the word vector representation of the user's query question. Specifically, calculate the cosine similarity between the word vector representation of the user's question and the vector representation of each sentence.

[0147] 3. Filter target sentences

[0148] Set similarity threshold: Preset a similarity threshold based on the application scenario and requirements. This threshold is used to filter sentences that are sufficiently relevant to the user's question. For example, the cosine similarity threshold can be set to 0.7.

[0149] Filter sentences: Filter sentences from the knowledge question and answer database whose similarity score is greater than the preset threshold. These sentences are considered to be highly relevant to the user's question and can be used as target sentences. If no sentence's similarity reaches the threshold, the threshold can be appropriately lowered or the top N most similar sentences can be returned.

[0150] 4. Statement combination

[0151] Semantic analysis: Perform semantic analysis on the selected target sentences to determine the semantic relationship between the target sentences. Natural language processing technology (such as dependency syntactic analysis, semantic role labeling, etc.) can be used to analyze the logical relationship between sentences.

[0152] Arrange in order: Arrange the target sentences in a logical order according to their semantic relationships. For example, you can arrange them according to the contextual relationship between the target sentences. During the arrangement process, you can add appropriate conjunctions or punctuation marks to enhance the coherence of the text.

[0153] 5. Generate target prompt text

[0154] Combine text: Combine the arranged target sentences into a coherent text as the target prompt text. Ensure that the generated text is semantically coherent and logically reasonable.

[0155] Optimize text: Optimize the generated target prompt text and adjust the order or structure of the sentences to improve the logic and readability of the text. You can polish the text to make it more in line with the expression habits of natural language.

[0156] Check and revise: Check the combined text to ensure that it is complete, coherent and grammatically correct. If any incoherent or repetitive parts are found in the text, make appropriate revisions or delete them.

[0157] Through the above process, this embodiment can efficiently filter out sentences that are highly relevant to user questions from the knowledge question and answer database, and generate coherent and accurate target prompt texts, providing support for the subsequent generation of high-quality answers by the large knowledge model.

[0158] Furthermore, in one embodiment, the knowledge question answering method, wherein the calculating the similarity between the word vector representation and the vector representation of each of the sentences, and filtering out the target sentence in the knowledge question answering database according to the calculation result, specifically comprises the steps of:

[0159] Pre-set similarity threshold;

[0160] Calculating the similarity score between the word vector representation and the vector representation of each of the sentences using a cosine similarity algorithm;

[0161] The target sentences having a similarity score greater than the similarity threshold are screened out from the knowledge question and answer database.

[0162] In specific implementation, the knowledge question and answer method of this embodiment calculates the similarity score between the word vector representation and the vector representation of each sentence by using the cosine similarity algorithm, and screens out the target sentence whose similarity score is greater than the preset similarity threshold. This implementation method realizes the accurate quantitative evaluation and efficient screening of the relevance between the sentence and the user's query question in the knowledge question and answer database. As a method widely used in text similarity calculation, the cosine similarity algorithm can effectively measure the cosine value of the angle between word vectors, thereby accurately reflecting the semantic relevance between the sentence and the query question. By presetting the similarity threshold, the system can accurately screen out sentences that are highly relevant to the user's query question, avoid the interference of irrelevant or low-relevance information, and significantly improve the accuracy and efficiency of information retrieval. This process not only ensures the high quality and high relevance of the answers generated subsequently, but also optimizes the performance of the system, so that the knowledge question and answer system can provide users with the required answers more quickly and accurately, greatly improving the user experience and the practicality of the system.

[0163] Furthermore, in one embodiment, the knowledge question answering method, wherein the combining of all the target sentences to obtain the target prompt text, specifically comprises the steps of:

[0164] Performing semantic analysis on all the selected target sentences to determine the semantic relationship between the target sentences;

[0165] According to the semantic relationship, the target sentences are arranged in a logical order to obtain the target prompt text.

[0166] During specific implementation, the knowledge question and answer method of this embodiment determines the semantic relationship between each target sentence through semantic analysis, and arranges each target sentence in order according to logic based on these relationships to generate a target prompt text. This process significantly improves the coherence and logic of generating the target prompt text. Specifically, semantic analysis can deeply understand the meaning of each target sentence and its relationship with each other, thereby providing a basis for the reasonable sorting of the target sentence. This arrangement method based on semantic relations not only ensures the semantic coherence of the generated target prompt text and avoids information fragmentation, but also can better reflect the logical structure in the original document, making the generated text easier to understand and use. The target prompt text finally obtained not only contains all the key information related to the user's query question, but also presents it in a logically clear and naturally expressed way, providing a higher quality input for the subsequent knowledge large model to generate high-quality answers, and further enhancing the performance and user experience of the entire knowledge question and answer system.

[0167] Furthermore, in one embodiment, the knowledge question answering method, wherein the step S400, taking the target prompt text as input and generating the target answer to the question to be queried through the pre-trained knowledge model, specifically comprises the steps of:

[0168] Acquire a fine-tuning dataset related to the target knowledge domain, and use the fine-tuning dataset to fine-tune the pre-trained large model to obtain the knowledge large model;

[0169] Formatting the target prompt text, inputting the formatted target prompt text into the knowledge macro model, and generating the target answer to the question to be queried;

[0170] The target answer is output or displayed.

[0171] In specific implementation, the knowledge question and answer method of this embodiment realizes an efficient, accurate and domain-adaptable knowledge question and answer function by fine-tuning the pre-trained large model, formatting the target prompt text, and generating and outputting the target answer. Specifically, the pre-trained large model is fine-tuned using a fine-tuning data set related to the target knowledge domain, wherein the pre-trained large model, i.e., the pre-trained large language model (LLM), can make the generated knowledge large model better adapt to the language style and knowledge structure of a specific domain (i.e., the target knowledge domain), thereby improving the question and answer performance of the knowledge large model in this domain. The target prompt text is formatted to ensure that the input data meets the input requirements of the knowledge large model, further improving the processing efficiency of the knowledge large model and the accuracy of the generated answers. Finally, the generated target answer is output or displayed, providing users with clear and accurate answers, meeting the user's needs for efficient and accurate knowledge questions and answers. This process not only makes full use of the powerful language generation ability of the pre-trained large model, but also ensures the quality and practicality of the answers generated by the knowledge large model obtained after fine-tuning through domain fine-tuning and text formatting, significantly improving the performance and user experience of the entire knowledge question and answer system.

[0172] The specific implementation process of the steps in this embodiment is roughly as follows:

[0173] 1. Obtain a fine-tuning dataset related to the target knowledge domain

[0174] Data collection: Collect annotated data related to the target knowledge domain, which should include questions and their corresponding answers or relevant background information. Data sources can include professional books, academic papers, industry reports, user manuals, etc.

[0175] Data labeling: Label the collected data, clearly mark the questions and their corresponding answers or related background information. Domain experts can be introduced in the labeling process to ensure the accuracy and professionalism of the labeling.

[0176] Data preprocessing: Clean and normalize the labeled data to ensure the neatness and consistency of the data. For example, remove HTML tags, extra spaces, special characters, etc., and unify the text format and encoding method.

[0177] 2. Fine-tune the pre-trained large model

[0178] Select a pre-trained model: According to the characteristics and requirements of the target knowledge domain, select a suitable pre-trained large model (such as BERT, GPT, etc.). Make sure that the selected large model has been fully pre-trained and has the ability to handle knowledge question-answering tasks.

[0179] Fine-tune the model: Use the prepared fine-tuning dataset to fine-tune the pre-trained large model. During the fine-tuning process, adjust the parameters of the model to better fit the target knowledge domain. During the fine-tuning process, use the validation set to evaluate the model performance and adjust the hyperparameters (such as learning rate, batch size, number of training rounds, etc.) to optimize the model performance.

[0180] Save the fine-tuned model: Save the fine-tuned model as a large knowledge model for subsequent use. Record the key parameters and performance indicators during the fine-tuning process to provide a reference for the continuous optimization of the model.

[0181] 3. Format the target prompt text

[0182] Text preprocessing: Perform necessary formatting on the target prompt text to ensure that it meets the input requirements of the knowledge model. Formatting may include text encoding, length adjustment, adding special tags, etc. For example, convert the text to UTF8 encoding to ensure that the text length meets the maximum input length limit of the model.

[0183] Add context information (optional): If necessary, you can add additional context information to the target prompt text to help the knowledge model better understand the background of the query. For example, you can attach information such as the category of the question, keywords in related fields, etc. to the target prompt text.

[0184] 4. Generate target answers

[0185] Input target prompt text: Input the formatted target prompt text into the knowledge model. The model will generate the target answer to the query question based on the input prompt text.

[0186] Answer generation: The model generates the target answer based on the input prompt text. The generated answer should have high accuracy and readability and be able to directly answer the user's question.

[0187] Answer post-processing (optional): Perform necessary post-processing on the generated target answer to ensure that it meets the requirements of the actual application. For example, the answer can be grammatically checked, duplicate content can be removed, and the format can be adjusted to improve the quality and readability of the answer.

[0188] 5. Output or display the target answer

[0189] Answer output: Output or display the generated target answer to the user or system, such as web page display, mobile application interface display, API return, etc.

[0190] Result record: The generated target answer and its corresponding original question are stored in the system to form a question and answer record. Timestamp, user ID and other information are added to each question and answer record to facilitate subsequent query and analysis.

[0191] Through the above process, this embodiment can efficiently use the pre-trained knowledge model to generate high-quality target answers, meeting the user's needs for efficient and accurate knowledge questions and answers.

[0192] It can be seen from the above method embodiments that the knowledge question and answer method provided by the present invention includes: constructing a knowledge question and answer database; obtaining the original prompt text corresponding to the question to be queried, performing word segmentation and vectorization processing on the original prompt text, and obtaining a word vector representation of the original prompt text; calculating the similarity between the word vector representation and each sentence in the knowledge question and answer database, and filtering out the content related to the word vector representation in the knowledge question and answer database according to the calculation result, and generating a target prompt text; using the target prompt text as input, and generating a target answer to the question to be queried through a pre-trained knowledge model. In this way, the method of the present invention can effectively process the original prompt text of the question to be queried and accurately understand the key information, and generate high-quality answers, which can meet the user's needs for efficient and accurate knowledge questions and answers.

[0193] It should be understood that, although the present application provides method operation steps as described in the embodiments or flowcharts, more or less operation steps may be included based on conventional or non-creative labor, and these operation steps are not necessarily performed in sequence according to the embodiment or flowchart. The order of steps listed in the embodiment or flowchart is only one way of executing the order of many steps, and does not represent the only execution order. It should be noted that there is not necessarily a certain order between the above steps. A person of ordinary skill in the art can understand from the description of the embodiment of the present invention that in different embodiments, the above steps may have different execution orders, that is, they may be executed in parallel, or they may be executed in exchange, etc. Moreover, at least a part of the steps in the embodiment or flowchart may include multiple sub-steps or multiple stages, and these sub-steps or stages are not necessarily executed at the same time, but may be executed at different times, and the execution order of these sub-steps or stages is not necessarily performed in sequence, but may be executed in turn, alternately or synchronously with other steps or at least a part of the sub-steps or stages of other steps.

[0194] Based on the above method embodiment, please refer to Figure 3 Another embodiment of the present invention further provides a knowledge question answering device, wherein the device comprises:

[0195] A construction module 11 is used to construct a knowledge question and answer database;

[0196] The acquisition module 12 is used to acquire the original prompt text corresponding to the query question, perform word segmentation and vectorization processing on the original prompt text, and obtain a word vector representation of the original prompt text;

[0197] A screening module 13 is used to calculate the similarity between the word vector representation and each sentence in the knowledge question and answer database, and screen out the content related to the word vector representation in the knowledge question and answer database according to the calculation result to generate a target prompt text;

[0198] The answer generation module 14 is used to take the target prompt text as input and generate the target answer to the question to be queried through a pre-trained knowledge model.

[0199] Furthermore, in one embodiment, in the knowledge question answering device, the construction module 11 is specifically used for:

[0200] Collecting several document data related to knowledge question answering from multiple data sources;

[0201] Cleaning and normalizing the document data, and organizing it into question-and-answer data including questions and corresponding answers;

[0202] Each of the question and answer data is categorized, and the knowledge question and answer database is constructed based on all the question and answer data after categorization.

[0203] Furthermore, in one embodiment, in the knowledge question answering device, the acquisition module 12 is specifically used to:

[0204] Acquire the original prompt text corresponding to the question to be queried input by the target user through the user interface;

[0205] Using a natural language processing tool, the original prompt text is segmented, and based on a predefined stop word list, stop words in the segmentation result are removed to obtain a segmented text;

[0206] The vectorization model is used to obtain the word vector of each word in the segmented text, and all the word vectors are aggregated to obtain the word vector representation of the original prompt text.

[0207] Furthermore, in one embodiment, in the knowledge question answering device, the screening module 13 is specifically used for:

[0208] Performing vectorization processing on each of the statements in the knowledge question and answer database;

[0209] Calculating the similarity between the word vector representation and the vector representation of each of the sentences, and filtering out the target sentence in the knowledge question and answer database according to the calculation result;

[0210] All of the target sentences are combined to obtain the target prompt text.

[0211] Furthermore, in one embodiment, the knowledge question answering device, wherein the calculating the similarity between the word vector representation and the vector representation of each of the sentences, and filtering out the target sentence in the knowledge question answering database according to the calculation result, specifically includes:

[0212] Pre-set similarity threshold;

[0213] Calculating the similarity score between the word vector representation and the vector representation of each of the sentences using a cosine similarity algorithm;

[0214] The target sentences having a similarity score greater than the similarity threshold are screened out from the knowledge question and answer database.

[0215] Furthermore, in one embodiment, the knowledge question answering device, wherein the combining all the target sentences to obtain the target prompt text specifically includes:

[0216] Performing semantic analysis on all the selected target sentences to determine the semantic relationship between the target sentences;

[0217] According to the semantic relationship, the target sentences are arranged in a logical order to obtain the target prompt text.

[0218] Furthermore, in one embodiment, in the knowledge question and answer device, the answer generation module 14 is specifically used to:

[0219] Acquire a fine-tuning dataset related to the target knowledge domain, and use the fine-tuning dataset to fine-tune the pre-trained large model to obtain the knowledge large model;

[0220] Formatting the target prompt text, inputting the formatted target prompt text into the knowledge macro model, and generating the target answer to the question to be queried;

[0221] The target answer is output or displayed.

[0222] It should be noted that in the embodiment of the device of the present invention, the information interaction, execution process and other contents between the above-mentioned modules are based on the same concept as the embodiment of the method of the present invention. Their specific functions and technical effects can be found in the aforementioned method embodiment part and will not be repeated here.

[0223] Based on the above method embodiment, another embodiment of the present invention further provides a computer device, which may be a server, and its internal structure diagram may be as follows: Figure 4 As shown. The computer device includes a processor, a memory, a network interface and a database connected via a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external terminal via a network connection. When the computer program is executed by the processor, it implements the functions or steps on the server side of the knowledge question and answer method in any of the above method embodiments.

[0224] Based on the above method embodiment, another embodiment of the present invention further provides a computer device, which may be a client, and its internal structure diagram may be as follows: Figure 5As shown. The computer device includes a processor, a memory, a network interface, a display screen and an input device connected via a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external terminal via a network connection. When the computer program is executed by the processor, it implements the functions or steps of the client side of the knowledge question and answer method in any of the above method embodiments.

[0225] Those skilled in the art will understand that Figure 4 and Figure 5 The structural schematic diagram shown in the figure is only a schematic diagram of a partial structure related to the solution of the present invention, and does not constitute a limitation on the computer device to which the solution of the present invention is applied. The specific computer device may include more components than those shown in the figure, or combine certain components, or have a different arrangement of components.

[0226] The processor may be a CPU, or other general-purpose processors, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor, or the processor may be any conventional processor, etc.

[0227] Among them, the memory includes a readable storage medium, an internal memory, etc., wherein the internal memory can be the memory of a computer device, and the internal memory provides an environment for the operation of an operating system and computer-readable instructions in the readable storage medium. The readable storage medium can be a hard disk of a computer device, and in other embodiments, it can also be an external storage device of a computer device, for example, a plug-in hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (Flash Card), etc. equipped on a computer device. Further, the memory can also include both an internal storage unit of a computer device and an external storage device. The memory is used to store an operating system, an application program, a boot loader (BootLoader), data, and other programs, such as the program code of a computer program, etc. The memory can also be used to temporarily store data that has been output or is to be output.

[0228] Based on the above method embodiments, another embodiment of the present invention further provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, wherein when the computer program is executed by a processor, the knowledge question answering method in any of the above method embodiments is implemented. The computer-readable storage medium may be non-volatile or volatile.

[0229] It should be noted that the above-mentioned functions or steps that can be implemented by the computer-readable storage medium or computer device, and the technical effects brought about by the functions / steps, can be found in the relevant descriptions in the aforementioned method embodiments. To avoid repetition, they will not be described one by one here.

[0230] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM). The disclosed memory components or memories of the operating environments described herein are intended to comprise one or more of these and / or any other suitable types of memory.

[0231] The technicians in the relevant field can clearly understand that for the convenience and simplicity of description, in the embodiment of the device of the present invention, only the division of the above-mentioned functional units and modules is used as an example. In practical applications, the above-mentioned function allocation can be completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiment can be integrated in a processing unit, or each unit can exist physically separately, or two or more units can be integrated in one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of software functional units. In addition, the specific names of the functional units and modules are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of the present invention. The specific working process of the units and modules in the above-mentioned device can refer to the corresponding process in the above-mentioned method embodiment, which will not be repeated here. If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium.

[0232] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention.

[0233] In the embodiments provided by the present invention, it should be understood that the disclosed devices / computer equipment and methods can be implemented in other ways. For example, the device / computer equipment embodiments described above are only schematic, for example, the division of modules or units is only a logical function division, and there may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0234] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0235] It should be noted that if software tools or components other than those of the Company appear in the embodiments of the present application, they are only used for illustration and do not represent actual use. The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the above embodiments, a person of ordinary skill in the art should understand that the technical solutions described in the above embodiments can still be modified, or some of the technical features thereof can be replaced by equivalents; and these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention, and should be included in the protection scope of the present invention.

Claims

1. A knowledge question answering method, characterized in that: include: Build a knowledge question-answering database; Obtaining the original prompt text corresponding to the question to be queried, performing word segmentation and vectorization processing on the original prompt text, and obtaining a word vector representation of the original prompt text; Calculate the similarity between the word vector representation and each sentence in the knowledge question and answer database, filter out the content related to the word vector representation in the knowledge question and answer database according to the calculation result, and generate a target prompt text; The target prompt text is used as input, and the target answer to the question to be queried is generated through a pre-trained knowledge model.

2. The knowledge question answering method according to claim 1, characterized in that: The constructing of the knowledge question and answer database comprises: Collecting several document data related to knowledge question answering from multiple data sources; Cleaning and normalizing the document data, and organizing it into question-and-answer data including questions and corresponding answers; Each of the question and answer data is categorized, and the knowledge question and answer database is constructed based on all the question and answer data after categorization.

3. The knowledge question answering method according to claim 1, characterized in that: The obtaining of the original prompt text corresponding to the question to be queried, performing word segmentation and vectorization processing on the original prompt text, and obtaining a word vector representation of the original prompt text includes: Acquire the original prompt text corresponding to the question to be queried input by the target user through the user interface; Using a natural language processing tool, the original prompt text is segmented, and based on a predefined stop word list, stop words in the segmentation result are removed to obtain a segmented text; The vectorization model is used to obtain the word vector of each word in the segmented text, and all the word vectors are aggregated to obtain the word vector representation of the original prompt text.

4. The knowledge question answering method according to claim 1, characterized in that: The calculating the similarity between the word vector representation and each sentence in the knowledge question and answer database, filtering out the content in the knowledge question and answer database related to the word vector representation according to the calculation result, and generating the target prompt text includes: Performing vectorization processing on each of the statements in the knowledge question and answer database; Calculating the similarity between the word vector representation and the vector representation of each of the sentences, and filtering out the target sentence in the knowledge question and answer database according to the calculation result; All of the target sentences are combined to obtain the target prompt text.

5. The knowledge question answering method according to claim 4, characterized in that: The calculating the similarity between the word vector representation and the vector representation of each of the sentences, and filtering out the target sentence in the knowledge question and answer database according to the calculation result, comprises: Pre-set similarity threshold; Calculating the similarity score between the word vector representation and the vector representation of each of the sentences using a cosine similarity algorithm; The target sentences having a similarity score greater than the similarity threshold are screened out from the knowledge question and answer database.

6. The knowledge question answering method according to claim 4, characterized in that: The step of combining all the target sentences to obtain the target prompt text includes: Performing semantic analysis on all the selected target sentences to determine the semantic relationship between the target sentences; According to the semantic relationship, the target sentences are arranged in a logical order to obtain the target prompt text.

7. The knowledge question answering method according to claim 1, characterized in that: The method of taking the target prompt text as input and generating the target answer to the question to be queried through the pre-trained knowledge model includes: Acquire a fine-tuning dataset related to the target knowledge domain, and use the fine-tuning dataset to fine-tune the pre-trained large model to obtain the knowledge large model; Formatting the target prompt text, inputting the formatted target prompt text into the knowledge macro model, and generating the target answer to the question to be queried; The target answer is output or displayed.

8. A knowledge question-answering device, characterized in that: include: A construction module, used to construct a knowledge question-answering database; An acquisition module is used to acquire the original prompt text corresponding to the question to be queried, perform word segmentation and vectorization processing on the original prompt text, and obtain a word vector representation of the original prompt text; A screening module, used to calculate the similarity between the word vector representation and each sentence in the knowledge question and answer database, and screen out the content related to the word vector representation in the knowledge question and answer database according to the calculation result, and generate a target prompt text; The answer generation module is used to take the target prompt text as input and generate the target answer to the question to be queried through a pre-trained knowledge model.

9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the knowledge question and answer method as described in any one of claims 1 to 7 is implemented.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the knowledge question and answer method as described in any one of claims 1 to 7 is implemented.

Citation Information

Cited By

  • Question answering method for compliance consultation and knowledge base generation method and device

    CN120653815A

  • Evidence-based medical question and answer method and device, electronic equipment and storage medium

    CN121144469A