A question-and-answer robot assisted by a large language model

By adopting large language model assistive technology in question-and-answer robots, building vector knowledge bases and matching data, the problem of limited Q&A capabilities and accuracy of traditional Q&A robots is solved, and more efficient and accurate Q&A services are achieved, reducing costs and improving scalability.

CN118349652BActive Publication Date: 2025-05-27BEIJING HUAYUN WORLD TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410497084.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-04-24
Publication Date
2025-05-27
Estimated Expiration
2044-04-24

AI Technical Summary

Technical Problem

Traditional Q&A robots rely on predefined rules and limited training data, resulting in limited question-answer capabilities and accuracy, and require a lot of manpower investment, which is costly and inefficient.

Method used

A question-and-answer robot based on a large language model is adopted to build a large language model, obtain application scenarios and generate question prompt data, determine the question-and-answer intent, generate similar text data, review whether the data meets the standards, convert it into text vectors, and build a vector knowledge base to match user question-and-answer data to provide answers.

Benefits of technology

It improves the accuracy and efficiency of the Q&A system, can better understand user questions and provide relevant answers, reduces labor costs, improves work efficiency, and has strong scalability. It is suitable for Q&A systems in various fields and scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118349652B_ABST
    Figure CN118349652B_ABST
Patent Text Reader

Abstract

The present invention belongs to the technical field of robot intelligent question answering, and specifically relates to a question answering robot assisted by a large language model. By generating a large number of similar questions and answers, and screening and correcting them according to the review criteria, the accuracy of the question answering system can be improved, enabling it to better answer the questions raised by users. Since the question answering system can more accurately understand the user's questions and give relevant answers, it can enhance user satisfaction and experience, and strengthen the user's trust in the product or service. Compared with the traditional manual question answering system, using a large language model to assist in question answering can save labor costs and improve work efficiency, especially when dealing with a large number of repetitive questions. Due to the generality and flexibility of the large language model, this method has strong scalability and can be applied to question answering systems in various different fields and scenarios to meet the needs of different user groups.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of robot intelligent question answering, and specifically relates to a question answering robot assisted by a large language model. Background Art

[0002] With the rapid development of artificial intelligence technology, question answering robots have been widely used in fields such as intelligent customer service, online education, and e-commerce. The core function of a question answering robot is to understand the question input by the user and match the corresponding answer from a preset knowledge base. The traditional intention matching methods of question answering robots mainly rely on two strategies: one is to pre-configure similar intentions manually; the other is to generate similar intentions by training models such as seq2seq.

[0003] In the first strategy, professionals in the field need to configure each possible intention one by one. This method is not only time-consuming and laborious, but also difficult to cover all possible scenarios, resulting in limited recognition rate of the question answering robot. In the second strategy, although intention recognition can be automatically performed through machine learning models, a large amount of labeled corpus is required for model training, which also requires human input and a professional algorithm engineer team for model tuning and maintenance.

[0004] Therefore, traditional question answering robots are often limited by predefined rules and limited training data, resulting in limited question answering ability and accuracy. Whether it is manually configuring intentions or training models, a large amount of time and human input are required, resulting in high costs. The method of manually configuring intentions is difficult to cover all possible scenarios, while the method of training models requires a large amount of labeled corpus and long-term training, resulting in low efficiency. Summary of the Invention

[0005] The purpose of the present invention is to provide a question answering robot assisted by a large language model, which can more accurately understand the user's question and provide accurate answers.

[0006] The technical solution adopted by the present invention is specifically as follows:

[0007] A question answering robot assisted by a large language model is applied to a question answering method assisted by a large language model. The question answering method assisted by a large language model includes:

[0008] Construct a large language model;

[0009] Obtain the application scenario of the question answering robot, generate question prompt data, and determine the question answering intention of the large language model according to the prompt data;

[0010] According to the question answering intention, the large language model generates similar text data with a specific intention;

[0011] Determine whether the similar text data meets the review criteria;

[0012] If it does not meet the review criteria, it indicates that the intent of the similar text data is inaccurate, and the redundant and incorrect text information in the similar text data needs to be removed;

[0013] If it meets the review criteria, it indicates that the intent of the similar text data is accurate;

[0014] Convert the text information in the similar text data that passes the review criteria into text vectors and build a vector knowledge base;

[0015] Obtain the Q&A data of the user, match the similar intent text information according to the vector knowledge base, and feedback the corresponding answers of the similar intent text information to the user.

[0016] In a preferred solution, the steps of obtaining the application scenario of the Q&A robot, generating question prompt data, and determining the Q&A intent of the large language model according to the prompt data include:

[0017] Obtain the application scenario of the Q&A robot;

[0018] Obtain the key Q&A business areas and common Q&A conversations according to the application scenario;

[0019] Obtain the Q&A information pairs corresponding to the key Q&A business areas and common Q&A conversations as training data according to the Q&A business areas and common Q&A scenarios;

[0020] Extract question features from the training data and generate question prompt data for training the large language model;

[0021] Train the large language model according to the generated question prompt data to obtain the Q&A intent corresponding to the application scenario.

[0022] In a preferred solution, the steps of generating similar text data with a specific intent by the large language model according to the Q&A intent include:

[0023] Obtain the Q&A text related to the Q&A intent in the internal knowledge base of the large language model according to the Q&A intent;

[0024] Obtain the Q&A similar text that matches the Q&A intent in the Q&A text;

[0025] Generate similar text data with a specific intent according to the Q&A similar text.

[0026] In a preferred solution, the steps of determining whether the similar text data meets the review criteria include:

[0027] Obtain the text review standard evaluation interval;

[0028] Extract semantic feature information and question keyword information from similar text data;

[0029] According to the semantic feature information and question keyword information, determine whether the similar text data belongs to the text review standard evaluation range;

[0030] If it does not meet the review standard, it indicates that the intention of the similar text data is inaccurate, and redundant and incorrect text information in the similar text data needs to be removed;

[0031] If it meets the review standard, it indicates that the intention of the similar text data is accurate.

[0032] In a preferred embodiment, the step of converting the text information passing the review standard in the similar text data into a text vector and constructing a vector knowledge base includes:

[0033] Obtain the text information in the similar text data passing the review standard;

[0034] Extract multiple feature information blocks from the text information;

[0035] Convert the text information into a text vector according to the multiple feature information blocks;

[0036] Summarize the text vectors into a vector knowledge base.

[0037] In a preferred embodiment, the step of obtaining the Q&A data of the user, matching the text information with similar intentions according to the vector knowledge base, and feeding back the answers corresponding to the text information with similar intentions to the user includes:

[0038] Obtain the Q&A data of the user;

[0039] Extract the Q&A intention from the Q&A data and convert the Q&A data into a Q&A text vector;

[0040] Match according to the Q&A intention and the Q&A text vector with the vector knowledge base, and label the matching result as text information with similar intentions;

[0041] Extract the answers from the text information with similar intentions and feed them back to the user.

[0042] In a preferred embodiment, after the step of obtaining the Q&A data of the user, matching the text information with similar intentions according to the vector knowledge base, and feeding back the answers corresponding to the text information with similar intentions to the user, it further includes:

[0043] Generate an inquiry instruction according to the answer corresponding to the text information with similar intentions fed back to the user;

[0044] The user selects whether it is the required answer according to the inquiry instruction;

[0045] If so, it indicates that the answer corresponding to the similar intention text information is correct, and the result is directly output;

[0046] If not, it indicates that the answer corresponding to the similar intention text information is incorrect, and the answer corresponding to the similar intention text information is re-matched.

[0047] In a preferred embodiment, after the step of re-matching the answer corresponding to the similar intention text information, it further includes:

[0048] Mark the incorrect answer as a questionable answer and record the marking times;

[0049] Mark the re-matched correct answer as a suspected answer and record the marking times;

[0050] When the marking times of the suspected answer are greater than those of the questionable answer, mark the suspected answer as the correct answer and the questionable answer as the incorrect answer.

[0051] The present invention also provides a question-answering system assisted by a large language model for the above-mentioned question-answering robot assisted by a large language model, including:

[0052] A language module for constructing a large language model;

[0053] An intention module for obtaining the application scenario of the question-answering robot and generating question prompt data, and determining the question-answering intention of the large language model according to the prompt data;

[0054] A text generation module for generating similar text data with a specific intention according to the question-answering intention and the large language model;

[0055] An audit module for judging whether the similar text data meets the audit criteria;

[0056] If it does not meet the audit criteria, it indicates that the intention of the similar text data is inaccurate, and the redundant and incorrect text information in the similar text data needs to be removed;

[0057] If it meets the audit criteria, it indicates that the intention of the similar text data is accurate;

[0058] A vector module for converting the text information passing the audit criteria in the similar text data into text vectors and constructing a vector knowledge base;

[0059] A matching and retrieval module for obtaining the question-answering data of the user and matching the similar intention text information according to the vector knowledge base, and feeding back the answer corresponding to the similar intention text information to the user.

[0060] And, a question-answering terminal assisted by a large language model, including:

[0061] One or more processors;

[0062] A storage device on which one or more programs are stored;

[0063] When the one or more programs are executed by the one or more processors, the one or more processors implement the large language model-assisted question and answer robot.

[0064] The technical effects achieved by the present invention are as follows:

[0065] In the present invention, by generating a large number of similar questions and answers and screening and correcting them according to the review criteria, the accuracy of the question and answer system can be improved, enabling it to better answer the questions raised by users. Since the question and answer system can more accurately understand user questions and give relevant answers, it can enhance user satisfaction and experience, and strengthen users' trust in the product or service. Compared with traditional manual question and answer systems, using a large language model to assist in question and answer can save labor costs and improve work efficiency, especially when dealing with a large number of repetitive questions. Due to the versatility and flexibility of the large language model, this method has strong scalability and can be applied to question and answer systems in various different fields and scenarios to meet the needs of different user groups. Description of the Drawings

[0066] Figure 1 is a flowchart of the method provided by the present invention;

[0067] Figure 2 is a system module diagram provided by the present invention. Detailed Embodiments

[0068] In order to make the above objects, features, and advantages of the present invention more obvious and understandable, the following will describe the detailed embodiments of the present invention in conjunction with the drawings of the specification.

[0069] In the following description, many specific details are set forth to fully understand the present invention. However, the present invention can also be implemented in other ways different from those described herein. Those skilled in the art can make similar extensions without departing from the connotation of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed below.

[0070] Secondly, the so-called "one embodiment" or "embodiment" herein refers to a specific feature, structure, or characteristic that can be included in at least one implementation manner of the present invention. The "in a preferred embodiment" that appears in different places in this specification does not necessarily refer to the same embodiment, nor is it a separate or selectively exclusive embodiment from other embodiments.

[0071] Thirdly, the present invention is described in detail in conjunction with the schematic diagrams. When describing the embodiments of the present invention in detail, for the sake of convenience of explanation, the schematic diagrams are only examples and should not limit the scope of protection of the present invention herein.

[0072] Please refer to the attached Figure 1 As shown, a question-answering robot assisted by a large language model is provided, which is applied to a question-answering method assisted by a large language model. The large language model-assisted question-answering method includes:

[0073] S1. Build a large language model;

[0074] S2. Obtain the application scenario of the question-answering robot, generate question prompt data, and determine the question-answering intention of the large language model according to the prompt data;

[0075] S3. According to the question-answering intention, the large language model generates similar text data with a specific intention;

[0076] S4. Judge whether the similar text data meets the review criteria;

[0077] If it does not meet the review criteria, it means that the intention of the similar text data is inaccurate, and redundant and incorrect text information in the similar text data needs to be removed;

[0078] If it meets the review criteria, it means that the intention of the similar text data is accurate;

[0079] S5. Convert the text information passing the review criteria in the similar text data into text vectors and build a vector knowledge base;

[0080] S6. Obtain the question-answering data of the user, match the similar intention text information according to the vector knowledge base, and feedback the corresponding answers of the similar intention text information to the user.

[0081] In the above steps S1 to S6, first, a large language model needs to be trained, such as the GPT series models, to be able to understand and generate natural language text. Determine the application scenarios of the Q&A robot, such as the customer service field, the education field, the entertainment field, the historical museum field, etc. Generate question prompt data, which can be questions that users may ask, used to guide the model to generate relevant answers. Utilize the large language model to generate similar questions and answers based on the question prompt data. These similar text data should cover various questions that users may ask and the corresponding answers. Review the generated similar text data to determine whether it meets the pre-set review criteria. If it does not meet the criteria, the similar text data needs to be corrected to remove redundant and incorrect text information. Manual review and screening can be adopted, and the process of manual review and screening is also more efficient and accurate, avoiding complex algorithms and models in traditional methods, simplifying the system structure and maintenance difficulty, quickly analyzing and understanding the questions input by users, improving the speed of system launch, ensuring intention accuracy. If it meets the criteria, it means that the intention of the similar text data is accurate, and the subsequent steps can be continued. Convert the similar text data that passes the review criteria into text vectors for vector matching. Build a vector knowledge base to store these text vectors for subsequent matching with the questions asked by users. Obtain the Q&A data of users, convert the questions asked by users into text vectors, utilize the vector knowledge base to match the text information similar to the user's questions, and then feedback the corresponding answers to users. By generating a large number of similar questions and answers and screening and correcting them according to the review criteria, the accuracy of the Q&A system can be improved, enabling it to better answer the questions asked by users. Since the Q&A system can more accurately understand the user's questions and give relevant answers, it can improve the user's satisfaction and experience, and enhance the user's trust in the product or service. Compared with traditional manual Q&A systems, using a large language model to assist in Q&A can save labor costs and improve work efficiency, especially when dealing with a large number of repetitive questions. Due to the versatility and flexibility of the large language model, this method has strong scalability and can be applied to Q&A systems in various different fields and scenarios to meet the needs of different user groups.

[0082] The steps of obtaining the application scenarios of the Q&A robot and generating question prompt data, and determining the Q&A intention of the large language model according to the prompt data include:

[0083] S201. Obtain the application scenarios of the Q&A robot;

[0084] S202. Obtain the key Q&A business fields and common Q&A conversations according to the application scenarios;

[0085] S203. Obtain the Q&A information pairs corresponding to the key Q&A business fields and common Q&A conversations as training data according to the Q&A business fields and common Q&A scenarios;

[0086] S204. Extract question features from the training data and generate question prompt data for training the large language model;

[0087] S205. Train the large language model according to the generated question prompt data to obtain the Q&A intent corresponding to the application scenario.

[0088] In the above steps S201 to S205, first, it is necessary to determine the application scenario of the Q&A robot, which may be the customer service field, the education field, the medical field, the historical museum field, etc. After determining the application scenario, it is necessary to deeply understand the key Q&A business areas and common Q&A conversations in this scenario. For example, in the customer service field, it may involve product usage questions, order inquiries, etc. Collect real Q&A information pairs related to the key Q&A business areas and common Q&A conversations as training data. These data can come from existing documents, historical conversation records, etc. Process the collected training data, extract the features of the questions, and generate question prompt data for training the large language model. These data can include question text, question type, question topic, etc. Use the generated question prompt data to train the large language model so that it can understand and generate Q&A intents corresponding to the application scenario. In this way, the model can accurately answer relevant questions according to the questions raised by users. By deeply understanding the specific application scenario and business requirements, the Q&A service can be customized to provide more accurate and effective solutions. Since the training data is obtained from real scenarios, the trained large language model is more targeted and accurate, and can better understand and answer users' questions. The Q&A model trained for specific scenarios and business requirements can respond to users' questions more quickly and give accurate and useful answers, thereby improving the user experience. By collecting and processing training data in a targeted manner, the time and cost required for training the model can be reduced, and the training efficiency can be improved.

[0089] The steps for the large language model to generate similar text data with specific intents according to the Q&A intent include:

[0090] S301. According to the Q&A intent, obtain the Q&A texts related to the Q&A intent in the internal knowledge base of the large language model;

[0091] S302. Obtain the Q&A similar texts that match the Q&A intent in the Q&A texts;

[0092] S303. Generate similar text data with specific intents according to the Q&A similar texts.

[0093] In the above steps S301 to S303, first, according to a specific question-and-answer intention, relevant question-and-answer text data is retrieved from the internal knowledge base of the large language model. This data may include historical conversation records, relevant documents, question-and-answer information on the Internet, etc. After obtaining the question-and-answer text data related to the question-and-answer intention, it is necessary to screen out the question-and-answer similar texts that match the specific intention. These texts may have similar semantics, structures, or themes to the question-and-answer intention. Finally, using the screened question-and-answer similar text data, similar text data for a specific intention can be further generated. These data can be generated by rewriting, reorganizing, or expanding the original text, etc., to meet the specific intention requirements. By using the internal knowledge base of the large language model to obtain question-and-answer text data related to a specific intention and generate similar text data, the training data of the question-and-answer system can be greatly enriched, thereby improving the system's understanding ability and coverage. Generating similar text data for a specific intention helps the training model better understand and generalize different expressions of a specific intention, enabling the question-and-answer system to more accurately respond to various questions raised by users. The rich training data and similar text data for a specific intention can improve the recognition and answer accuracy of the question-and-answer system for a specific intention, thereby enhancing the overall performance of the system and user satisfaction. By continuously generating and updating similar text data for a specific intention, the iteration and optimization process of the model can be accelerated, enabling the question-and-answer system to continuously adapt to user needs and scenario changes and maintain a competitive advantage.

[0094] The steps for judging whether the similar text data meets the review criteria include:

[0095] S401. Obtain the evaluation range of the text review criteria;

[0096] S402. Extract semantic feature information and question keyword information from the similar text data;

[0097] S403. According to the semantic feature information and question keyword information, judge whether the similar text data belongs to the evaluation range of the text review criteria;

[0098] If it does not meet the review criteria, it indicates that the intention of the similar text data is inaccurate, and redundant and incorrect text information in the similar text data needs to be removed;

[0099] If it meets the review criteria, it indicates that the intention of the similar text data is accurate.

[0100] As in the above steps S401 to S403, first, it is necessary to determine the standard evaluation range for text review, which can be a set of criteria set according to text accuracy, business requirements, laws and regulations, ethical guidelines, etc., for evaluating the accuracy and legality of text data. Analyze similar text data, extract semantic feature information and question keyword information from it, and these information can help judge the content and intention of the text data. According to the extracted semantic feature information and question keyword information, evaluate the similar text data to determine whether it meets the standard evaluation range of text review. If the similar text data does not meet the review criteria, it may contain incorrect or redundant information and needs to be corrected or deleted. If the similar text data meets the review criteria, it indicates that its intention is accurate and can be used for subsequent training or applications. By reviewing similar text data, the accuracy and legality of the data can be ensured, and the performance and credibility of the question-answering system can be prevented from being affected by incorrect or inappropriate information. Deleting similar text data that does not meet the review criteria can improve the quality of the training data, thereby improving the training effect and performance of the model. Establishing a strict text review mechanism can enhance users' trust in the question-answering system and make users more willing to use and rely on the services provided by the system. Through the review mechanism, inappropriate or sensitive text content can be effectively filtered out, protecting users' privacy and security and enhancing the user experience.

[0101] The steps of converting the text information passing the review criteria in the similar text data into text vectors and constructing a vector knowledge base include:

[0102] S501. Obtain the text information in the similar text data passing the review criteria;

[0103] S502. Extract multiple feature information blocks from the text information;

[0104] S503. Convert the text information into text vectors according to the multiple feature information blocks;

[0105] S504. Summarize the text vectors into a vector knowledge base.

[0106] As in the above steps S501 to S504, text information that meets the review criteria is extracted from the reviewed similar text data. These text information are the basis for constructing the vector knowledge base. The text information is processed to extract multiple feature information blocks, which can be words, phrases, sentences, etc., and reflect the semantics and content of the text information. Using text vectorization technology, the multiple extracted feature information blocks are converted into corresponding text vectors. Commonly used text vectorization methods include the bag-of-words model, TF-IDF, Word2Vec, BERT, etc. The converted text vectors are aggregated to construct a vector knowledge base. This vector knowledge base stores the vector representations of the similar text data that meet the review criteria, providing a basis for subsequent text matching and question-and-answer services. After converting the text information into text vectors, the similarity between vectors can be used to quickly match the questions raised by users and similar text information, thereby improving the matching efficiency and accuracy. Text vectors are more compact than the original text information, can save storage space, improve the storage efficiency and processing speed of the system. The vectorized text information can support various matching requirements, including similarity-based matching, semantic matching, etc., so as to better meet the question-and-answer needs in different scenarios. Constructing a vector knowledge base enables the system to retrieve and match relevant text information more quickly, thereby enhancing the performance of the system and the user experience.

[0107] The steps of obtaining the question-and-answer data of the user and matching the similar-intention text information according to the vector knowledge base and feeding back the answers corresponding to the similar-intention text information to the user include:

[0108] S601. Obtain the question-and-answer data of the user;

[0109] S602. Extract the question-and-answer intention from the question-and-answer data and convert the question-and-answer data into question-and-answer text vectors;

[0110] S603. Match according to the question-and-answer intention and the question-and-answer text vectors with the vector knowledge base, and label the matching result as similar-intention text information;

[0111] S604. Extract the answers from the similar-intention text information and feed back the answers to the user.

[0112] In the above steps S601 to S604, first, it is necessary to obtain the Q&A data proposed by the user. These data can come from the interaction records between the user and the Q&A system, the questions submitted by the user, etc. Process the user's Q&A data, extract the Q&A intent therein, and convert the Q&A text data into text vectors. This can use the same text vectorization technology as when constructing the vector knowledge base. Match the user's Q&A text vectors with the text vectors in the vector knowledge base to find text information similar to the user's intent. The matching can be based on similarity metrics between vectors, such as cosine similarity, etc. The matching result can be labeled as similar intent text information. Extract the answer content from the obtained similar intent text information, and then feedback the answer to the user. This answer can be pre-prepared or dynamically generated. By matching the similar intent text information, personalized answers can be provided according to the user's Q&A intent, so as to better meet the user's needs. Using the vector knowledge base for matching can improve the response speed of the system to the user's questions, thereby enhancing the user experience. The matching similar intent text information often contains accurate answer content, so the answer feedback to the user is also more accurate and credible. By continuously matching the user's Q&A data and providing relevant answers, the system can continuously learn and optimize, improving its own intelligence level and adaptability.

[0113] After the steps of obtaining the user's Q&A data, matching the similar intent text information according to the vector knowledge base, and feeding back the answer corresponding to the similar intent text information to the user, it further includes:

[0114] S605. Generate an inquiry instruction according to the answer corresponding to the similar intent text information fed back to the user;

[0115] S606. The user selects whether it is the required answer according to the inquiry instruction.

[0116] If so, it indicates that the answer corresponding to the similar intent text information is correct, and the result is directly output;

[0117] If not, it indicates that the answer corresponding to the similar intent text information is incorrect, and the answer corresponding to the similar intent text information is re-matched.

[0118] As in the above steps S605 to S606, according to the answer corresponding to the similar intent text information fed back to the user, a corresponding inquiry instruction is generated to prompt the user to confirm whether it is the required answer. The user selects whether to approve the provided answer according to the inquiry instruction generated by the system. If the user confirms that it is the required answer, the result is directly output. If the user denies or chooses to re-match the answer, the system re-matches the answer corresponding to the similar intent text information and feeds it back to the user again. By allowing the user to choose whether to approve the answer, the user's participation and initiative in the answer are enhanced, and the user experience is improved. If the user denies the provided answer, the system can re-match the answer corresponding to the similar intent text information, thereby correcting the wrong answer and providing more accurate information. By the user selecting the answer according to the inquiry instruction, the interaction and communication between the user and the system are enhanced, which helps to better meet the user's needs. Through the user's feedback information, the system can continuously learn and optimize, improve the system's accurate understanding and satisfaction of the user's needs, and thus improve the overall accuracy and efficiency of the system.

[0119] After the step of re-matching the answer corresponding to the similar intent text information, it further includes:

[0120] S6061. Mark the wrong answer as a questioned answer and record the marking times;

[0121] S6062. Mark the re-matched correct answer as a suspected answer and record the marking times;

[0122] S6063. When the marking times of the suspected answer are greater than those of the questioned answer, mark the suspected answer as the correct answer and the questioned answer as the wrong answer.

[0123] In the above steps S6061 to S6063, the wrong answers denied by the user are marked as suspected answers, and the marking times are recorded. This can help the system track and count the user's feedback on the answers. The correct answers obtained by re-matching are marked as suspected answers, and the marking times are recorded. This can compare the performance of the wrong answers and the re-matched answers, so as to make more accurate judgments in the future. When the marking times of the suspected answers are greater than those of the suspected answers, the suspected answers are marked as correct answers, and the suspected answers are marked as wrong answers. By calibrating and comparing the wrong answers and the re-matched answers, the system can dynamically adjust the correctness of the answers, improve the accuracy and reliability of the system. The system can automatically adjust and optimize the answers according to the user's feedback information, reduce the user's intervention and operation, and improve the user experience. The system continuously records and compares the marking times of the answers, and can judge the reliability of the answers according to the changes in the marking times, so as to improve the intelligence level and adaptive ability of the system. By dynamically adjusting and optimizing the answers, the performance and efficiency of the system can be improved, the occurrence of wrong answers can be reduced, and the credibility and user satisfaction of the system can be improved.

[0124] Please refer to the attached Figure 2 As shown in the figure, the present invention also provides a question-and-answer system assisted by a large language model for the above-mentioned question-and-answer robot assisted by a large language model, including:

[0125] A language module for constructing a large language model;

[0126] An intention module for obtaining the application scenario of the question-and-answer robot and generating question prompt data, and determining the question-and-answer intention of the large language model according to the prompt data;

[0127] A text generation module for generating similar text data with a specific intention according to the question-and-answer intention and the large language model;

[0128] An audit module for judging whether the similar text data meets the audit criteria;

[0129] If it does not meet the audit criteria, it indicates that the intention of the similar text data is inaccurate, and redundant and incorrect text information in the similar text data needs to be removed;

[0130] If it meets the audit criteria, it indicates that the intention of the similar text data is accurate;

[0131] A vector module for converting the text information passing the audit criteria in the similar text data into a text vector and constructing a vector knowledge base;

[0132] A matching and retrieval module for obtaining the question-and-answer data of the user and matching similar intention text information according to the vector knowledge base, and feeding back the answers corresponding to the similar intention text information to the user.

[0133] As described above, the language module is responsible for building large language models, usually using pre-trained large language models as a basis, such as GPT series models. These models have been pre-trained on large-scale text data and have powerful language understanding and generation capabilities. The intent module is responsible for obtaining the application scenarios of the question-answering robot and generating question prompt data, and then determining the question-answering intent of the large language model. By deeply understanding the application scenarios and collecting relevant data, the question-answering intent can be accurately determined, enabling the system to focus more on question-answering services in specific fields and improving accuracy and efficiency. The text generation module uses the large language model to generate similar text data with specific intents based on the determined question-answering intent. These similar text data can be used to train the model, expand the training data set, or be used as candidate answers for subsequent question-answering processes. The review module is used to determine whether the generated similar text data meets the review criteria. Through review, the accuracy and legality of the similar text data can be ensured, improving the reliability and trustworthiness of the system. The vector module converts the similar text data that passes the review criteria into text vectors and constructs a vector knowledge base. This can speed up the text matching and retrieval speed, improving the response efficiency and accuracy of the system. The matching and retrieval module is responsible for obtaining the user's question-answering data and matching similar intent text information according to the vector knowledge base, and feeding back the corresponding answers of the similar intent text information to the user. By matching similar intent text information, the system can provide accurate answers according to the user's questions, improving the performance and user experience of the question-answering system. The system uses the large language model and the vector knowledge base to be able to more accurately understand the user's questions and provide relevant answers, thereby improving the accuracy of the question-answering system. Through the application of the pre-trained large language model and the vector knowledge base, the system can quickly match similar intent text information, speed up the retrieval and feedback speed of answers, and improve the efficiency of the system. The system can perform intelligent matching based on the user's question-answering data and the similar text information generated by the intent, realizing a more natural and intelligent question-answering interaction and enhancing the user experience. The system can continuously optimize the training process according to the user's feedback information, improve the accuracy and performance of the model, and achieve continuous optimization and improvement.

[0134] And, a question-answering terminal assisted by a large language model, comprising:

[0135] One or more processors;

[0136] A storage device on which one or more programs are stored;

[0137] When the one or more programs are executed by the one or more processors, the one or more processors implement a question-answering robot assisted by a large language model.

[0138] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention. The structures, devices, and operation methods not specifically described and explained in the present invention are implemented by conventional means in the art without special description and limitation.

Claims

1. A question-answering robot based on the assistance of a large language model, applied to a question-answering method assisted by a large language model, characterized in that: The large language model-assisted question answering method includes: Build large language models; Obtain the application scenario of the question-answering robot and generate question prompt data, and determine the question-answering intention of the large language model based on the prompt data; Based on the question-answering intent, the large language model generates similar text data for specific intents; Determine whether similar text data meets the review standards; If it does not meet the review standards, it means that the intention of the similar text data is inaccurate, and the redundant and erroneous text information in the similar text data needs to be removed; If the review criteria are met, it indicates that the intent of similar text data is accurate; Convert text information that passes the review criteria in similar text data into text vectors and build a vector knowledge base; Obtain the user's question and answer data, match text information with similar intent based on the vector knowledge base, and feedback the corresponding answers to the text information with similar intent to the user; The step of obtaining the application scenario of the question-answering robot, generating question prompt data, and determining the question-answering intention of the large language model according to the prompt data includes: Get the application scenarios of question-answering robots; Obtain key Q&A business areas and frequently asked questions and answers based on application scenarios; According to the question-and-answer business domain and common question-and-answer scenarios, obtain question-and-answer information pairs corresponding to key question-and-answer business domains and common question-and-answer conversations as training data; Extract question features from training data and generate question prompt data for training large language models; Train a large language model based on the generated question prompt data to obtain the question-answering intent corresponding to the application scenario; After the step of obtaining the user's question and answer data, matching similar intent text information according to the vector knowledge base, and feeding back the answer corresponding to the similar intent text information to the user, the method further includes: Generate inquiry instructions based on the corresponding answers to the text information with similar intentions fed back to the user; The user selects whether it is the desired answer according to the inquiry instruction; If so, it indicates that the answer corresponding to the text information with similar intent is correct, and the result is output directly; If not, it indicates that the answer corresponding to the text information with similar intent is wrong, and the answer corresponding to the text information with similar intent is re-matched; After the step of re-matching the corresponding answers of the text information with similar intentions, the method further includes: Mark the wrong answers as question answers and record the number of times they were marked; Mark the re-matched correct answer as the suspected answer, and record the number of markings; When the number of times a suspected answer is marked is greater than the number of times a question answer is marked, the suspected answer is marked as the correct answer and the question answer is marked as the wrong answer.

2. The question-answering robot based on the assistance of a large language model according to claim 1, characterized in that: The step of generating similar text data of a specific intent by a large language model according to the question-answering intent includes: According to the question-answering intent, obtain the question-answering text related to the question-answering intent from the internal knowledge base of the large language model; Obtain question-and-answer similar texts that match the question-and-answer intent in the question-and-answer text; Generate similar text data for specific intents based on question-answer similar texts.

3. The question-answering robot based on the assistance of a large language model according to claim 1, characterized in that: The step of determining whether the similar text data meets the review criteria includes: Get the text review standard evaluation interval; Extract semantic feature information and question keyword information from similar text data; Based on the semantic feature information and question keyword information, determine whether the similar text data belongs to the text review standard evaluation range; If it does not meet the review standards, it means that the intention of the similar text data is inaccurate, and the redundant and erroneous text information in the similar text data needs to be removed; If the review criteria are met, it indicates that the intent of similar text data is accurate.

4. The question-answering robot based on the assistance of a large language model according to claim 1, characterized in that: The step of converting text information that passes the review standard in similar text data into text vectors and constructing a vector knowledge base includes: Obtain text information in similar text data that passes the review criteria; Extracting multiple feature information blocks from text information; According to the multiple feature information blocks, the text information is converted into a text vector; Aggregate text vectors into a vector knowledge base.

5. The question-answering robot based on the assistance of a large language model according to claim 1, characterized in that: The step of obtaining the user's question and answer data, matching similar intent text information according to the vector knowledge base, and feeding back the answer corresponding to the similar intent text information to the user includes: Get the user's question and answer data; Extract question-answering intent from question-answering data and convert the question-answering data into question-answering text vectors; Match the question-answering intent and question-answering text vector with the vector knowledge base, and mark the matching results as text information with similar intent; Extract answers from text messages with similar intent and feed them back to the user.

6. A question-answering system assisted by a large language model, applied to the question-answering robot assisted by a large language model as claimed in any one of claims 1 to 5, characterized in that: include: Language module, used to build large language models; The intention module is used to obtain the application scenarios of the question-answering robot and generate question prompt data, and determine the question-answering intention of the large language model based on the prompt data; The text generation module is used to generate similar text data with specific intents based on the question-answering intent and the large language model; The review module is used to determine whether similar text data meets the review standards; If it does not meet the review standards, it means that the intention of the similar text data is inaccurate, and the redundant and erroneous text information in the similar text data needs to be removed; If the review criteria are met, it indicates that the intent of similar text data is accurate; The vector module is used to convert text information that passes the review standards in similar text data into text vectors and build a vector knowledge base; The matching retrieval module is used to obtain the user's question and answer data, match the text information with similar intent according to the vector knowledge base, and feed back the corresponding answers of the text information with similar intent to the user.

7. A question-answering terminal based on the assistance of a large language model, characterized in that: include: one or more processors; a storage device having one or more programs stored thereon; When the one or more programs are executed by the one or more processors, the one or more processors implement the question-answering robot based on the assistance of a large language model as described in any one of claims 1-5.

Citation Information

Patent Citations

  • Remote supervision question and answer method and system for product knowledge base

    CN117668175A

  • Construction method of mold professional question-answering system based on LLM model

    CN117909458A