Dialogue method based on large model and computing device
By filtering, chunking and compressing effective dialogue blocks, combined with knowledge base and accuracy detection, the problem of redundancy and hallucination of large language models in multiple rounds of dialogue is solved, improving the accuracy and efficiency of dialogue.
Patent Information
- Application Number
- CN202510191954.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-20
- Publication Date
- 2025-07-08
AI Technical Summary
Large language models are prone to generate redundant information and false information in multiple rounds of dialogue scenarios, affecting the accuracy and efficiency of dialogue content.
By filtering out effective dialogue blocks related to user input, chunking and compression processing, generating reply information in combination with relevant knowledge text in the knowledge base, and introducing enhanced propt and accuracy detection mechanisms to reduce hallucination phenomena.
It improves the processing efficiency and accuracy of dialogue information, reduces the probability of hallucination, and enhances the intelligence and user experience of dialogue systems.
Smart Images

Figure CN120277181A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular, to a dialogue method and a computing device based on a large model. Background Art
[0002] With the rapid development of artificial intelligence technology, especially the emergence of large pre-trained models in the fields of deep learning and natural language processing, human-computer dialogue systems have been able to achieve unprecedented complex interactions. These large language models (LLMs) have obtained extensive language understanding and generation capabilities through unsupervised learning on massive text data, and can show excellent performance in various tasks.
[0003] In practical applications, especially in multi-turn dialogue scenarios, the above large models face a significant problem - "hallucination". The so-called "hallucination" refers to the situation where the model generates inaccurate or false information during the dialogue process, which may affect the accuracy of the dialogue content. Summary of the Invention
[0004] A dialogue method and a computing device based on a large model provided by an embodiment of this application reduce the interference of redundant information, improve the processing efficiency of dialogue information and the accuracy of dialogue content, and to a certain extent reduce the probability of the large model having "hallucinations".
[0005] In a first aspect, an embodiment of this application provides a dialogue method based on a large model. The large model implements dialogue based on a knowledge base. The method includes: obtaining a user input; screening out valid dialogues related to the user input from historical dialogues; splitting the valid dialogues to obtain valid dialogue chunks; compressing each valid dialogue chunk to generate compressed valid dialogue chunks; recalling knowledge texts related to the user input from the knowledge base; and generating a reply message for the user input according to the compressed valid dialogue chunks and the knowledge texts.
[0006] Compared with the existing method of directly combining historical dialogues and a knowledge base to form the input information of a large model, an embodiment of this application not only screens out multi-turn historical dialogue segments most relevant to the current dialogue as valid dialogues, but also introduces valid dialogue processing steps. By performing chunking and compression processing on the valid dialogues, redundant information is reduced, making the valid dialogues shorter while still retaining all key information and core points. It reduces the interference of redundant information, improves the processing efficiency of dialogue information, as well as the accuracy and reliability of dialogue content, and to a certain extent avoids the hallucination phenomenon of the large model.
[0007] In a possible implementation, before segmenting the valid dialogue to obtain valid dialogue blocks, the valid dialogue is enhanced based on a preset enhanced prompt to obtain an enhanced valid dialogue.
[0008] In the above embodiment, the valid dialogue is enhanced through a carefully designed prompt. The enhancement process of the valid dialogue mainly strengthens the special processing for the phenomenon leading to hallucinations, which belongs to the enhancement of factual basis, thus avoiding the hallucination phenomenon to a certain extent.
[0009] In a possible implementation, the generating of the response information for the user input according to the compressed valid dialogue block and the knowledge text includes: performing chunking processing on the knowledge text to generate knowledge text chunks; compressing each of the knowledge text chunks to generate compressed text chunks; and generating the response information for the user input according to the compressed valid dialogue block and the compressed text chunks.
[0010] In the above embodiment, the knowledge text recalled from the knowledge base is divided into multiple chunks, and each chunk undergoes independent content compression processing, thereby removing redundant parts while retaining key information. In this way, the accuracy of the information is further ensured and the information processing speed is improved.
[0011] In a possible implementation, the method further includes: obtaining the accuracy of the response information; and in response to the accuracy being lower than a preset threshold, reconstructing the user input according to the valid dialogue and the knowledge text.
[0012] In the above embodiment, whether there is a hallucination is detected by detecting the accuracy of the response information, and the processing operation after hallucination occurs is added, thereby reducing the probability of hallucination, enabling the system to more intelligently handle complex dialogue scenarios, providing more accurate services, and greatly improving the performance and user experience of the multi-round dialogue system.
[0013] In a possible implementation, the obtaining of the accuracy of the response information includes: obtaining the accuracy of the response information according to a preset comparison rule.
[0014] In a possible implementation, the obtaining of the accuracy of the response information includes: obtaining the accuracy of the response information according to a preset comparison prompt.
[0015] In a possible implementation, screening out valid conversations related to the user input from the historical conversations includes: calculating the similarity between the historical conversations and the user input; selecting the historical conversations corresponding to the similarities greater than a preset similarity threshold as valid conversations, or sorting the similarities according to the rule of decreasing values, and selecting the top N historical conversations corresponding to the similarities as valid conversations.
[0016] In a possible implementation, segmenting the valid conversations to obtain valid conversation blocks includes: segmenting the question-and-answer pairs in the valid conversations according to a preset number of questions and answers.
[0017] In a possible implementation, compressing each of the valid conversation blocks includes: compressing each of the valid conversation blocks based on a preset compression prompt.
[0018] In a second aspect, an embodiment of the present application provides a computing device, including: a processor and a memory: the processor and the memory are coupled; the memory is used to store computer program instructions; the processor is used to execute the computer program instructions stored in the memory to implement the large model-based dialogue method in the first aspect and its various possible implementations above.
[0019] In a third aspect, an embodiment of the present application provides a computer-readable storage medium, on which computer program instructions are stored. When the computer program instructions run on a computing device, the computing device is enabled to implement the large model-based dialogue method in the first aspect and its various possible implementations above.
[0020] In a fourth aspect, an embodiment of the present application provides a computer program product, which includes computer program instructions. When the computer program instructions run on a computing device, the computing device is enabled to implement the large model-based dialogue method in the first aspect and its various possible implementations.
[0021] The technical effects obtained in the second, third, and fourth aspects above are similar to the technical effects obtained by the corresponding technical means in the first aspect, and will not be elaborated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] Figure 1 It is a schematic hardware structure diagram of a computing device for a large model-based dialogue method provided by an embodiment of the present application;
[0023] Figure 2 It is a schematic flowchart of a large model-based dialogue method provided by an embodiment of the present application;
[0024] Figure 3Schematic flowchart of processing knowledge text in a dialogue method based on a large model provided by an embodiment of this application. Detailed implementation manners
[0025] The terms "first", "second", "third", etc. in the specification, claims and drawings of this application are used to distinguish different objects, rather than to limit a specific order.
[0026] In the embodiments of this application, words such as "exemplary" or "for example" are used to represent examples, illustrations or explanations. Using words such as "exemplary" or "for example" aims to present relevant concepts in a specific way.
[0027] For the sake of clear and concise description of the following embodiments, a brief introduction to the implementation scenarios of the embodiments is given first:
[0028] Large Language Model (LLM): It is a natural language processing model based on the transformer architecture, with a large number of parameters and trained on a large amount of text. Therefore, the LLM can approximate the human language cognition and generation process, and it has super strong language understanding ability, generation ability, logical reasoning ability and multi-round dialogue ability. Common large language models such as ChatGPT model, PaLM (Pathways Language Model) model, Pangu large model, ERNIE Bot large model, Spark Descartes large model, etc.
[0029] Retrieval-augmented Generation, abbreviated as RAG, is one of the current popular cutting-edge technologies of large models. Retrieval-augmented Generation combines language models and information retrieval technologies. Specifically, when the model needs to generate text or answer questions, it will first retrieve relevant information from a large collection of knowledge base documents, and then use this retrieved information to guide the generation of text, thereby improving the quality and accuracy of the answer.
[0030] Knowledge base: It refers to a collection of a series of knowledge, which can be presented in the form of a structured database (such as: MySQL), or in the form of a set of unstructured document systems (such as: files, pictures, audio, video, etc.), or even a comprehensive form that combines both.
[0031] Prompt: It refers to the input text used to guide or stimulate the model to perform a specific task, and it is a kind of instruction. The purpose is to help the model understand the type of task the user wants to perform or the required output format. It can be a text description, such as "Please recommend a popular music for me" input by the user during a conversation, or a parameter description in a certain format, such as asking the large model to draw a picture according to a certain format, and relevant drawing parameters need to be described.
[0032] Recall: It refers to retrieving several paragraphs related to the question from the knowledge graph or document library.
[0033] The method and device provided by the embodiments of the present application can be applied to the scenario of human-machine dialogue in the Natural Language Processing (NLP) technology. Specifically, the embodiments of the present application can be applied to the scenario of providing dialogue services to users.
[0034] Exemplarily, Figure 1 It is a schematic diagram of the hardware structure of a computing device that can implement the dialogue method based on a large model provided by the embodiments of the present application. Taking this computing device as a server as an example, from the perspective of form, the server can be a high-density server, a rack server or a whole cabinet server; from the perspective of performance, it can be a general server, a GPU (graphics processing unit) server, etc. or an artificial intelligence (AI) server.
[0035] Among them, the hardware part of the computing device includes a processor and a memory. The processor can be a central processing unit (CPU) for performing logical operations on data. An OS management unit runs on the processor, and processor firmware is set inside the processor. The memory, also known as internal memory or main memory, is installed in the memory slot on the motherboard of the computing device for data storage management.
[0036] The above server can be communicatively connected to a mobile terminal. The user inputs a question that requires an answer from the large model (such as text input) through the information input and output device of the mobile terminal (such as a display screen with a virtual keyboard). The mobile terminal sends this question as user input to the server, and then through the large model running on the server, generates dialogue information (reply information for the user input) according to this user input, and feeds back this dialogue information to the information input and output device of the mobile terminal, finally realizing the dialogue between the user and the large model. Here, the mobile terminal can be an electronic device such as a smart phone, a tablet computer, or a personal computer.
[0037] It should be noted that the system architecture and application scenarios described in the embodiments of the present application are for more clearly explaining the technical solutions of the embodiments of the present application, and do not constitute a limitation to the technical solutions provided by the embodiments of the present application. Those of ordinary skill in the art know that with the evolution of the system architecture and the emergence of new business scenarios, the technical solutions provided by the embodiments of the present application are equally applicable to similar technical problems. The following combination Figures 2 - 3, introduce in detail a dialogue method based on a large model provided by an embodiment of the present application.
[0038] As Figure 2 shown, an embodiment of the present application provides a dialogue method based on a large model. The large model realizes dialogue based on a knowledge base, and the dialogue method may include steps S11 - S16. Specifically
[0039] In step S11, obtain user input. For example, obtain a question in text form or a question in voice form input by the user (e.g., What are the popular English learning apps?). Among them, a question in voice form needs to be converted into text information that the large model can process.
[0040] In step S12, screen out valid dialogues related to the above user input from the historical dialogues between the large model and the user. Here, a valid dialogue refers to a dialogue with a relatively high relevance or similarity to the current user input. Specifically, in an application scenario, a pre - trained model (e.g., Sentence - BERT) can be used to transform the historical dialogues into a vector space, and then the cosine similarity formula is used to calculate the similarity between the aforementioned user input and each question in all question - answer pairs (a question - answer pair is a question and its answer) in the historical dialogues. Then, the cosine similarities are sorted according to the rule of decreasing numerical values, and finally, the first N dialogue pairs corresponding to the cosine similarities are selected as valid dialogues. Among them, similarity represents the degree of similarity between two questions. Cosine similarity, also known as cosine similarity measure, is used to evaluate the similarity between two vectors by calculating the cosine value of the included angle between them. Cosine similarity can plot vectors into a vector space according to their coordinate values and then calculate the result. The cosine similarity formula is:
[0041]
[0042] where x i represents the i - th sub - vector in the word - frequency vector of the user input, and y i represents the i - th sub - vector in the word - frequency vector of a certain question in the historical dialogue. Among them, i is a positive integer, n is a positive integer, 1 <= i <= n, and n is the length of the word - frequency vectors of the user input and a certain question in the historical dialogue for calculating similarity.
[0043] For example, the historical dialogue is:
[0044] "Question - answer pair 1:
[0045] Q: How to improve English listening?
[0046] Answer: The key to improving listening skills doesn't lie in how much you listen, but in your reading foundation (including vocabulary and grammar). The better your reading foundation, the more you can understand, and the smoother the listening process will be. The goal of listening training is to strive to understand what you can read. For things you can't understand, you can only solve it by increasing your vocabulary. You should choose appropriate listening materials according to your actual vocabulary. Therefore, to improve listening, you can start from the following points: 1. Memorize more words; 2. Combine listening, speaking, reading, and writing; 3. Choose appropriate listening materials.
[0047] Q&A Pair 2:
[0048] Question: What are the listening materials for improving English listening?
[0049] Answer: First of all, we need to clarify the purpose of improving our listening level. What is the reason for wanting to improve our listening level? If there is a current exam requirement, then our intensive listening materials should be the corresponding exam real questions.
[0050] If it is for being able to understand what foreigners say in daily life communication, then we can listen to more materials related to life and spoken language.
[0051] If you don't have a clear need, then you can choose some materials you are interested in for practice. If you want to know about current affairs, then you can choose to listen to some hot news materials. The pronunciation of these materials is clear and standard, and the words used are regular and the grammar is formal, which are good materials for listening practice.
[0052] The following are generally recognized good listening materials:
[0053] 1. CNN Student News, 2. 60-Second Science, 3. NPR News, 4. AP News Minute, 5. BBC World Service.
[0054] Q&A Pair 3:
[0055] Question: Can the French movie "Amélie" be used to improve English listening?
[0056] Answer: "Amélie" is a French movie, and its help in improving English listening is limited.
[0057] Q&A Pair 4:
[0058] Question: Is the movie "Goodbye Mr. Loser" starring Nicole Kidman suitable for improving English listening?
[0059] Answer: The movie "Goodbye Mr. Loser" starring Nicole Kidman is an English movie and can be used to effectively improve English listening.
[0060] Q&A Pair 5:
[0061] Q: Recommend some nice English songs.
[0062] A: 1. Stairway_To_Heaven 2. How You Remind Me 3. here without you 4. fade to black 5. 18and life 6. Day Of Your Beliefs 7. Her mantle so green 8. For the Love of God 9. no luckygrass 10. Don't you forget about me.
[0063] User input: What are some good English movies for improving English listening?
[0064] In the embodiment of this application, taking the calculation of the similarity between the user input (hereinafter referred to as sentence 1) and the question of the first Q&A pair (hereinafter referred to as sentence 2) as an example, the calculation method of the similarity between the user input and any one of the n questions in the historical conversation is described. The method for calculating the similarity between sentence 1 and sentence 2 includes steps S121 - S123.
[0065] S121. Perform word segmentation on sentence 1 and sentence 2 to obtain a first set of words.
[0066] Among them, word segmentation refers to splitting sentence 1 and sentence 2 into one or more independent words, which can be nouns, verbs, adjectives, or words of any other part of speech. The first set of words refers to the set of words obtained by splitting the user input and the question of the first Q&A pair, where the first set of words does not include duplicate words.
[0067] For example, sentence 1: What are some good English movies for improving English listening. Perform word segmentation on sentence 1 to obtain {improve, English, listening, of, good, English, movies, what are there}. Sentence 2: How to improve English listening, perform word segmentation on sentence 2 to obtain {how, improve, English, listening}. At this time, the first set of words that can be obtained is {how, improve, English, listening, of, good, English, movies, what are there}.
[0068] S122. Represent the word frequency vectors of sentence 1 and sentence 2 in the space composed of the first set of words.
[0069] Among them, the word frequency of each word in the first word set in sentence 1 refers to the number of times each word in the first word set appears in the user input. For example, the word frequency of each word in the first word set in sentence 1 is {how 0, improve 1, English 1, listening 1, nice-looking 1, movie 1}, and the word frequency vector of sentence 1 is (0, 1, 2, 1, 1, 1).
[0070] The word frequency of each word in the first word set in sentence 2 refers to the number of times each word in the first word set appears in sentence 2. For example, the word frequency of each word in the first word set in sentence 2 is {how 1, improve 1, English 1, listening 1, nice-looking 0, movie 0}, and the word frequency vector of sentence 2 is (1, 1, 1, 1, 0, 0).
[0071] S123. Use the cosine similarity formula to calculate that the cosine similarity between sentence 1 and sentence 2 is approximately 0.7.
[0072]
[0073] Next, calculate the cosine similarity between the user input and the questions of other Q&A pairs (the above 2nd - 5th Q&A pairs) in the historical conversation according to the above steps. The cosine similarities between the questions of Q&A pairs 2, 3, 4, and 5 and the user input are approximately: 0.75, 0.5, 0.53, and 0.35 respectively. In one implementation, this similarity can also be obtained by other methods of calculating similarity. For example, Jaccard distance and Dice coefficient, that is, the similarity between two questions can be calculated according to the number of identical words, which is not limited here.
[0074] When selecting the top N cosine similarity corresponding dialogue pairs as valid dialogues, and when the aforementioned N is fixed at 2, dialogue pair 1 and dialogue pair 2 can be selected as valid dialogues; when selecting the cosine similarity corresponding question-and-answer pairs greater than the preset similarity threshold as valid dialogues, and when the similarity threshold is 0.7, dialogue pair 1 and dialogue pair 2 can still be selected as valid dialogues. Through screening, question-and-answer pairs 3-5 with insufficient relevance to the user input are excluded. When the similarity threshold is 0.5, question-and-answer pair 5 with weak relevance to the user input can be excluded. In an application scenario, the method of inverse ranking fusion and custom vector scoring weighting can also be used to sort and screen historical dialogues. For the screening method, no limitation is made here. Then, based on the screened valid dialogues, the true intention of the user input is inferred, and then an answer is given to the user input. Obviously, the screened valid dialogues are mainly related to English listening learning. Therefore, subsequent answers will also focus on this theme. The aforementioned N can be a fixed value (such as 2), or a certain proportion (such as 30%, rounded to an integer) of the total number of question-and-answer pairs in the historical dialogue. When the number of question-and-answer pairs in the historical dialogue is 1, N is 1. In an embodiment, the cosine similarity corresponding question-and-answer pairs greater than the preset similarity threshold (such as 0.7) can also be selected as valid dialogues. This screening operation fully considers the logical relationship between rounds of dialogues, thereby avoiding problems such as context confusion or information omission.
[0075] In an embodiment, the valid dialogue can be enhanced based on a preset enhanced prompt to obtain an enhanced valid dialogue. The specific factual basis enhanced Prompt used can be:
[0076] "You are a Q&A assistant dedicated to providing accurate and reliable information, capable of answering questions with the help of context knowledge. You need to answer questions according to the following rules:
[0077] 1. If the correct relevant answer is included in the context, you need to answer accurately according to the context. However, please note that the information in the context may contain factual errors. If there is content inconsistent with the facts in the document, please answer according to the actual correct information.
[0078] 2. If the answer is not included in the context, just say you don't know, rather than trying to guess or fabricate an answer.
[0079] 3. You need to give a detailed answer according to the context, don't try to be lazy, and you must answer as detailed as possible. When referring to people, places or other entities in the context, use appropriate pronouns or names to enhance understanding and coherence, making the answer more natural and fluent."
[0080] For example, when the above similarity threshold is 0.5, the valid conversations are the above Q&A pairs 1-4. After being enhanced by the above Prompt, it can identify the context inconsistency problem of "French movies... improve English listening" in Q&A pair 3 and the factual error of "The movie 'Goodbye Mr. Loser' starring Nicole Kidman" in Q&A pair 4, and then generate corrective answers. The corrected Q&A pairs 3 and 4 are respectively:
[0081] "Q&A pair 3:
[0082] Question: Can the French movie 'Amélie' be used to improve English listening?
[0083] Answer: 'Amélie' is a French movie and is of limited help in improving English listening.
[0084] Q&A pair 4:
[0085] Question: Is the movie 'Goodbye Mr. Loser' starring Nicole Kidman suitable for improving English listening?
[0086] Answer: Nicole Kidman did not star in the movie 'Goodbye Mr. Loser'. 'Goodbye Mr. Loser' is a Chinese movie and is of limited help in improving English listening."
[0087] Above, through the above carefully designed Prompt, the special processing for context inconsistency or non-factual phenomena is strengthened, so it belongs to fact-based enhancement. From the enhanced answers, it can be known that 'Amélie' and 'Goodbye Mr. Loser' are of limited help in improving English listening. Therefore, in subsequent answers regarding "improving English listening", the content of 'Amélie' and 'Goodbye Mr. Loser' will be excluded, improving the accuracy and reliability of the conversation content, and to a certain extent avoiding the hallucination phenomenon caused by context inconsistency or information errors.
[0088] In step S13, the above valid conversations are segmented to obtain valid conversation blocks. Specifically, multiple Q&A pairs in the valid conversations are segmented according to a preset length (the number of Q&A pairs). For example, they are segmented every 2 or every 5 Q&A pairs. These 2 or 5 Q&A pairs form a valid conversation block, and then the next processing is performed for each valid conversation block. For example, when the above N is fixed at 2, the valid conversations selected from the above historical conversations are Q&A pair 1 and Q&A pair 2, then Q&A pair 1 and Q&A pair 2 can be used as a valid conversation block. This processing method can avoid the problem that the model is difficult to effectively process all information due to information overload (for example, as the number of conversation rounds increases, the historical conversation information may become too large), thereby affecting the quality of the answer.
[0089] In step S14, each of the foregoing valid dialogue blocks is compressed to generate a compressed valid dialogue block. For example, the compression can be achieved through a preset Prompt for context compression, and the specific Prompt can be:
[0090] "You will be dealing with a series of question-and-answer pairs from a historical conversation record. Your goal is to compress these question-and-answer pairs so that they are shorter but still retain all the key information and core points. The compressed version should be easy to understand and complete in information for the reader.
[0091] Requirements:
[0092] 1. Maintain the Q&A structure: Ensure that each question-and-answer pair still maintains the original question and answer structure after compression.
[0093] 2. Retain key information: Ensure that all important questions and answers are accurately retained, especially those parts containing specific facts, data points, or conclusions.
[0094] 3. Remove redundancy: Remove duplicate information and details that add no value, including overly long explanations or unnecessary background information.
[0095] 4. Simplify language: Use simpler and more direct language to express complex concepts without sacrificing accuracy or changing the original meaning.
[0096] 5. Logical coherence: Ensure that the compressed content is still logically coherent without jumps or missing steps.
[0097] 6. Appropriate length: Try to compress each question-and-answer pair to the specified word count range without sacrificing content quality for the sake of reaching the word count.
[0098] 7. Context relevance: If there are multiple consecutive question-and-answer pairs, ensure that the compressed version can still reflect the necessary context relationships."
[0099] For example, after compressing the valid dialogue block formed by the above question-and-answer pair 1 and question-and-answer pair 2, specifically:
[0100] "Question-and-answer pair 1:
[0101] Q: How to improve English listening
[0102] A: 1. Memorize more words; 2. Combine listening, speaking, reading, and writing; 3. Choose suitable listening materials.
[0103] Question-and-answer pair 2:
[0104] Q: Listening materials for improving English listening
[0105] Answer: 1. CNN Student News, 2. 60-Second Science, 3. NPR News, 4. AP News Minute, 5. BBC World Service.
[0106] By compressing the context through a specific prompt, redundant information in the effective dialogue is reduced, so that the key information "improve English listening" in the effective dialogue can be quickly extracted, thus improving the processing efficiency of the effective dialogue and further enhancing the coherence of the dialogue.
[0107] In step S15, recall the knowledge text related to the above user input from the aforementioned knowledge base. This step is independent of steps S12 - S14 and can be executed before S12 or processed in parallel with S12 - S14. In an application scenario, this step can be achieved by dynamically retrieving information from the knowledge base (such as an external knowledge source) through RAG. For the sake of understanding, the working process of RAG is introduced here. The working process of RAG mainly includes the following three steps:
[0108] First is indexing: The construction of the text index includes the following steps: First, document parsing, text chunking, Embedding vectorization, and index creation. Specifically, first parse and convert the original files in different formats in the knowledge base into plain text, and then split the text into smaller text chunks. Second, generate a vector representation for each text chunk through Embedding (vector mapping, which refers to a method of "representing" an object with a numerical vector) to calculate the similarity between the text vector and the vector of the user input. Third, create an index to store the original text chunks and Embedding vectors in the form of key-value pairs for fast and frequent searches in the future.
[0109] Next is retrieval: Use the Embedding model to convert the user input (for example: What are some good English movies for improving English listening?) into a vector, calculate the similarity between the Embedding vector of the user input and the Embedding vectors of the text chunks in the corpus, and select the top K text chunks with the highest similarity as the knowledge text related to the above user input. For example, when K is 2, the first 2 text chunks are respectively:
[0110] (1) More and more adults are starting to learn English, and their biggest English problems are in English listening and speaking. To solve these two problems, many people also choose to watch movies as a form of entertainment to practice English listening and speaking. This is just one of the methods. Today, the editor will share 5 movies suitable for adults to practice English speaking and listening.
[0111] 1. "Zootopia"
[0112] This animated film has a high score of 9.2 on Douban and is a box office and word-of-mouth success. Besides being great to watch, it's also very suitable for learning English. The dialogue is very clear, making it a very popular movie for learning English. Everyone should start learning from it!
[0113] 2. "The Intern"
[0114] This workplace movie starring Anne Hathaway is also very suitable for learning English. The "intern" in the title actually refers to a retired old man who used to be a VP and becomes a junior intern in Anne's company. This huge age difference brings a lot of freshness to the movie. The extremely high emotional intelligence of the elderly intern will surely bring inspiration to everyone, so this movie is also a must-watch for learning.
[0115] 3. "Me Before You"
[0116] This movie tells the story of a wealthy young man with a spinal cord injury and his caregiver. The plot may seem clichéd, but for girls, they will cry buckets after watching it. It's a very touching love movie, not to mention great for learning English.
[0117] 4. "Like Sunday, Like Rain"
[0118] This movie doesn't have the kind of life-and-death separation like "Me Before You". Instead, it's a story about two people opening up to each other and growing together. The dialogue in the movie mainly focuses on daily life and feelings, making it also a very suitable movie for learning English.
[0119] 5. "The Martian"
[0120] This movie is basically a one-man show by Matt Damon. It is based on some existing theories and some future technologies that are really possible to achieve, rather than a pure fictional science fiction movie like superhero movies. So this movie is really very suitable for science students to study. There are also a large number of words and dialogues related to various disciplines such as computer, Internet, physics, and agriculture in the movie, making it a very good movie for expanding vocabulary.
[0121] (2) The 10 Most Popular European and American Movies, the Best Choice for Practicing English Listening
[0122] 1. "Friends" [No need to say more, it's a classic and definitely a must-watch.]
[0123] 2. "Everybody Loves Raymond"
A 24-minute family comedy. To be honest, it can't compare with "Friends", but it has typical American life English, which is more useful than that in "Friends".
[0124] 3. "Joey"
It's still very funny, but it didn't continue. Of course, it's also not as good as "Friends".
[0125] 4. "Prison Break"
Basically, "PB" is the stepping stone for everyone to get into American TV series. Everyone has to watch it, but whether to watch it from the second season on doesn't really matter.
[0126] 5. "Lost"
If we were to rate American TV series, so far, the Douban score of "Lost" is 9.9.
[0127] 6. "Veronica Mars"
The first and second seasons are okay. It's campus + detective.
[0128] 7. "Heroes"
No need to say much. The first season is great, and the second season is a disaster.
[0129] 8. "The 4400" 【4400 people have abilities. Since it came out earlier than "Heroes", although the themes are similar, it's still good.】
[0130] 9. "CSI Las Vegas"
"LV" is the originator of the "CSI" series. Naturally, it's the best-looking. It has now reached the eighth season, and its ratings are always among the top one or two. It's very classic and definitely worth watching.
[0131] 10. "CSI: New York"
The leading actor in "NY" is definitely not as handsome as that in "LV", and the rhythm is slower.
[0132] Finally, there is Generation: The top K (a preset value) text blocks retrieved from the knowledge base and the user input are sent into the large model together, allowing the large model to answer the user's question based on the given text blocks. For example, the information retrieved in step S15 is used as the knowledge text and sent to the large model, and the large model generates a reply message for the user input based on this knowledge text and the above valid dialogue blocks.
[0133] In one implementation, after recalling the knowledge text (for example, using the information retrieved in step S15 as the knowledge text) and before sending this knowledge text to the large model, it also includes steps S151 - S152, as Figure 3As shown, in step S151, the foregoing knowledge text is chunked to generate knowledge text chunks. For example, the 2 text chunks included in the above knowledge text are directly used as 2 independent knowledge text chunks. In step S152, each knowledge text chunk is compressed to generate a compressed text chunk, so that redundant parts in the knowledge text can be removed while key information is retained. In this way, the accuracy of the information is further ensured and the processing speed of the knowledge text is improved. For example, the 2 compressed knowledge text chunks are respectively as follows:
[0134] (1) 1. "Zootopia"
[0135] An animated film that has achieved double success at the box office and in terms of word-of-mouth. The dialogue is very clear.
[0136] 2. "The Intern"
[0137] This fresh and clean workplace movie is also very suitable for learning English.
[0138] 3. "Me Before You"
[0139] It is a very touching love movie, and needless to say, it is even better for learning English.
[0140] 4. "Like Sunday, Like Rain"
[0141] The dialogue in the movie mainly focuses on daily life and feelings, etc., and it is also a very suitable movie for learning English.
[0142] 5. "The Martian"
[0143] This movie is really very suitable for science students to study. There are a lot of words and dialogues related to various disciplines such as computer, Internet, physics, agriculture, etc. in the movie, and it is a very good movie for expanding vocabulary.
[0144] (2) The 10 most popular European and American movies, the best choice for practicing English listening
[0145] 1. "Friends"
Classic, must-see
[0146] 2. "Everybody Loves Raymond"
Typical American life English, more useful than that in Friends
[0147] 3. "Joey"
Not as good as Friends
[0148] 4. "Prison Break"
The stepping stone for everyone to watch American TV series
[0149] 5. "Lost" 【9.9 points】
[0150] 6. "Veronica Mars"
Campus + Detective
[0151] 7. "Heroes"
The first season is great
[0152] 8. "The 4400"
Released earlier than "Heroes", with similar themes
[0153] 9. "CSI Las Vegas"
Very classic, definitely worth watching
[0154] 10. "CSI: New York"
Slower pace
[0155] In step S16, generate a response message for the user input based on the compressed effective dialogue block and the compressed text block. Since the key information in the effective dialogue is "improve English listening", after combining the compressed effective dialogue block, the content in the compressed text block that is only related to "improve English listening" can be used as the response message, deleting the content that is related to both "spoken and listening", or placing the content that is related to both "spoken and listening" in the second place of the response message.
[0156] Therefore, the response message can be:
[0157] "The following are 10 movies recommended for you to practice English listening:
[0158] 1. "Friends"
Classic, must-watch
[0159] 2. "Everybody Loves Raymond"
Typical American life English, more useful than that in "Friends"
[0160] 3. "Joey"
Not as good as "Friends"
[0161] 4. "Prison Break"
Is the stepping stone for everyone's American TV series
[0162] 5. "Lost" 【9.9 points】
[0163] 6. "Veronica Mars"
Campus + Detective
[0164] 7. "Heroes"
The first season is great
[0165] 8. "The 4400"
Released earlier than "Heroes", with similar themes
[0166] 9. "CSI Las Vegas"
Very classic, definitely worth watching
[0167] 10. 《CSI: New York》
Slower rhythm
[0168] Obviously, after the processing of steps S11 - S16, the obtained dialogue information is more accurate and refined. On mobile terminal devices with limited screen sizes, such as mobile phones, this kind of answer is more convenient for users to view.
[0169] In a possible implementation, in order to verify the accuracy of the reply information generated during the conversation between the large model and the user, in other words, to verify whether there are hallucinations during the conversation of the large model, after the large model generates reply information for the user input, the above - mentioned dialogue method may further include: obtaining the accuracy of the reply information. In different implementation scenarios, the accuracy of the reply information can be obtained according to a preset comparison prompt, or the accuracy of the reply information can be obtained according to a preset comparison rule.
[0170] Specifically, the above - mentioned preset comparison prompt can be:
[0171] "You are a grader, evaluating whether the content generated by a large language model (LLM) is based on or supported by a set of retrieved facts. To complete this evaluation, please follow the following steps and criteria:
[0172] 1. During the evaluation process, please carefully read and consider the conversation content in the context and the relevant knowledge obtained through retrieval. This information will be used as the basis for judgment.
[0173] 2. Give a binary score of "yes" or "no":
[0174] If the LLM's answer is completely based on or supported by a set of confirmed facts, the score is "yes". This means that the model's answer is accurate and has sufficient factual basis.
[0175] If the LLM's answer is not based on facts, or contains information that does not match known facts, the score is "no". This may be because the model's answer lacks factual support, contains incorrect information, or has an incorrect understanding and application of facts."
[0176] Obtaining the accuracy of the reply information according to the above - mentioned preset comparison rule mainly involves training the large model to enable the large model to have the ability to distinguish real answers from hallucinations, which can specifically include the following steps:
[0177] First, create a contrastive learning framework with the goal of training a deep learning model that can distinguish between correct answers and hallucinated content, and the model should be able to capture semantic similarities while being sensitive to false information. The specific steps can include: I. Select a basic model architecture: Adopt a pre-trained language model (such as BERT, ROBERTA, or T5) as the base. These models have been extensively pre-trained on a large amount of text data and thus have good language understanding capabilities. II. Fine-tuning strategy: Fine-tune the above basic model according to specific task requirements to ensure its adaptation to the task of hallucination detection. III. Construct a contrastive learning environment: Design a loss function to encourage the model to make correct judgments when facing positive and negative sample pairs. IV. Use CONTRASTIVE LOSS to measure the distance between samples, minimizing the distance between positive samples and maximizing the distance between negative samples.
[0178] Next, prepare positive and negative sample pairs. The specific steps can include: I. Data collection: Collect conversation records from different fields, ensuring coverage of a wide range of topics and scenarios. This includes but is not limited to fields such as education, healthcare, technology, history, etc., to increase the generality of the model. It is also possible to only collect data from a specific field to enable the large model to have expertise in that specific field. II. Annotation process: Have professionals annotate each conversation fragment to clearly indicate which parts are true statements and which may be hallucinations. For each true answer (positive sample), try to construct a reasonable but incorrect answer as a negative sample (hallucinated content), which can be achieved by modifying key facts, adding misleading information, etc. III. Sample balancing: Ensure that the number of positive and negative samples is roughly equal to avoid the model biasing towards a certain type of output. If the natural distribution is unbalanced, oversampling or undersampling methods can be used to adjust the ratio.
[0179] Finally, perform binary classification on the results and output whether hallucinations occur. The specific steps can include the following: I. Feature extraction: Use the fine-tuned language model to extract feature representations from the input text, including context embeddings, sentence-level representations, etc. These features will be used for subsequent classification tasks. II. Classifier design: Add a simple classification layer (such as a fully connected layer plus a softmax activation function) on top of the language model to predict which class the input belongs to (true vs. hallucination). The classifier should be well-trained to achieve high precision and recall on the test set. III. Decision threshold setting: Set a decision threshold based on the probability value output by the classifier. When the probability exceeds this threshold, it is considered that the input contains hallucinated components. The choice of the threshold can be flexibly adjusted according to the requirements of the actual application. For example, in some scenarios, more attention may be paid to reducing the false positive rate. IV. Feedback mechanism: Establish a user feedback channel to allow users to report potential misjudgments. And regularly collect feedback and update the training data to continuously optimize the model performance.
[0180] By means of a preset comparison prompt, the large model can judge the accuracy of the response information in a wide range of scenarios without training; by constructing a comparison framework according to the preset comparison rules, since fixed-scenario or specific-domain data is used for training, the large model can judge the accuracy of the response information for fixed scenarios or specific domains (such as medical assistants), and the accuracy of the judgment result is higher than that of the judgment result by the comparison prompt method.
[0181] After obtaining the accuracy of the above response information, judge whether the accuracy is lower than the preset threshold. If it is lower, reconstruct the user input according to the aforementioned valid conversation and knowledge text. Reconstructing the user input means making certain corrections or adjustments to the user input, aiming to make the user input more in line with the user's true meaning and not deviate from the user's intention, so as to make the response information more accurate. There are various ways to reconstruct the user input, such as multiple queries, query decomposition, Step-back prompting, and HyDE (Hypothetical Document Embeddings). In addition, the valid conversation is a screened conversation with higher accuracy. Therefore, the user input reconstructed according to the valid conversation and knowledge text is more accurate.
[0182] Among them, multiple queries can be understood as follows: predicting that there is an inaccurate answer to a single question, using the large model to turn a single question into a bunch of questions, and then conducting multiple retrievals. This method attempts to supplement the original question from multiple perspectives, turning it into multiple questions to search for results as comprehensively as possible from the vector database. The goal is to refine the query to make it more relevant to the topic, so as to retrieve more relevant documents from the database.
[0183] Query decomposition is a strategy to improve the question-and-answer effect by decomposing a question into sub-questions, and there are two implementation paths:
[0184] (1) Sequential solution, throwing the answer to the previous sub-question and the current question together to the LLM to generate an answer, and then giving the currently generated answer and the next sub-question to the LLM to generate an answer until the final answer is generated for the last sub-question;
[0185] (2) Independently answering questions in parallel, and then merging the multiple-channel answers into the final answer.
[0186] Step-back prompting is a prompting method that enables the LLM to perform abstract operations, thus allowing for accurate answers. Since the question is made more abstract and summarized, a wider range of retrieval is provided. HyDE first generates a hypothesized answer to the question using a large model. This method assumes that the hypothesized answer may be closer to the answer retrieved from the document. It is suitable for cases where the original question is generally short, and the generated hypothesized document may better align with the indexed documents. Generally speaking, the optimization strategies for question reconstruction include four directions: rewriting the question into multiple ones, summarizing the question, decomposing the question, and generating an answer to the question using a large model and then retrieving it.
[0187] In the above embodiments, a step for hallucination detection of the response information is added, and a processing operation after hallucination occurs is added, enabling the system to more intelligently handle complex dialogue scenarios, provide more accurate and personalized services, and greatly improve the performance and user experience of the multi-turn dialogue system.
[0188] In addition, the embodiment of the present application further provides a computing device. The server includes a processor and a memory. The processor is coupled to the memory, and the memory stores computer-executable instructions. When the processor executes the computer-executable instructions, the dialogue method based on the large model in the above embodiments is implemented.
[0189] The embodiment of the present application further provides a computer-readable storage medium. A computer program is stored on the computer-readable storage medium. When the computer program runs on a computer, the computer is enabled to execute the dialogue method based on the large model in the above embodiments. For the explanations and beneficial effects of the relevant content in any of the above-provided computer-readable storage media, reference can be made to the corresponding embodiments above, and details are not repeated here.
[0190] The embodiments of the present application also provide a computer program product containing instructions. When the instructions run on a computer, the computer is caused to execute any one of the above-mentioned large model-based dialogue methods in the embodiments. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions according to the embodiments of the present application are generated in whole or in part. The computer may be a general-purpose computer, an artificial intelligence computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from a website, a computer, a server, or a data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wirelessly (such as infrared, wireless, microwave, etc.). The computer-readable storage medium may be any available medium that can be accessed by the computer or a data storage device such as a server or a data center that includes one or more media integrated therewith. The available medium may be a magnetic medium (such as a floppy disk, a hard disk, a magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as an SSD), etc.
[0191] It should be noted that the devices for storing computer instructions or computer programs provided in the embodiments of the present application, such as but not limited to, the above-mentioned memory, computer-readable storage medium, and communication chip, etc., are all non-transitory.
[0192] Although the present application has been described in conjunction with various embodiments herein, however, in the process of implementing the claimed present application, those skilled in the art can understand and achieve other variations of the disclosed embodiments by viewing the drawings, the disclosure content, and the appended claims. In the claims, the word "comprising" does not exclude other components or steps, and "a" or "one" does not exclude a plurality. A single processor or other unit may implement several functions recited in the claims. Certain measures are recited in mutually different dependent claims, but this does not mean that these measures cannot be combined to produce good results.
[0193] Although the present application has been described in connection with specific features and their embodiments, it will be apparent that various modifications and combinations can be made without departing from the spirit and scope of the present application. Accordingly, the present specification and the drawings are merely exemplary illustrations of the present application as defined by the appended claims, and are considered to cover any and all modifications, variations, combinations, or equivalents within the scope of the present application. Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalent technologies, the present application is also intended to include these changes and modifications.
Claims
1. A dialogue method based on a large model, where the large model realizes dialogue based on a knowledge base, characterized in that, The method includes: Obtaining user input; Filtering conversations related to the user input from historical conversations as valid conversations; Segmenting the valid conversations to obtain valid conversation blocks; Compressing each of the valid conversation blocks to generate compressed valid conversation blocks; Determining knowledge texts related to the user input from the knowledge base; Generating response information for the user input based on the compressed valid conversation blocks and the knowledge texts.
2. The method according to claim 1, characterized in that Before segmenting the valid conversations to obtain valid conversation blocks, the method further includes: Enhancing the valid conversations based on a preset enhanced prompt to obtain enhanced valid conversations.
3. The method according to claim 1 or 2, characterized in that, The generating response information for the user input based on the compressed valid conversation blocks and the knowledge texts includes: Chunking the knowledge texts to generate knowledge text chunks; Compressing each of the knowledge text chunks to generate compressed text chunks; Generating response information for the user input based on the compressed valid conversation blocks and the compressed text chunks.
4. The method according to any one of claims 1 to 3, characterized in that, The method further includes: Obtaining the accuracy of the response information; In response to the accuracy being lower than a preset threshold, reconstructing the user input based on the valid conversations and the knowledge texts.
5. The method according to claim 4, wherein The obtaining the accuracy of the response information includes: Obtaining the accuracy of the response information according to a preset comparison rule.
6. The method according to any one of claims 1-5, characterized in that, The filtering conversations related to the user input from historical conversations as valid conversations includes: Calculating the similarity between the historical conversations and the user input; Selecting the historical conversations corresponding to the similarities greater than a preset similarity threshold as valid conversations, or Sorting the similarities according to the rule of decreasing values, and selecting the top N historical conversations corresponding to the similarities as valid conversations, where N is a positive integer.
7. The method according to any one of claims 1-6, characterized in that, The segmenting the valid conversations to obtain valid conversation blocks includes: Segmenting the question-and-answer pairs in the valid conversations according to a preset number of questions and answers.
8. The method according to any one of claims 1 to 7, characterized in that The compressing each of the valid conversation blocks includes: Compressing each of the valid conversation blocks based on a preset compression prompt.
9. A computing device, characterized in that, Includes: A processor and a memory: The processor and the memory are coupled; The memory is used to store computer program instructions; The processor is used to execute the computer program instructions stored in the memory so that the computing device implements the large model-based conversation method according to any one of claims 1-8.