Mobile phone retail store knowledge answer generation method based on LIama3 and retrieval enhancement

By building a local text vector knowledge base and using fine-tuned Whisper model and RH-FAISS Merge hybrid search technology, the problem of insufficient knowledge in specific fields of large language models is solved, and accurate answers to mobile phone retail-related questions and improved resource efficiency.

CN119988533APending Publication Date: 2025-05-13XIAN UNIV OF TECH
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202411809958.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-10
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The existing large language models lack sufficient domain-specific knowledge during training, resulting in the inability to accurately understand and answer questions related to mobile phone retail.

Method used

Using LIama3 and search enhancement methods, we use the method to collect and process information related to mobile retail store sales scenarios, build a local text vector knowledge base, and use the fine-tuned Whisper model and RH-FAISS Merge hybrid search technology to identify speech problems and document retrieval to generate accurate answers.

Benefits of technology

It realizes accurate answers to big models in specific fields, reduces "illusion" problems, and avoids the huge resource demand for big models retraining, and is suitable for all types of sales stores.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119988533A_ABST
    Figure CN119988533A_ABST
Patent Text Reader

Abstract

The invention discloses a method for generating knowledge answers to a mobile phone retail store based on LIama3 and retrieval enhancement. The method comprises the following steps: collecting related information of a sales scene of the mobile phone retail store and performing text processing to form a data set; converting texts in the data set into vectors through a BGE model, and storing the vectors into a local vector knowledge base; recognizing a voice problem uploaded by a user by using the finely-adjusted Whisper model, and converting the voice problem into a problem vector; documents related to problem vectors are retrieved in a local vector knowledge base through RH-FAISS Merge mixed search, and a final knowledge document is selected through a Rerank method; generating cue words by using a cue word template Prompt according to the final knowledge document; and submitting the generated prompt word to the Llama3 model to generate an answer, and displaying the answer to the user. According to the method, the problem that related problems of mobile phone retail cannot be accurately understood and answered due to the fact that an existing large language model lacks enough specific domain knowledge during training is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of artificial intelligence application technology, and in particular relates to a method for generating knowledge answers of mobile phone retail stores based on LIama3 and retrieval enhancement. Background Art

[0002] In the sales scenario of mobile phone retail stores, sales staff and customers need to communicate about the detailed parameters, functions, prices, promotions, and other information of mobile phones. The traditional question-and-answer method relies on the professional knowledge of sales staff, but with the update of mobile phone models and functions, the amount of information that sales staff need to master has increased significantly, posing a great challenge. At the same time, customers hope to obtain the required information quickly and accurately in order to make a purchase decision. Traditional methods have problems in question-and-answer knowledge extraction, such as insufficient feature extraction and poor answer generation.

[0003] Existing technologies have introduced question-answering methods based on large language models (LLMs), which process common questions in the public domain and generate corresponding answers by training large language models. However, when applied to specific vertical fields (such as mobile phone retail), LLMs often produce inaccurate or fabricated information (i.e., "hallucinations"), especially for questions about the latest mobile phones. This is because large language models lack sufficient domain-specific knowledge during training, resulting in an inability to accurately understand and answer relevant questions.

[0004] To solve this problem, existing solutions mainly adopt two approaches: one is to fine-tune the large language model in a targeted manner to make it better adapt to the knowledge in a specific field; the other is to use prompt engineering to guide the model to generate answers by designing prompt word templates. Although these methods are effective, they still have certain limitations: fine-tuning requires a lot of data and computing resources and may lead to overfitting; and although prompt word design can reduce "hallucinations", it depends on the quality and design of the prompt words. Summary of the invention

[0005] The purpose of the present invention is to provide a method for generating knowledge answers for mobile phone retail stores based on LIama3 and retrieval enhancement, which solves the problem that the existing large language model lacks sufficient specific domain knowledge during training, resulting in the inability to accurately understand and answer mobile phone retail related questions.

[0006] The technical solution adopted by the present invention is: a method for generating mobile phone retail store knowledge answers based on LIama3 and retrieval enhancement, comprising the following steps: Step 1: Collect information related to mobile phone retail store sales scenarios; Step 2: Perform text processing on the collected information to form a data set; Step 3: Convert the text in the dataset into vectors through the BGE model and store them in the local vector knowledge base; Step 4: Use the fine-tuned Whisper model to recognize the voice questions uploaded by users, convert them into text, and then vectorize them to generate question vectors; Step 5: Retrieve documents related to the question vector in the local vector knowledge base through RH-FAISS Merge hybrid search, and select the final knowledge document through the Rerank method; Step 6: Generate prompt words using the prompt word template Prompt according to the final knowledge document; Step 7: Submit the generated prompt words to the Llama3 model to generate answers and display them to the user through the user interface.

[0007] The present invention is also characterized in that: The mobile phone retail store sales scenario-related information in step 1 includes official standard mobile phone parameter information for each model, official standard training scripts, and store clerks' daily sales scripts.

[0008] Step 2 specifically includes the following steps: Step 2.1: Clean the collected text data information to remove useless formats, special characters, extra spaces and line breaks; Step 2.2, remove stop words from text data; Step 2.3: Standardize the text data, including unifying capitalization, abbreviations, and terminology.

[0009] Step 3 specifically includes the following steps: Step 3.1: Read the knowledge text in the dataset through the unstructured loader provided by LangChain, and use the CharacterTextSplitter to split the text into text blocks; Step 3.2: Convert the segmented text blocks into vector representations through the pre-trained BGE model and store them in the Milvus vector database, i.e., the local vector knowledge base; Step 3.3: Selectively build Rhnsw and FAISS indexes based on specific query scenarios or performance requirements.

[0010] Step 4 specifically includes the following steps: Step 4.1. Use the public AIShell dataset to fine-tune the Whisper model using the LoRA fine-tuning technique with timestamped speech-text pairs as input. Step 4.2: Automatically recognize the voice questions uploaded by users through the fine-tuned Whisper model to generate corresponding text; Step 4.3: Convert the recognized text into a question vector through the BGE model.

[0011] Step 5 specifically includes the following steps: Step 5.1, transfer the question vector into the local vector knowledge base; Step 5.2: Use the two vector retrieval methods, RHNSW and FAISS, to simultaneously perform similarity matching between the question vector and the local vector knowledge base, and return two sets of similar text paragraphs; Step 5.3: Rerank the two groups of similar text paragraphs returned using the Rerank method, and select the top k most relevant text paragraphs as the final knowledge documents based on the weighted comprehensive scores.

[0012] Step 6 is as follows: Through the ChatPromptTemplate template provided by LangChain, combined with the user role, question and final knowledge document, a prompt word template is dynamically generated. The prompt word template adjusts the prompt word content in real time according to the user's specific question and the retrieved final knowledge document to guide Llama3 to generate answers related to the user's question.

[0013] Step 7 is as follows: define the Llama3 task as "You are a mobile intelligent question-answering assistant. Use the retrieved final knowledge document to answer questions. If you are not sure about the answer, just reply 'I don'".

[0014] The beneficial effects of the present invention are: the mobile phone retail store knowledge answer generation method based on LIama3 and retrieval enhancement of the present invention, by constructing a local text knowledge base, provides accurate data support for the Llama3 model, so that the large model can answer professional questions in specific fields, generate contextual, accurate and efficient answers, and can effectively respond to user questions and provide valuable knowledge, avoiding the hallucination problem of large model answers and the huge resource demand for retraining the large model, and is suitable for various types of sales stores. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Figure 1 It is a flow chart of the method for generating knowledge answers of mobile retail stores based on LIama3 and retrieval enhancement of the present invention; Figure 2 It is a logical schematic diagram of the method for generating knowledge answers of mobile phone retail stores based on LIama3 and retrieval enhancement of the present invention; Figure 3 It is a flow chart of using LoRA technology to fine-tune the Whisper large model in the method for generating knowledge answers of mobile retail stores based on LIama3 and retrieval enhancement of the present invention; Figure 4 It is a schematic diagram of text vectorization in the method for generating knowledge answers of mobile retail stores based on LIama3 and retrieval enhancement of the present invention; Figure 5It is a schematic diagram of the present invention being used for conducting relevant professional knowledge question and answer on a PC side. DETAILED DESCRIPTION

[0016] The present invention is described in detail below with reference to the accompanying drawings and specific embodiments.

[0017] Example 1 The present invention provides a method for generating knowledge answers for mobile phone retail stores based on LIama3 and retrieval enhancement. First, the Transformer pre-trained model BGE (BAAI General Embedding) is used to build a local text vector knowledge base for mobile phone retail stores. During the construction process, by collecting high-quality information such as official mobile phone introductions and sales pitches, the text is cleaned, screened and annotated in combination with NLTK (Natural Language Toolkit) technology to form a data set for building a knowledge base. Next, the CharacterTextSplitter module in the LangChain library is used to split the long text content into smaller text blocks for subsequent processing. Then, the BGE model is used to convert these text blocks into vectors to build a local vector knowledge base.

[0018] When a user asks a question by voice, the system first performs speech recognition through a fine-tuned Whisper large language model to convert the voice question into text. Then, the BGE model is used to convert the text into vectors, and the RH-FAISS Merge vector retrieval method is used. The system searches for the top k documents related to the question in the local vector knowledge base, sorts them by similarity score, and merges the results. Finally, the final answer is generated by combining the Llama3 and LangChain frameworks to provide users with intelligent question-answering services.

[0019] Through the above-mentioned methods, the method provided by the present invention is applicable to various types of sales stores. By constructing a local text knowledge base, it provides accurate data support for the Llama3 model, helps store managers accurately predict and grasp key operating indicators, and ensures the reliability of decision support. The system improves service quality by evaluating the semantic similarity between the sales talk of the store clerks and the standard requirements, and provides personalized sales strategy suggestions to optimize the sales process. The intelligent verification mechanism ensures the accuracy and reliability of business decisions. At the same time, the intelligent management of the system reduces manual dependence and improves operational efficiency and market response speed. Ultimately, through accurate and efficient service and management, customer satisfaction and loyalty are enhanced, thereby improving overall business performance.

[0020] Example 2 The present invention provides a method for generating mobile phone retail store knowledge answers based on LIama3 and retrieval enhancement, such as Figure 1As shown in the figure, it includes the following steps: S1. Collect high-quality information such as official standard mobile phone introductions and sales pitches; S2. Clean, screen, and label the collected information to form a dataset; S3. Use the BGE pre-trained model to convert the cleaned data into vectors and build a local text vector knowledge base; S4. When the user asks a question by voice, the system uses the fine-tuned Whisper model for speech recognition, converts the question into text, and converts the text question into a vector through the BGE model; S5. Use the RH-FAISS Merge vector retrieval method to retrieve the top k documents most relevant to the user's question from the local knowledge base; S6. Design a prompt template for Llama3 to guide it to generate corresponding answers; S7. Pass the retrieval results and prompts to Llama3 through LangChain to generate the final answer.

[0021] Embodiment 3 The present invention provides a method for generating knowledge answers for mobile phone retail stores based on Llama3 and retrieval enhancement, as Figure 2 shown in the figure, including the following steps: S1. Collect parameter information of various models of mobile phones currently promoted by VIVO, daily sales pitches of store clerks, and official standard training pitches.

[0022] S2. Perform text processing on the collected information, including text cleaning, stop word removal, and standardization, to form a dataset for building a knowledge base. Specifically, it includes: S2.1. First, clean the obtained text data to remove useless formats, special characters, extra spaces, and line breaks; S2.2. On the basis of text cleaning, further remove stop words, which are usually words that frequently appear in the language but contribute little to the meaning of the text, such as "de", "he", "shi", etc.; S2.3. Finally, perform standardization processing on the text, including unifying case, standardizing abbreviations, and term usage, etc.

[0023] S3. Build a local knowledge base system, use the above dataset to convert text into vectors through the BGE model, and store them in the local text vector knowledge base; among them, building a local text vector knowledge base using the dataset includes: reading the dataset and splitting it into text blocks, converting each text block into a high-dimensional vector representation through the pre-trained BGE (Embedding-based Generation) model. The BGE model will generate semantically rich vectors according to the text content, and these vectors can effectively capture the semantic information in the text. Subsequently, the generated vector data is stored in the local Milvus vector database. Milvus, as a vector database, supports efficient storage, management, and retrieval of large-scale vector data, ensuring low latency and high throughput during the query process. Specifically, it includes: S3.1. The dataset includes files in txt, word or pdf format. The knowledge text is read through the unstructured loader provided by LangChain, and the text is split into text blocks using the text splitter CharacterTextSplitter. S3.2. Each text block is converted into a vector representation through the pre-trained BGE model. The generated vector will be stored in the local Milvus vector database. Milvus provides efficient storage and indexing functions to ensure that vector data can be quickly accessed and retrieved. S3.3. After the vector data is stored, Rhnsw (Randomized Hierarchical Navigable Small World) and FAISS (Facebook AI SimilaritySearch) indexes are selectively constructed to enhance vector retrieval performance, based on specific query scenarios or performance requirements.

[0024] S4: For the voice questions uploaded by users, we first use the Whisper model fine-tuned by LoRA to perform voice recognition, convert them into text, and then vectorize them to generate question vectors, which are convenient for similarity matching in the local knowledge base. The specific steps include: S4.1, if Figure 3 As shown in the figure, using the public AIShell dataset and the LoRA fine-tuning technology, the Whisper model is fine-tuned with timestamped speech-text pairs as input, effectively improving its performance in Mandarin speech recognition tasks; S4.2. Automatically recognize the voice questions uploaded by users through the fine-tuned Whisper model and convert them into text form; S4.3. Convert the recognized text into question vectors through the BGE model so as to perform similarity matching in the local knowledge base.

[0025] S5. Build RH-FAISS Merge hybrid search technology to retrieve documents related to user questions in the local vector library, and select the top k (k is usually less than 5) relevant documents as the final knowledge documents through the Rerank method for the large language model to generate answers. Specifically include: S5.1, transfer the question vector in S4 to the local vector database and perform similarity matching; S5.2. Use two vector retrieval methods (RHNSW and FAISS) to perform similarity matching between the question vector and the text vector library, and return two sets of similar text paragraphs. RHNSW accelerates retrieval through graph structure and uses small-world graph traversal to find the text closest to the question vector; while FAISS uses quantization and hierarchical indexing methods to accelerate the matching of high-dimensional vectors and return the text most similar to the question vector. S5.3. Rerank the two groups of similar text paragraphs returned, and select the top N most relevant text paragraphs based on the weighted comprehensive score to ensure that the final selected text is superior to other text paragraphs in terms of accuracy and relevance.

[0026] S6. Design a dynamic prompt template (Prompt) for Llama3 to automatically generate prompts related to the question based on the user's question and the relevant documents retrieved from the knowledge base, ensuring that the generated prompts are highly relevant and accurate, so as to guide Llama3 to generate accurate answers. Specifically include: S6.1. First, specify the role of Llama3 as “intelligent question-answering assistant” and define its task as “You are an intelligent question-answering assistant about VIVO mobile phones. Use the following retrieved context fragments to answer the user’s questions. If you are not sure about the answer, just reply ‘I don’t know’”; S6.2. Through the ChatPromptTemplate template provided by LangChain, combined with the user question (Question) and N text fragments (Context) retrieved from the local knowledge base, a prompt word template is dynamically generated according to the query content and context information to guide Llama3 to generate answers related to the user's question. Specifically, the prompt word template will flexibly adjust its content according to different query contexts to ensure that the generated answer is highly relevant and accurate.

[0027] S7. Connect the searcher, prompt word template and Llama3 through LangChain. First, the searcher retrieves relevant information from the knowledge base, generates prompt words using the prompt word template, and then submits the generated prompt words to the Llama3 model to finally generate answers and display them to the user through the user interface. The LangChain module links the searcher, prompt word template and large language model, and passes the designed dynamic prompt word template to Llama3. Llama3 generates accurate, rich and relevant answers to user questions based on the information and context in the template, combined with its powerful generation capabilities. Specifically, Llama3 uses the key information in the input search results and dynamic prompt word templates to generate accurate and efficient answers that meet the query context, ensuring that the generated answers can effectively respond to user questions and provide valuable knowledge.

[0028] Example 4 The present invention provides a method for generating knowledge answers for mobile phone retail stores based on Llama3 and retrieval enhancement, including the following steps: S1. Collect the parameter information of various models of mobile phones currently promoted by VIVO, the daily sales promotion words of store clerks, and the official standard training words.

[0029] S2. Perform text processing on the collected information, including text cleaning, stop word removal, and standardization, to form a data set that can be used to build a knowledge base. The specific steps include: S2.1. First, clean the obtained text data to remove useless formats, special characters, extra spaces, and line breaks; S2.2. On the basis of text cleaning, further remove stop words, which are usually common words in the language and contribute less to the text, such as "de", "he", "shi", etc.; S2.3. Finally, perform standardization processing on the text, including unifying case, using standard abbreviations and terms, etc.

[0030] S3. Build a local knowledge base system, convert the text into vectors through the BGE model using the above data set, and store it in the local text vector knowledge base. The specific steps include: S3.1. The data set includes files in txt, word, or pdf format. Read the knowledge text through the unstructured loader provided by LangChain, and use CharacterTextSplitter to split the text into text blocks; S3.2. Convert the text blocks into vectors through the BGE model, build RHNSW and FAISS indexes, and finally store them in the local Milvus vector database.

[0031] S4. For the voice questions uploaded by users, first use the Whisper model fine-tuned by LoRA for speech recognition. After converting it into text, further vectorize it to generate question vectors for similarity matching in the local knowledge base. The specific steps include: S4.1. Using the public AIShell dataset, the Whisper model is optimized through the LoRA fine-tuning technology, which significantly reduces GPU memory usage and accelerates fine-tuning. Freeze the original pre-trained parameters and introduce low-rank adapters for adjustment. Using timestamped speech-text pairs as input, combined with DeepSpeed ​​and mixed precision training, the performance of the model in Mandarin speech recognition tasks is improved. In this step, the LoRA (Low-Rank Adaptation) fine-tuning technology is used to optimize the Whisper model to improve its accuracy in Mandarin speech recognition tasks. The core principle of LoRA fine-tuning is to reduce the number of trainable parameters by inserting low-rank adapters without changing the original model structure. The specific operation process is as follows: Initialize the model: First, load the pre-trained parameters of the Whisper model to ensure that the model has basic speech processing capabilities.

[0032] Introduction of low-rank adapters: Add low-rank adapters to key layers of the model (such as the attention mechanism of the Transformer layer). The adapter exists in the form of a low-rank matrix, and only a small number of additional parameters are required to adjust the representation ability of the model without retraining the entire model.

[0033] Freeze original parameters: During fine-tuning, the original pre-trained parameters of the Whisper model are frozen and only the parameters of the low-rank adapter are updated. This significantly reduces the computational resources required for training and avoids overfitting the model.

[0034] Data input: Use speech-text pairs with timestamps as input. Timestamp information helps the model better understand the temporal characteristics of speech, thereby improving the accuracy of speech recognition.

[0035] Training acceleration: DeepSpeed ​​and mixed-precision training technologies are combined. DeepSpeed ​​accelerates the fine-tuning process through distributed training and efficient memory management, while mixed-precision training uses half-precision and single-precision calculation methods to reduce video memory usage and increase calculation speed. S4.2. Automatically recognize the voice questions uploaded by users through the fine-tuned Whisper model and convert them into corresponding text; S4.3. Convert the recognized text into question vectors through the BGE model so as to perform similarity matching in the local knowledge base.

[0036] S5. Build RH-FAISS Merge hybrid search technology to retrieve documents related to user questions in the local vector library, and select the top k (k is usually less than 5) relevant documents as the final knowledge documents through the Rerank method for the large language model to generate answers. The specific steps include: S5.1. The user question vector is transferred into the local vector database and a preliminary similarity matching is performed; S5.2, use two vector retrieval methods (RHNSW, FAISS) to simultaneously perform similarity matching on the question vector and the text vector knowledge base, and return two sets of similar text paragraphs respectively; S5.3. Rerank the two groups of similar text paragraphs returned, sort the two groups of search results according to the weighted scores, and select the top k most relevant documents as the final candidate knowledge documents for the large language model to generate answers.

[0037] S6. Design a dynamic prompt template (Prompt) for Llama3. According to the user's question and the relevant documents retrieved from the knowledge base, automatically generate prompts related to the question, and ensure that the generated prompts are highly relevant and accurate, so as to guide Llama3 to generate accurate answers. The specific steps include: S6.1. First, specify the role of Llama3 as "intelligent question-answering assistant" and define its task as "You are an intelligent question-answering assistant about VIVO phones. Use the following retrieved context snippets to answer the user's questions. If you are not sure about the answer, just reply 'I don'". This task description will be dynamically adjusted according to the actual query to ensure that Llama3 can generate more appropriate answers based on different question types and contexts; S6.2. The prompt word template is dynamically generated by combining the user role (user), question (question) and N text fragments (context) retrieved from the local knowledge base through the ChatPromptTemplate template provided by LangChain. Specifically, the template will adjust the prompt word content in real time according to the user's specific question and the retrieved context information to ensure that the prompt word is highly relevant to the query context and can guide Llama3 to generate accurate, rich and relevant answers.

[0038] S7. Connect the retriever, prompt word template and Llama3 through LangChain. First, retrieve relevant information from the knowledge base through the retriever, generate prompt words using the prompt word template, and then submit the generated prompt words to the Llama3 model. Finally, generate the answer and display it to the user through the user interface.

[0039] Example 5 As a powerful tool set, the LangChain framework makes it easy for developers to build end-to-end applications based on language models. Through the tools, components and interfaces it provides, it integrates large language models with external data and promotes the interaction between models and the operating environment, making it a compelling solution to build a vertical question-answering system in combination with a knowledge base.

[0040] Llama3 is an advanced reasoning framework that has performed well in tests on multiple Chinese and English public datasets with its efficient dynamic reasoning and memory optimization technology. When loading the Llama3 large language model in the LangChain framework, you can use the LLM wrapper to construct an ollama model class. In order to construct a custom LLM class in LangChain, you need to use the input and output modules of the framework and complete two main tasks: one is to load the user-defined LLM pre-trained model file in the initialization method through the model_path parameter; the second is to implement the _call method so that it can accept a prompt string and return the corresponding response string.

[0041] In order to achieve efficient storage of vector data and fast query of semantic vectors, this system uses Milvus vector database to store text vector data. Milvus has excellent scalability and can easily cope with the growing vector data in the knowledge base of mobile phone retail stores. Whether it is a massive amount of mobile phone product information vectors or a large number of user evaluation vectors, it can ensure that the system maintains excellent performance when the data scale expands. Its advanced indexing algorithms (such as HNSW-based indexing) and GPU acceleration technology have greatly improved the speed and accuracy of vector queries, especially when processing high-dimensional vectors (such as vectors containing multiple characteristics of mobile phones), showing significant advantages, and can quickly provide users with accurate mobile phone product recommendations or information query results.

[0042] Both FAISS and RHNSW are libraries designed for efficient vector search, and can demonstrate powerful performance when processing large-scale data sets. FAISS achieves data compression through vector quantization technology (such as PQ), and combines multiple index types (such as IndexPQ, IndexHNSWFlat) and exact and approximate search strategies, with GPU acceleration, and can process billions of vector data. RHNSW builds a hierarchical graph structure based on the HNSW algorithm, and traverses the top-level nodes according to similarity selection to quickly find the approximate nearest neighbor vector. RHNSW is particularly suitable for complex queries (such as multi-attribute matching of mobile phones), and after integration with machine learning models, it can improve semantic understanding and search accuracy, especially when processing mobile phone-related text data.

[0043] The present invention provides a method for generating mobile phone retail store knowledge answers based on LIama3 and retrieval enhancement, comprising the following steps: S1. Collect parameter information of various mobile phone models currently promoted by VIVO, daily sales pitches of store clerks and official standard training pitches.

[0044] S2. Perform text processing on the collected information, including text cleaning, removal of stop words and standardization, to form a data set that can be used to build a knowledge base. The specific steps include: First, text cleaning is performed to remove irrelevant characters and format errors. Next, stop words are removed to reduce data noise and highlight key information. Then, text standardization is performed to unify the format and case to ensure data consistency. In addition, feature extraction is performed, such as calculating word frequency and inverse document frequency (TF-IDF), to enhance the expressiveness of the data.

[0045] S3. Build a local knowledge base system, use the above data set to convert text into vectors through the BGE model, and store them in the local text vector knowledge base. The specific steps include: S3.1. The dataset includes files in txt, word or pdf format. The knowledge text is read through the unstructured loader provided by LangChain, and the CharacterTextSplitter is used to split the text into text blocks. S3.2. Select and prepare a pre-trained text embedding model, such as the BGE model, which supports similarity calculation and heterogeneous text retrieval between Chinese and English texts; S3.3. In the LangChain framework, load the pre-trained file of the BGE model through the provided interface to prepare for text vectorization processing; S3.4, using the loaded BGE model, convert the text to be processed into numerical vectors, which can effectively capture the semantic information of the text; S3.5. Store the generated text vector data into the Milvus vector database to ensure effective management and access of the data.

[0046] S4: For the voice questions uploaded by users, we first use the Whisper model fine-tuned by LoRA to perform voice recognition, convert them into text, and then vectorize them to generate question vectors, which are convenient for similarity matching in the local knowledge base. The specific steps include: S4.1. Using the public AIShell dataset, we fine-tune the Whisper model using timestamped speech-text pairs as input through LoRA fine-tuning technology to improve its performance in Mandarin speech recognition tasks. S4.2. Automatically recognize the voice questions uploaded by users into text through the fine-tuned Whisper model; S4.3. Convert the recognized text into a question vector through the BGE model to facilitate similarity matching in the local knowledge base.

[0047] S5. Build RH-FAISS Merge hybrid search technology to retrieve documents related to user questions in the local vector library, and select the top k (k is usually less than 5) relevant documents as the final knowledge documents through the Rerank method for the large language model to generate answers. The specific steps include: S5.1. In the Milvus vector database, use the question vector of the user's question as a query to retrieve the document vector in the database that is most similar to the question vector. S5.2. Extract a set of candidate answers from the search results. These candidate answers are sorted based on the document vectors that are most similar to the question vector to ensure the relevance of the candidate answers. S5.3. Re-score the extracted candidate answers using the Rerank model. The model evaluates the relevance and accuracy of each candidate answer based on the similarity between the user question and the candidate answer and other contextual information; S5.4. From the re-ranked candidate answers, select the top k (usually less than 5) most relevant answers as the final result, ensuring that the generated answers meet the requirements in terms of relevance and accuracy.

[0048] S6. Design a dynamic prompt template (Prompt) for Llama3. According to the user's question and the relevant documents retrieved from the knowledge base, automatically generate prompts related to the question, and ensure that the generated prompts are highly relevant and accurate, so as to guide Llama3 to generate accurate answers. The specific steps include: S6.1. First, specify the role of Llama3 as "intelligent question-answering assistant" and define its task as "You are an intelligent question-answering assistant about VIVO mobile phones. Use the following retrieved context snippets to answer the user's questions. If you are not sure about the answer, just reply 'I don'". This task description is dynamically adjusted according to the user's question and query context to ensure that Llama3 can generate the most relevant answers according to different situations; S6.2. Then, through the ChatPromptTemplate template provided by LangChain, combined with the user role (User), question (Question) and N text fragments (Context) retrieved from the local knowledge base, a prompt word template is dynamically generated. The content of the template is flexibly adjusted according to the actual query and retrieval results to ensure that the answer generated by Llama3 is highly relevant to the user's question and can provide accurate and detailed answers.

[0049] S7. Connect the searcher, prompt word template and Llama3 through LangChain. First, use the searcher to retrieve relevant information from the knowledge base, use the prompt word template to generate prompt words, and then submit the generated prompt words to the Llama3 model to finally generate answers and display them to the user through the user interface. Specifically, it includes: passing the designed dynamic prompt word template to Llama3, using the generation ability of the large language model, combining the information in the prompt template (such as user questions, search results and context fragments), to generate accurate, rich and highly relevant outputs. Specifically, Llama3 will generate answers that meet the query context based on the input prompt template and context information, ensuring that the output is not only accurate but also highly relevant to meet user needs.

[0050] Example 6 The present invention develops a question-answering system based on knowledge retrieval enhancement, such as Figure 4 and Figure 5 As shown in the figure, the system fully utilizes the powerful functions of the LangChain framework to achieve rapid construction of vector knowledge bases and efficient data retrieval. By combining the excellent generation capabilities of the open source and lightweight Llama3 pre-trained model, the development cycle of the question-answering system in a specific field is significantly shortened, while ensuring that users can easily build personalized intelligent question-answering systems on standard computing resources and consumer-grade graphics processors. This design not only improves the wide applicability and practical value of the system, but also provides accurate and reliable solutions to the "hallucination" problem that may occur in large language models in vertical fields, greatly facilitating users and promoting the further development and widespread application of intelligent question-answering technology.

[0051] This one-click, end-to-end voice interaction service combined with a question-and-answer solution customized for vertical fields has become a key trend in the future development of artificial intelligence. By providing seamless voice command execution and accurate field-specific answers, these technologies will greatly enrich the application scenarios of artificial intelligence and significantly improve the user's interactive experience.

Claims

1. A method for generating answers to mobile phone retail store knowledge based on LIama3 and retrieval enhancement, characterized in that: The following steps are involved: Step 1: Collect information related to mobile phone retail store sales scenarios; Step 2: Perform text processing on the collected information to form a data set; Step 3: Convert the text in the dataset into vectors through the BGE model and store them in the local vector knowledge base; Step 4: Use the fine-tuned Whisper model to recognize the voice questions uploaded by users, convert them into text, and then vectorize them to generate question vectors; Step 5: Retrieve documents related to the question vector in the local vector knowledge base through RH-FAISS Merge hybrid search, and select the final knowledge document through the Rerank method; Step 6: Generate prompt words using the prompt word template Prompt according to the final knowledge document; Step 7: Submit the generated prompt words to the Llama3 model to generate answers and display them to the user through the user interface.

2. The method for generating mobile phone retail store knowledge answers based on LIama3 and retrieval enhancement as claimed in claim 1, characterized in that: The mobile phone retail store sales scenario-related information in step 1 includes official standard mobile phone parameter information of various models, official standard training scripts, and store clerks' daily sales scripts.

3. The method for generating mobile phone retail store knowledge answers based on LIama3 and retrieval enhancement as claimed in claim 1, characterized in that: The step 2 specifically includes the following steps: Step 2.1: Clean the collected text data information to remove useless formats, special characters, extra spaces and line breaks; Step 2.2, remove stop words from text data; Step 2.3: Standardize the text data, including unifying capitalization, abbreviations, and terminology.

4. The method for generating mobile phone retail store knowledge answers based on LIama3 and retrieval enhancement as claimed in claim 1, characterized in that: The step 3 specifically comprises the following steps: Step 3.1: Read the knowledge text in the dataset through the unstructured loader provided by LangChain, and use the CharacterTextSplitter to split the text into text blocks; Step 3.2: Convert the segmented text blocks into vector representations through the pre-trained BGE model and store them in the Milvus vector database, i.e., the local vector knowledge base; Step 3.3: Selectively build Rhnsw and FAISS indexes based on specific query scenarios or performance requirements.

5. The method for generating mobile phone retail store knowledge answers based on LIama3 and retrieval enhancement as claimed in claim 1, characterized in that: The step 4 specifically comprises the following steps: Step 4.

1. Use the public AIShell dataset to fine-tune the Whisper model using the LoRA fine-tuning technique with timestamped speech-text pairs as input. Step 4.2: Automatically recognize the voice questions uploaded by users through the fine-tuned Whisper model to generate corresponding text; Step 4.3: Convert the recognized text into a question vector through the BGE model.

6. The method for generating mobile phone retail store knowledge answers based on LIama3 and retrieval enhancement as claimed in claim 1, characterized in that: The step 5 specifically comprises the following steps: Step 5.1, transfer the question vector into the local vector knowledge base; Step 5.2: Use the two vector retrieval methods, RHNSW and FAISS, to simultaneously perform similarity matching between the question vector and the local vector knowledge base, and return two sets of similar text paragraphs; Step 5.3: Rerank the two groups of similar text paragraphs returned using the Rerank method, and select the top k most relevant text paragraphs as the final knowledge documents based on the weighted comprehensive scores.

7. The method for generating mobile phone retail store knowledge answers based on LIama3 and retrieval enhancement as claimed in claim 1, characterized in that: The step 6 is specifically as follows: through the ChatPromptTemplate template provided by LangChain, a prompt word template is dynamically generated in combination with the user role, question and final knowledge document. The prompt word template adjusts the prompt word content in real time according to the user's specific question and the retrieved final knowledge document to guide Llama3 to generate answers related to the user's question.

8. The method for generating mobile phone retail store knowledge answers based on LIama3 and retrieval enhancement as claimed in claim 1, characterized in that: The step 7 is specifically as follows: define the Llama3 task as "You are a mobile intelligent question-answering assistant, use the retrieved final knowledge document to answer questions, and directly reply 'I don't know' if the answer is not clear."

Citation Information

Cited By

  • Method for generating fine-tuning question and answer data set based on retrieval enhancement

    CN120745855A

  • Cognitive reconstruction psychological intervention system and device based on large language model

    CN121328740A