Knowledge base-based large language model intelligent customer service system and implementation method thereof

By replacing the Faiss vector library with the ElasticSearch database in the intelligent customer service system, real-time index updates and multiple query methods are realized, solving the problem of low Q&A and low efficiency of the intelligent customer service system, and improving the system's query efficiency and accuracy.

CN120144714APending Publication Date: 2025-06-13FUDAN UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510249900.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-04
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

The existing Langchain-ChatGLM2-6B large language model intelligent customer service system based on local knowledge base has low correlation when searching for Q&A. Moreover, due to the limitations of the Faiss vector library, index updates are not real-time, resulting in search query delays and low efficiency.

Method used

The ElasticSearch database is used to replace the Faiss vector library in Langchain, realizing real-time index updates and multiple query methods, including full-text search, precise matching, mixed query and range query.

Benefits of technology

It improves the accuracy and efficiency of Q&A in the intelligent customer service system, realizes real-time index updates, avoids the need to rebuild indexes, and solves the problems of search query latency and low efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120144714A_ABST
    Figure CN120144714A_ABST
Patent Text Reader

Abstract

The invention provides a knowledge base-based big language model intelligent customer service system and an implementation method thereof, and belongs to the technical field of big language models and data processing. According to the method, a new ElasticSearch full-text retrieval database is deployed and accessed to replace a default Faiss vector library in Langchain to serve as a new knowledge storage library, and the question and answer accuracy and efficiency requirements of an intelligent customer service system are met. The Elasticsearch can realize a plurality of query modes, meet the requirements under different scenes, and solve the problem that Fiss is limited to be used for vector similarity search and is only suitable for a specific vector retrieval scene. According to the Elasticsearch, real-time index updating can be achieved, the consistency of data is kept, and the problem that data changes can be reflected only when an index needs to be reconstructed for Faiss is solved. The local knowledge base can be updated only by uploading the document for Elasticsearch storage, so that the problems of query delay and low efficiency are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of large language models and data processing, and specifically relates to a large language model intelligent customer service system based on a knowledge base and its implementation method. Background Art

[0002] When the existing Langchain-ChatGLM2-6B large language model intelligent customer service system based on a local knowledge base conducts a search and answer in the local knowledge base, the relevance of the generated question answers is low, and the essence of the question is not answered completely and accurately. Secondly, when updating the local knowledge base and uploading documents for data vectorization, new vectors need to be frequently inserted, old vectors need to be deleted, and the query index needs to be rebuilt, resulting in search query delays for customer service staff and low efficiency.

[0003] The above business problems are caused by the limitations of the Faiss vector library itself:

[0004] (1) Faiss is mainly suitable for processing floating-point vector data, has limited support for other types of data, is limited to vector similarity search, and is only applicable to specific vector retrieval scenarios. Therefore, the similarity of the search query question answers is not high, resulting in low accuracy.

[0005] (2) The index structure of Faiss is static after construction and does not support dynamic updates. If vectorized data needs to be frequently inserted, deleted, or updated, the index needs to be rebuilt and then used for similarity search queries. Therefore, real-time index updates cannot be achieved, resulting in low search query efficiency. Summary of the Invention

[0006] The present invention is made to solve the above problems, and aims to provide a large language model intelligent customer service system based on a knowledge base and its implementation method.

[0007] The present invention provides a large language model intelligent customer service system based on a knowledge base, which has the following characteristics: it is deployed on a server and uses Langchain as the main framework, including: a client for a customer service staff to input a question and Langchain vectorizes it into a question content vector; a knowledge base for uploading documents through its management function and Langchain vectorizes them into text vector data; an ElasticSearch database for replacing the Faiss vector library in Langchain, receiving the text vector data, automatically updating the index in real time and storing it, wherein Langchain returns several text paragraphs most relevant to the question content vector in the text vector data through similarity analysis and generates a prompt word according to the question content vector; a large language model for generating an answer according to the prompt word, and then the document interaction function module of Langchain returns it to the client for display.

[0008] In the large language model intelligent customer service system based on the knowledge base provided by the present invention, it may also have the following characteristics: Among them, the types of documents include any one or more of txt, word, and pdf.

[0009] In the large language model intelligent customer service system based on the knowledge base provided by the present invention, it may also have the following characteristics: Among them, the large language model includes ChatGLM-6B and / or Qwen-72B.

[0010] In the large language model intelligent customer service system based on the knowledge base provided by the present invention, it may also have the following characteristics: Among them, after Langchain converts the document into plain text, it divides it into text paragraphs, and then calls the embedding model to convert it into embedding vectors to obtain text vector data.

[0011] The present invention also provides an implementation method of the large language model intelligent customer service system based on the knowledge base described above, having the following characteristics, including the following steps: S10, replacing the Faiss vector library in Langchain with an ElasticSearch database, deploying and starting Langchain, the ElasticSearch database, and the large language model, and linking the knowledge base, the client, and the large language model through Langchain; S20, uploading documents through the management function of the knowledge base, and after Langchain processes the documents into plain text, performing text division to obtain text paragraphs; S30, after Langchain vectorizes the text paragraphs into text vector data, storing them in ElasticSearch and automatically updating the index in real time; S40, after the customer service inputs a question to the client, Langchain vectorizes the question into a question content vector; S50, setting the query method and similarity parameters of the ElasticSearch database, performing similarity analysis on the question content vector and the text vector data, and then returning several similar text paragraphs; S60, the Langchain framework generates prompt words based on the text paragraphs and the question content vector; S70, the large language model receives the prompt words and outputs an answer corresponding to the question, and then the document interaction function module of Langchain returns the answer to the client for display.

[0012] In the implementation method of the large language model intelligent customer service system based on the knowledge base provided by the present invention, it may also have the following characteristics: Among them, in step S50, the query method includes approximate query, hybrid query, or exact query.

[0013] In the implementation method of the large language model intelligent customer service system based on the knowledge base provided by the present invention, it may also have the following characteristics: Among them, in step S50, the similarity parameter selects the default parameter.

[0014] In the implementation method of the large language model intelligent customer service system based on the knowledge base provided by the present invention, the following features may also be included: Among them, in step S50, the number of text paragraphs returned is 3.

[0015] Functions and effects of the invention

[0016] According to a large language model intelligent customer service system based on a knowledge base and its implementation method involved in the present invention, because a new ElasticSearch full-text retrieval database is provided for deployment and access to replace the Faiss vector library stored by default in the Langchain framework as a new knowledge repository, therefore, a large language model intelligent customer service system based on a knowledge base and its implementation method of the present invention solve the accuracy and efficiency requirements of question answering in the intelligent customer service system.

[0017] Elasticsearch can implement various query methods such as full-text search, exact matching, hybrid query, and range query, which can meet the requirements in different scenarios. It solves the problem that Faiss is limited to vector similarity search and is only applicable to specific vector retrieval scenarios.

[0018] Elasticsearch can achieve real-time index update, which can update the index in a timely manner when the data changes to maintain data consistency. It solves the problem that Faiss needs to rebuild the index to reflect the data changes. That is, when updating the local knowledge base, only need to upload the document for Elasticsearch storage, and there is no need to insert, delete or update the vectorized data anymore. Elasticsearch will automatically update the index in real time, solving the problems of search query delay and low efficiency. Brief description of the drawings

[0019] Figure 1 It is a demonstration screenshot of the large language model intelligent customer service system based on the knowledge base in the embodiment of the present invention;

[0020] Figure 2 It is a flowchart of the implementation method of the large language model intelligent customer service system based on the knowledge base in the embodiment of the present invention. Detailed implementation manners

[0021] In order to make the technical means, creative features, achieved purposes and effects of the present invention easy to understand, the following embodiments will specifically describe a large language model intelligent customer service system based on a knowledge base and its implementation method of the present invention in conjunction with the accompanying drawings.

[0022] Before further elaborating on the technical details of this embodiment, it is necessary to explain several key concepts to better understand the technical background and implementation mechanism of this embodiment.

[0023] Langchain: A powerful framework designed to assist developers in building end-to-end applications using language models. It provides a set of tools, components, and interfaces that simplify the process of creating applications supported by large language models and chat models. It can manage interactions with language models, link multiple components together, and integrate additional resources such as APIs and databases.

[0024] ChatGLM2-6B, Qwen-72B: A text-generation-based large language model for dialogue that enables free-form question answering as well as knowledge base question answering. It consists of a neural network with numerous parameters (usually billions of weights or more) and is trained using self-supervised learning or semi-supervised learning on a large amount of unlabeled text to generate a language model.

[0025] ElasticSearch: A highly scalable distributed full-text search engine, similar to a distributed database. It can store and retrieve data almost in real time, has good scalability itself, can be scaled to hundreds of servers, and handle PB-level data. Its main feature is that it can perform hybrid queries and exact queries in two ways, namely, text-oriented and vector-oriented.

[0026] Knowledge base: A special database for knowledge management to facilitate the collection, organization, and extraction of domain knowledge. The knowledge in the knowledge base comes from the experience and lessons of domain experts or practitioners. It is a collection of domain knowledge required to solve problems, including basic facts, rules, and other relevant information.

[0027] Faiss: A vector database for efficient similarity search and clustering, mainly used to process large-scale vector data such as images, text, and audio. It provides a series of index structures and search algorithms that can quickly perform similarity search and clustering operations in large-scale datasets.

[0028] <Example>

[0029] This example provides a large language model intelligent customer service system based on a knowledge base, which is deployed on a server and uses Langchain as the main framework.

[0030] The large language model intelligent customer service system based on the knowledge base in this example includes a client, a knowledge base, an ElasticSearch database, and a large language model.

[0031] The client is used for the customer service to input questions, and Langchain vectorizes them into question content vectors.

[0032] The knowledge base is used to upload documents (any one or more of txt, word, and pdf) through its management function, and after Langchain converts them into plain text, they are divided into text paragraphs, and then the embedding model is called to convert them into embedding vectors to obtain text vector data.

[0033] The ElasticSearch database is used to replace the Faiss vector library in Langchain, receive text vector data, and automatically update the index in real time and store it.

[0034] Among them, Langchain returns several text paragraphs in the text vector data that are most relevant to the question content vector through similarity analysis and generates prompt words according to the question content vector.

[0035] Large language models (ChatGLM-6B and Qwen-72B in this embodiment) are used to generate answers according to the prompt words, and then the document interaction function module of Langchain returns them to the client for display.

[0036] Figure 1 is a demonstration screenshot of the large language model intelligent customer service system based on the knowledge base in the embodiment of the present invention; Figure 2 is a flow block diagram of the implementation method of the large language model intelligent customer service system based on the knowledge base in the embodiment of the present invention.

[0037] As Figure 1 and 2 shown, this embodiment also provides an implementation method of the foregoing large language model intelligent customer service system based on the knowledge base, including the following steps:

[0038] S10, use the ElasticSearch database to replace the Faiss vector library in Langchain, deploy and start Langchain, the ElasticSearch database, and the large language model, and link the knowledge base, the client, and the large language model through Langchain.

[0039] S20, upload documents through the management function of the knowledge base, and after Langchain processes the documents into plain text, perform text division to obtain text paragraphs.

[0040] S30, Langchain vectorizes the text paragraphs into text vector data and then stores them in ElasticSearch and automatically updates the index in real time.

[0041] S40, after the customer service inputs a question to the client, Langchain vectorizes the question into a question content vector.

[0042] S50, set the query method of the ElasticSearch database (the query method supports approximate query, hybrid query or exact query, and exact query is specifically selected in this embodiment) and the similarity parameter (the default parameter is specifically selected in this embodiment), and perform similarity analysis on the problem content vector and the text vector data, and then return the three most relevant text paragraphs.

[0043] S60, the Langchain framework generates prompt words based on the text paragraphs and the problem content vector.

[0044] S70, the large language model receives the prompt words and outputs the answer corresponding to the problem, and then the document interaction function module of Langchain returns the answer to the client for display.

[0045] Functions and effects of the embodiment

[0046] According to a large language model intelligent customer service system and its implementation method based on a knowledge base provided in this embodiment, because a new ElasticSearch full-text retrieval database is provided for deployment and access to replace the Faiss vector library stored by default in the Langchain framework as a new knowledge repository, therefore, the large language model intelligent customer service system and its implementation method based on a knowledge base in this embodiment can implement hybrid query and exact query in both text and vector modes, and solve the requirements for the accuracy and efficiency of question answering in the intelligent customer service system.

[0047] Elasticsearch can implement various query methods such as full-text search, exact matching, hybrid query, and range query, which can meet the requirements in different scenarios. It solves the problem that Faiss is limited to vector similarity search and is only applicable to specific vector retrieval scenarios.

[0048] Elasticsearch can achieve real-time index update, which can update the index in time when the data changes and maintain data consistency. It solves the problem that Faiss needs to rebuild the index to reflect the data changes. That is, when updating the local knowledge base, only need to upload the document for Elasticsearch storage, and there is no need to insert, delete or update the vectorized data again. Elasticsearch will automatically update the index in real time, solving the problems of search query delay and low efficiency.

[0049] Those skilled in the art of this industry should understand that the present invention is not limited by the above embodiments. What is described in the above embodiments and the specification only illustrates the principle of the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed. The scope of protection of the present invention is defined by the appended claims and their equivalents.

Claims

1. A large language model intelligent customer service system based on a knowledge base, characterized in that: Deployed on the server and using Langchain as the main framework, including: The client is used for customer service to input questions and Langchain will vectorize them into question content vectors; Knowledge base, which is used to upload documents through its management function and vectorize them into text vector data by Langchain; ElasticSearch database, used to replace the Faiss vector library in Langchain, receive the text vector data, automatically update the index in real time and store it, wherein Langchain returns several text paragraphs in the text vector data that are most relevant to the question content vector through similarity analysis and generates prompt words according to the question content vector; and The large language model is used to generate an answer based on the prompt word, and the document interaction function module of Langchain returns it to the client for display.

2. The large language model intelligent customer service system based on the knowledge base according to claim 1 is characterized by: in, The document type includes any one or more of txt, word and pdf.

3. The large language model intelligent customer service system based on the knowledge base according to claim 1 is characterized by: in, The large language model includes ChatGLM-6B and / or Qwen-72B.

4. The large language model intelligent customer service system based on the knowledge base according to claim 1 is characterized by: in, After Langchain converts the document into plain text, it is divided into text paragraphs, and an embedding model is called to convert it into an embedding vector to obtain the text vector data.

5. A method for implementing a large language model intelligent customer service system based on a knowledge base as described in any one of claims 1 to 4, characterized in that: The following steps are involved: S10, using the ElasticSearch database to replace the Faiss vector library in Langchain, deploying and starting Langchain, the ElasticSearch database and the large language model, and linking the knowledge base, the client and the large language model through Langchain; S20, uploading a document through the management function of the knowledge base, Langchain processes the document into plain text, and then divides the text into text paragraphs; S30, Langchain vectorizes the text paragraph into text vector data, stores it in ElasticSearch, and automatically updates the index in real time; S40, after the customer service inputs a question to the client, Langchain vectorizes the question into a question content vector; S50, setting the query mode and similarity parameters of the ElasticSearch database, performing similarity analysis on the question content vector and the text vector data, and returning a number of similar text paragraphs; S60, the Langchain framework generates prompt words according to the text paragraph and the question content vector; S70, the large language model receives the prompt word and outputs an answer corresponding to the question, and then the document interaction function module of Langchain returns the answer to the client for display.

6. The method for implementing the large language model intelligent customer service system based on the knowledge base according to claim 5 is characterized by: in, In step S50, the query method includes approximate query, mixed query or exact query.

7. The method for implementing the large language model intelligent customer service system based on the knowledge base according to claim 5 is characterized by: in, In step S50, the similarity parameter uses a default parameter.

8. The method for implementing the large language model intelligent customer service system based on the knowledge base according to claim 5 is characterized by: in, In step S50, the number of text paragraphs returned is 3.