Customer service knowledge base optimization method and device, equipment and medium

By automatically processing user questions and answers in historical dialogue records, new customer service knowledge points are generated and added to the knowledge base, the problems of inefficient knowledge base construction and poor answer quality in the existing technology are solved, and efficient construction and optimization of the knowledge base are achieved.

CN120045668APending Publication Date: 2025-05-27GUANGZHOU SHANGYUN NETWORK TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510122541.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-24
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

In the prior art, the construction of knowledge bases relies on manual operations, resulting in inefficiency and easy to miss long-tail knowledge points, and there are problems of duplication of knowledge points and poor answer quality.

Method used

By obtaining the Q&A pairs in the historical dialogue records, calculate the similarity between the user's question text and the customer service knowledge points in the knowledge base, filter and classify the user's question text, generate new customer service knowledge points and add them to the knowledge base, and optimize the content and structure of the knowledge base.

Benefits of technology

It significantly improves the construction efficiency of the knowledge base and the quality of answers, avoids the omission and duplication of knowledge points, and improves the comprehensiveness and query efficiency of the knowledge base.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120045668A_ABST
    Figure CN120045668A_ABST
Patent Text Reader

Abstract

The invention relates to a customer service knowledge base optimization method, a customer service knowledge base optimization device, customer service knowledge base optimization equipment and a medium in the technical field of e-commerce. Determining a question matching degree between the user question text in the question and answer pair and a similar question text set in each customer service knowledge point in a preset knowledge base, and constructing a new customer service knowledge point and adding the new customer service knowledge point to the knowledge base on the basis of the user question text of which the question matching degree is lower than a first preset threshold value; adding the user question text of which the question matching degree exceeds a second preset threshold value and is lower than a third preset threshold value to a similar question text set in the customer service knowledge points matched with the user question text, wherein the second preset threshold value is greater than the first preset threshold value; and optimizing a standard reply text in the corresponding customer service knowledge point by adopting the customer service reply text of the user question text of which the question matching degree exceeds a third preset threshold value in the question and answer pair. According to the method, the knowledge base can be continuously and automatically optimized by fully and reasonably utilizing the historical dialogue records.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of e-commerce technologies, and in particular, to a method for optimizing a customer service knowledge base, as well as a corresponding device, a computer device, and a computer-readable storage medium. Background Art

[0002] In the current era of rapid development of informatization, the application of intelligent customer service systems has become increasingly common. Such systems process customer inquiries in an automated manner, not only significantly improving the efficiency of customer service, but also effectively reducing the labor cost for enterprises. In an intelligent customer service system, the knowledge base plays a core role. It is responsible for storing and managing various types of information closely related to user inquiries, in order to enable the customer service system to quickly and accurately respond to the consulting needs of users.

[0003] Currently, the construction of the knowledge base mainly relies on manual operations. In traditional technologies, common questions and their corresponding answers are usually manually sorted out by analyzing user inquiry records and the responses of customer service staff therein. Although this method ensures the reliability of the knowledge base to a certain extent, it also exposes several obvious defects:

[0004] Highly dependent on manual labor: The manual construction of the knowledge base requires a large amount of human resources input. Especially in enterprises with a large volume of customer inquiries, the efficiency of manual review and update often fails to keep up with the pace of actual needs. As the number of user inquiries continues to increase, the pressure to maintain the knowledge base also grows day by day. This process is not only time-consuming and laborious, but also prone to omitting some potentially important information in the collection of knowledge points.

[0005] Omission of long-tail knowledge points: Manual processing usually mines data from chat records regularly and summarizes knowledge points. Those frequently occurring problems are easily noticed, but for long-tail problems that occur less frequently, they are easily overlooked during the manual collection process.

[0006] Duplication of knowledge points: During the process of manually constructing the knowledge base, due to different expressions of similar questions, there are often redundancies and duplications of knowledge points. Different customer service staff may express the same question in different ways, resulting in multiple knowledge points with basically the same content being stored in the knowledge base. This not only increases the storage cost, but also affects the query efficiency of the knowledge base.

[0007] Insufficient quality of answers: Over time, the answers in the knowledge base may need to be updated. However, during the update process, knowledge base operators may miss some, resulting in incomplete or untimely answers.

[0008] Through in-depth analysis of the prior art, this application aims to propose an improved method to solve the above problems, improve the construction efficiency of the knowledge base and the quality of answers, so as to provide users with more accurate and efficient intelligent customer service. Summary of the Invention

[0009] The primary object of this application is to provide a method for optimizing a customer service knowledge base, as well as a corresponding device, computer device, and computer program product, to solve at least one of the above problems.

[0010] To meet the various objectives of this application, the following technical solutions are adopted:

[0011] A method for optimizing a customer service knowledge base provided to meet one of the objectives of this application includes the following steps:

[0012] Obtain each question-and-answer pair in the historical conversation records. For each question-and-answer pair, determine the user's question text in the question-and-answer pair, and respectively determine the question matching degree between the user's question text and the set of similar question texts in each customer service knowledge point in the preset knowledge base.

[0013] Filter out the user's question text with a question matching degree lower than the first preset threshold. Use the user's question texts with the same semantics as similar question texts to construct a set of similar question texts. For each set of similar question texts, determine the corresponding intention description text and standardized answer text according to the set of similar question texts and the customer service answer texts of each similar question text in the question-and-answer pair, and construct new customer service knowledge points to append to the knowledge base.

[0014] Filter out the user's question text with a question matching degree exceeding the second preset threshold and lower than the third preset threshold. Use the user's question text as a similar question text and append it to the set of similar question texts in the customer service knowledge point that matches the user's question text. The second preset threshold is greater than the first preset threshold.

[0015] Filter out the user's question text with a question matching degree exceeding the third preset threshold. Optimize the standardized answer text in the customer service knowledge point that matches the user's question text according to the customer service answer text of the user's question text in the question-and-answer pair.

[0016] On the other hand, a customer service knowledge base optimization device provided to meet one of the purposes of this application includes a matching degree determination module, a first addition module, a second addition module, and a reply optimization module. Among them, the matching degree determination module is used to obtain each question-and-answer pair in the historical conversation record. For each question-and-answer pair, determine the user question text in the question-and-answer pair, and respectively determine the question matching degree between the user question text and the set of similar question texts in each customer service knowledge point in the preset knowledge base; the first addition module is used to screen the user question texts with a question matching degree lower than the first preset threshold, use the user question texts with the same semantics as the similar question texts to construct a set of similar question texts, and for each set of similar question texts, determine the corresponding intention description text and standard reply text according to the set of similar question texts and the customer service reply texts of each similar question text in the question-and-answer pair, so as to construct new customer service knowledge points and add them to the knowledge base; the second addition module is used to screen the user question texts with a question matching degree exceeding the second preset threshold and lower than the third preset threshold, and add the user question texts as similar question texts to the set of similar question texts in the customer service knowledge point that matches the user question text, and the second preset threshold is greater than the first preset threshold; the reply optimization module is used to screen the user question texts with a question matching degree exceeding the third preset threshold, and optimize the standard reply text in the customer service knowledge point that matches the user question text according to the customer service reply text of the user question text in the question-and-answer pair.

[0017] On the other hand, a computer device provided to meet one of the purposes of this application includes a central processing unit and a memory. The central processing unit is used to call and run the computer program stored in the memory to execute the steps of the customer service knowledge base optimization method described in this application.

[0018] On the other hand, a computer program product provided to meet another purpose of this application includes computer programs / instructions. When the computer programs / instructions are executed by a processor, the steps of the method described in any embodiment of this application are implemented.

[0019] The technical solution of this application has many advantages, including but not limited to the following aspects:

[0020] This application first obtains each Q&A pair from the historical conversation records and calculates the question matching degree between the user's question text and the set of similar question texts of the customer service knowledge points in the knowledge base. During the screening process, the user's question text is classified according to different matching degree thresholds. For the user's question text with a matching degree lower than the first preset threshold, questions with the same semantics can be identified and classified into the set of similar question texts. Based on these sets of similar question texts and their corresponding customer service reply texts, intention description texts and standardized reply texts are generated, new customer service knowledge points are constructed and appended to the knowledge base, effectively filling the gaps in the long-tail knowledge points in the knowledge base, avoiding the problem of knowledge point omission caused by omission, and significantly improving the comprehensiveness and integrity of the knowledge base.

[0021] For the user's question text with a question matching degree between the second preset threshold and the third preset threshold, it is used as a similar question text and appended to the set of similar question texts in the customer service knowledge point that matches it, which can effectively enrich the different expression forms of similar questions in the knowledge base, avoid the phenomenon of knowledge point duplication caused by appending basically the same expressions of the same question in the knowledge base, significantly improve the standardization and query efficiency of the knowledge base, and at the same time reduce the storage cost.

[0022] In addition, for the user's question text with a question matching degree exceeding the third preset threshold, the standardized reply text in the matching customer service knowledge point is optimized according to its corresponding customer service reply text, which can ensure that the replies in the knowledge base always remain in the optimal state, avoid the problems of incomplete answers or lack of timeliness that may occur in the traditional knowledge base update process, and thus improve the quality and reliability of the answers in the knowledge base.

[0023] To sum up, based on the multi-dimensional matching and screening mechanism, the automatic continuous update and optimization of the knowledge base are realized, which not only significantly improves the optimization efficiency of the knowledge base without relying on manual work, but also effectively solves the problems of long-tail knowledge point omission, knowledge point duplication, and poor reply quality existing in the traditional knowledge base construction process, so as to be able to provide users with more accurate, efficient and comprehensive customer service, and provide a brand-new high-value solution for the customer service knowledge base management in the e-commerce field. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] The above and / or additional aspects and advantages of this application will become apparent and easy to understand from the following description of the embodiments in conjunction with the drawings, where:

[0025] Figure 1 is the network architecture of an exemplary e-commerce platform of this application;

[0026] Figure 2 is the flowchart of a typical embodiment of the customer service knowledge base optimization method of this application;

[0027] Figure 3 It is the principle block diagram of the customer service knowledge base optimization device of the present application;

[0028] Figure 4 It is the structural schematic diagram of a computer device adopted by the present application. Specific embodiments

[0029] The embodiments of the present application will be described in detail below. Examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals represent the same or similar elements or elements with the same or similar functions throughout. The embodiments described below by referring to the accompanying drawings are exemplary and are only used to explain the present application, and cannot be construed as a limitation to the present application.

[0030] Those skilled in the art of the present technology can understand that, unless specifically stated otherwise, the singular forms "a", "an", "the" and "said" used herein may also include the plural forms. It should be further understood that the term "including" used in the specification of the present application means the presence of the described features, integers, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or their groups. It should be understood that when we say an element is "connected" or "coupled" to another element, it can be directly connected or coupled to other elements, or there may also be intermediate elements. In addition, the "connection" or "coupling" used herein may include wireless connection or wireless coupling. The phrase "and / or" used herein includes all or any unit and all combinations of one or more related listed items.

[0031] Those skilled in the art of the present technology can understand that, unless otherwise defined, all terms (including technical terms and scientific terms) used herein have the same meaning as the general understanding of those of ordinary skill in the art to which the present application belongs. It should also be understood that terms such as those defined in a general dictionary should be understood to have a meaning consistent with the meaning in the context of the prior art, and will not be interpreted with an idealized or overly formal meaning unless specifically defined as here.

[0032] As Figure 1 shown in the network architecture, the e-commerce platform 82 is deployed in the Internet to provide corresponding services to its users. Similarly, the devices 80 of the merchant users and the devices 81 of the consumer users of the e-commerce platform 82 are also connected to the Internet to use the services provided by the e-commerce platform.

[0033] The exemplary e-commerce platform 82 matches the supply and demand of products and / or services for the general public by means of the Internet infrastructure. In the e-commerce platform 82, the products and / or services are provided as commodity information. For the sake of simplicity in description, in this application, concepts such as commodities and products are used to refer to the products and / or services in the e-commerce platform 82, which may specifically be physical products, digital products, tickets, service subscriptions, other offline-performed services, etc.

[0034] All parties in reality can access the e-commerce platform 82 in the capacity of users and use various online services provided by the e-commerce platform 82 to achieve the purpose of participating in the business activities realized by the e-commerce platform 82. These entities may be natural persons, legal persons or social organizations, etc. Corresponding to the two types of entities, namely merchants and consumers, in business activities, there are two major types of users in the e-commerce platform 82, namely merchant users and consumer users. All parties in the product circulation chain in business activities, including manufacturers, sellers, retailers, logistics providers, etc., can use the online services in the e-commerce platform 82 in the capacity of merchant users, while consumers in business activities, including actual or potential consumers, can use the online services in the e-commerce platform 82 in the capacity of their corresponding consumer users. In actual business activities, the same entity can act both in the capacity of a merchant user and in the capacity of a consumer user, and this should be understood flexibly.

[0035] The infrastructure for deploying the e-commerce platform 82 mainly includes a back-end architecture and front-end devices. The back-end architecture runs various online services through a service cluster, including middleware or front-end services for the platform side, services for consumers, services for merchants, etc., to enrich and improve its service functions; the front-end devices mainly cover the terminal devices used by users to access the e-commerce platform 82 as clients, including but not limited to various mobile terminals, personal computers, point-of-sale devices, etc. By way of example, a merchant user can use his terminal device 80 to enter commodity information for his online store or generate his commodity information using the interfaces opened by the e-commerce platform; a consumer user can access the web page of the online store realized by the e-commerce platform 82 through his terminal device 81, trigger the shopping process by means of the shopping buttons provided on the web page, and call various online services provided by the e-commerce platform 82 during the shopping process, so as to achieve the purpose of placing an order for shopping.

[0036] In some embodiments, the e-commerce platform 82 can be implemented by a processing facility including a processor and a memory. The processing facility stores a set of instructions that, when executed, cause the e-commerce platform 82 to perform the e-commerce and support functions involved in this application. The processing facility can be part of a server, a client, a network infrastructure, a mobile computing platform, a cloud computing platform, a fixed computing platform, or other computing platforms, and provides the electronic components of the e-commerce platform 82, merchant devices, payment gateways, application developers, marketing channels, shipping providers, customer devices, point-of-sale devices, etc.

[0037] The e-commerce platform 82 can be implemented as online services such as cloud computing services, software as a service (SaaS), infrastructure as a service (IaaS), platform as a service (PaaS), desktop as a service (DaaS), managed software as a service, mobile backend as a service (MBaaS), information technology management as a service (ITMaaS), etc. In some embodiments, the various functional components of the e-commerce platform 82 can be implemented to be suitable for operating on various platforms and operating systems. For example, for an online store, its administrator user enjoys the same or similar functions in various embodiments such as iOS, Android, HomonyOS, or the web.

[0038] The e-commerce platform 82 can implement corresponding independent websites for each merchant to run their corresponding online stores, and provide corresponding business management engine instances for the merchants to establish, maintain, and run one or more online stores in one or more independent websites. The business management engine instance can be used for content management, task automation, and data management of one or more online stores, and can configure various specific business processes of the online store through interfaces or built-in components, etc., to support the realization of business activities. The independent website is the infrastructure of the e-commerce platform 82 with cross-border service functions. Merchants can relatively centrally and autonomously maintain their online stores based on the independent website. The independent website usually has a domain name and storage space dedicated to the merchant. Different independent websites are relatively independent. The e-commerce platform 82 can provide standardized or personalized technical support for a large number of independent websites, enabling merchant users to customize their own suitable business management engine instances and use this business management engine instance to maintain one or more online stores they own.

[0039] An online store can implement back-end configuration and maintenance by a merchant user logging in to its business management engine instance as an administrator. With the support of various online services provided by the infrastructure of the e-commerce platform 82, the merchant user can, in its administrator capacity, configure various functions in its online store, view various data, etc. For example, the merchant user can manage all aspects of its online store, such as viewing recent activities of the online store, updating the product catalog of the online store, managing orders, recent access activities, total order activities, etc.; the merchant user can also view more detailed information about the business and visitors to the merchant's online store by obtaining reports or metrics, such as showing a sales summary of the merchant's overall business, specific sales and participation data of the active sales and marketing channels, etc.

[0040] The e-commerce platform 82 can provide communication facilities and associated merchant interfaces for providing electronic communication and marketing. For example, it uses an electronic message aggregation facility to collect and analyze communication interactions among merchants, consumers, merchant devices, customer devices, point-of-sale devices, etc., aggregate and analyze the communication, such as for increasing the potential of providing product sales. For example, a consumer may have a question about a product, which may generate a conversation between the consumer and the merchant (or an automated processor-based agent on behalf of the merchant), where the communication facility is responsible for the interaction and provides the merchant with an analysis of how to increase the probability of a sale.

[0041] In some embodiments, application programs suitable for installation on terminal devices can be provided to serve the access needs of different users, so that various users can, in the terminal device, access the e-commerce platform 82 by running the application program, such as the merchant back-end module of the online store in the e-commerce platform 82. In the process of implementing business activities through these functions, the e-commerce platform 82 can implement various functions related to supporting business activities as middleware or online services and open corresponding interfaces, and then implant a toolkit corresponding to the interface access function into the application program to achieve function expansion and task implementation. The business management engine can include a series of basic functions and expose these functions to online services and / or application programs for calling through an API. The online services and application programs use the corresponding functions by remotely calling the corresponding API.

[0042] With the support of the various components of the business management engine instance, the e-commerce platform 82 can provide an online shopping function, enabling merchants to establish contact with customers in a flexible and transparent manner. Consumer users can select items online, create a merchandise order, provide the delivery address of the goods in the merchandise order, and complete the payment confirmation of the merchandise order. Then, the merchant can review and complete or cancel the order.

[0043] An optimization method for a customer service knowledge base of the present application can be programmed as a computer program product and implemented by running on a client or a server. For example, in an exemplary application scenario of the present application, it can be implemented by deploying it on the server of an e-commerce customer service platform. Thereby, by accessing the interface opened after the computer program product runs, human-computer interaction can be performed with the process of the computer program product through a graphical user interface to execute this method.

[0044] Please refer to Figure 2 , in a typical embodiment of the optimization method for the customer service knowledge base of the present application, the following steps are included:

[0045] Step S1100: Obtain each question-and-answer pair in the historical conversation record. For each of the question-and-answer pairs, determine the user's question text in the question-and-answer pair, and respectively determine the question matching degree between the user's question text and the set of similar question texts in each customer service knowledge point in the preset knowledge base.

[0046] It can be understood that users on the e-commerce platform can use the consulting customer service function provided by the platform to trigger the creation of a session between the user and a human customer service or a machine customer service for the user. In the session, the user can have a conversation with the human customer service or the machine customer service. Thus, the specific conversation content generated in each session is stored as a complete historical conversation record for future invocation.

[0047] In order to extract the question-and-answer pairs in the historical conversation record, it is necessary to analyze and filter the historical conversation record. The question-and-answer pair refers to the conversation content of one question and one answer, including the user's question text and its corresponding customer service reply text. Specifically, for each historical conversation record, after identifying and filtering the invalid conversation content in the historical conversation record, the question-and-answer pairs composed of the text sent by the user, that is, the user's question text, and the text sent by the customer service to reply to this text, that is, the customer service reply text, can be segmented from the remaining valid conversation content. In this process, when the user sends multiple texts to express a question, these texts can be concatenated into one user's question text in the order from the earliest to the latest according to the sending time. The user's question text expresses a single complete question, and the customer service reply text accurately answers this question.

[0048] The invalid conversation content includes, but is not limited to, non-business-related conversations between the user and the customer service, repeated texts, and conversation fragments that do not form a question-and-answer structure. The non-business-related conversations are usually greetings or insignificant chats between the user and the customer service. The repeated texts are usually texts sent by the user or the customer service multiple times with the same content. The conversation fragments that do not form a question-and-answer structure are usually those where the customer service does not reply to the question sent by the user, or the reply from the customer service cannot accurately answer the user's question.

[0049] In one embodiment, for identifying non-business-related conversations of users or customer service representatives in invalid conversation content, a text classification model can be used to classify each piece of text in the historical conversation record to determine whether each text is business-related. The text classification model is pre-supervised and trained to a convergent state to acquire the classification ability to determine whether the input text belongs to business-related or non-business-related. The supervised training can be flexibly implemented by those skilled in the art; for identifying duplicate conversations in invalid conversation content, a text encoding model can be first used to encode each piece of text in the historical conversation record to determine the corresponding text encoding vector. Then, any similarity algorithm such as cosine similarity algorithm, Euclidean distance algorithm, Pearson correlation coefficient algorithm, Jaccard coefficient algorithm, etc. can be used to calculate the similarity between the text encoding vectors as the similarity between the corresponding texts. For every two texts arranged in the order of sending time from earliest to latest, if the similarity of these two texts exceeds a preset threshold, it is determined that these two texts are duplicates. The preset threshold can be set by those skilled in the art according to business requirements, such as 0.98. The text encoding model is pre-supervised and trained to a convergent state to acquire the ability to map text to vectors in a high-dimensional space, so that texts with similar semantics are closer in the vector space; for identifying conversation fragments that do not form a question-and-answer structure in invalid conversation content, after determining each piece of text sent by the user arranged in the order of sending time from earliest to latest in the historical conversation record, it is determined whether there is text sent by the customer service representative. When there is none, it means that the customer service representative has not replied to the question sent by the user, that is, a conversation fragment that does not form a question-and-answer structure. When there is, the Bert model is used to determine whether the text sent by the customer service representative with the earliest sending time can accurately answer the text sent by the user. If it can accurately answer, a question-and-answer structure is formed; if it cannot accurately answer, a conversation fragment that does not form a question-and-answer structure is formed. The Bert model is pre-supervised and trained to a convergent state to acquire the ability to determine whether the text sent by the customer service representative can accurately answer the text sent by the user. The supervised training can be flexibly implemented by those skilled in the art.

[0050] For each question-and-answer pair segmented from the valid conversation content, the valid conversation content can be embedded into a preset prompt template to obtain the corresponding prompt text and input it into the large language model to obtain each question-and-answer pair output by the large language model.

[0051] The prompt template includes a task description and the conversation content to be embedded. The task description can be: "Please split the following conversation content into question-and-answer pairs in the form of one question and one answer. Extract each complete question-and-answer pair, which contains a user question text and a customer service reply text. The user question text should express a complete question, and the customer service reply text should accurately answer the question. If the user sends multiple texts to express one question, these texts should be concatenated into a complete user question text in the order of sending time. During the process of extracting question-and-answer pairs, special attention should be paid to possible expression problems in the user question text, such as misspelled words, missing words, and incoherent expression order. Please identify whether there are such expression problems in the user question text. If so, the user question text should be corrected when extracting the question-and-answer pair to ensure the accuracy and readability of the user question text. The corrected user question text should maintain the original meaning while the language expression is more standardized and clear. Conversation content: {the conversation content to be embedded}."

[0052] To improve the matching efficiency, it can be understood that the text encoding model can be used in advance to encode each similar question text in the set of similar question texts of each customer service knowledge point in the knowledge base to obtain corresponding text semantic vectors, and associate each text semantic vector with its corresponding similar question text for storage. When calculating the similarity between the user question text and the similar question text, there is no need to repeatedly encode the similar question texts in the knowledge base, thus saving computing resources and improving efficiency. Therefore, for the user question text in each of the question-and-answer pairs, after inputting the user question text into the text encoding model to encode it into a corresponding text encoding vector, calculate the similarity between the text encoding vector and the text encoding vectors corresponding to each of the pre-encoded similar question texts in the knowledge base, and use it as the similarity between the user question text and each similar question text. Then, for each set of similar question texts of each customer service knowledge point in the knowledge base, select the maximum value among the similarities between each similar question text in the set and the user question text as the question matching degree between the set and the user question text.

[0053] Step S1200: Screen the user question texts with a question matching degree lower than the first preset threshold, use the user question texts with the same semantics as similar question texts to construct a set of similar question texts. For each set of similar question texts, determine the corresponding intention description text and standard reply text according to the set of similar question texts and the customer service reply texts of each similar question text in the question-and-answer pairs, and construct new customer service knowledge points to append to the knowledge base.

[0054] The first preset threshold is used to determine whether there is a significant difference between the user question text and the set of similar question texts in the existing customer service knowledge points in the knowledge base. It is recommended to set this threshold to 0.8.

[0055] When the question matching degree is lower than the first preset threshold, it means that the similarity between the user's question text and the set of similar question texts in any of the customer service knowledge points is relatively low, and it is not sufficient to consider that their semantics are the same or highly similar. In this regard, the user question texts with a question matching degree lower than the first preset threshold are screened out. Such user question texts and their corresponding customer service reply texts in the question-and-answer pairs can be used to construct new customer service knowledge points. Specifically, first, a text encoding model is used to encode each of the user question texts into vectors in a high-dimensional semantic space, obtaining corresponding text encoding vectors for each. These vectors can capture the semantic information of the text, making texts with similar semantics closer in the vector space and texts with different semantics farther apart. Then, a clustering algorithm is used to cluster the user question texts based on the text encoding vectors representing the semantics of each of the user question texts numerically, so that user question texts with similar semantics are grouped into the same cluster. For each cluster, each of the user question texts in it is used as a similar question text and grouped into a set of similar question texts. The clustering algorithm can be the K-means algorithm, the GMM Gaussian mixture model clustering algorithm, the DBSCAN algorithm, the meanshift algorithm, the spectral clustering algorithm, etc. Those skilled in the art can choose one to implement according to their needs.

[0056] For each set of similar question texts, first obtain the customer service knowledge point in the knowledge base corresponding to the set of similar question texts with the highest question matching degree from the question matching degrees of each similar question text in the set of similar question texts. Then, embed the customer service knowledge point, the single set of similar question texts it targets, and the corresponding customer service reply texts of each similar question text in the question-and-answer pair into a preset prompt template to obtain the corresponding prompt text, which is input into the large language model, and the customer service knowledge point generated by the model is appended to the knowledge base.

[0057] The prompt template includes a task description and the customer service knowledge point, the set of similar question texts, and the customer service reply texts to be embedded. The task description can be "Please refer to the existing customer service knowledge points demonstrated below. According to the set of similar question texts provided below and the corresponding customer service reply texts of each similar question text in it, construct a new customer service knowledge point. This knowledge point includes an intention description text, the set of similar question texts provided, and a standard reply text.

[0058] Set of similar question texts: {Set of similar question texts to be embedded};

[0059] Customer service reply texts: {Corresponding customer service reply texts of each similar question text in the set of similar question texts to be embedded};

[0060] Demonstrated existing customer service knowledge points: {Customer service knowledge points to be embedded};

[0061] Requirements:

[0062] 1. Since the provided set of similar question texts is constructed based on an unsupervised clustering algorithm, it may lead to a large semantic deviation between some individual similar question texts in the set and the vast majority of similar question texts. If there are such individual similar question texts, you need to delete them from the set and not use the corresponding customer service reply texts during the process of constructing new customer service knowledge points.

[0063] 2. The intention description text should accurately summarize the common intention of all questions in the set of similar question texts.

[0064] 3. The standardized reply text should comprehensively consider all customer service reply texts and provide a unified and accurate reply that can be applied to any question in the set of similar question texts.

[0065] 4. If there are inconsistent or redundant information in the provided multiple customer service reply texts, appropriate integration and optimization are required to ensure the accuracy and practicality of the standardized reply text.

[0066] 5. Please ensure that the intention description text and the standardized reply text are clearly and normatively expressed and easy to understand.

[0067] Step S1300: Screen the user question texts whose question matching degree exceeds the second preset threshold and is lower than the third preset threshold, and use the user question texts as similar question texts and append them to the set of similar question texts in the customer service knowledge points that match the user question texts. The second preset threshold is greater than the first preset threshold;

[0068] The second preset threshold is used to determine whether the user question text is highly similar to the set of similar question texts in the existing customer service knowledge points in the knowledge base, and is also used to determine whether the user question text is highly similar to the intention description text in the customer service knowledge points in the knowledge base. It is recommended to set this threshold to 0.9.

[0069] The third preset threshold is used to determine whether the user question text is extremely similar to the set of similar question texts in the existing customer service knowledge points in the knowledge base, and is also used to determine whether there is an extremely high similarity between the customer service question text and the standardized question text in the customer service knowledge points in the knowledge base. It is recommended to set this threshold to 0.98.

[0070] When the question matching degree exceeds the second preset threshold and is lower than the third preset threshold, it indicates that the user question text corresponding to the question matching degree has a relatively high similarity with the set of similar question texts, and it is sufficient to consider that their semantics are highly similar and express the same intention, but they adopt different expressions. That is, it indicates that the user question text is of the same nature as the similar question texts in the question text set. Therefore, the user question text is used as a new similar question text and appended to the set of similar question texts.

[0071] Step S1400: Screen the user question texts whose question matching degree exceeds the third preset threshold, and optimize the standard reply text in the customer service knowledge points that match the user question text according to the customer service reply text in the question-and-answer pair.

[0072] When the question matching degree exceeds the third preset threshold, it indicates that the user question text corresponding to the question matching degree has an extremely high similarity with the set of similar question texts, and it is sufficient to consider that their semantics are extremely similar and express the same intention, and they adopt basically the same expression. That is, it indicates that there are similar question texts identical to the user question text in the set of similar question texts. Therefore, the user question text is not suitable as an additional similar question text to be appended to the set of similar question texts. However, the customer service reply text of the user question text in the question-and-answer pair can be used to optimize the standard reply text in the customer service knowledge points where the set of similar question texts is located. Specifically, the customer service reply text and the customer service knowledge points are embedded into a preset prompt template to obtain the corresponding prompt text, which is input into the large language model to obtain the standard reply text output by the model.

[0073] The prompt template includes a task description and the customer service knowledge points and customer service reply text to be embedded. The task description can be: "Please evaluate and possibly optimize the existing standard reply text in the following provided customer service knowledge points using the following provided customer service reply text. If optimization is required, output the optimized standard reply text; if no optimization is required, output the original text of the standard reply text.

[0074] Customer service reply text: {Customer service reply text to be embedded};

[0075] Customer service knowledge points: {Customer service knowledge points to be embedded};

[0076] Requirements:

[0077] 1. First, evaluate the differences between the customer service reply text and the existing standard reply text to determine whether optimization is required. If the customer service reply text provides more detailed or updated information, it should be integrated into the standard reply text to ensure the timeliness and integrity of the answer; if there is no significant improvement or new information in the customer service reply text compared with the existing standard reply text, no optimization is required.

[0078] 2. If, after comparing the expressions of the customer service reply text with those of the existing standard reply text, it is found that the standard reply text needs to be edited and reorganized as necessary, please optimize it in this way while maintaining the original meaning to improve its readability and attractiveness.

[0079] 3. Check and eliminate any possible inconsistencies or redundant information in the existing standard reply text to improve the clarity and professionalism of the reply.

[0080] 4. Ensure that the optimized reply text is not only accurate and comprehensive, but also able to answer the user's questions in a clear and understandable manner.

[0081] According to the typical embodiments of the present application, it can be known that the technical solution of the present application has many advantages, including but not limited to the following aspects:

[0082] The present application first obtains each question-and-answer pair from the historical conversation records, and calculates the question matching degree between the user's question text and the set of similar question texts of the customer service knowledge points in the knowledge base. During the screening process, the user's question text is classified according to different matching degree thresholds. For the user's question text with a matching degree lower than the first preset threshold, questions with the same semantics can be identified and classified into the set of similar question texts. Based on these sets of similar question texts and their corresponding customer service reply texts, an intention description text and a standard reply text are generated, a new customer service knowledge point is constructed and appended to the knowledge base, effectively filling the gap of the long-tail knowledge points in the knowledge base, avoiding the problem of knowledge point loss caused by omission, and significantly improving the comprehensiveness and integrity of the knowledge base.

[0083] For the user's question text with a matching degree between the second preset threshold and the third preset threshold, it is used as a similar question text and appended to the set of similar question texts in the customer service knowledge point that matches it, which can effectively enrich the different expression forms of similar questions in the knowledge base, avoid the phenomenon of repeated knowledge points caused by different expressions of the same question, significantly improve the standardization and query efficiency of the knowledge base, and at the same time reduce the storage cost.

[0084] In addition, for the user's question text with a matching degree exceeding the third preset threshold, the standard reply text in the matching customer service knowledge point is optimized according to its corresponding customer service reply text, which can ensure that the reply in the knowledge base always maintains the optimal state, avoid the problems of incomplete answers or lack of timeliness that may occur during the update process of the traditional knowledge base, and thus improve the quality and reliability of the answers in the knowledge base.

[0085] In summary, based on the multi-dimensional matching and screening mechanism, the automated continuous update and optimization of the knowledge base are realized, which not only significantly improves the optimization efficiency of the knowledge base without relying on manual labor, but also effectively solves the problems such as the omission of long-tail knowledge points, the duplication of knowledge points, and the poor quality of answers existing in the construction of traditional knowledge bases. Therefore, it can provide users with more accurate, efficient and comprehensive customer service, and provide a brand-new and highly valuable solution for the customer service knowledge base management in the e-commerce field.

[0086] In a further embodiment, before step S1100 of obtaining each Q&A pair in the historical conversation record, the following steps are included:

[0087] Step S1000: Obtain the customer service knowledge document library, and for each customer service knowledge document therein, perform semantic segmentation on the customer service knowledge document to obtain each knowledge block corresponding to different semantics;

[0088] For the customer service knowledge document with a clear chapter structure, such as including headings at all levels and paragraphs belonging to the heading, the semantic segmentation process is relatively straightforward. In this regard, each knowledge block can be composed of a heading and its subordinate paragraphs. The heading serves as the theme of the knowledge block, while the paragraphs provide detailed explanations and / or information, ensuring that each knowledge point organizes content around a central theme.

[0089] For the customer service knowledge document without a clear chapter structure, the segmentation process needs to be more meticulous. In this regard, it is necessary to identify the paragraph boundaries in the document, that is, those pairs of sentences that are semantically incoherent. Since a paragraph usually consists of a series of semantically coherent sentences, when the coherence degree of the context between two sentences is lower than a preset threshold, it can be considered that they do not belong to the same paragraph. Thus, by identifying these pairs of sentences in the document, the document can be split into paragraphs composed of semantically coherent sentences as knowledge blocks. To achieve this goal, a pre-trained and converged Bert model can be used to determine whether each pair of adjacent sentences in the document forms an upper and lower sentence relationship. When it can form an upper and lower sentence relationship, it means that the corresponding two sentences belong to the same paragraph; when it cannot form an upper and lower sentence relationship, it means that the corresponding two sentences do not belong to the same paragraph. From this, the pairs of two sentences in the document that cannot form an upper and lower sentence relationship are used as paragraph boundaries. In this way, by splitting these pairs of sentences, the document can be split into paragraphs composed of semantically coherent sentences as knowledge blocks.

[0090] Step S1010: Use a large language model to generate corresponding customer service knowledge points according to each knowledge block, and the customer service knowledge points include intention description text, similar question text set, and standardized answer text;

[0091] For each knowledge block, embed the knowledge block into a preset prompt template to obtain the corresponding prompt text, input it into the large language model, and obtain the customer service knowledge points generated by the model.

[0092] The prompt template includes a task description and the knowledge block to be embedded. The task description can be: "Please generate corresponding customer service knowledge points based on the following provided knowledge block content. The customer service knowledge points should include an intention description text, a set of similar question texts, and a standard response text. The generation process should follow the following steps:

[0093] Step 1: Since the knowledge block provides detailed content for answering user questions, please analyze and identify the core theme of the knowledge block and the scope that can support the answer, and reverse-infer the possible user consultation intentions, and accordingly provide a corresponding accurate, clear, and concise intention description text to summarize the information or help that the user may seek.

[0094] Step 2: Based on the intention description text, think about various ways of expressing questions that the user may adopt. Generate a series of similar question texts that are semantically closely related to the intention description. Ensure that this set of question texts is diverse and representative, and can reflect the user's possible question habits.

[0095] Step 3: Combine the knowledge block content and the set of similar question texts to generate a standard response text. The response text should be comprehensive, accurate, and easy to understand, and can provide a satisfactory answer to any question in the set of similar question texts. Please ensure that the content of the response text is complete, the language expression is standard, and it is easy for the user to understand.

[0096] Customer service knowledge points: {to-be-embedded customer service knowledge points}

[0097] Requirement: Please ensure that the generated customer service knowledge points have a clear structure and accurate content, and can provide high-quality customer service for users.

[0098] Step S1020: Aggregate all the customer service knowledge points to build a knowledge base.

[0099] Initialize and create an empty knowledge base, and collect all the customer service knowledge points and add them to this knowledge base.

[0100] In this embodiment, first, the customer service knowledge document is fully semantically segmented to obtain each knowledge block, ensuring that each knowledge block organizes content around a core theme, thereby providing an accurate and rich information source for generating customer service knowledge points, and solving the technical problems of unclear information organization, incomplete coverage of knowledge points, and low user query matching degree commonly found in traditional knowledge base construction. Further, when using a large language model to generate customer service knowledge points, it can accurately reverse-infer the user's consultation intention, and accordingly generate a more representative and comprehensive set of similar question texts and corresponding standard reply texts, not only improving the construction efficiency of the knowledge base, but also enhancing the response ability and accuracy of the knowledge base to user queries.

[0101] In a further embodiment, after step S1100, for each of the question-and-answer pairs, after determining the user question text in the question-and-answer pair and the question matching degrees between the user question text and the sets of similar question texts in each customer service knowledge point in the preset knowledge base respectively, the following steps are included:

[0102] Step S1500: Screen the user question texts with question matching degrees exceeding the first preset threshold and lower than the second preset threshold, and determine the reply matching degree between the customer service reply text of the user question text in the question-and-answer pair and the standard reply text in the customer service knowledge point matching the user question text.

[0103] When the question matching degree exceeds the first preset threshold and is lower than the second preset threshold, it means that it is not sufficient to ensure that the user question text corresponding to the question matching degree is highly similar to the similar question texts in the set of similar question texts. For this, calculate the reply matching degree between the customer service reply text of the user question text in the question-and-answer pair and the standard reply text of the set of similar question texts in the customer service knowledge point. Specifically, use the text encoding model to encode the customer service reply text and the standard reply text to obtain the corresponding text encoding vectors, and use any one of the similarity algorithms such as the cosine similarity algorithm, Euclidean distance algorithm, Pearson correlation coefficient algorithm, and Jaccard coefficient algorithm to calculate the similarity between these two text encoding vectors as the reply matching degree between the customer service reply text and the standard reply text.

[0104] Step S1510: When the reply matching degree exceeds the third preset threshold, append the user question text as a similar question text to the set of similar question texts in the customer service knowledge point matching the user question text.

[0105] That is, it indicates that the reply matching degree at this time is high enough. Based on this, it can be determined that the nature of the user's question text is the same as that of the similar question texts in the similar question text set, and the same standardized reply text can be used for reply. Only the question expression method is different from any of the similar question texts in the similar question text set. Therefore, the user's question text is added to the similar question text set as a similar question text.

[0106] In this embodiment, in practical applications, the expression methods of user questions may be diverse. Even if the question matching degree does not reach the standard of high similarity, the corresponding customer service reply text may still be highly consistent with the standardized reply text in the knowledge base. At this time, by calculating the reply matching degree, the relevance between the user's question text and the customer service knowledge points can be further confirmed, avoiding matching omissions caused by the diversity of question expressions.

[0107] In a further embodiment, in step S1200, the user question texts with the same semantics are used as similar question texts to construct a similar question text set, including the following steps:

[0108] Step S1210: Encode each of the user question texts using a preset text encoding model to determine the corresponding question semantic vectors;

[0109] In one embodiment, the text encoding model uses a pre-trained Bert model in a converged state to encode each of the user question texts into corresponding high-dimensional vectors, that is, question semantic vectors. The question semantic vectors can effectively capture the semantic information of the user question texts, making the vectors corresponding to semantically similar texts closer in the vector space. For example, if two user question texts are semantically similar (such as "How to query the order status?" and "How to view the order progress?"), although their literal expressions are different, the semantic vectors after Bert encoding will have a high similarity in the vector space.

[0110] The Bert model is a pre-trained language model based on the Transformer architecture. Its core principle is to pre-train through a large amount of unsupervised text data to learn the deep semantic information of the language. One of the significant features of the Bert model is its bidirectional encoding mechanism: it simultaneously considers the context information of each word in the text, that is, when encoding a word, it will consider the context content on both its left and right sides. This bidirectional context modeling method enables Bert to capture rich semantic relationships and semantic dependencies in the text, thereby generating vector representations with high semantic expression capabilities.

[0111] Step S1220: Cluster each of the question semantic vectors using a clustering algorithm to determine the clusters to which each question semantic vector belongs;

[0112] In one embodiment, the DBSCAN algorithm is used as the clustering algorithm to cluster all the query semantic vectors. Specifically, a preset neighborhood radius Eps and a threshold MinPts for the number of data object points in the neighborhood are set. All the query semantic vectors are aggregated to form a data set, where each query semantic vector is regarded as a single data object point. An arbitrary data object point is selected from the data set as the target data object point, and the vector distance between each data object point in the data set and the target object point is calculated. According to the Eps and MinPts, it is determined whether the target data object point is a core point. When the target data object point is a core point, all the data object points that are density-reachable from the target data object point are found to form a cluster. When the target object point is a noise point, another data object is selected as the target object point, and the foregoing process is iterated until all the data object points in the data set are processed. When the number of data object points within the neighborhood of any single data object point is greater than MinPts, this data object point is a core point, and the data object points within the neighborhood are boundary points, and the vector distance between the boundary points and the core point is less than or equal to EPS. In addition, the data object points that neither belong to the core points nor belong to the boundary points are noise points.

[0113] After all the query semantic vectors are clustered by the DBSCAN algorithm, multiple clusters can be determined, and the cluster to which each query semantic vector belongs can be determined. Specifically, there may be query semantic vectors that do not belong to any of the determined clusters.

[0114] In a further embodiment, in order to ensure the accuracy and reliability of the number of clusters determined by the clustering algorithm, the number of clusters is determined by any one of the elbow method, the silhouette coefficient method, the information criterion method, and the hierarchical clustering method. By way of example, the elbow method is used to determine the number of clusters by plotting the relationship between the clustering results and the clustering performance evaluation indicators (such as SSE, silhouette coefficient) under different numbers of clusters. When there is an obvious inflection point (shaped like an elbow) in the change trend of the clustering performance evaluation indicator, the number of clusters corresponding to this point can be used as a reasonable choice.

[0115] Step S1230: For each of the clusters, the user query texts corresponding to the query semantic vectors belonging to the cluster are used as similar query texts and aggregated into a set of similar query texts.

[0116] Taking a single cluster as an example, it is not difficult to understand that the user query texts corresponding to the query semantic vectors belonging to the cluster are semantically similar, so these texts can be used as similar query texts respectively, and a set of similar query texts is formed by aggregating them.

[0117] In this embodiment, a process of constructing a similar question text set is disclosed, during which user question texts with similar semantics but different expressions can be reasonably and efficiently classified.

[0118] In a further embodiment, after step S1400, based on the customer service reply text in the question-answer pair of the user question text, optimizing the standard reply text in the customer service knowledge point matching the user question text, the following steps are included:

[0119] Step S1600: respond to the user's inquiry request and obtain the user's question text carried in the request;

[0120] Users can use the customer service consultation function provided by the e-commerce platform to edit and upload user question text, thereby triggering the front-end of the e-commerce platform to construct a user consultation request containing the text and send it to the back-end server. The server then receives and responds to the request and obtains the user question text it carries.

[0121] Step S1610: determining the initial screening matching degree between the user question text and the intention description text in each customer service knowledge point in the knowledge base;

[0122] Similarly, in order to improve the service efficiency of online customer service, the text encoding model can be used in advance to encode the text encoding vector corresponding to the intent description text in each customer service knowledge point in the knowledge base for ready call.

[0123] First, the text encoding model is used to encode the text encoding vector of the user's question text as the question semantic vector, and then the text encoding vectors of each intent description text in the knowledge base are called, and the similarity between these text encoding vectors and the question semantic vector is calculated using any similarity algorithm such as the cosine similarity algorithm, Euclidean distance algorithm, Pearson correlation coefficient algorithm, Jaccard coefficient algorithm, etc., which corresponds to the initial screening matching degree between the user's question text and each intent description text.

[0124] Step S1620: Screen the intention description texts whose preliminary screening matching degree meets the first preset condition, and determine the fine screening matching degree between the user question text and the similar question text set of each intention description text in the customer service knowledge point;

[0125] Sort the intent description texts in descending order according to the initial screening matching degree, and select N intent description texts that are ranked high and exceed the second preset threshold. 1 intention description texts, these intention description texts are considered to meet the first preset condition, the N 1 It can be set by technicians in this field according to business needs.

[0126] Further, call the text encoding vectors corresponding to each similar question text in the set of similar question texts in the customer service knowledge points for each of the intention description texts, and calculate the similarity between these text encoding vectors and the question semantic vector using any one of the similarity algorithms such as the cosine similarity algorithm, Euclidean distance algorithm, Pearson correlation coefficient algorithm, Jaccard coefficient algorithm, etc. For each set of similar question texts, screen out the maximum value among the similarities corresponding to each similar question text in the set as the refined screening matching degree between the user question text and the set of similar question texts.

[0127] Step S1630: Screen out the set of similar question texts whose refined screening matching degree meets the second preset condition, and send the standard reply text of the set of similar question texts in the customer service knowledge points to the user who triggered the user consultation request.

[0128] Screen out the set of similar question texts with the highest refined screening matching degree. This set of similar question texts is regarded as meeting the second preset condition. Obtain the standard reply text of this set of similar question texts in the customer service knowledge points, which can answer the user question text, so it is sent to the user.

[0129] In this embodiment, through the dual screening mechanism of the preliminary screening matching degree and the refined screening matching degree, the most suitable reply content can be quickly located in the vast knowledge base. Among them, the preliminary screening matching degree uses the semantic similarity between the intention description text and the user question text to quickly narrow the matching range and avoid a comprehensive search of the entire knowledge base, thus significantly improving the matching efficiency; while the refined screening matching degree further ensures the accuracy of the matching result through the semantic similarity between the set of similar question texts and the user question text. This hierarchical screening mechanism not only improves the response speed of the request but also effectively reduces the probability of mis-matching, enabling a more accurate understanding of the user's intention and providing high-quality replies.

[0130] In a further embodiment, after step S1400: Optimize the standard reply text in the customer service knowledge points that matches the user question text according to the customer service reply text in the question-and-answer pair for the user question text, the following steps are included:

[0131] Step S1700: Respond to the knowledge point mining event, obtain the customer service knowledge document library, and for each customer service knowledge point in the knowledge base, associate the customer service knowledge document in the customer service knowledge document library that matches it;

[0132] In one embodiment, the server can monitor each time it responds to a user's consultation request. Eventually, when it fails to send a standardized response text to the user who triggered the request, it accumulates the number of such occurrences as the number of unanswered times. When the number of unanswered times reaches a predetermined number, it responds to the knowledge point mining event, thereby mining more new customer service knowledge points and adding them to the knowledge base, which helps reduce the situation of being unable to answer user consultations in the future. In another embodiment, the server can set up a scheduled task to respond to the knowledge point mining event every predetermined period of time, thereby regularly mining more new customer service knowledge points and adding them to the knowledge base, so that the customer service knowledge points in the knowledge base can be continuously enriched, the response scope of the knowledge base can be broadened, and thus better serve the consultation needs of users.

[0133] Since for each standardized response text in the knowledge base of customer service knowledge points, the content material of the response basis associated with this standardized response text can usually be found in the customer service knowledge document in the customer service knowledge document library. Accordingly, such a customer service knowledge document can be associated with this standardized response text for further mining more customer service knowledge points from this customer service knowledge document.

[0134] In one embodiment, taking a single customer service knowledge point as an example, for each customer service knowledge document, the intention description text and the set of similar question texts in this customer service knowledge point, as well as this customer service knowledge document, can be embedded into a preset prompt template to obtain the corresponding prompt text and input it into the large language model. The model determines whether it can extract content from this customer service document to answer the corresponding user question based on the intention description text and the set of similar question texts in this customer service knowledge point and the customer service knowledge document, thereby generating the corresponding customer service response text. If the customer service response text cannot be generated, it can be confirmed that this customer service knowledge document does not match this customer service knowledge point.

[0135] The prompt template includes a task description and the intention description text, the set of similar question texts, and the customer service knowledge document to be embedded. The task description can be: "Please review the provided intention description text, the set of similar question texts, and the customer service knowledge document, determine whether content that can answer the user's question can be extracted from the customer service knowledge document, and generate a standardized customer service response text accordingly.

[0136] Intention description text: {intention description text to be embedded}

[0137] Set of similar question texts: {set of similar question texts to be embedded}

[0138] Customer service knowledge document: {customer service knowledge document to be embedded}

[0139] Requirements:

[0140] 1. Understand the possible consultation intentions of users based on the provided intention description text and the set of similar question texts.

[0141] 2. Carefully read the customer service knowledge document, identify the content related to the user's consultation intention from it, and analyze whether this content contains enough information to generate an accurate and useful reply text to respond to the user's question. If it is sufficient, generate a standard reply text based on this. The reply should be detailed, clear, and able to directly and accurately solve the problem consulted by the user.

[0142] 3. If the document does not contain enough content or is not relevant to the user's consultation intention, output "The customer service knowledge document does not match the intention description text and the set of similar question texts."

[0143] When the large language model generates the customer service reply text, further submit the customer service reply text and the standard reply text in the customer service knowledge points to the large language model to confirm whether their semantics are approximately the same. Those skilled in the art can make flexible changes according to the disclosure here. When they are approximately the same, it is confirmed that the customer service knowledge document matches the customer service knowledge points; otherwise, they do not match.

[0144] It is not difficult to understand that in the above implementation, a two-step method is adopted to let the large language model independently judge whether it can generate a customer service reply text, and when it can generate, judge whether the self-generated reply is consistent with the actual existing reply. First, it ensures that the large language model independently extracts information from the customer service knowledge document and generates a reply without knowing the standard reply text at all, avoiding the model directly "copying" or relying on the existing standard reply text, thus ensuring the independence and autonomy of the generation process.

[0145] Step S1710: For each customer service knowledge document in the customer service knowledge document library, use the large language model to take each customer service knowledge point associated with the customer service knowledge document as a sample instance, and based on the customer service knowledge document and each sample instance, mine new customer service knowledge points from the customer service knowledge document and append them to the knowledge base.

[0146] Embed each sample instance and the customer service knowledge document into a preset prompt template to obtain the corresponding prompt text, input it into the large language model, and obtain the new customer service knowledge points mined by the model from the customer service Q&A and append them to the knowledge base.

[0147] The prompt template includes a task description and the sample instances and customer service knowledge document to be embedded. The task description can be: "Please analyze the following provided customer service knowledge document and each associated sample instance, and these sample instances are generated based on the customer service knowledge document. Therefore, based on the document content and sample instances, identify and extract new customer service knowledge points. These new customer service knowledge points should be able to enrich the knowledge base, broaden the reply scope and accuracy of the knowledge base to respond to user questions, and thus more effectively respond to the consultation needs of users.

[0148] Sample instance: {Each customer service knowledge point associated with the customer service knowledge document to be embedded}

[0149] Customer service knowledge document: {Customer service knowledge document to be embedded}

[0150] Requirements:

[0151] 1. Carefully read and understand the content of the customer service knowledge document, and identify the key information and themes therein.

[0152] 2. Analyze the sample instance, understand the intention description of each customer service knowledge point, the structure and content of the similar question text set and the standardized reply text, as well as the user questions that can be addressed and solved by this customer service knowledge point.

[0153] 3. Combine the content of the document and the sample instance, and identify the reply materials that may be included in the document but are not covered by the existing customer service knowledge points and the user questions they can solve.

[0154] 4. Generate new intention description text, similar question text set and standardized reply text based on the identified reply materials to form new customer service knowledge points.

[0155] 5. Ensure that the newly generated customer service knowledge points do not duplicate the solution of the same user questions as the existing knowledge points, and can provide accurate answers to new user consultations.

[0156] 6. Output the newly generated customer service knowledge points so that they can be appended to the knowledge base.

[0157] In this embodiment, by monitoring the number of times that user consultations cannot be answered or setting a timed task to trigger a knowledge point mining event, the optimization requirements of the knowledge base are actively responded to. It not only solves the technical problems of the traditional knowledge base being updated lagging behind and unable to cover new user questions in a timely manner, but also can dynamically mine and supplement new customer service knowledge points, continuously broaden the reply scope of the knowledge base, so as to more effectively meet the consultation needs of users.

[0158] In a further embodiment, in step S1700, for each customer service knowledge point in the knowledge base, associating the customer service knowledge point with the customer service knowledge document in the customer service knowledge document library that matches it includes the following steps:

[0159] Step S1701, for each customer service knowledge point in the knowledge base, use a large language model to determine whether each customer service knowledge document in the customer service knowledge document library belongs to the reply basis document of the standardized reply text in this customer service knowledge point;

[0160] Taking a single customer service knowledge point as an example, for each customer service knowledge document, embed the standardized response text in the customer service knowledge point and the customer service knowledge document into a preset prompt template to obtain the corresponding prompt text, and input it into the large language model. The model determines whether the customer service knowledge text belongs to the response basis document of the standardized response text in the customer service knowledge point and outputs a judgment result.

[0161] The prompt template includes a task description and the customer service knowledge document and standardized response text to be embedded. The task description can be: "Please analyze the provided customer service knowledge document and standardized response text below to determine whether the content in the document directly supports or provides the basis for generating the standardized response text. Specifically, it is necessary to evaluate whether there is content in the document that is closely related to the answer provided in the standardized response text and whether it can be used as a reliable source for answering users."

[0162] Customer service knowledge document: {Content of the customer service knowledge document to be embedded}

[0163] Standardized response text: {Standardized response text in the customer service knowledge point to be embedded}

[0164] Requirements:

[0165] 1. Carefully read and understand the content of the customer service knowledge document, including all details and context information. After that, compare the content of the document with the standardized response text and pay attention to the relevance between the compared content and the standardized response text.

[0166] 3. Judge whether the content in the document is sufficiently relevant to the standardized response text. For this purpose, the extraction-based summarization method can be used to extract all the content in the document that is closely related to the standardized response text, and then based on these contents, judge whether the text fully contains all the bases for answering the user's question in the standardized response text. When it fully contains, there is sufficient evidence to prove that the document directly supports or provides the basis for generating the standardized response text.

[0167] 4. Output a clear judgment result indicating whether the customer service knowledge document belongs to the response basis document of the standardized response text.

[0168] Step S1702: When the customer service knowledge document belongs to the response basis document, confirm its match with the customer service knowledge point and establish an association relationship between the two.

[0169] It is not difficult to understand that at this time, it means that the customer service knowledge document matches the customer service knowledge point, so the two are associated.

[0170] In this embodiment, the present embodiment analyzes the customer service knowledge document by using a large language model to determine whether it is a reply basis document for the standardized reply text of the customer service knowledge point, and accordingly establishes the association relationship between the customer service knowledge document and the customer service knowledge point, which not only solves the problem of the disconnection between the knowledge point and the original knowledge document in the traditional knowledge base, but also provides a reliable data source association for the dynamic update and expansion of the knowledge base.

[0171] Please refer to Figure 3 , a customer service knowledge base optimization device provided to meet one of the purposes of the present application, is a functional embodiment of the customer service knowledge base optimization method of the present application. On the other hand, a customer service knowledge base optimization device provided to meet one of the purposes of the present application includes a matching degree determination module 1100, a first addition module 1200, a second addition module 1300, and a reply optimization module 1400. Among them, the matching degree determination module 1100 is used to obtain each question-and-answer pair in the historical conversation record. For each question-and-answer pair, determine the user question text in the question-and-answer pair, and respectively determine the question matching degree between the user question text and the similar question text sets in each customer service knowledge point in the preset knowledge base; the first addition module 1200 is used to screen the user question texts with the question matching degree lower than the first preset threshold, use the user question texts with the same semantics as the similar question texts to construct a similar question text set, and for each similar question text set, according to the similar question text set and the customer service reply text of each similar question text in the question-and-answer pair, determine the corresponding intention description text and standardized reply text, and construct new customer service knowledge points to be added to the knowledge base; the second addition module 1300 is used to screen the user question texts with the question matching degree exceeding the second preset threshold and lower than the third preset threshold, and add the user question texts as similar question texts to the similar question text sets in the customer service knowledge points matching the user question texts, and the second preset threshold is greater than the first preset threshold; the reply optimization module 1400 is used to screen the user question texts with the question matching degree exceeding the third preset threshold, and optimize the standardized reply text in the customer service knowledge point matching the user question text according to the customer service reply text of the user question text in the question-and-answer pair.

[0172] In a further embodiment, before the matching degree determination module 1100, it includes: a semantic segmentation sub-module, which is used to obtain the customer service knowledge document library, perform semantic segmentation on each customer service knowledge document in the library, and obtain each knowledge block corresponding to different semantics; a knowledge point generation sub-module, which is used to generate corresponding customer service knowledge points by using a large language model according to each knowledge block, and the customer service knowledge points include intention description text, similar question text set, and standardized reply text; a knowledge base construction sub-module, which is used to construct a knowledge base by aggregating all the customer service knowledge points.

[0173] In a further embodiment, after the matching degree determination module 1100, it includes: a first matching degree determination sub-module, configured to screen user question texts whose question matching degree exceeds a first preset threshold and is lower than a second preset threshold, determine the customer service reply text in the question-answer pair corresponding to the user question text, and the reply matching degree between the standard reply text in the customer service knowledge point matching the user question text; a third append sub-module, configured to, when the reply matching degree exceeds a third preset threshold, use the user question text as a similar question text and append it to the set of similar question texts in the customer service knowledge point matching the user question text.

[0174] In a further embodiment, the matching degree determination module 1100 includes: a text encoding sub-module, configured to encode each of the user question texts using a preset text encoding model to determine corresponding question semantic vectors; a vector clustering sub-module, configured to cluster each of the question semantic vectors using a clustering algorithm to determine the clusters to which each of the question semantic vectors belongs; a text collection sub-module, configured to, for each of the clusters, use the user question texts corresponding to the question semantic vectors belonging to the cluster as similar question texts and collect them into a set of similar question texts.

[0175] In a further embodiment, after the reply optimization module 1400, it includes: a request response sub-module, configured to respond to a user consultation request and obtain the user question text carried by the request; a second matching degree sub-module, configured to determine the initial screening matching degree between the user question text and the intention description texts in each of the customer service knowledge points in the knowledge base; a third matching degree sub-module, configured to screen the intention description texts whose initial screening matching degree meets a first preset condition, and determine the refined screening matching degree between the user question text and the sets of similar question texts in the customer service knowledge points corresponding to each of the intention description texts; a user reply sub-module, configured to screen the sets of similar question texts whose refined screening matching degree meets a second preset condition, and send the standard reply text in the sets of similar question texts in the customer service knowledge points to the user who triggered the user consultation request.

[0176] In a further embodiment, after the reply optimization module 1400, it includes: an event response sub-module, configured to respond to a knowledge point mining event, obtain a customer service knowledge document library, and for each of the customer service knowledge points in the knowledge base, associate the customer service knowledge document in the customer service knowledge document library that matches it; a knowledge point append sub-module, configured to, for each of the customer service knowledge documents in the customer service knowledge document library, use the customer service knowledge points associated with the customer service knowledge document as sample instances, and use a large language model to mine new customer service knowledge points from the customer service knowledge document according to the customer service knowledge document and each of the sample instances and append them to the knowledge base.

[0177] In a further embodiment, the event response sub-module includes: a reply basis sub-module, which is used to determine, for each customer service knowledge point in the knowledge base, whether each customer service knowledge document in the customer service knowledge document library belongs to the reply basis document of the standard reply text in the customer service knowledge point by using a large language model; an association establishment sub-module, which is used to establish an association relationship between the customer service knowledge document and the customer service knowledge point when the customer service knowledge document belongs to the reply basis document.

[0178] To solve the above technical problems, an embodiment of the present application also provides a computer device. As Figure 4 shown, it is a schematic internal structure diagram of the computer device. The computer device includes a processor, a computer-readable storage medium, a memory, and a network interface connected through a system bus. Among them, the computer-readable storage medium of the computer device stores an operating system, a database, and computer-readable instructions. The database can store a control information sequence. When the computer-readable instructions are executed by the processor, the processor can implement a customer service knowledge base optimization method. The processor of the computer device is used to provide computing and control capabilities to support the operation of the entire computer device. The memory of the computer device can store computer-readable instructions. When the computer-readable instructions are executed by the processor, the processor can execute the customer service knowledge base optimization method of the present application. The network interface of the computer device is used to connect and communicate with the terminal. Those skilled in the art can understand that Figure 4 the structure shown in the figure is only a block diagram of some structures related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.

[0179] In this embodiment, the processor is used to execute Figure 3 the specific functions of each module and its sub-module. The memory stores the program code and various types of data required to execute the above modules or sub-modules. The network interface is used for data transmission between the user terminal or the server. The memory in this embodiment stores the program code and data required to execute all modules / sub-modules in the customer service knowledge base optimization device of the present application. The server can call the program code and data of the server to execute the functions of all sub-modules.

[0180] The present application also provides a storage medium storing computer-readable instructions. When the computer-readable instructions are executed by one or more processors, one or more processors are caused to execute the steps of the customer service knowledge base optimization method of any embodiment of the present application.

[0181] Those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments of the present application can be completed by instructing relevant hardware through a computer program. This computer program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. Among them, the aforementioned storage medium can be a computer-readable storage medium such as a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM), etc.

[0182] In summary, the present application can fully and reasonably utilize historical conversation records to continuously and automatically optimize the knowledge base.

[0183] Those skilled in the art of the present technology can understand that the steps, measures, and solutions in the various operations, methods, and processes discussed in the present application can be alternated, changed, combined, or deleted. Further, the other steps, measures, and solutions in the various operations, methods, and processes discussed in the present application can also be alternated, changed, rearranged, decomposed, combined, or deleted. Further, the steps, measures, and solutions in the prior art that are the same as those in the various operations, methods, and processes open-sourced in the present application can also be alternated, changed, rearranged, decomposed, combined, or deleted.

[0184] The above are only some implementation manners of the present application. It should be noted that for those of ordinary skill in the art of the present technology, without departing from the principle of the present application, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present application.

Claims

1. A customer service knowledge base optimization method, characterized in that: The steps include: Obtain each question-answer pair in the historical conversation record, and for each question-answer pair, determine the question matching degree between the user question text in the question-answer pair and the similar question text set in each customer service knowledge point in the preset knowledge base; Filtering the user question texts whose question matching degree is lower than a first preset threshold, taking the user question texts with the same semantics as similar question texts, constructing a similar question text set, and for each similar question text set, determining the corresponding intent description text and the standard reply text according to the similar question text set and the customer service reply text of each similar question text in the question-answer pair, so as to construct a new customer service knowledge point and append it to the knowledge base; Filtering the user question texts whose question matching degree exceeds a second preset threshold and is lower than a third preset threshold, taking the user question texts as similar question texts, and adding them to a similar question text set in a customer service knowledge point that matches the user question texts, wherein the second preset threshold is greater than the first preset threshold; The user question texts whose question matching degree exceeds a third preset threshold are screened, and based on the customer service reply texts of the user question texts in the question-answer pair, the standard reply texts in the customer service knowledge points that match the user question texts are optimized.

2. The customer service knowledge base optimization method according to claim 1, characterized in that: Before obtaining each question-answer pair in the historical conversation record, the following steps are included: Obtain a customer service knowledge document library, and semantically segment each customer service knowledge document therein to obtain knowledge blocks corresponding to different semantics; A large language model is used to generate corresponding customer service knowledge points according to each knowledge block, wherein the customer service knowledge points include intention description text, similar question text set, and standard answer text; Gather all the customer service knowledge points to build a knowledge base.

3. The customer service knowledge base optimization method according to claim 1, characterized in that: For each question-answer pair, after determining the question matching degree between the user question text in the question-answer pair and the similar question text sets in each customer service knowledge point in the preset knowledge base, the following steps are included: Screening user question texts whose question matching degree exceeds a first preset threshold and is lower than a second preset threshold, and determining a reply matching degree between a customer service reply text in a question-answer pair of the user question text and a standard reply text in a customer service knowledge point matching the user question text; When the answer matching degree exceeds a third preset threshold, the user question text is taken as a similar question text and added to a similar question text set in the customer service knowledge point that matches the user question text.

4. The customer service knowledge base optimization method according to claim 1, characterized in that: The user question texts with the same semantics are used as similar question texts to construct a similar question text set, including the following steps: Encode each of the user's question texts using a preset text encoding model to determine each corresponding question semantic vector; Using a clustering algorithm to cluster each of the question semantic vectors to determine the cluster to which each question semantic vector belongs; For each of the clusters, the user question texts corresponding to each question semantic vector belonging to the cluster are taken as similar question texts and grouped into a similar question text set.

5. The customer service knowledge base optimization method according to claim 1, characterized in that: After optimizing the standard answer text in the customer service knowledge point matching the user question text according to the customer service answer text in the question-answer pair, the following steps are included: Respond to the user's inquiry request and obtain the user's question text carried in the request; Determine the initial screening matching degree between the user question text and the intention description text in each customer service knowledge point in the knowledge base; Screening the intention description texts whose preliminary screening matching degree meets the first preset condition, and determining the fine screening matching degree between the user question text and the similar question text set of each intention description text in the customer service knowledge point; The similar question text set whose fine screening matching degree meets the second preset condition is screened, and the standard reply text of the similar question text set in the customer service knowledge point is sent to the user who triggered the user consultation request.

6. The customer service knowledge base optimization method according to claim 1, characterized in that: After optimizing the standard answer text in the customer service knowledge point matching the user question text according to the customer service answer text in the question-answer pair, the following steps are included: In response to a knowledge point mining event, a customer service knowledge document library is obtained, and for each customer service knowledge point in the knowledge library, the customer service knowledge point is associated with a customer service knowledge document in the customer service knowledge document library that matches the customer service knowledge point; For each customer service knowledge document in the customer service knowledge document library, each customer service knowledge point associated with the customer service knowledge document is used as a sample instance, and a large language model is used to mine new customer service knowledge points from the customer service knowledge document and each sample instance and append them to the knowledge base.

7. The customer service knowledge base optimization method according to claim 1, characterized in that: For each customer service knowledge point in the knowledge base, associating the customer service knowledge point with a customer service knowledge document in a customer service knowledge document base that matches the customer service knowledge point includes the following steps: For each customer service knowledge point in the knowledge base, a large language model is used to determine whether each customer service knowledge document in the customer service knowledge document library belongs to the reply basis document of the standard reply text in the customer service knowledge point; When the customer service knowledge document belongs to the reply basis document, it is confirmed that it matches the customer service knowledge point and an association relationship is established between the two.

8. A customer service knowledge base optimization device, characterized in that: include: A matching degree determination module is used to obtain each question-answer pair in the historical conversation record, and for each question-answer pair, determine the question matching degree between the user question text in the question-answer pair and the similar question text set in each customer service knowledge point in the preset knowledge base; A first appending module is used to screen user question texts whose question matching degree is lower than a first preset threshold, and use user question texts with the same semantics as similar question texts to construct a similar question text set, and for each similar question text set, determine the corresponding intention description text and standard reply text according to the similar question text set and the customer service reply text of each similar question text in the question-answer pair, so as to construct a new customer service knowledge point and append it to the knowledge base; A second appending module is used to screen the user question texts whose question matching degree exceeds a second preset threshold and is lower than a third preset threshold, and to append the user question texts as similar question texts to a similar question text set in a customer service knowledge point that matches the user question texts, wherein the second preset threshold is greater than the first preset threshold; The answer optimization module is used to screen the user question texts whose question matching degree exceeds the third preset threshold, and optimize the standard answer text in the customer service knowledge points that match the user question text according to the customer service reply text in the question and answer pair.

9. A computer device comprising a central processing unit and a memory, characterized in that: The central processing unit is used to call and run the computer program stored in the memory to execute the steps of the method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: It stores a computer program implemented according to the method described in any one of claims 1 to 7 in the form of computer-readable instructions, and when the computer program is called and executed by a computer, the steps included in the corresponding method are executed.

Citation Information

Cited By

  • Intelligent customer service knowledge base automatic updating method based on large language model

    CN121501919A