Intelligent customer service reply processing method, device and system, electronic equipment, storage medium and program product

By preprocessing and extracting the long text data of the bank's intelligent customer service system, and clustering to determine the target topic terms, the problem of insufficient accuracy in replying to smart customer service when handling long text consultations is solved, and accurate response to user problems and improvement of service experience is achieved.

CN120046604APending Publication Date: 2025-05-27INDUSTRIAL AND COMMERCIAL BANK OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510167071.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-14
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

When handling long text consultations from users, the bank's intelligent customer service system is not good at refining topic words, resulting in insufficient reply accuracy. Users need to simplify the problem by themselves, increase communication costs and reduce service experience.

Method used

By preprocessing the long text data entered by the user, feature phrases are extracted, and subject extraction is performed, target topic words are determined, and corresponding reply text is found based on the reply knowledge base, and feedback is given to the user.

Benefits of technology

It realizes accurate responses to user problems, improves the accuracy of responses of intelligent customer service, reduces user operation steps and communication costs, and improves service experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120046604A_ABST
    Figure CN120046604A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides an intelligent customer service reply processing method, device and system, electronic equipment, a storage medium and a program product, and relates to the field of artificial intelligence. The method comprises the following steps: in response to an input instruction of a user, preprocessing long text data indicated by the input instruction to obtain feature phrases; wherein the input instruction is used for indicating long text data, and the feature word group comprises part of feature words in the long text data; performing subject extraction processing on feature words in the feature word group to obtain a subject word group; wherein the subject word group comprises subject words corresponding to the feature words; according to the feature word group and the subject word group, performing clustering processing on the subject word, and determining a target subject word; the target subject term represents a unique subject term corresponding to the long text data; based on a reply knowledge base, determining a reply text corresponding to the target subject term; and feeding back the obtained reply text to the user. The effect of accurately replying a user question is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence, and in particular to a method, device, system, electronic device, storage medium and program product for intelligent customer service response processing. Background Art

[0002] When dealing with long text inquiries from users, the bank's intelligent customer service system is not good at refining keywords, resulting in inaccurate responses. In order to make the intelligent customer service understand the problem accurately, users often simplify and refine the problem into short text, but this is particularly difficult for complex or lengthy problems, which leads to more operation steps or inaccurate responses.

[0003] In this case, users often turn to manual customer service, which not only increases communication costs but also reduces satisfaction with the service experience.

[0004] Therefore, there is an urgent need for an intelligent customer service response processing solution that can accurately answer user questions. Summary of the invention

[0005] The present application provides an intelligent customer service reply processing method, device, system, electronic device, storage medium and program product to achieve the effect of accurately replying to user questions.

[0006] In a first aspect, the present application provides a method for processing a reply of an intelligent customer service, comprising:

[0007] In response to a user's input instruction, preprocessing the long text data indicated by the input instruction to obtain a feature phrase group; wherein the input instruction is used to indicate the long text data, and the feature phrase group includes some feature words in the long text data;

[0008] Performing a topic extraction process on the feature words in the feature word group to obtain a topic word group; wherein the topic word group includes topic words corresponding to the feature words; performing a clustering process on the topic words according to the feature word group and the topic word group to determine a target topic word; the target topic word represents a unique topic word corresponding to the long text data;

[0009] Based on a reply knowledge base, a reply text corresponding to the target subject word is determined; wherein the reply knowledge base contains reply texts corresponding to the subject words; and the obtained reply text is fed back to the user.

[0010] In a second aspect, the present application provides a reply processing device for intelligent customer service, comprising:

[0011] An acquisition module, for responding to an input instruction of a user, preprocessing the long text data indicated by the input instruction to obtain a feature phrase; wherein the input instruction is used to indicate the long text data, and the feature phrase includes a part of the feature words in the long text data;

[0012] An extraction module is used to perform a topic extraction process on the feature words in the feature word group to obtain a topic word group; wherein the topic word group includes a topic word corresponding to the feature word; based on the feature word group and the topic word group, the topic word is clustered to determine a target topic word; the target topic word represents a unique topic word corresponding to the long text data;

[0013] The reply module is used to determine the reply text corresponding to the target subject word based on the reply knowledge base; wherein the reply knowledge base contains the reply text corresponding to the subject word; and feed back the obtained reply text to the user.

[0014] In a third aspect, an embodiment of the present application provides a reply processing system for intelligent customer service, wherein the system is used to implement the method provided in the first aspect and any implementation method of the first aspect.

[0015] In a fourth aspect, an embodiment of the present application provides an electronic device, including: a memory, a processor;

[0016] The memory stores computer executable instructions.

[0017] The processor executes the computer-executable instructions stored in the memory, so that the processor performs the method provided in the first aspect and any embodiment of the first aspect.

[0018] In a fifth aspect, an embodiment of the present application provides a computer-readable storage medium, in which computer-executable instructions are stored. When the computer-executable instructions are executed by a processor, they are used to implement the first aspect above and / or various possible implementations of the first aspect.

[0019] In a sixth aspect, an embodiment of the present application provides a computer program product, wherein the computer program product includes a computer program, and when the computer program is executed by a processor, the method provided in the first aspect and any implementation method of the first aspect is implemented.

[0020] The present embodiment provides a method, device, system, electronic device, storage medium and program product for intelligent customer service reply processing. The long text data indicated by the user's input instruction is preprocessed. This process aims to extract key feature words in the long text data to form a feature phrase. Then, the feature words in the feature phrase are subjected to topic extraction processing to obtain a theme phrase group containing the feature words corresponding to the theme words. This process utilizes the topic modeling technology, which can deeply explore the potential connection between the feature words, so as to accurately identify the theme structure in the long text data. Subsequently, according to the feature phrase group and the theme phrase group, the method clusters the theme words to determine the target theme word. As the only theme word corresponding to the long text data, the target theme word can accurately reflect the main content and intention of the long text data. The implementation of this step depends on an efficient clustering algorithm, which can ensure the accuracy and stability of the clustering results. After determining the target theme word, the reply text corresponding to the target theme word is searched based on the reply knowledge base. The reply knowledge base is a database containing theme words and their corresponding reply texts, which enables the intelligent customer service to quickly and accurately find replies that match the user's questions. Finally, the obtained reply text is fed back to the user, achieving an accurate response to the user's question. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.

[0022] Figure 1 A schematic diagram of a process for processing a reply of an intelligent customer service provided in an embodiment of the present application Figure 1 ;

[0023] Figure 2 A schematic diagram of a process for processing a reply of an intelligent customer service provided in an embodiment of the present application Figure 2 ;

[0024] Figure 3 A schematic diagram of the structure of an intelligent customer service reply processing device provided in an embodiment of the present application Figure 1 ;

[0025] Figure 4 A schematic diagram of the structure of an intelligent customer service reply processing device provided in an embodiment of the present application Figure 2 ;

[0026] Figure 5 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application.

[0027] The above drawings have shown clear embodiments of the present application, which will be described in more detail later. These drawings and text descriptions are not intended to limit the scope of the present application in any way, but to illustrate the concept of the present application to those skilled in the art by referring to specific embodiments. DETAILED DESCRIPTION

[0028] Exemplary embodiments will be described in detail herein, examples of which are shown in the accompanying drawings. When the following description refers to the drawings, the same numbers in different drawings represent the same or similar elements unless otherwise indicated. The implementations described in the following exemplary embodiments do not represent all implementations consistent with the present application. Instead, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.

[0029] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, storage, use, processing, transmission, provision, disclosure and application of the relevant data comply with the relevant laws, regulations and standards of the relevant countries and regions, take necessary confidentiality measures, do not violate public order and good morals, and provide corresponding operation entrances for users to choose to authorize or refuse.

[0030] In addition, this application involves conducting big data analysis of user information (including but not limited to personal biometrics, identity data, consumption data, asset data, electronic terminal operation data, etc.), and using artificial intelligence technology to make automated decisions, and making technical solutions that have a significant impact on personal rights and interests based on the results of automated decisions. The application provides users with corresponding operation entrances for them to choose to agree or reject the results of automated decisions; if the user chooses to reject, the expert decision-making process will be entered.

[0031] It should be noted that the reply processing method, device, system, electronic device, storage medium and product for intelligent customer service provided in this application can be used in the field of artificial intelligence, and can also be used in any field other than artificial intelligence. The application field of the reply processing method, device, system, electronic device, storage medium and product for intelligent customer service in this application is not limited.

[0032] The bank's intelligent customer service automatic reply system aims to understand the user's intentions and needs by identifying the text information entered by the user, and then provide appropriate replies. To achieve this goal, the system mainly uses advanced technologies such as natural language processing and machine learning. The core of these technologies is to discover the regularity of text appearance and the connection between text and semantics and grammar, that is, to extract valuable information from unstructured data and convert it into structured data.

[0033] Existing systems have limited processing capabilities for long text messages, especially when users find it difficult to refine the core of the problem on their own. When faced with complex or multi-step problem descriptions, intelligent customer service often cannot effectively parse the user's true intentions, resulting in inaccurate or insufficiently relevant answers. Because intelligent customer service has difficulty handling complex, long-text inquiries, many users have to simplify their questions or even ask them multiple times to ensure that they are correctly understood. This practice not only increases the user's communication costs, but may also affect the overall service efficiency and satisfaction because more rounds of interaction are required. When intelligent customer service cannot meet user needs, users may turn to manual customer service for help. This not only increases operating costs, but also means that intelligent customer service fails to fully utilize its automation advantages, reducing the overall benefits of the system.

[0034] The intelligent customer service response processing method, device, system, electronic device, storage medium and product provided in this application are intended to solve the above technical problems of the prior art.

[0035] The technical solution of the present application and how the technical solution of the present application solves the above-mentioned technical problems are described in detail below with specific embodiments. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of the present application will be described below in conjunction with the accompanying drawings.

[0036] Figure 1 A schematic diagram of a process for processing a reply of an intelligent customer service provided in an embodiment of the present application Figure 1 ,like Figure 1 As shown, the method includes:

[0037] S101. In response to a user input instruction, preprocess the long text data indicated by the input instruction to obtain a feature phrase group; wherein the input instruction is used to indicate the long text data, and the feature phrase group includes some feature words in the long text data.

[0038] Exemplarily, the input instruction is long text data directly sent by the user to the intelligent customer service system and requires the intelligent customer service to reply and process. Preprocessing the long text data refers to performing a series of operations on the long text data to improve the data quality and make it more suitable for subsequent feature extraction and analysis. The specific preprocessing includes: removing stop words, word segmentation, removing numbers and symbols, data cleaning, data transformation, etc. Stop words are words that frequently appear in the text but have no actual meaning, such as "of", "is", "in", etc. By removing these stop words, noise interference can be reduced and the efficiency of subsequent processing can be improved; word segmentation is to cut continuous long text data into a sequence of meaningful words. Word segmentation is a basic step in text processing and is crucial for subsequent feature extraction; removing numbers and symbols is to reduce noise interference and requires removing the numbers and symbols in the long text data and retaining the words with actual meaning; data cleaning is to detect and correct errors, incomplete or inaccurate data in the data set, such as handling missing values and outliers, etc.; data transformation is to ensure that the data set meets the requirements of analysis and modeling, including data standardization, normalization, etc. By preprocessing the long text data, a feature word group containing some feature words is obtained.

[0039] S102. Perform theme extraction processing on the feature words in the feature word group to obtain a theme word group; among them, the theme word group includes the theme words corresponding to the feature words; according to the feature word group and the theme word group, perform clustering processing on the theme words to determine the target theme word; the target theme word represents the unique theme word corresponding to the long text data.

[0040] For example, in text processing, feature words represent keywords or phrases that can represent long text data; subject words are the abstraction of feature words at a higher level, representing the theme or concept pointed to by each feature word; clustering is an unsupervised learning method used to group similar subject words together, so that subject words in the same group are similar to each other, while subject words in different groups are different. Based on the feature word group, a corresponding subject word is generated for each feature word to form a subject word group. The subject model is applied to the feature word group for topic modeling. The subject model will assign feature words to different topics, and each topic is represented by a group of related feature words; although multiple feature words may be mapped to the same topic, they are not strictly one-to-one corresponding, but independent of each other, that is, multiple feature words can point to a broader topic together. The subject words corresponding to all feature words are combined to form a subject word group. Based on the feature word group and the subject word group, the subject words are clustered to determine the target subject word. Select a suitable clustering algorithm (such as K-means, hierarchical clustering, DBSCAN, etc.); the subject words in the subject word group are used as the input of the clustering algorithm for clustering processing. Clustering algorithms divide subject words into different clusters based on their similarities. One or more representative subject words are selected from the clustering results as target subject words. This is usually determined by calculating the average score of the subject words in the cluster, selecting the cluster center, or based on other evaluation indicators. The target subject word should be able to uniquely represent the theme of the long text data, that is, it is the core or summary of the text content.

[0041] S103, based on the reply knowledge base, determining the reply text corresponding to the target subject word; wherein the reply knowledge base contains the reply text corresponding to the subject word; and feeding back the obtained reply text to the user.

[0042] Exemplarily, before this, it is necessary to build a reply knowledge base, which contains a series of subject words and their corresponding reply texts. The reply text is pre-written to answer or respond to questions or requests related to the subject words. After the target subject word is determined, the reply text matching the subject word is searched from the reply knowledge base. The matching process can be a simple string matching or a matching based on semantic similarity (such as using cosine similarity, word embedding and other technologies). The found reply text is formatted and processed as necessary to ensure that it is suitable for display to the user. The processed reply text is fed back to the user through appropriate channels (such as web pages, application program interfaces, etc.). The user can perform further interactions or operations based on the content of the reply text.

[0043] The present embodiment provides a method for processing a reply of an intelligent customer service. The long text data indicated by the user's input instruction is preprocessed. This process aims to extract key feature words in the long text data to form a feature phrase. Then, the feature words in the feature phrase are subjected to a topic extraction process to obtain a topic phrase group containing the feature words corresponding to the subject words. This process utilizes the topic modeling technology, which can deeply explore the potential connection between the feature words, thereby accurately identifying the subject structure in the long text data. Subsequently, according to the feature phrase group and the subject phrase group, the method clusters the subject words to determine the target subject word. As the only subject word corresponding to the long text data, the target subject word can accurately reflect the main content and intention of the long text data. The implementation of this step depends on an efficient clustering algorithm, which can ensure the accuracy and stability of the clustering results. After determining the target subject word, the reply text corresponding to the target subject word is searched based on the reply knowledge base. The reply knowledge base is a database containing subject words and their corresponding reply texts, which enables the intelligent customer service to quickly and accurately find a reply that matches the user's question. Finally, the obtained reply text is fed back to the user, achieving an accurate response to the user's question.

[0044] Figure 2 A schematic diagram of a process for processing a reply of an intelligent customer service provided in an embodiment of the present application Figure 2 ,like Figure 2 As shown, the embodiment of the present application is Figure 1 Based on the embodiment, a method for processing a reply of an intelligent customer service is described in detail, and the method includes:

[0045] S201. Obtain historical question and answer data, and summarize the historical question and answer data to obtain a preliminary knowledge base; create a question and answer knowledge base based on the knowledge graph and the preliminary knowledge base.

[0046] Exemplarily, historical question and answer data is obtained and preprocessed, including: removing irrelevant content, formatting, etc. Removing irrelevant content means removing those parts of the conversation that are not related to the business, such as small talk and greetings; formatting ensures that all data is stored in a unified format for subsequent analysis and use. Natural language processing technology is used to extract keywords from historical question and answer data, and the question and answer data is classified into different topics according to the keywords to obtain a preliminary knowledge base.

[0047] Exemplarily, a knowledge graph is a structured knowledge base that represents entities, concepts, and the relationships between them in the form of a graph. Based on the knowledge graph and the preliminary knowledge base, the first step in creating a question-answering knowledge base is to create a node for each entity. These nodes will serve as the basis of the graph to represent different entities. The node creation process usually involves the identification and extraction of entities, which can be achieved through natural language processing technology. After the nodes are created, edges need to be added between the nodes based on the relationships between the entities. These edges represent the associations between entities and are an important part of the graph structure. The addition of relationships can be achieved by analyzing text data, using existing knowledge bases, or combining user input. When adding relationships, it is necessary to ensure the accuracy and consistency of the relationships to avoid errors and redundancy in the graph. In the process of graph construction, the graph also needs to be optimized. This includes operations such as denoising and merging duplicate nodes to improve the quality and usability of the graph. Denoising refers to removing irrelevant information or noise in the graph, such as duplicate edges, invalid nodes, etc. Merging duplicate nodes is to merge multiple nodes representing the same entity into one to reduce the complexity and redundancy of the graph. These optimization operations help improve the accuracy and efficiency of the graph, enabling it to better serve various application scenarios.

[0048] The knowledge graph represents entities, concepts and their relationships in the form of a graph, making the association between keywords clearer and more intuitive. This structured representation greatly improves the efficiency of knowledge retrieval, and the system can quickly retrieve relevant keywords and reply texts.

[0049] S202, performing word segmentation processing on the long text data to obtain word segmentation groups; wherein the word segmentation groups include at least one word; and deleting repeated and meaningless words in the word segmentation groups to obtain feature word groups.

[0050] Exemplarily, the purpose of word segmentation for long text data is to divide the continuous long text into independent vocabulary units (i.e., words or phrases) for subsequent processing and analysis. The word segmentation algorithm can be a rule-based method, a statistical method, or a combination of the two. Rule-based methods usually rely on predefined dictionaries and grammatical rules, while statistical methods rely on the statistical characteristics of a large amount of text data. Rule-based and statistical methods combine the advantages of rules and statistics, first preliminarily segment the text by rules, and then use statistical models to optimize the results. After word segmentation, a word segmentation group is obtained, which contains all the words in the long text data. For example: given a Chinese text "I like to eat Beijing roast duck.", after word segmentation, ["I", "like", "eat", "Beijing", "roast duck"] is obtained.

[0051] Exemplarily, the purpose of deleting the repeatedly occurring and meaningless words in the word segmentation group is to remove the same words that appear multiple times within the same long text data and reduce redundant information. Traverse the word segmentation group to identify and delete the repeatedly occurring words. This can be achieved through data structures such as hash tables or sets to efficiently detect duplicates. Identify and delete meaningless words according to a preset stop word list (including common meaningless words such as "de", "le", "zai", etc.) or based on the statistical characteristics of words (such as word frequency, context relevance, etc.). By deleting the repeatedly occurring and meaningless words in the word segmentation group, a feature word group is obtained, where the feature word group contains feature words, and the feature words are valuable words that have been screened.

[0052] Through word segmentation processing, the long text data is transformed into smaller units, reducing the complexity of subsequent processing. Deleting duplicate and meaningless words reduces data redundancy and improves processing efficiency.

[0053] S203: Randomly assign initial topic words to the feature words; for each feature word, temporarily remove the randomly assigned initial topic word, calculate the probability that the feature word belongs to each topic word, and re-determine the topic word corresponding to the feature word according to the probability; repeat the process of re-assigning the corresponding topic words for each feature word in the feature word group; until the Gibbs sampling converges or reaches the preset number of iterations, to obtain the topic word group corresponding to the feature word group.

[0054] Exemplarily, for each feature word in the feature word group, a subject word is randomly selected as its initial assignment. The number of topics K to be identified is predetermined, and then all feature words are ensured to be assigned to one of the K topics. For example: if there are 10 feature words and 5 preset topics, each feature word will be randomly assigned to one of the 5 topics. By iteratively updating the subject words corresponding to each feature word, the real subject words are gradually approached. In each iteration, for each feature word in the feature word group, the subject words currently assigned to it are temporarily removed. After removing the current subject words of the feature words, the probability that the feature word belongs to each subject word needs to be recalculated. The calculation of this probability is usually based on the conditional probability at the text level and the conditional probability at the subject level. The conditional probability at the text level is to calculate the conditional probability that the feature word belongs to a certain topic in other feature words in the current text; the conditional probability at the subject level is the conditional probability that the feature word belongs to a certain topic in the entire document collection. The product of these two probabilities is used to update the probability that the feature word belongs to each topic, and the subject words corresponding to the feature word are re-determined according to the probability. According to the above probability distribution, a subject word is randomly selected again as the new assignment of the feature word. For example, if a feature word "roast duck" was previously randomly assigned to topic 3, we now remove this association and decide which topic it should more likely belong to based on the new probability distribution, such as topic 1 or topic 4. Repeat the above process for all feature words in the long text data until the topic words of all feature words are reassigned.

[0055] Exemplarily, the above process is repeated for multiple iterations. After each iteration, the subject words corresponding to each feature word are updated. As the iteration proceeds, the subject words corresponding to the feature words will gradually stabilize, that is, the subject word allocation corresponding to the feature words will no longer change significantly, and the model is considered to have converged. If the model does not converge before reaching the preset number of iterations, the iteration will also stop, and the result of the last iteration will be used as the final subject word allocation.

[0056] By calculating the probability that a feature word belongs to each topic word and reallocating topic words based on these probabilities, the potential topic structure in the text can be captured more accurately. Compared with simple word frequency statistics or rule-based assignment methods, this method considers the global distribution of feature words in the document collection and the potential association between feature words and topics, thereby improving the accuracy of topic word assignment.

[0057] S204, calculating term distribution according to the feature word group and the subject word group, wherein the term distribution represents the probability of each feature word appearing under different subject words; clustering the subject word group according to the term distribution to determine the target subject word.

[0058] In one example, the importance of feature words in a feature phrase is calculated to obtain a weight vector of the feature words; wherein the weight vector represents the importance of the feature words in the feature phrase; based on the feature phrases and the subject phrases, a subject-feature co-occurrence frequency matrix is ​​generated; wherein the subject-feature co-occurrence frequency matrix represents the number of times the subject words co-occur with different feature words; and based on the weight vector and the subject-feature co-occurrence frequency matrix, the term distribution is determined.

[0059] In one example, based on the term distribution, the cosine similarity is used to calculate the similarity between subject words, and the two subject words with the highest similarity are merged; the new subject word after the merger is determined by the weight of the feature words of the original two subject words, until only one subject word remains; based on the silhouette coefficient, the optimal threshold of hierarchical clustering is determined; based on the optimal threshold, the clustering result is determined; wherein the optimal threshold represents the number of subject words in the subject word group obtained by hierarchical clustering; the clustering result represents that the subject words in the subject word group correspond to the feature words in multiple feature word groups; based on the clustering result, the topic distribution is calculated, and the topic distribution represents the probability distribution of the subject words in the subject word group in the long text data; based on the topic distribution, the frequency of occurrence of the subject words corresponding to the feature words in the feature word group is counted, and the subject word with the highest frequency of occurrence is the target subject word.

[0060] Exemplarily, term distribution is calculated based on feature word groups and subject word groups, wherein the term distribution represents the probability of each feature word appearing under different subject words; based on the term distribution, the subject word groups are clustered to determine the target subject words.

[0061] Exemplarily, firstly, the weights of the feature words are calculated using statistical methods or machine learning algorithms (such as TF-IDF, word embedding models, etc.). These weights reflect the frequency of occurrence of feature words in long text data and their importance. Then, based on the feature word group and the subject word group, the number of times each subject word co-occurs with different feature words is counted. This involves traversing the document collection and recording the frequency of occurrence of feature words under each subject word. Construct a subject-feature co-occurrence frequency matrix, in which the rows represent subject words, the columns represent feature words, and the elements in the matrix represent the number of co-occurrences of subject words and feature words. Finally, the weight vector and the subject-feature co-occurrence frequency matrix are combined to calculate the probability of occurrence of each feature word under each subject word using a probability formula or a normalization method. This usually involves dividing the number of co-occurrences by the total number of occurrences of all feature words under the subject word to obtain a probability value. A term distribution matrix is ​​obtained, in which the rows represent feature words, the columns represent subject words, and the elements in the matrix represent the probability of occurrence of feature words under the corresponding subject words.

[0062] By calculating the weight of feature words, we can distinguish the importance of different feature words in long text data, so as to more accurately reflect the topic structure of the text. Combining the topic-feature co-occurrence frequency matrix and the term distribution determined by the weight vector can more accurately quantify the relationship between feature words and subject words, and improve the accuracy of topic modeling.

[0063] Exemplarily, based on the term distribution, the cosine similarity formula is used to calculate the similarity between each subject term. Cosine similarity measures the degree of similarity between two vectors in direction and is suitable for measuring the similarity between subject terms. The two subject terms with the highest cosine similarity are merged, and the new subject term is determined by the weight of the feature terms of the original two subject terms (for example, it can be summed by probability weight). Update the term distribution matrix, remove the two subject term columns before the merge, add the new subject term column, and recalculate the probability of each feature term under the new subject term. Repeat the above steps until only one subject term is left or the preset stop condition (such as the number of merges, the minimum number of subject terms, etc.) is reached. In the process of merging subject terms, the clustering results after each merge and the corresponding clustering level (that is, the order of merging) are recorded. This is equivalent to executing a bottom-up hierarchical clustering process.

[0064] Exemplarily, the silhouette coefficient is an indicator to measure the effectiveness of clustering, and its value range is between -1 and 1. The larger the silhouette coefficient, the better the clustering effect. For each clustering level (i.e., the result after each merger), calculate its silhouette coefficient. Find the clustering level with the largest silhouette coefficient, and the number of subject words corresponding to this level is the optimal threshold. The optimal threshold represents the number of subject words in the subject word group obtained by hierarchical clustering, and is an important basis for the rationality of the clustering results. According to the optimal threshold, select the corresponding clustering level from the hierarchical clustering results. The subject word group at this level is the final clustering result. The clustering result represents the relationship between the subject words in the subject word group and the feature words in multiple feature word groups. Each subject word contains a set of feature words associated with it and their weights (probabilities).

[0065] Exemplarily, given the clustering results and long text data, the probability distribution of each subject word in the long text data is calculated. This usually involves traversing the long text data, counting the frequency of occurrence of feature words under each subject word, and normalizing it to probability. For each feature word in the feature phrase group, the frequency of occurrence of its corresponding subject word in the long text data is counted. This helps to understand the distribution of feature words under different topics. For each feature word in the feature phrase group, the subject word with the highest frequency of occurrence is selected as the target subject word. The target subject word represents the most important subject attribution of the feature word in the long text data.

[0066] By accurately calculating the similarity between subject words and merging similar subject words, the number of subject words can be significantly reduced while retaining the main information in the text data, making the results of topic modeling more concise and easy to understand. Secondly, using the silhouette coefficient to determine the optimal threshold can ensure the rationality and effectiveness of the clustering results. The selection of the optimal threshold not only makes the number of subject words in the clustering results moderate, but also can accurately reflect the topic structure of the text data.

[0067] Figure 3 A schematic diagram of the structure of an intelligent customer service reply processing device provided in an embodiment of the present application Figure 1 ,like Figure 3 As shown, the present embodiment provides an intelligent customer service reply processing device 30 including:

[0068] The acquisition module 301 is used to pre-process the long text data indicated by the input instruction in response to the user's input instruction to obtain a feature phrase; wherein the input instruction is used to indicate the long text data, and the feature phrase includes some feature words in the long text data;

[0069] Extraction module 302, used for subject extraction processing of feature words in feature word group to obtain subject word group; wherein the subject word group includes subject words corresponding to the feature words; clustering processing is performed on the subject words according to the feature word group and the subject word group to determine the target subject word; the target subject word represents the unique subject word corresponding to the long text data;

[0070] The reply module 303 is used to determine the reply text corresponding to the target subject word based on the reply knowledge base; wherein the reply knowledge base contains the reply text corresponding to the subject word; and feed back the obtained reply text to the user.

[0071] The device provided in this embodiment can execute the method provided in the above method embodiment. Its implementation principle and technical effect are similar, and this embodiment will not be described in detail here.

[0072] Figure 4 A schematic diagram of the structure of an intelligent customer service reply processing device provided in an embodiment of the present application Figure 2 ,like Figure 4 As shown, the present embodiment provides an intelligent customer service reply processing device 40 including:

[0073] The acquisition module 401 is used to pre-process the long text data indicated by the input instruction in response to the user's input instruction to obtain a feature phrase; wherein the input instruction is used to indicate the long text data, and the feature phrase includes some feature words in the long text data;

[0074] Extraction module 402, used for subject extraction processing of feature words in feature word groups to obtain subject word groups; wherein the subject word groups include subject words corresponding to the feature words; clustering processing is performed on the subject words according to the feature word groups and the subject word groups to determine target subject words; the target subject words represent unique subject words corresponding to the long text data;

[0075] The reply module 403 is used to determine the reply text corresponding to the target subject word based on the reply knowledge base; wherein the reply knowledge base contains the reply text corresponding to the subject word; and feed back the obtained reply text to the user.

[0076] In one example, the acquisition module 401 is specifically used for:

[0077] The long text data is segmented to obtain segmented word groups, wherein the segmented word groups include at least one word; repeated and meaningless words in the segmented word groups are deleted to obtain feature word groups.

[0078] In one example, the extraction module 402 is specifically configured to:

[0079] Randomly assign initial topic words to feature words;

[0080] For each feature word, temporarily remove the randomly assigned initial subject words, calculate the probability that the feature word belongs to each subject word, and re-determine the subject word corresponding to the feature word based on the probability; repeat the process of re-assigning the corresponding subject word to each feature word in the feature word group;

[0081] Until Gibbs sampling converges or reaches a preset number of iterations, the subject phrases corresponding to the feature phrases are obtained.

[0082] In one example, the extraction module 402 further includes:

[0083] A calculation module 4021 is used to calculate term distribution according to the feature word group and the subject word group, wherein the term distribution represents the probability of each feature word appearing under different subject words;

[0084] The determination module 4022 is used to cluster the subject word groups according to the word distribution and determine the target subject word.

[0085] In one example, the calculation module 4021 is specifically used for:

[0086] The importance of the feature words in the feature word group is calculated to obtain the weight vector of the feature word; wherein the weight vector represents the importance of the feature word in the feature word group;

[0087] According to the feature word group and the subject word group, a subject-feature co-occurrence frequency matrix is ​​generated; wherein the subject-feature co-occurrence frequency matrix represents the number of times the subject word and different feature words co-occur;

[0088] Based on the weight vector and the topic-feature co-occurrence frequency matrix, the term distribution is determined.

[0089] In one example, the determination module 4022 is specifically configured to:

[0090] According to the distribution of terms, the cosine similarity is used to calculate the similarity between the subject words, and the two subject words with the highest similarity are merged; the new subject word after the merger is determined by the weight of the feature words of the original two subject words, until only one subject word is left;

[0091] According to the silhouette coefficient, the optimal threshold of hierarchical clustering is determined; according to the optimal threshold, the clustering result is determined; wherein the optimal threshold represents the number of subject words in the subject word group obtained by hierarchical clustering; the clustering result represents that the subject words in the subject word group correspond to the feature words in multiple feature word groups;

[0092] According to the clustering results, the topic distribution is calculated. The topic distribution represents the probability distribution of the subject words in the topic phrase in the long text data. According to the topic distribution, the frequency of occurrence of the subject words corresponding to the feature words in the feature phrase is counted, and the subject word with the highest frequency is the target subject word.

[0093] In one example, before obtaining module 401, the following steps are also included:

[0094] A creation module 400 is used to obtain historical question and answer data and summarize the historical question and answer data to obtain a preliminary knowledge base; a question and answer knowledge base is created based on the knowledge graph and the preliminary knowledge base.

[0095] The device provided in this embodiment can execute the method provided in the above method embodiment. Its implementation principle and technical effect are similar, and this embodiment will not be described in detail here.

[0096] An embodiment of the present application also provides an intelligent customer service reply processing system, which is used to implement the above method.

[0097] Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. Figure 5 As shown, the electronic device 50 provided in this embodiment includes: at least one processor 501 and a memory 502. Optionally, the device 50 also includes a communication component 503. The processor 501, the memory 502 and the communication component 503 are connected via a bus 504.

[0098] In a specific implementation process, at least one processor 501 executes the computer-executable instructions stored in the memory 502, so that at least one processor 501 executes the above method.

[0099] The specific implementation process of the processor 501 can be found in the above method embodiment, and its implementation principle and technical effect are similar, so this embodiment will not be repeated here.

[0100] In the above embodiments, it should be understood that the processor can be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), etc. A general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in the invention can be directly implemented as a hardware processor, or can be implemented by a combination of hardware and software modules in the processor.

[0101] The memory may include a high-speed memory (Random Access Memory, RAM), and may also include a non-volatile memory (NVM), such as at least one disk storage.

[0102] The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, the bus in the drawings of this application is not limited to only one bus or one type of bus.

[0103] The present application also provides a computer program product, including a computer program, which implements the above method when executed by a processor.

[0104] The present application also provides a computer-readable storage medium, in which computer-executable instructions are stored. When a processor executes the computer-executable instructions, the above method is implemented.

[0105] The above-mentioned readable storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk. The readable storage medium can be any available medium that can be accessed by a general or special-purpose computer.

[0106] An exemplary readable storage medium is coupled to a processor so that the processor can read information from the readable storage medium and write information to the readable storage medium. Of course, the readable storage medium can also be a component of the processor. The processor and the readable storage medium can be located in an application specific integrated circuit (Application Specific Integrated Circuits, referred to as: ASIC). Of course, the processor and the readable storage medium can also exist in the device as discrete components.

[0107] The division of units is only a logical function division, and there may be other divisions in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interface, device or unit, which can be electrical, mechanical or other forms.

[0108] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0109] In addition, each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.

[0110] If the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions for a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the methods of each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, etc. Various media that can store program codes.

[0111] Those skilled in the art can understand that all or part of the steps of implementing the above-mentioned method embodiments can be completed by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When the program is executed, the steps of the above-mentioned method embodiments are executed; and the aforementioned storage medium includes: ROM, RAM, disk or optical disk and other media that can store program codes.

[0112] It should be noted that, for the aforementioned method embodiments, for the sake of simplicity, they are all expressed as a series of action combinations, but those skilled in the art should be aware that the present application is not limited by the described order of actions, because according to the present application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all optional embodiments, and the actions and modules involved are not necessarily required by the present application.

[0113] It should be further noted that, although the various steps in the flowchart are displayed in sequence according to the indication of the arrows, these steps are not necessarily executed in sequence in the order indicated by the arrows. Unless there is a clear description in this article, the execution of these steps is not strictly limited in order, and these steps can be executed in other orders. Moreover, at least a portion of the steps in the flowchart may include multiple sub-steps or multiple stages, and these sub-steps or stages are not necessarily executed at the same time, but can be executed at different times, and the execution order of these sub-steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a portion of the sub-steps or stages of other steps.

[0114] It should be understood that the above-mentioned device embodiments are only illustrative, and the device of the present application can also be implemented in other ways. For example, the division of units / modules in the above-mentioned embodiments is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units, modules or components can be combined, or can be integrated into another system, or some features can be ignored or not executed.

[0115] In addition, unless otherwise specified, each functional unit / module in each embodiment of the present application may be integrated into one unit / module, each unit / module may exist physically separately, or two or more units / modules may be integrated together. The above-mentioned integrated unit / module may be implemented in the form of hardware or in the form of a software program module.

[0116] If the integrated unit / module is implemented in the form of hardware, the hardware may be a digital circuit, an analog circuit, etc. The physical implementation of the hardware structure includes but is not limited to transistors, memristors, etc. Unless otherwise specified, the processor may be any appropriate hardware processor, such as a CPU, a GPU, an FPGA, a DSP, an ASIC, etc. Unless otherwise specified, the storage unit may be any appropriate magnetic storage medium or magneto-optical storage medium, such as a resistive random access memory (RRAM), a dynamic random access memory (DRAM), a static random access memory (SRAM), an enhanced dynamic random access memory (EDRAM), a high-bandwidth memory (HBM), a hybrid memory cube (HMC), etc.

[0117] If the integrated unit / module is implemented in the form of a software program module and sold or used as an independent product, it can be stored in a computer-readable memory. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a memory, including a number of instructions to enable a computer device (which can be a personal computer, server or network device, etc.) to execute all or part of the steps of the various embodiments of the present application. The aforementioned memory includes: U disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), mobile hard disk, disk or optical disk and other media that can store program codes.

[0118] In the above embodiments, the description of each embodiment has its own emphasis. For the part not described in detail in a certain embodiment, please refer to the relevant description of other embodiments. The technical features of the above embodiments can be combined arbitrarily. In order to make the description concise, all possible combinations of the technical features in the above embodiments are not described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0119] Those skilled in the art will readily appreciate other embodiments of the present application after considering the specification and practicing the invention disclosed herein. The present application is intended to cover any modification, use or adaptation of the present application, which follows the general principles of the present application and includes common knowledge or customary techniques in the art that are not disclosed in the present application. The specification and examples are intended to be exemplary only, and the true scope and spirit of the present application are indicated by the following claims.

[0120] It should be understood that the present application is not limited to the precise structures that have been described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof. The scope of the present application is limited only by the appended claims.

Claims

1. A method for processing responses of intelligent customer service, characterized in that: include: In response to a user's input instruction, preprocessing the long text data indicated by the input instruction to obtain a feature phrase group; wherein the input instruction is used to indicate the long text data, and the feature phrase group includes some feature words in the long text data; Performing a topic extraction process on the feature words in the feature word group to obtain a topic word group; wherein the topic word group includes topic words corresponding to the feature words; performing a clustering process on the topic words according to the feature word group and the topic word group to determine a target topic word; the target topic word represents a unique topic word corresponding to the long text data; Based on a reply knowledge base, a reply text corresponding to the target subject word is determined; wherein the reply knowledge base contains reply texts corresponding to the subject words; and the obtained reply text is fed back to the user.

2. The method according to claim 1, characterized in that Preprocessing the long text data indicated by the input instruction to obtain characteristic phrases includes: Performing word segmentation processing on the long text data to obtain word segmentation groups; wherein the word segmentation groups include at least one word; The repeated and meaningless words in the word groups are deleted to obtain the characteristic word groups.

3. The method according to claim 1, characterized in that Performing topic extraction processing on the feature words in the feature phrase group to obtain a topic phrase group includes: Randomly assigning initial subject words to the feature words; For each feature word, temporarily remove the randomly assigned initial subject words, calculate the probability that the feature word belongs to each subject word, and re-determine the subject word corresponding to the feature word based on the probability; repeat the process of re-assigning the corresponding subject word to each feature word in the feature word group; Until Gibbs sampling converges or reaches a preset number of iterations, the subject phrases corresponding to the feature phrases are obtained.

4. The method according to claim 1, characterized in that: According to the characteristic word group and the subject word group, clustering processing is performed on the subject words to determine target subject words, including: Calculating term distribution according to the feature word group and the subject word group, wherein the term distribution represents the probability of each feature word appearing under different subject words; According to the distribution of terms, the subject word group is clustered to determine the target subject word.

5. The method according to claim 4, characterized in that Determining term distribution according to the characteristic phrases and the subject phrases includes: Calculating the importance of the feature words in the feature word group to obtain a weight vector of the feature word; wherein the weight vector represents the importance of the feature word in the feature word group; Generate a subject-feature co-occurrence frequency matrix based on the feature word group and the subject word group; wherein the subject-feature co-occurrence frequency matrix represents the number of times the subject word and different feature words co-occur; The term distribution is determined based on the weight vector and the topic-feature co-occurrence frequency matrix.

6. The method according to claim 4, characterized in that According to the distribution of terms, clustering is performed on the subject word group to determine the target subject word, including: According to the term distribution, the similarity between the subject words is calculated using cosine similarity, and the two subject words with the highest similarity are merged; the merged new subject word is determined by the weight of the feature words of the original two subject words, until only one subject word is left; According to the silhouette coefficient, an optimal threshold of hierarchical clustering is determined; according to the optimal threshold, a clustering result is determined; wherein the optimal threshold represents the number of subject words in the subject word group obtained by hierarchical clustering; the clustering result represents that the subject words in the subject word group correspond to the feature words in multiple feature word groups; According to the clustering results, the topic distribution is calculated, and the topic distribution represents the probability distribution of the subject words in the subject phrase group in the long text data; according to the topic distribution, the frequency of occurrence of the subject words corresponding to the feature words in the feature phrase group is statistically analyzed, and the subject word with the highest occurrence frequency is the target subject word.

7. The method according to any one of claims 1 to 6, characterized in that Before responding to the user's input instruction, the method further includes: Acquire historical question-and-answer data, and summarize the historical question-and-answer data to obtain a preliminary knowledge base; The question-and-answer knowledge base is created based on the knowledge graph and the preliminary knowledge base.

8. An intelligent customer service reply processing device, characterized in that: include: An acquisition module, for responding to an input instruction of a user, preprocessing the long text data indicated by the input instruction to obtain a feature phrase; wherein the input instruction is used to indicate the long text data, and the feature phrase includes some feature words in the long text data; An extraction module is used to perform a topic extraction process on the feature words in the feature word group to obtain a topic word group; wherein the topic word group includes a topic word corresponding to the feature word; based on the feature word group and the topic word group, the topic word is clustered to determine a target topic word; the target topic word represents a unique topic word corresponding to the long text data; The reply module is used to determine the reply text corresponding to the target subject word based on the reply knowledge base; wherein the reply knowledge base contains the reply text corresponding to the subject word; and feed back the obtained reply text to the user.

9. An intelligent customer service reply processing system, characterized in that: The system is used to implement the method according to any one of claims 1 to 7.

10. An electronic device, characterized in that: include: A processor, and a memory communicatively connected to the processor; The memory stores computer-executable instructions; The processor executes the computer-executable instructions stored in the memory to implement the method according to any one of claims 1 to 7.

11. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-executable instructions, which are used to implement the method according to any one of claims 1 to 7 when executed by a processor.

12. A computer program product, characterized in that The invention comprises a computer program, which implements the method according to any one of claims 1 to 7 when being executed by a processor.