E-commerce intelligent customer service knowledge base automatic expansion method based on deep learning
Patent Information
- Application Number
- CN202610340590.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-03-19
- Publication Date
- 2026-09-22
- Estimated Expiration
- 2046-03-19
AI Technical Summary
例如,一些方案仅通过简单的相似度匹配从现有知识条目中选择答案,缺乏对用户咨询数据中潜在知识关系的挖掘能力;另一些方案虽然能够从交互数据中提取新问题,但缺少系统化的知识归并与规则挖掘机制,难以形成结构化的扩展知识条目
[0075]首先,本发明通过对电商客服交互数据进行语义建模,并构建双塔语义匹配网络实现问题表达与知识条目的统一语义表示,使系统能够识别语义相同但表达不同的用户问题。同时,通过对问题向量进行聚类划分形成问题意图簇,使系统能够自动发现用户咨询中的潜在问题模式,为知识发现提供基础。
Smart Images

Figure CN122220466B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of knowledge base expansion technology, and in particular to a method for automatically expanding the knowledge base of e-commerce intelligent customer service based on deep learning. Background Technology
[0002] With the rapid development of e-commerce platforms, online customer service systems have become a crucial channel for communication between these platforms and users. To improve customer service efficiency, an increasing number of e-commerce platforms are adopting intelligent customer service systems to automatically respond to user inquiries. In existing technologies, intelligent customer service systems typically rely on pre-built knowledge bases. By semantically matching user questions, they retrieve corresponding answers from the knowledge base and provide responses, thereby reducing the workload of human customer service representatives. However, as business scenarios and user needs continuously evolve, the knowledge entries in the knowledge base need constant updating and expansion; otherwise, inaccurate or unanswerable questions may arise.
[0003] Currently, the common methods for maintaining e-commerce intelligent customer service knowledge bases mainly rely on manual compilation and periodic updates. Technical personnel typically need to manually summarize and organize knowledge items based on customer service logs, user inquiry records, and business rules, and then add new Q&A content to the knowledge base. This method is not only labor-intensive but also has a long update cycle, making it difficult to reflect changes in user inquiry content in a timely manner. When the way users express their questions constantly changes, traditional systems based on keyword matching or simple semantic matching also struggle to accurately identify the question intent, thus affecting the response effectiveness of the intelligent customer service system.
[0004] In recent years, some technical solutions have begun to incorporate deep learning models to semantically represent user questions, aiming to improve question comprehension. However, shortcomings remain in knowledge base expansion. For example, some solutions simply select answers from existing knowledge entries through simple similarity matching, lacking the ability to mine potential knowledge relationships within user consultation data. Other solutions, while able to extract new questions from interaction data, lack systematic knowledge merging and rule mining mechanisms, making it difficult to form structured expanded knowledge entries. Furthermore, existing methods typically lack effective mechanisms for integrating newly generated knowledge entries and lack technical means for incremental model updates based on new knowledge, hindering the formation of a continuous optimization loop between the knowledge base and the semantic model.
[0005] Therefore, how to provide an automatic expansion method for the knowledge base of e-commerce intelligent customer service based on deep learning is a problem that urgently needs to be solved by those skilled in the art. Summary of the Invention
[0006] One objective of this invention is to propose an automatic expansion method for e-commerce intelligent customer service knowledge base based on deep learning. This invention fully utilizes a dual-tower semantic matching network, K-means clustering algorithm, k-nearest neighbor search algorithm, and FP-Growth association rule mining algorithm to perform semantic modeling, question intent recognition, knowledge item retrieval, and association rule mining on e-commerce customer service interaction data. It describes in detail the technical process of intelligently realizing automatic knowledge base discovery, knowledge merging and generation, and incremental updating of semantic models. It has the advantages of strong automatic knowledge expansion capability, high semantic matching accuracy, and high knowledge base update efficiency.
[0007] The method for automatically expanding the knowledge base of e-commerce intelligent customer service based on deep learning according to an embodiment of the present invention includes the following steps:
[0008] S1. Obtain and preprocess the interaction data of e-commerce customer service to generate a semantic dataset;
[0009] S2. A semantic coding model is constructed using a dual-tower semantic matching network as the semantic coding framework. Synonym merging and knowledge matching are performed on the question expressions in the semantic dataset. Comparative training samples are constructed and the semantic coding model is trained through comparative learning.
[0010] S3. Based on the trained semantic encoding model, the semantic vector mapping method is used to semantically encode the interactive question statements to generate question vectors, and the K-means clustering algorithm is used to cluster the question vectors to generate question intent clusters;
[0011] S4. Using the center vector of the question intent cluster as the retrieval vector, the k-nearest neighbor search algorithm is used to perform semantic retrieval in the knowledge base data, and the hierarchical clustering algorithm is used to cluster and divide the related knowledge items. Candidate knowledge units are generated through semantic merging.
[0012] S5. Construct a knowledge transaction set based on candidate knowledge units, use an association rule mining algorithm to perform frequent itemset mining on the knowledge transaction set, and filter high-confidence rules through support calculation and confidence calculation to generate extended knowledge items.
[0013] S6. Integrate the extended knowledge entries with the knowledge base data, construct incremental training samples, and incrementally update the semantic coding model using the stochastic gradient descent method.
[0014] Optionally, the preprocessing specifically includes:
[0015] The acquired interactive data is cleaned by removing HTML tags, emojis, special characters and repeated spaces, and the character encoding format is standardized to obtain standardized text data.
[0016] Standardized text data is processed by sentence segmentation and word segmentation, and a stop word filtering method is used to remove stop words that do not contribute semantics, generating an initial question expression sequence;
[0017] The initial problem expression sequence is subjected to spell correction and synonym standardization. Synonyms of different expression forms are replaced with unified expressions by dictionary matching to obtain a standardized problem expression sequence.
[0018] Part-of-speech tagging and keyword extraction are performed on the normalized problem expression sequence. The TF-IDF algorithm is used to calculate the term weights and select keywords with weights higher than a set threshold to generate a semantic feature sequence.
[0019] Question expression samples are constructed based on semantic feature sequences, and semantic alignment is performed on the knowledge entry texts in the knowledge base data to generate a semantic dataset.
[0020] Optionally, S2 specifically includes:
[0021] S21. Convert the question expressions and knowledge entry texts in the semantic dataset into word vector sequences, specifically including:
[0022] The problem expression and knowledge entry text are segmented into words, and the segmentation results are mapped into fixed-dimensional word vectors according to a preset vocabulary. Word vector sequences with insufficient length are padded.
[0023] The first and second encoding branches are constructed based on the dual-tower semantic matching network. The word vector sequence corresponding to the question expression is input into the first encoding branch and the word vector sequence corresponding to the knowledge item text is input into the second encoding branch.
[0024] A semantic encoding layer with shared parameters is set in the first and second encoding branches to perform context-related encoding on adjacent word vectors in the word vector sequence, thereby generating a context semantic feature sequence.
[0025] Average pooling is performed on the context semantic feature sequence to generate question expression vectors and knowledge entry vectors, respectively, and a semantic encoding model is constructed.
[0026] S22. Based on the question expression vector, calculate the cosine similarity between each question expression in the semantic dataset, and perform clustering and merging of semantically similar question expressions according to the set similarity threshold to form a set of synonymous question expressions;
[0027] S23. Perform semantic matching calculations on the question expression vectors and knowledge item vectors in the synonym question expression set, mark the sample pairs with similarity higher than the matching threshold as positive sample pairs, and use the random negative sampling method to select semantically mismatched question expressions and knowledge items from the semantic dataset to construct negative sample pairs, and generate a set of comparative training samples.
[0028] S24. Based on the comparison training sample set, calculate the similarity difference between positive sample pairs and negative sample pairs, construct the comparison learning loss function, and use the stochastic gradient descent algorithm to iteratively update the parameters of the dual-tower semantic matching network to complete the training of the semantic coding model.
[0029] Optionally, S24 specifically includes:
[0030] S241. Divide the comparison training sample set into batches according to the preset batch size, and input the positive sample pairs and negative sample pairs in each batch into the two encoding branches of the dual-tower semantic matching network respectively. Perform forward encoding on the question expression and knowledge item, extract the corresponding vector representation, and perform normalization processing on the vector representation.
[0031] S242. Based on the normalized vector representation, calculate the cosine similarity between the question expression vector and the knowledge item vector of the positive sample pair in each batch, and the cosine similarity between the question expression vector and the knowledge item vector of the negative sample pair, and construct the similarity matrix corresponding to the current batch.
[0032] S243. Sort the negative sample pairs corresponding to each positive sample pair based on the similarity matrix, and select the negative sample pairs whose similarity ranking is within a preset range as difficult negative samples.
[0033] S244. Construct a contrastive learning loss function based on the similarity of positive sample pairs, the similarity of difficult negative sample pairs, and a preset interval value, and sum the loss terms of all samples in the current batch to obtain the batch loss value.
[0034] S245. Based on the batch loss value, the backpropagation method is used to calculate the parameter gradient of the dual-tower semantic matching network, and the stochastic gradient descent algorithm is used to iteratively update the network parameters. After the current batch parameter update is completed, the next batch of samples is called to repeat the forward encoding, similarity calculation, hard negative sample screening, loss calculation and parameter update.
[0035] S246. After each training round, calculate the average loss value of all batches. When the change in the average loss value is lower than the preset convergence threshold or the training rounds reach the preset upper limit, stop updating the parameters and complete the training of the semantic coding model.
[0036] Optionally, S3 specifically includes:
[0037] S31. Input the interactive question statement into the trained semantic encoding model, perform semantic feature mapping on the interactive question statement through the semantic encoding model, and calculate the corresponding question vector set.
[0038] S32. Calculate the cosine similarity between any two question vectors based on the question vector set, construct a question vector similarity matrix, and analyze the semantic proximity relationship between question vectors based on the similarity matrix;
[0039] S33. Based on the preset number of clusters, the K-means clustering algorithm is used to perform initial clustering partitioning on the problem vector set, specifically including:
[0040] Randomly select one problem vector from the set of problem vectors as the first initial cluster center, and establish an initial cluster center set;
[0041] For the remaining question vectors in the question vector set excluding the selected cluster centers, calculate the squared Euclidean distance from each question vector to the nearest cluster center in the initial cluster center set, and construct a center selection probability distribution based on each squared Euclidean distance.
[0042] Based on the probability distribution of center selection, new problem vectors are selected from the remaining problem vectors and added to the initial cluster center set. The process of calculating the nearest cluster center distance, updating the squared Euclidean distance, and constructing the probability distribution is repeated until the number of cluster centers in the initial cluster center set reaches the preset number of clusters.
[0043] Calculate the Euclidean distance from each question vector in the question vector set to each cluster center in the initial cluster center set, and construct the sample center distance matrix;
[0044] Perform a minimum search on each row of the sample center distance matrix to determine the target cluster center for each question vector, and assign each question vector to the category corresponding to the target cluster center to generate the initial clustering result;
[0045] S34. After completing the current round of problem vector allocation, recalculate the mean vector for each category of problem vectors to update the cluster center. Repeat the problem vector allocation and cluster center update process until the change in cluster center is lower than the preset convergence threshold or the preset number of iterations is reached, and generate the problem intent cluster.
[0046] Optionally, S32 specifically includes:
[0047] S321. Assign sample indices to each question vector in the question vector set, and construct a two-dimensional question vector similarity matrix according to the order of the sample indices, where the row index and column index of the similarity matrix correspond to the sample indices in the question vector set, respectively.
[0048] S322. Read the question vectors corresponding to any two sample indices in sequence, calculate the dot product of the two question vectors, and calculate the magnitude of the two question vectors respectively.
[0049] S323. Divide the dot product by the product of the magnitudes of the two problem vectors to obtain the cosine similarity of the corresponding sample pair, and write the cosine similarity into the corresponding position in the problem vector similarity matrix.
[0050] S324. Utilizing the symmetry of cosine similarity, assign the same similarity value to the symmetrical positions in the question vector similarity matrix, and assign the preset similarity base value to the positions with the same sample index, thereby generating a complete question vector similarity matrix.
[0051] S325. Perform threshold comparison on each cosine similarity value in the question vector similarity matrix, mark sample pairs with similarity higher than the preset proximity threshold as semantically close sample pairs, and mark sample pairs with similarity lower than the preset difference threshold as semantically different sample pairs, thus forming a semantic proximity relationship between question vectors.
[0052] Optionally, S4 specifically includes:
[0053] S41. Extract the center vector of the question intent cluster as the retrieval vector, and perform semantic feature mapping on the knowledge entry text in the knowledge base data through a semantic encoding model to generate a set of knowledge vectors corresponding to the knowledge entry.
[0054] S42. Input the retrieval vector into the knowledge vector set and perform k-nearest neighbor search to obtain the k knowledge vectors that are semantically closest to the retrieval vector, and extract the corresponding knowledge entry text as a candidate set of associated knowledge entries.
[0055] S43. Perform word segmentation on the text of knowledge entries in the candidate associated knowledge entry set, and calculate the weight value of each term in the candidate associated knowledge entry set using the TF-IDF algorithm. Construct the semantic feature representation of the knowledge entry based on the term weight.
[0056] S44. Based on the semantic feature representation of knowledge items, a hierarchical clustering algorithm is executed to gradually merge knowledge items with similar semantic features through a bottom-up aggregation method, generating multiple knowledge item clusters.
[0057] S45. Perform semantic merging processing on the knowledge entry texts in each knowledge entry cluster, and unify and organize knowledge entries with similar semantic expressions and consistent semantic directions to generate candidate knowledge units.
[0058] Optionally, S43 specifically includes:
[0059] S431. Perform word segmentation on the text of knowledge entries in the candidate related knowledge entry set, and remove invalid terms by combining the stop word list and low information word filtering rules to obtain the effective term sequence of candidate knowledge entries;
[0060] S432. Based on the distribution of each effective term sequence in different knowledge entry texts, construct a domain term set corresponding to the candidate associated knowledge entry set, and establish feature indexes for the terms in the domain term set;
[0061] S433. Using the TF-IDF weight allocation mechanism, the importance of each term in the domain term set is evaluated, and high-weight terms that can represent the topic content of the knowledge item are extracted to form a keyword set for the knowledge item.
[0062] S434. Perform feature filtering processing based on the keyword set of knowledge items, remove words with high cross-item repetition and low distinguishability, and retain feature words that can distinguish the semantic content of different knowledge items.
[0063] S435. Encode the retained feature terms according to the feature index to form a sparse feature representation for each knowledge entry, and use each sparse feature representation as the semantic feature representation of the knowledge entry.
[0064] Optionally, S5 specifically includes:
[0065] S51. Construct candidate knowledge units in a transactional manner, extract the term set in each candidate knowledge unit, and organize the corresponding term set into independent transaction records according to the semantic boundary of the candidate knowledge unit to generate a knowledge transaction set.
[0066] S52. The FP-Growth algorithm is used to mine frequent itemsets in the knowledge transaction set. A frequent pattern tree is constructed based on the occurrence frequency of each term in the knowledge transaction set, and frequent itemsets that meet the preset support threshold are recursively extracted along the conditional path of the frequent pattern tree.
[0067] S53. Generate association rules based on frequent itemsets, divide each frequent itemset into rule antecedent and rule consequent, and calculate the support and confidence of each association rule in combination with the co-occurrence relationship of itemsets in the knowledge transaction set, and screen high confidence rules that meet the preset confidence threshold.
[0068] S54. Perform rule mapping on high-confidence rules, determine the subject terms and constraint terms corresponding to the rule antecedents as knowledge triggering conditions, determine the result terms corresponding to the rule consequents as knowledge response content, and generate extended knowledge entries according to the preset knowledge organization format.
[0069] Optionally, S6 specifically includes:
[0070] S61. Perform knowledge alignment processing on the extended knowledge entries and knowledge base data, identify knowledge content in the extended knowledge entries that is semantically consistent or semantically similar to the knowledge base data, and merge and organize the corresponding knowledge content according to the semantic structure of the knowledge entries to generate a fused knowledge entry set.
[0071] S62. Based on the fused knowledge item set, extract the correspondence between the question expression and the knowledge item, and construct an incremental training sample set according to the pairing results of the question expression and the corresponding knowledge item;
[0072] S63. Perform semantic encoding processing on the question expressions in the incremental training sample set, and establish a semantic matching relationship based on the encoding results and the semantic representation of the corresponding knowledge items to form semantic matching samples;
[0073] S64. Based on semantic matching samples, the parameters of the semantic coding model are iteratively updated using the stochastic gradient descent method, enabling the semantic coding model to learn the semantic features of the fused knowledge items and complete the incremental update of the semantic coding model.
[0074] The beneficial effects of this invention are:
[0075] First, this invention performs semantic modeling on e-commerce customer service interaction data and constructs a dual-tower semantic matching network to achieve a unified semantic representation of question expressions and knowledge items, enabling the system to identify user questions that have the same semantics but different expressions. Simultaneously, by clustering question vectors to form question intent clusters, the system can automatically discover potential question patterns in user inquiries, providing a foundation for knowledge discovery.
[0076] Secondly, this invention performs semantic retrieval in the knowledge base data using the k-nearest neighbor search algorithm, and combines it with a hierarchical clustering algorithm to cluster and merge related knowledge items to generate candidate knowledge units. On this basis, the FP-Growth algorithm is used to mine frequent itemsets of the knowledge transaction set and generate association rules.
[0077] Finally, this invention integrates extended knowledge entries with knowledge base data, constructs incremental training samples based on the integrated knowledge entries, and incrementally updates the semantic encoding model, enabling the model to continuously learn new knowledge information and achieve coordinated updates of knowledge base expansion and semantic model optimization. Attached Figure Description
[0078] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings:
[0079] Figure 1 This is a flowchart of the method for automatically expanding the knowledge base of intelligent e-commerce customer service based on deep learning proposed in this invention.
[0080] Figure 2This is a flowchart of the training of the dual-tower semantic matching network and the construction of the semantic encoding model for the automatic expansion method of the e-commerce intelligent customer service knowledge base based on deep learning proposed in this invention.
[0081] Figure 3 This is a flowchart illustrating the incremental update process of the knowledge fusion and semantic encoding model for the automatic expansion method of the e-commerce intelligent customer service knowledge base based on deep learning proposed in this invention. Detailed Implementation
[0082] The present invention will now be described in further detail with reference to the accompanying drawings. These drawings are simplified schematic diagrams, illustrating only the basic structure of the invention, and therefore only show the components relevant to the invention.
[0083] refer to Figures 1-3 The method for automatically expanding the knowledge base of e-commerce intelligent customer service based on deep learning includes the following steps:
[0084] S1. Obtain and preprocess the interaction data of e-commerce customer service to generate a semantic dataset;
[0085] S2. A semantic coding model is constructed using a dual-tower semantic matching network as the semantic coding framework. Synonym merging and knowledge matching are performed on the question expressions in the semantic dataset. Comparative training samples are constructed and the semantic coding model is trained through comparative learning.
[0086] S3. Based on the trained semantic encoding model, the semantic vector mapping method is used to semantically encode the interactive question statements to generate question vectors, and the K-means clustering algorithm is used to cluster the question vectors to generate question intent clusters;
[0087] S4. Using the center vector of the question intent cluster as the retrieval vector, the k-nearest neighbor search algorithm is used to perform semantic retrieval in the knowledge base data, and the hierarchical clustering algorithm is used to cluster and divide the related knowledge items. Candidate knowledge units are generated through semantic merging.
[0088] S5. Construct a knowledge transaction set based on candidate knowledge units, use an association rule mining algorithm to perform frequent itemset mining on the knowledge transaction set, and filter high-confidence rules through support calculation and confidence calculation to generate extended knowledge items.
[0089] S6. Integrate the extended knowledge entries with the knowledge base data, construct incremental training samples, and incrementally update the semantic coding model using the stochastic gradient descent method.
[0090] In this embodiment, the preprocessing specifically includes:
[0091] The acquired interactive data is cleaned by removing HTML tags, emojis, special characters and repeated spaces, and the character encoding format is standardized to obtain standardized text data.
[0092] Standardized text data is processed by sentence segmentation and word segmentation, and a stop word filtering method is used to remove stop words that do not contribute semantics, generating an initial question expression sequence;
[0093] The initial problem expression sequence is subjected to spell correction and synonym standardization. Synonyms of different expression forms are replaced with unified expressions by dictionary matching to obtain a standardized problem expression sequence.
[0094] Part-of-speech tagging and keyword extraction are performed on the normalized problem expression sequence. The TF-IDF algorithm is used to calculate the term weights and select keywords with weights higher than a set threshold to generate a semantic feature sequence.
[0095] Question expression samples are constructed based on semantic feature sequences, and semantic alignment is performed on the knowledge entry texts in the knowledge base data to generate a semantic dataset.
[0096] In this embodiment, S2 specifically includes:
[0097] S21. Convert the question expressions and knowledge entry texts in the semantic dataset into word vector sequences, specifically including:
[0098] The problem expression and knowledge entry text are segmented into words, and the segmentation results are mapped into fixed-dimensional word vectors according to a preset vocabulary. Word vector sequences with insufficient length are padded.
[0099] The first and second encoding branches are constructed based on the dual-tower semantic matching network. The word vector sequence corresponding to the question expression is input into the first encoding branch and the word vector sequence corresponding to the knowledge item text is input into the second encoding branch.
[0100] A semantic encoding layer with shared parameters is set in the first and second encoding branches to perform context-related encoding on adjacent word vectors in the word vector sequence, thereby generating a context semantic feature sequence.
[0101] Average pooling is performed on the context semantic feature sequence to generate question expression vectors and knowledge entry vectors, respectively, and a semantic encoding model is constructed.
[0102] S22. Based on the question expression vector, calculate the cosine similarity between each question expression in the semantic dataset, and perform clustering and merging of semantically similar question expressions according to the set similarity threshold to form a set of synonymous question expressions;
[0103] S23. Perform semantic matching calculations on the question expression vectors and knowledge item vectors in the synonym question expression set, mark the sample pairs with similarity higher than the matching threshold as positive sample pairs, and use the random negative sampling method to select semantically mismatched question expressions and knowledge items from the semantic dataset to construct negative sample pairs, and generate a set of comparative training samples.
[0104] S24. Based on the comparison training sample set, calculate the similarity difference between positive sample pairs and negative sample pairs, construct the comparison learning loss function, and use the stochastic gradient descent algorithm to iteratively update the parameters of the dual-tower semantic matching network to complete the training of the semantic coding model.
[0105] In this embodiment, S24 specifically includes:
[0106] S241. Divide the comparison training sample set into batches according to the preset batch size, and input the positive sample pairs and negative sample pairs in each batch into the two encoding branches of the dual-tower semantic matching network respectively. Perform forward encoding on the question expression and knowledge item, extract the corresponding vector representation, and perform normalization processing on the vector representation.
[0107] S242. Based on the normalized vector representation, calculate the cosine similarity between the question expression vector and the knowledge item vector of the positive sample pair in each batch, and the cosine similarity between the question expression vector and the knowledge item vector of the negative sample pair, and construct the similarity matrix corresponding to the current batch.
[0108] S243. Sort the negative sample pairs corresponding to each positive sample pair based on the similarity matrix, and select the negative sample pairs whose similarity ranking is within a preset range as difficult negative samples.
[0109] S244. Construct a contrastive learning loss function based on the similarity of positive sample pairs, the similarity of difficult negative sample pairs, and a preset interval value, and sum the loss terms of all samples in the current batch to obtain the batch loss value.
[0110] S245. Based on the batch loss value, the backpropagation method is used to calculate the parameter gradient of the dual-tower semantic matching network, and the stochastic gradient descent algorithm is used to iteratively update the network parameters. After the current batch parameter update is completed, the next batch of samples is called to repeat the forward encoding, similarity calculation, hard negative sample screening, loss calculation and parameter update.
[0111] S246. After each training round, calculate the average loss value of all batches. When the change in the average loss value is lower than the preset convergence threshold or the training rounds reach the preset upper limit, stop updating the parameters and complete the training of the semantic coding model.
[0112] In this embodiment, S3 specifically includes:
[0113] S31. Input the interactive question statement into the trained semantic encoding model, perform semantic feature mapping on the interactive question statement through the semantic encoding model, and calculate the corresponding question vector set.
[0114] S32. Calculate the cosine similarity between any two question vectors based on the question vector set, construct a question vector similarity matrix, and analyze the semantic proximity relationship between question vectors based on the similarity matrix;
[0115] S33. Based on the preset number of clusters, the K-means clustering algorithm is used to perform initial clustering partitioning on the problem vector set, specifically including:
[0116] Randomly select one problem vector from the set of problem vectors as the first initial cluster center, and establish an initial cluster center set;
[0117] For the remaining question vectors in the question vector set excluding the selected cluster centers, calculate the squared Euclidean distance from each question vector to the nearest cluster center in the initial cluster center set, and construct a center selection probability distribution based on each squared Euclidean distance.
[0118] Based on the probability distribution of center selection, new problem vectors are selected from the remaining problem vectors and added to the initial cluster center set. The process of calculating the nearest cluster center distance, updating the squared Euclidean distance, and constructing the probability distribution is repeated until the number of cluster centers in the initial cluster center set reaches the preset number of clusters.
[0119] Calculate the Euclidean distance from each question vector in the question vector set to each cluster center in the initial cluster center set, and construct the sample center distance matrix;
[0120] Perform a minimum search on each row of the sample center distance matrix to determine the target cluster center for each question vector, and assign each question vector to the category corresponding to the target cluster center to generate the initial clustering result;
[0121] S34. After completing the current round of problem vector allocation, recalculate the mean vector for each category of problem vectors to update the cluster center. Repeat the problem vector allocation and cluster center update process until the change in cluster center is lower than the preset convergence threshold or the preset number of iterations is reached, and generate the problem intent cluster.
[0122] In this embodiment, S32 specifically includes:
[0123] S321. Assign sample indices to each question vector in the question vector set, and construct a two-dimensional question vector similarity matrix according to the order of the sample indices, where the row index and column index of the similarity matrix correspond to the sample indices in the question vector set, respectively.
[0124] S322. Read the question vectors corresponding to any two sample indices in sequence, calculate the dot product of the two question vectors, and calculate the magnitude of the two question vectors respectively.
[0125] S323. Divide the dot product by the product of the magnitudes of the two problem vectors to obtain the cosine similarity of the corresponding sample pair, and write the cosine similarity into the corresponding position in the problem vector similarity matrix.
[0126] S324. Utilizing the symmetry of cosine similarity, assign the same similarity value to the symmetrical positions in the question vector similarity matrix, and assign the preset similarity base value to the positions with the same sample index, thereby generating a complete question vector similarity matrix.
[0127] S325. Perform threshold comparison on each cosine similarity value in the question vector similarity matrix, mark sample pairs with similarity higher than the preset proximity threshold as semantically close sample pairs, and mark sample pairs with similarity lower than the preset difference threshold as semantically different sample pairs, thus forming a semantic proximity relationship between question vectors.
[0128] In this embodiment, S4 specifically includes:
[0129] S41. Extract the center vector of the question intent cluster as the retrieval vector, and perform semantic feature mapping on the knowledge entry text in the knowledge base data through a semantic encoding model to generate a set of knowledge vectors corresponding to the knowledge entry.
[0130] S42. Input the retrieval vector into the knowledge vector set and perform a k-nearest neighbor search to obtain the k knowledge vectors that are semantically closest to the retrieval vector, and extract the corresponding knowledge entry text as a candidate set of associated knowledge entries, specifically including:
[0131] A vector index structure is constructed for the knowledge vector set according to the vector feature dimension, and each knowledge vector is written into the vector index structure to form a vector retrieval space;
[0132] The search vector is input into the vector search space to perform a nearest neighbor search operation, and the set of candidate vectors that are adjacent to the search vector is located in the vector index structure.
[0133] In the candidate vector set, nearest neighbor filtering is performed according to the vector similarity, and the k knowledge vectors with the highest semantic similarity to the retrieved vector are determined as the target nearest neighbor vectors;
[0134] Extract the corresponding knowledge entry text based on the index position of the target nearest neighbor vector in the knowledge vector set, and collect and organize the extraction results to generate a candidate related knowledge entry set;
[0135] S43. Perform word segmentation on the text of knowledge entries in the candidate associated knowledge entry set, and calculate the weight value of each term in the candidate associated knowledge entry set using the TF-IDF algorithm. Construct the semantic feature representation of the knowledge entry based on the term weight.
[0136] S44. Based on the semantic feature representation of knowledge items, a hierarchical clustering algorithm is executed to gradually merge knowledge items with similar semantic features through a bottom-up aggregation method, generating multiple knowledge item clusters.
[0137] S45. Perform semantic merging processing on the knowledge entry texts in each knowledge entry cluster, and unify and organize knowledge entries with similar semantic expressions and consistent semantic directions to generate candidate knowledge units.
[0138] In this embodiment, S43 specifically includes:
[0139] S431. Perform word segmentation on the text of knowledge entries in the candidate related knowledge entry set, and remove invalid terms by combining the stop word list and low information word filtering rules to obtain the effective term sequence of candidate knowledge entries;
[0140] S432. Based on the distribution of each effective term sequence in different knowledge entry texts, construct a domain term set corresponding to the candidate associated knowledge entry set, and establish feature indexes for the terms in the domain term set;
[0141] S433. Using the TF-IDF weight allocation mechanism, the importance of each term in the domain term set is evaluated, and high-weight terms that can represent the topic content of the knowledge item are extracted to form a keyword set for the knowledge item.
[0142] S434. Perform feature filtering processing based on the keyword set of knowledge items, remove words with high cross-item repetition and low distinguishability, and retain feature words that can distinguish the semantic content of different knowledge items.
[0143] S435. Encode the retained feature terms according to the feature index to form a sparse feature representation for each knowledge entry, and use each sparse feature representation as the semantic feature representation of the knowledge entry.
[0144] In this embodiment, S5 specifically includes:
[0145] S51. Construct candidate knowledge units in a transactional manner, extract the term set in each candidate knowledge unit, and organize the corresponding term set into independent transaction records according to the semantic boundary of the candidate knowledge unit to generate a knowledge transaction set.
[0146] S52. The FP-Growth algorithm is used to mine frequent itemsets in the knowledge transaction set. A frequent pattern tree is constructed based on the frequency of occurrence of each term in the knowledge transaction set, and frequent itemsets that meet the preset support threshold are recursively extracted along the conditional path of the frequent pattern tree. Specifically, this includes:
[0147] The transaction records in the knowledge transaction set are traversed and statistically analyzed to identify terms whose frequency in the knowledge transaction set reaches a preset support threshold, and the terms that meet the support condition are determined as the frequent term set.
[0148] The frequent terms set is sorted according to the frequency of occurrence of terms in the knowledge transaction set, and the term sequence in each transaction record is rearranged according to the sorting result to form a transaction sequence for constructing a frequent pattern tree.
[0149] A frequent pattern tree is constructed based on the transaction sequence. Multiple transaction records are compressed and represented in the same tree structure by sharing a common prefix path, forming a tree structure that reflects the co-occurrence relationship of terms.
[0150] Starting from each leaf node of the frequent pattern tree, backtrack along the direction of the parent node to extract the conditional path, construct the corresponding conditional pattern base, and recursively execute the frequent item mining process based on the conditional pattern base to generate a set of frequent itemsets that meet the preset support threshold.
[0151] S53. Generate association rules based on frequent itemsets, divide each frequent itemset into rule antecedent and rule consequent, and calculate the support and confidence of each association rule in combination with the co-occurrence relationship of itemsets in the knowledge transaction set, and screen high confidence rules that meet the preset confidence threshold.
[0152] S54. Perform rule mapping on high-confidence rules, determine the subject terms and constraint terms corresponding to the rule antecedents as knowledge triggering conditions, determine the result terms corresponding to the rule consequents as knowledge response content, and generate extended knowledge entries according to the preset knowledge organization format.
[0153] In this embodiment, S6 specifically includes:
[0154] S61. Perform knowledge alignment processing on the extended knowledge entries and knowledge base data, identify knowledge content in the extended knowledge entries that is semantically consistent or semantically similar to the knowledge base data, and merge and organize the corresponding knowledge content according to the semantic structure of the knowledge entries to generate a fused knowledge entry set.
[0155] S62. Based on the fused knowledge item set, extract the correspondence between the question expression and the knowledge item, and construct an incremental training sample set according to the pairing results of the question expression and the corresponding knowledge item;
[0156] S63. Perform semantic encoding processing on the question expressions in the incremental training sample set, and establish a semantic matching relationship based on the encoding results and the semantic representation of the corresponding knowledge items to form semantic matching samples;
[0157] S64. Based on semantic matching samples, the parameters of the semantic coding model are iteratively updated using the stochastic gradient descent method, enabling the semantic coding model to learn the semantic features of the fused knowledge items and complete the incremental update of the semantic coding model.
[0158] Example 1: To verify the feasibility of this invention in practice, it was applied to an intelligent customer service system of an e-commerce platform. This system has accumulated a large number of user inquiry records during its long-term operation. User questions mainly involve product inquiries, order status, logistics progress, and after-sales processing. Due to the diverse ways users express themselves, such as "when will it be shipped?", "how long will it take to send it out?", and "how long will it take for the order to be shipped?", different expressions all point to the same meaning. Traditional customer service knowledge bases based on keywords or simple semantic matching struggle to accurately identify the intent of the questions, resulting in some questions failing to find suitable answers and requiring human intervention. Furthermore, updates to the existing knowledge base mainly rely on manual compilation, which is inefficient and makes it difficult to promptly add new knowledge content.
[0159] In this scenario, customer service interaction data is first preprocessed to construct a semantic dataset. A semantic encoding model is then established using a dual-tower semantic matching network to provide a unified semantic representation of user question expressions. Subsequently, cluster analysis is performed on the question vectors to form multiple question intent clusters. For example, in actual data processing, approximately 120,000 user inquiries were divided into more than 200 question intent clusters, among which the question intent cluster related to shipping contained approximately 8,000 semantically similar question expressions.
[0160] Next, the center vector of the question intent cluster is extracted, semantic retrieval is performed in the knowledge base to obtain candidate related knowledge items, and these are merged into candidate knowledge units through hierarchical clustering. Based on this, the FP-Growth algorithm is used to mine frequent itemsets and association rules, generating more than 270 new extended knowledge items from approximately 4200 candidate knowledge units, which are then integrated with the original knowledge base. At the same time, incremental training samples are constructed to update the semantic encoding model.
[0161] After a period of operation, the knowledge base expanded from approximately 1,850 knowledge entries to approximately 2,120. The newly added knowledge entries mainly focused on high-frequency consultation scenarios such as order processing, shipping rules, and after-sales procedures. Testing with 5,000 user inquiries revealed that the system's automatic matching success rate increased from 78.4% to approximately 91.2%, while the proportion of human customer service intervention decreased from approximately 22% to approximately 9%. This indicates that the method of this invention can effectively improve the automatic expansion capability of the knowledge base and the response efficiency of the intelligent customer service system. Table 1 shows the comparison results of the effects before and after applying the method of this invention.
[0162] Table 1. Comparison of the Automatic Expansion Effects of the Intelligent Customer Service Knowledge Base
[0163] Number of knowledge items 1850 2120 items Automatic matching success rate 78.40% 91.20% Average response time 1.6 seconds 1.2 seconds Human customer service intervention rate 22% 9% Number of new knowledge entries generated 0 270 items
[0164] As can be seen from the data in Table 1, after adopting the method of the present invention, the number of knowledge items in the intelligent customer service system increased significantly, the automatic matching success rate increased from 78.4% to 91.2%, and the proportion of human customer service intervention decreased from 22% to 9%. This indicates that the present invention can effectively improve the automatic expansion capability of the knowledge base and the system's efficiency in processing user inquiries.
[0165] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.
Claims
1. A method for automatically expanding the knowledge base of e-commerce intelligent customer service based on deep learning, characterized in that, Includes the following steps: S1. Obtain and preprocess the interaction data of e-commerce customer service to generate a semantic dataset; S2. A semantic coding model is constructed using a dual-tower semantic matching network as the semantic coding framework. Synonym merging and knowledge matching are performed on the question expressions in the semantic dataset. Comparative training samples are constructed and the semantic coding model is trained through comparative learning. S3. Based on the trained semantic encoding model, the semantic vector mapping method is used to semantically encode the interactive question statements to generate question vectors, and the K-means clustering algorithm is used to cluster the question vectors to generate question intent clusters; S4. Using the center vector of the question intent cluster as the retrieval vector, the k-nearest neighbor search algorithm is used to perform semantic retrieval in the knowledge base data, and the hierarchical clustering algorithm is used to cluster and divide the related knowledge items. Candidate knowledge units are generated through semantic merging. S5. Construct a knowledge transaction set based on candidate knowledge units, use an association rule mining algorithm to perform frequent itemset mining on the knowledge transaction set, and filter high-confidence rules through support calculation and confidence calculation to generate extended knowledge items. S6. Integrate the extended knowledge entries with the knowledge base data, construct incremental training samples, and incrementally update the semantic coding model using the stochastic gradient descent method. S2 specifically includes: S21. Convert the question expressions and knowledge entry texts in the semantic dataset into word vector sequences, specifically including: The problem expression and knowledge entry text are segmented into words, and the segmentation results are mapped into fixed-dimensional word vectors according to a preset vocabulary. Word vector sequences with insufficient length are padded. The first and second encoding branches are constructed based on the dual-tower semantic matching network. The word vector sequence corresponding to the question expression is input into the first encoding branch and the word vector sequence corresponding to the knowledge item text is input into the second encoding branch. A semantic encoding layer with shared parameters is set in the first and second encoding branches to perform context-related encoding on adjacent word vectors in the word vector sequence, thereby generating a context semantic feature sequence. Average pooling is performed on the context semantic feature sequence to generate question expression vectors and knowledge entry vectors, respectively, and a semantic encoding model is constructed. S22. Based on the question expression vector, calculate the cosine similarity between each question expression in the semantic dataset, and perform clustering and merging of semantically similar question expressions according to the set similarity threshold to form a set of synonymous question expressions; S23. Perform semantic matching calculations on the question expression vectors and knowledge item vectors in the synonym question expression set, mark the sample pairs with similarity higher than the matching threshold as positive sample pairs, and use the random negative sampling method to select semantically mismatched question expressions and knowledge items from the semantic dataset to construct negative sample pairs, and generate a set of comparative training samples. S24. Based on the comparison training sample set, calculate the similarity difference between positive sample pairs and negative sample pairs, construct the comparison learning loss function, and use the stochastic gradient descent algorithm to iteratively update the parameters of the dual-tower semantic matching network to complete the training of the semantic coding model. S24 specifically includes: S241. Divide the comparison training sample set into batches according to the preset batch size, and input the positive sample pairs and negative sample pairs in each batch into the two encoding branches of the dual-tower semantic matching network respectively. Perform forward encoding on the question expression and knowledge item, extract the corresponding vector representation, and perform normalization processing on the vector representation. S242. Based on the normalized vector representation, calculate the cosine similarity between the question expression vector and the knowledge item vector of the positive sample pair in each batch, and the cosine similarity between the question expression vector and the knowledge item vector of the negative sample pair, and construct the similarity matrix corresponding to the current batch. S243. Sort the negative sample pairs corresponding to each positive sample pair based on the similarity matrix, and select the negative sample pairs whose similarity ranking is within a preset range as difficult negative samples. S244. Construct a contrastive learning loss function based on the similarity of positive sample pairs, the similarity of difficult negative sample pairs, and a preset interval value, and sum the loss terms of all samples in the current batch to obtain the batch loss value. S245. Based on the batch loss value, the backpropagation method is used to calculate the parameter gradient of the dual-tower semantic matching network, and the stochastic gradient descent algorithm is used to iteratively update the network parameters. After the current batch parameter update is completed, the next batch of samples is called to repeat the forward encoding, similarity calculation, hard negative sample screening, loss calculation and parameter update. S246. After each training round, calculate the average loss value of all batches. When the change in the average loss value is lower than the preset convergence threshold or the training rounds reach the preset upper limit, stop updating the parameters and complete the training of the semantic coding model.
2. The method for automatically expanding the knowledge base of e-commerce intelligent customer service based on deep learning according to claim 1, characterized in that, The preprocessing specifically includes: The acquired interactive data is cleaned by removing HTML tags, emojis, special characters and repeated spaces, and the character encoding format is standardized to obtain standardized text data. Standardized text data is processed by sentence segmentation and word segmentation, and a stop word filtering method is used to remove stop words that do not contribute semantics, generating an initial question expression sequence; The initial problem expression sequence is subjected to spell correction and synonym standardization. Synonyms of different expression forms are replaced with unified expressions by dictionary matching to obtain a standardized problem expression sequence. Part-of-speech tagging and keyword extraction are performed on the normalized problem expression sequence. The TF-IDF algorithm is used to calculate the term weights and select keywords with weights higher than a set threshold to generate a semantic feature sequence. Question expression samples are constructed based on semantic feature sequences, and semantic alignment is performed on the knowledge entry texts in the knowledge base data to generate a semantic dataset.
3. The method for automatically expanding the knowledge base of e-commerce intelligent customer service based on deep learning according to claim 1, characterized in that, S3 specifically includes: S31. Input the interactive question statement into the trained semantic encoding model, perform semantic feature mapping on the interactive question statement through the semantic encoding model, and calculate the corresponding question vector set. S32. Calculate the cosine similarity between any two question vectors based on the question vector set, construct a question vector similarity matrix, and analyze the semantic proximity relationship between question vectors based on the similarity matrix; S33. Based on the preset number of clusters, the K-means clustering algorithm is used to perform initial clustering partitioning on the problem vector set, specifically including: Randomly select one problem vector from the set of problem vectors as the first initial cluster center, and establish an initial cluster center set; For the remaining question vectors in the question vector set excluding the selected cluster centers, calculate the squared Euclidean distance from each question vector to the nearest cluster center in the initial cluster center set, and construct a center selection probability distribution based on each squared Euclidean distance. Based on the probability distribution of center selection, new problem vectors are selected from the remaining problem vectors and added to the initial cluster center set. The process of calculating the nearest cluster center distance, updating the squared Euclidean distance, and constructing the probability distribution is repeated until the number of cluster centers in the initial cluster center set reaches the preset number of clusters. Calculate the Euclidean distance from each question vector in the question vector set to each cluster center in the initial cluster center set, and construct the sample center distance matrix; Perform a minimum search on each row of the sample center distance matrix to determine the target cluster center for each question vector, and assign each question vector to the category corresponding to the target cluster center to generate the initial clustering result; S34. After completing the current round of problem vector allocation, recalculate the mean vector for each category of problem vectors to update the cluster center. Repeat the problem vector allocation and cluster center update process until the change in cluster center is lower than the preset convergence threshold or the preset number of iterations is reached, and generate the problem intent cluster.
4. The method for automatically expanding the knowledge base of e-commerce intelligent customer service based on deep learning according to claim 3, characterized in that, Specifically, S32 includes: S321. Assign sample indices to each question vector in the question vector set, and construct a two-dimensional question vector similarity matrix according to the order of the sample indices, where the row index and column index of the similarity matrix correspond to the sample indices in the question vector set, respectively. S322. Read the question vectors corresponding to any two sample indices in sequence, calculate the dot product of the two question vectors, and calculate the magnitude of the two question vectors respectively. S323. Divide the dot product by the product of the magnitudes of the two problem vectors to obtain the cosine similarity of the corresponding sample pair, and write the cosine similarity into the corresponding position in the problem vector similarity matrix. S324. Utilizing the symmetry of cosine similarity, assign the same similarity value to the symmetrical positions in the question vector similarity matrix, and assign the preset similarity base value to the positions with the same sample index, thereby generating a complete question vector similarity matrix. S325. Perform threshold comparison on each cosine similarity value in the question vector similarity matrix, mark sample pairs with similarity higher than the preset proximity threshold as semantically close sample pairs, and mark sample pairs with similarity lower than the preset difference threshold as semantically different sample pairs, thus forming a semantic proximity relationship between question vectors.
5. The method for automatically expanding the knowledge base of e-commerce intelligent customer service based on deep learning according to claim 1, characterized in that, S4 specifically includes: S41. Extract the center vector of the question intent cluster as the retrieval vector, and perform semantic feature mapping on the knowledge entry text in the knowledge base data through a semantic encoding model to generate a set of knowledge vectors corresponding to the knowledge entry. S42. Input the retrieval vector into the knowledge vector set and perform k-nearest neighbor search to obtain the k knowledge vectors that are semantically closest to the retrieval vector, and extract the corresponding knowledge entry text as a candidate set of associated knowledge entries. S43. Perform word segmentation on the text of knowledge entries in the candidate associated knowledge entry set, and calculate the weight value of each term in the candidate associated knowledge entry set using the TF-IDF algorithm. Construct the semantic feature representation of the knowledge entry based on the term weight. S44. Based on the semantic feature representation of knowledge items, a hierarchical clustering algorithm is executed to gradually merge knowledge items with similar semantic features through a bottom-up aggregation method, generating multiple knowledge item clusters. S45. Perform semantic merging processing on the knowledge entry texts in each knowledge entry cluster, and unify and organize knowledge entries with similar semantic expressions and consistent semantic directions to generate candidate knowledge units.
6. The method for automatically expanding the knowledge base of e-commerce intelligent customer service based on deep learning according to claim 5, characterized in that, Specifically, S43 includes: S431. Perform word segmentation on the text of knowledge entries in the candidate related knowledge entry set, and remove invalid terms by combining the stop word list and low information word filtering rules to obtain the effective term sequence of candidate knowledge entries; S432. Based on the distribution of each effective term sequence in different knowledge entry texts, construct a domain term set corresponding to the candidate associated knowledge entry set, and establish feature indexes for the terms in the domain term set; S433. Using the TF-IDF weight allocation mechanism, the importance of each term in the domain term set is evaluated, and high-weight terms that can represent the topic content of the knowledge item are extracted to form a keyword set for the knowledge item. S434. Perform feature filtering processing based on the keyword set of knowledge items, remove words with high cross-item repetition and low distinguishability, and retain feature words that can distinguish the semantic content of different knowledge items. S435. Encode the retained feature terms according to the feature index to form a sparse feature representation for each knowledge entry, and use each sparse feature representation as the semantic feature representation of the knowledge entry.
7. The method for automatically expanding the knowledge base of e-commerce intelligent customer service based on deep learning according to claim 1, characterized in that, S5 specifically includes: S51. Construct candidate knowledge units in a transactional manner, extract the term set in each candidate knowledge unit, and organize the corresponding term set into independent transaction records according to the semantic boundary of the candidate knowledge unit to generate a knowledge transaction set. S52. The FP-Growth algorithm is used to mine frequent itemsets in the knowledge transaction set. A frequent pattern tree is constructed based on the occurrence frequency of each term in the knowledge transaction set, and frequent itemsets that meet the preset support threshold are recursively extracted along the conditional path of the frequent pattern tree. S53. Generate association rules based on frequent itemsets, divide each frequent itemset into rule antecedent and rule consequent, and calculate the support and confidence of each association rule in combination with the co-occurrence relationship of itemsets in the knowledge transaction set, and screen high confidence rules that meet the preset confidence threshold. S54. Perform rule mapping on high-confidence rules, determine the subject terms and constraint terms corresponding to the rule antecedents as knowledge triggering conditions, determine the result terms corresponding to the rule consequents as knowledge response content, and generate extended knowledge entries according to the preset knowledge organization format.
8. The method for automatically expanding the knowledge base of e-commerce intelligent customer service based on deep learning according to claim 1, characterized in that, S6 specifically includes: S61. Perform knowledge alignment processing on the extended knowledge entries and knowledge base data, identify knowledge content in the extended knowledge entries that is semantically consistent or semantically similar to the knowledge base data, and merge and organize the corresponding knowledge content according to the semantic structure of the knowledge entries to generate a fused knowledge entry set. S62. Based on the fused knowledge item set, extract the correspondence between the question expression and the knowledge item, and construct an incremental training sample set according to the pairing results of the question expression and the corresponding knowledge item; S63. Perform semantic encoding processing on the question expressions in the incremental training sample set, and establish a semantic matching relationship based on the encoding results and the semantic representation of the corresponding knowledge items to form semantic matching samples; S64. Based on semantic matching samples, the parameters of the semantic coding model are iteratively updated using the stochastic gradient descent method, enabling the semantic coding model to learn the semantic features of the fused knowledge items and complete the incremental update of the semantic coding model.
Citation Information
Patent Citations
Government affair department knowledge base retrieval method based on semantic enhancement
CN121233791A