Dialogue processing method, device, equipment and storage medium
Through the analysis and clustering of shopping guide questions in dialogue records, the problem of low accuracy of structured data in the existing technology is solved, and more accurate and intelligent product recommendations are achieved.
Patent Information
- Application Number
- CN202211256001.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-10-13
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2042-10-13
AI Technical Summary
The prior art may combine keywords without association when extracting structured data from conversation records, resulting in a decrease in accuracy of structured data.
By obtaining shopping guide questions related to user demands in the conversation, using the representation vector to determine the keywords corresponding to multiple preset categories, and clustering them, selecting the benchmark cluster and the target cluster. After meeting the similarity conditions, the keywords in the target cluster are fused into the benchmark cluster.
It improves the accuracy of structured data, ensures that there is a semantic relationship between keywords, avoids the combination of unrelated keywords, and improves the accuracy and intelligence of product recommendations.
Smart Images

Figure CN115658889B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of information technology, and in particular to a conversation processing method, device, equipment and storage medium. Background Art
[0002] Currently, by analyzing the conversation records between real customer service and consumers, structured data can be extracted from the conversation records, thereby providing consumers with more accurate and intelligent product recommendation capabilities.
[0003] However, the existing technology may combine unrelated keywords in the conversation records to form structured data, thereby reducing the accuracy of the structured data. Summary of the invention
[0004] In order to solve the above technical problems or at least partially solve the above technical problems, the present disclosure provides a dialogue processing method, device, equipment and storage medium to improve the accuracy of structured data.
[0005] In a first aspect, an embodiment of the present disclosure provides a method for processing a conversation, including:
[0006] Obtain one or more shopping guide questions related to the user's demands in one or more conversations;
[0007] For each of the one or more shopping guide questions, determining keywords in the shopping guide question corresponding to at least one preset category among a plurality of preset categories according to a representation vector of each text unit in the shopping guide question, wherein the plurality of preset categories include a target category, and the target category is used to establish a semantic relationship between keywords of different preset categories in other preset categories among the plurality of preset categories except the target category;
[0008] For each preset category of the plurality of preset categories, clustering the one or more keywords corresponding to the preset category in the one or more shopping guide questions according to the representation vectors respectively corresponding to the one or more keywords corresponding to the preset category to obtain one or more cluster clusters;
[0009] Selecting a reference cluster from one or more clusters corresponding to the target category, determining one or more target clusters from the multiple clusters corresponding to the other preset categories, and the similarity between the target cluster and the reference cluster meets a preset condition;
[0010] The keywords in the one or more target clusters are merged into the reference clusters.
[0011] In a second aspect, an embodiment of the present disclosure provides a dialog processing device, including:
[0012] An acquisition module, used to acquire one or more shopping guide questions related to user demands in one or more conversations;
[0013] A first determining module is configured to determine, for each of the one or more shopping guide questions, keywords in the shopping guide question corresponding to at least one preset category among a plurality of preset categories according to a representation vector of each text unit in the shopping guide question, wherein the plurality of preset categories include a target category, and the target category is used to establish a semantic relationship between keywords of different preset categories in other preset categories among the plurality of preset categories except the target category;
[0014] A clustering module, configured to cluster the one or more keywords corresponding to the preset category in the one or more shopping guide questions according to the representation vectors respectively corresponding to the one or more keywords corresponding to the preset category, for each preset category of the multiple preset categories, to obtain one or more clusters;
[0015] A selection module, used for selecting a reference cluster from one or more clusters corresponding to the target category;
[0016] A second determination module, configured to determine one or more target clusters from the plurality of clusters corresponding to the other preset categories, wherein the similarity between the target cluster and the reference cluster meets a preset condition;
[0017] A fusion module is used to fuse the keywords in the one or more target clusters into the reference cluster.
[0018] In a third aspect, an embodiment of the present disclosure provides an electronic device, including:
[0019] Memory;
[0020] Processor; and
[0021] Computer programs;
[0022] The computer program is stored in the memory and is configured to be executed by the processor to implement the method as described in the first aspect.
[0023] In a fourth aspect, an embodiment of the present disclosure provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program is executed by a processor to implement the method described in the first aspect.
[0024] The dialogue processing method, device, equipment and storage medium provided by the embodiments of the present disclosure obtain one or more shopping guide questions related to user demands in one or more dialogues, and for each of the one or more shopping guide questions, determine the keywords in the shopping guide question corresponding to at least one preset category in multiple preset categories according to the representation vector of each text unit in the shopping guide question. Since the multiple preset categories include a target category, the target category is used to establish a semantic relationship between keywords of different preset categories in other preset categories other than the target category in the multiple preset categories. Therefore, for each preset category in the multiple preset categories, according to the representation vectors corresponding to the one or more keywords in the one or more shopping guide questions corresponding to the preset category, the one or more keywords are clustered, and after obtaining one or more cluster clusters, a reference cluster cluster can be selected from the one or more cluster clusters corresponding to the target category, and one or more target cluster clusters can be determined from the multiple cluster clusters corresponding to the other preset categories, so that the similarity between the target cluster cluster and the reference cluster cluster meets the preset condition. Furthermore, the keywords in the one or more target clusters are merged into the benchmark cluster, so that keywords of different preset categories are connected together through the target category data structure, ensuring the existence of a semantic relationship between the connected keywords, avoiding combining keywords with no correlation to form structured data, thereby improving the accuracy of the structured data. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present disclosure and, together with the description, serve to explain the principles of the present disclosure.
[0026] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.
[0027] Figure 1 A flow chart of a method for processing a conversation provided by an embodiment of the present disclosure;
[0028] Figure 2 A schematic diagram of an application scenario provided by an embodiment of the present disclosure;
[0029] Figure 3 A schematic diagram of a dialogue provided for an embodiment of the present disclosure;
[0030] Figure 4 A flow chart of a conversation processing method provided by another embodiment of the present disclosure;
[0031] Figure 5 A schematic diagram of a joint task provided for another embodiment of the present disclosure;
[0032] Figure 6 A schematic diagram of clustering provided by another embodiment of the present disclosure;
[0033] Figure 7 A flow chart of a conversation processing method provided by another embodiment of the present disclosure;
[0034] Figure 8 A schematic diagram of homogeneous matching and heterogeneous matching provided for another embodiment of the present disclosure;
[0035] Fig. 9 A schematic diagram of the structure of a dialog processing device provided in an embodiment of the present disclosure;
[0036] Fig.10 A schematic diagram of the structure of an electronic device embodiment provided by the present disclosure. DETAILED DESCRIPTION
[0037] In order to more clearly understand the above-mentioned objectives, features and advantages of the present disclosure, the scheme of the present disclosure will be further described below. It should be noted that the embodiments of the present disclosure and the features in the embodiments can be combined with each other without conflict.
[0038] In the following description, many specific details are set forth to facilitate a full understanding of the present disclosure, but the present disclosure may also be implemented in other ways different from those described herein; it is obvious that the embodiments in the specification are only part of the embodiments of the present disclosure, rather than all of the embodiments.
[0039] At present, by analyzing the conversation records between real customer service and consumers, structured data can be extracted from the conversation records, thereby providing consumers with more accurate and intelligent product recommendation capabilities. However, the existing technology may combine keywords that have no correlation in the conversation records to form structured data, thereby reducing the accuracy of the structured data. To address this problem, the embodiments of the present disclosure provide a conversation processing method, which is introduced below in conjunction with specific embodiments.
[0040] Figure 1 This is a flow chart of the conversation processing method provided in the embodiment of the present disclosure. The method can be executed by a conversation processing device, which can be implemented in software and / or hardware. The device can be configured in an electronic device, such as a server or a terminal, wherein the terminal specifically includes a mobile phone, a computer or a tablet computer. In addition, the conversation processing method described in this embodiment can be applied to Figure 2 The application scenario shown in Figure 1. Figure 2As shown, the application scenario includes a terminal 21 and a server 22, wherein the terminal 21 can be used to record a real person-to-person conversation, which can be a conversation between a user and a shopping guide. Further, the terminal 21 can send the person-to-person conversation in text form or voice form to the server 22. It can be understood that the person-to-person conversation provided by the terminal 21 to the server 22 is not limited to one conversation, for example, it can also be a multi-conversation conversation, the multi-conversation conversation can be a conversation between the same user and the shopping guide at different times, or the multi-conversation conversation can be a conversation between multiple users and the shopping guide at different times or at the same time. Among them, the shopping guides at different times may be the same person or may not be the same person. In addition, the server 22 can not only receive the person-to-person conversation provided by the terminal 21, for example, it can also receive the person-to-person conversation provided by other terminals or other servers. Thus, the server 22 can obtain at least one person-to-person conversation. Further, the server 22 can process or analyze the at least one person-to-person conversation using the method described in this embodiment, so as to extract structured data from the at least one person-to-person conversation. It is understandable that, in some embodiments, the terminal 21 may also process or analyze the at least one person-to-person conversation obtained by it using the method described in this embodiment, so as to extract structured data from the at least one person-to-person conversation. The structured data may be one or more groups. If there are multiple groups of structured data, each group of structured data may correspond to a category, and different structured data may correspond to different categories. Figure 2 This method is described in detail. Figure 1 As shown, the specific steps of this method are as follows:
[0041] S101. Obtain one or more shopping guide questions related to user demands in one or more conversations.
[0042] For example, Figure 2 Taking the server 22 as an example, the server 22 may receive at least one person-to-person conversation from the terminal 21, other terminals or other servers. Further, the server 22 may obtain one or more shopping guide questions related to the user's demands from the at least one person-to-person conversation. Figure 3 The figure shows a dialogue between a user and a shopping guide. In the dialogue, "YYYYY", "MMMMM", and "VVVVV" are shopping guide questions, but some of them may be shopping guide questions related to the user's demands, and the rest are shopping guide questions unrelated to the user's demands. For example, "YYYYY" is "What brand of air conditioner do you want?", "MMMMM" is "How many horsepower do you want?", and "VVVVV" is "What is your last name?", among which "YYYYY" and "MMMMM" are shopping guide questions related to the user's demands, such as buying an air conditioner, and "VVVVV" is a shopping guide question unrelated to the user's demands.
[0043] It is understandable that the server 22 can not only extract one or more shopping guide questions related to the user's demands from one conversation, but if the server 22 obtains multiple conversations, the server 22 can extract one or more shopping guide questions related to the user's demands from each conversation of the multiple conversations.
[0044] S102. For each of the one or more shopping guide questions, determine keywords in the shopping guide question that respectively correspond to at least one preset category among a plurality of preset categories according to a representation vector of each text unit in the shopping guide question, wherein the plurality of preset categories include a target category, and the target category is used to establish a semantic relationship between keywords of different preset categories in other preset categories among the plurality of preset categories except the target category.
[0045] For example, after the server 22 extracts one or more shopping guide questions related to the user's demands from each of the multiple conversations, for each of the one or more shopping guide questions, the server 22 may first determine a representation vector of each text unit in the shopping guide question. In this embodiment, a text unit may be a word, and in other embodiments, a text unit may also be a segmentation, a phrase, a character, a letter, a word, etc. In this embodiment, a text unit may be recorded as a token.
[0046] In addition, in this embodiment, multiple preset categories may be provided, such as attribute names, attribute values, and attribute pairs. Among them, the keywords, participles, or phrases belonging to the attribute names may be various aspects of interactive inquiries, such as skin quality, skin type, oily status, etc. The keywords, participles, or phrases belonging to the attribute values may be various options that may be applicable under a certain attribute, such as dry skin, mixed type, oily skin, etc. Taking keywords as an example, if a keyword belonging to the attribute name and a keyword belonging to the attribute value appear in a shopping guide question at the same time, then the keyword belonging to the attribute name and the keyword belonging to the attribute value may constitute an attribute pair. For example, (skin quality, dryness, oiliness) extracted from the field of beauty and makeup constitute an attribute pair, wherein "skin quality" is a keyword belonging to the attribute name, and "dryness" and "oiliness" are keywords belonging to the attribute value, respectively.
[0047] In this embodiment, for each shopping guide question related to the user's demand, the keywords in the shopping guide question corresponding to at least one of the multiple preset categories can be determined according to the representation vector of each word in the shopping guide question. That is to say, for the same shopping guide question, the shopping guide question may include keywords of one or more categories of attribute names, attribute values, and attribute pairs, that is, different shopping guide questions may include keywords of different categories. For example, shopping guide question A only includes keywords corresponding to attribute names, that is, keywords belonging to attribute names, and shopping guide question B includes keywords corresponding to attribute names, attribute values, and attribute pairs at the same time.
[0048] In addition, in this embodiment, since the attribute pair is composed of keywords belonging to the attribute name and keywords belonging to the attribute value in the same shopping guide question, that is, the attribute pair establishes a semantic relationship between keywords of different preset categories in other preset categories except the attribute pair, therefore, the attribute pair can be recorded as a target category among the multiple preset categories.
[0049] S103. For each preset category of the multiple preset categories, cluster the one or more keywords corresponding to the preset category in the one or more shopping guide questions according to the representation vectors respectively corresponding to the one or more keywords, so as to obtain one or more clusters.
[0050] For example, after determining the keywords corresponding to the attribute name, attribute value, and attribute pair from each shopping guide question related to the user's demands, all keywords belonging to the attribute name can constitute an attribute name set, all keywords belonging to the attribute value can constitute an attribute value set, and all attribute pairs constitute an attribute pair set. Further, for each preset category in the multiple preset categories, word clustering is performed within the category. For example, for the attribute name set, all keywords belonging to the attribute name are clustered according to the representation vector corresponding to each keyword in the attribute name set, and one or more cluster clusters under the attribute name category are obtained. Similarly, for the attribute value set, all keywords belonging to the attribute value are clustered according to the representation vector corresponding to each keyword in the attribute value set, and one or more cluster clusters under the attribute value category are obtained. Similarly, for the attribute pair set, since the attribute pair set may include one or more attribute pairs, for example, taking multiple attribute pairs as an example, since each attribute pair includes multiple keywords, the representation vectors of multiple keywords in an attribute pair can be averaged to obtain the representation vector of the attribute pair. Furthermore, according to the representation vector of each attribute pair in the attribute pair set, all attribute pairs in the attribute pair set are clustered to obtain one or more clusters under the attribute pair category. In other words, each cluster represents an aggregation of a specific attribute name, attribute value or attribute pair.
[0051] S104, selecting a reference cluster from one or more clusters corresponding to the target category, and determining one or more target clusters from the multiple clusters corresponding to other preset categories, wherein the similarity between the target cluster and the reference cluster meets a preset condition.
[0052] For example, there are one or more clusters under the category of attribute pair, and the server 22 in this embodiment can randomly select a cluster from the one or more clusters as the reference cluster. Further, based on the reference cluster, a target cluster is determined from one or more clusters under the category of attribute name, and / or a target cluster is determined from one or more clusters under the category of attribute value, so that the similarity between the target cluster and the reference cluster meets the preset condition.
[0053] S105: Merge the keywords in the one or more target clusters into the reference cluster.
[0054] For example, keywords in the target cluster under the category of attribute name and / or keywords in the target cluster under the category of attribute value are merged into the benchmark cluster, so that the keywords in the benchmark cluster are continuously expanded and increased, and the updated benchmark cluster is obtained. It can be understood that if there is a cluster under the category of attribute pair, then the cluster is used as the benchmark cluster, and when the updated benchmark cluster is obtained, the updated benchmark cluster can be used as a structured data. If there are multiple clusters under the category of attribute pair, each cluster in the multiple clusters can be used as a benchmark cluster respectively, so that each cluster in the multiple clusters is updated respectively. In this case, each updated benchmark cluster can be used as a structured data, so as to obtain multiple structured data. In addition, in other embodiments, the update process can also be continuously iterated. For example, when multiple clusters under the category of attribute pair are updated as a benchmark cluster respectively, each updated cluster can also be used as a benchmark cluster in turn, so as to continue to perform steps S104 and S105.
[0055] In the embodiment of the present disclosure, one or more shopping guide questions related to user demands in one or more conversations are obtained, and for each of the one or more shopping guide questions, the keywords in the shopping guide question corresponding to at least one preset category in a plurality of preset categories are determined according to the representation vector of each text unit in the shopping guide question. Since the plurality of preset categories include a target category, and the target category is used to establish a semantic relationship between keywords of different preset categories in other preset categories other than the target category in the plurality of preset categories, therefore, for each of the plurality of preset categories, the one or more keywords corresponding to the preset category in the one or more shopping guide questions are clustered according to the representation vectors respectively corresponding to the one or more keywords corresponding to the preset category, and after obtaining one or more cluster clusters, a reference cluster cluster can be selected from the one or more cluster clusters corresponding to the target category, and one or more target cluster clusters can be determined from the multiple cluster clusters corresponding to the other preset categories, so that the similarity between the target cluster cluster and the reference cluster cluster meets a preset condition. Furthermore, the keywords in the one or more target clusters are merged into the benchmark cluster, so that keywords of different preset categories are connected together through the target category data structure, ensuring the existence of a semantic relationship between the connected keywords, avoiding combining keywords with no correlation to form structured data, thereby improving the accuracy of the structured data.
[0056] In addition, some existing technologies are that operators extract the content to be inquired from the conversation records, and provide complete and comprehensive options corresponding to the content for consumers to choose, thereby using the content to be inquired and the complete and comprehensive options as structured data. However, if the operator does not have professional knowledge of the industry category, it will be difficult for the operator to extract complete and comprehensive structured data from the conversation records. In addition, extracting structured data manually will also result in low extraction efficiency. Therefore, compared to this type of existing technology, the present embodiment can automatically extract structured data from the conversation records, improve the extraction efficiency, and save labor costs. In addition, it can also solve the problem that it is difficult for operators to extract complete and comprehensive structured data from conversation records because they have no professional knowledge of the industry category.
[0057] Figure 4 A flow chart of a dialog processing method provided by another embodiment of the present disclosure. In this embodiment, according to the representation vector of each text unit in the shopping guide question, determining the keywords in the shopping guide question corresponding to at least one of the multiple preset categories includes the following steps:
[0058] S401: Determine the sentence category of the shopping guide question.
[0059] For example Figure 5 As shown in the figure, "Is your skin dry or oily?" is a shopping guide question. Taking this shopping guide question as an example, the sentence category of this shopping guide question can be determined first. For example, in this embodiment, the sentence category includes mixed questions, selection questions, open questions and others.
[0060] Optionally, determining the sentence category of the shopping guide question includes: adding a preset character at the starting position of the shopping guide question; inputting the preset character and the shopping guide question into an encoder, so that the encoder outputs a representation vector of the preset character in the context of the shopping guide question and a representation vector of each text unit in the shopping guide question in the context; and determining the sentence category of the shopping guide question based on the representation vector of the preset character in the context of the shopping guide question.
[0061] like Figure 5 As shown, a preset character such as [CLS] is added to the beginning of "Is your skin dry or oily?", and [CLS] can be recorded as a token. Then, [CLS] and each token (such as a word) in "Is your skin dry or oily?" are input into the bidirectional encoding representation algorithm (Bidirectional Encoder Representations from Transformers, BERT) encoder based on the Transformer algorithm. The BERT encoder first embeds the input into the embedding layer, and then passes through multiple layers of superimposed Transformer layers. The implementation mechanism inside the Transformer layer is the self-attention mechanism. Through the self-attention mechanism, the representation vector of each token in [CLS] and "Is your skin dry or oily?" in the context of the shopping guide question is learned.
[0062] like Figure 5 As shown, the representation vector of [CLS] in the context of the shopping guide question is further input into a multilayer perceptron (MLP), so that the MLP performs sentence classification based on the representation vector of [CLS] in the context of the shopping guide question, that is, determines the sentence category of the shopping guide question. In this embodiment, the sentence classification can be recorded as task 1. In addition, as Figure 5 The processing shown in the figure also includes Task 2, which is a sequence labeling task. The so-called sequence labeling is a major task in the field of natural language processing at the sentence level, that is, predicting the labels that need to be labeled in the sequence based on a given text sequence. For example, Task 2 includes: Figure 5The representation vector of each word in the context of the shopping guide question "Is your skin dry or oily?" is input into the bidirectional long short-term memory neural network (Bi-directional Long Short Term Memory Network, Bi-LSTM), and the output of the Bi-LSTM is further used as the input of the conditional random field (Conditional Random Field, CRF), so that the CRF can output the label of each word in "Is your skin dry or oily?", for example, the label of "skin" is B-attribute name, the label of "quality" is i-attribute name, and the label of "is" is o, where B means the beginning, i means the middle, and o means others. Therefore, it is determined that the keyword "skin quality" belongs to the attribute name. Similarly, it can be determined that "dry" and "oily" belong to attribute values respectively.
[0063] In addition, in this embodiment, Task 1 and Task 2 may be two joint tasks of natural language understanding. By jointly learning the loss functions of Task 1 and Task 2, the accuracy of the two tasks can be improved at the same time.
[0064] S402: Determine at least one preset category corresponding to the keyword in the shopping guide question according to the sentence category.
[0065] Since the sentence category of the shopping guide question and the keywords corresponding to the attribute names or attribute values to be extracted from the shopping guide question are mutually constrained and related, according to the sentence category of the shopping guide question, it can be determined which one or more preset categories of keywords are included in the shopping guide question.
[0066] Optionally, the multiple preset categories include attribute names, attribute values, and attribute pairs, and the attribute pairs include attribute names and attribute values; at least one preset category corresponding to the keywords in the shopping guide question is determined based on the sentence category, including: if the sentence category is a mixed question, then determining that the shopping guide question contains keywords corresponding to the attribute names and the attribute values respectively.
[0067] For example, if the sentence category of a shopping guide question is a mixed question, it means that there is a high probability that keywords corresponding to attribute names and keywords corresponding to attribute values exist in the shopping guide question at the same time.
[0068] If the sentence category of the shopping guide question is an open-ended question, for example, "What is your skin type?", then the shopping guide question is likely to only include keywords corresponding to the attribute names.
[0069] If the sentence category of the shopping guide question is a selection question, for example, "Do you have dry or oily skin?", then the shopping guide question is likely to only include keywords corresponding to attribute values.
[0070] S403: Determine keywords in the shopping guide question that respectively correspond to the at least one preset category according to the representation vector of each text unit in the shopping guide question.
[0071] For example, Figure 5 The representation vector of each word in the context of the shopping guide question "Is your skin dry or oily?" is input into the bidirectional long short-term memory neural network (Bi-directional Long Short Term Memory Network, Bi-LSTM), and the output of the Bi-LSTM is further used as the input of the conditional random field (Conditional Random Field, CRF), so that the CRF can output the label of each word in "Is your skin dry or oily?", for example, the label of "skin" is B-attribute name, the label of "quality" is i-attribute name, and the label of "is" is o, where B means the beginning, i means the middle, and o means others. Therefore, it is determined that the keyword "skin quality" belongs to the attribute name. Similarly, it can be determined that "dry" and "oily" belong to the attribute value respectively. It can be seen that since "Is your skin dry or oily?" is a mixed question, the keywords corresponding to the attribute name and the keywords corresponding to the attribute value can be analyzed and obtained from the shopping guide question.
[0072] This embodiment transforms a supervised traditional entity recognition task into a task of extracting keyword sequences of specific categories through sequence labeling, thereby reducing the manual labeling cost of entity category definition and facilitating cross-domain expansion of mining algorithms.
[0073] For example, after determining the keywords corresponding to the attribute name, attribute value, and attribute pair from each shopping guide question related to the user's demand, all keywords belonging to the attribute name can constitute an attribute name set, all keywords belonging to the attribute value can constitute an attribute value set, and all attribute pairs constitute an attribute pair set.
[0074] Further, if Figure 6As shown, each keyword in the attribute name set is input into the word embeddings (Word2vec) model to obtain the representation vector of each keyword in the attribute name set. Similarly, each keyword in the attribute value set is input into the Word2vec model to obtain the representation vector of each keyword in the attribute value set. Each attribute pair in the attribute pair set is input into the Word2vec model to obtain the representation vector of each attribute pair. Since an attribute pair includes multiple keywords, the representation vector of an attribute pair can be the average value of the representation vectors of the keywords included in the attribute pair. Further, the K-means clustering algorithm is used to cluster the keywords under each category (such as attribute name, attribute value, attribute pair) respectively. For example, the K-means clustering algorithm is used to cluster all the keywords in the attribute name set to obtain cluster clusters 61, cluster cluster 62, and cluster cluster 63, wherein, according to the representation vectors of different keywords in each cluster cluster, the distance between different keywords can be calculated, and the distance between different keywords in the same cluster cluster is less than or equal to the preset value. Similarly, after K-means clustering of all keywords in the attribute value set, clusters 71, 72, and 73 are obtained. After K-means clustering of all attribute pairs in the attribute pair set, clusters 81, 82, and 83 are obtained. Any cluster in clusters 81, 82, and 83 includes one or more attribute pairs. It can be understood that Figure 6 This is just a schematic illustration and does not limit the number of clusters under each category. In addition, a keyword in each cluster can be recorded as a node.
[0075] In addition, in some embodiments, Figure 6 The clusters under each category shown can also evaluate the quality of each cluster according to dimensions such as the size of the cluster, the average value of the similarity between each keyword in the cluster, etc., so as to retain the clusters with a confidence level greater than or equal to a preset threshold.
[0076] Figure 7 This is a flow chart of a conversation processing method provided by another embodiment of the present disclosure. In this embodiment, one or more target clusters are determined from the multiple clusters corresponding to the other preset categories, and the similarity between the target cluster and the reference cluster meets the preset conditions, including the following steps:
[0077] S701. Take each of the multiple clusters corresponding to the other preset categories as a candidate cluster, and determine the first similarity between the candidate cluster and the benchmark cluster according to the representation vectors of all keywords in the candidate cluster and the representation vectors of keywords in the benchmark cluster that have the same preset category as the candidate cluster.
[0078] For example Figure 6 As shown, in each cluster, the black solid dots are used to represent the keywords corresponding to the attribute name, and the hollow dots are used to represent the keywords corresponding to the attribute value. Since the attribute pair is composed of the keywords corresponding to the attribute name and the keywords corresponding to the attribute value, cluster 81, cluster 82, and cluster 83 include black solid dots and hollow dots, respectively. Further, any cluster, such as cluster 82, is selected from cluster 81, cluster 82, and cluster 83 as a reference cluster. Cluster 61, cluster 62, cluster 63, cluster 71, cluster 72, and cluster 73 are sequentially selected as candidate clusters. For example, taking cluster 61 as an example, the first similarity between cluster 61 and cluster 82 is determined based on the representation vectors of all keywords in cluster 61 and the representation vectors of keywords in cluster 82 that have the same preset category as cluster 61. Among them, since the preset category corresponding to cluster 61 is the attribute name, the keyword belonging to the attribute name in cluster 82 is "skin quality".
[0079] Optionally, the first similarity between the candidate cluster and the benchmark cluster is determined based on the representation vectors of all keywords in the candidate cluster and the representation vectors of keywords in the benchmark cluster that have the same preset category as the candidate cluster, including: calculating a first average value of the representation vectors of all keywords in the candidate cluster; calculating a second average value of the representation vectors of keywords in the benchmark cluster that have the same preset category as the candidate cluster; and determining the first similarity between the candidate cluster and the benchmark cluster based on the first average value and the second average value.
[0080] For example, based on the representation vector of each keyword in cluster 61, a first average value of the representation vectors of all keywords in cluster 61 is calculated. Further, a second average value of the representation vectors of keywords in cluster 82 that have the same preset category as cluster 61 is calculated. Since there is only one keyword belonging to the attribute name in cluster 82, namely "skin quality", in this case, the second average value is the representation vector of "skin quality". Further, based on the first average value and the second average value, a first similarity between cluster 61 and cluster 82 is determined.
[0081] S702: Determine a second similarity between the candidate cluster and the reference cluster according to the representation vectors of all keywords in the candidate cluster and the representation vectors of all keywords in the reference cluster.
[0082] In addition, if Figure 6As shown, the second similarity between cluster 61 and cluster 82 may also be determined based on the representation vectors of all keywords in cluster 61 and the representation vectors of all keywords in cluster 82 .
[0083] Optionally, the second similarity between the candidate cluster and the benchmark cluster is determined based on the representation vectors of all keywords in the candidate cluster and the representation vectors of all keywords in the benchmark cluster, including: calculating a first average value of the representation vectors of all keywords in the candidate cluster; calculating a third average value of the representation vectors of all keywords in the benchmark cluster; and determining the second similarity between the candidate cluster and the benchmark cluster based on the first average value and the third average value.
[0084] For example, based on the representation vector of each keyword in cluster 61, a first average value of the representation vectors of all keywords in cluster 61 is calculated. Further, a third average value of the representation vectors corresponding to all keywords in cluster 82, i.e., "skin quality", "dryness", and "oilyness", is calculated. Further, based on the first average value and the third average value, a second similarity between cluster 61 and cluster 82 is determined.
[0085] S703: When the first similarity satisfies a first preset condition and / or the second similarity satisfies a second preset condition, use the candidate cluster as the target cluster.
[0086] In this embodiment, the process of calculating the first similarity between the candidate cluster and the reference cluster is recorded as homogeneous matching, and the process of calculating the second similarity between the candidate cluster and the reference cluster is recorded as heterogeneous matching. In other words, homogeneous matching is to match the candidate cluster with keywords in the reference cluster that are of the same category as the candidate cluster, and heterogeneous matching is to match the candidate cluster with the entire reference cluster.
[0087] In a feasible implementation, since the same candidate cluster may correspond to a first similarity and a second similarity, for the same candidate cluster, if the first similarity is greater than the first threshold, and / or the second similarity is greater than the second threshold, then the candidate cluster is used as the target clustering cluster, thereby merging all keywords in the candidate cluster into the benchmark clustering cluster.
[0088] In another possible implementation, Figure 6As shown, cluster 61, cluster 62, and cluster 63 are clusters under the category of attribute name, respectively. Each cluster in cluster 61, cluster 62, and cluster 63 corresponds to a first similarity and a second similarity. If the first similarity and / or the second similarity of a cluster is the largest, then the cluster is used as the target cluster. In other words, a target cluster is selected from multiple clusters under the same category. Similarly, a target cluster can be selected from cluster 71, cluster 72, and cluster 73.
[0089] That is to say, based on the benchmark cluster, the other two different categories of clusters (attribute name cluster and attribute value cluster) can be aligned with it to achieve the goal of heterogeneous cluster fusion. Among them, the attribute name cluster refers to the cluster under the attribute name category, and the attribute value cluster refers to the cluster under the attribute value category.
[0090] In addition, in other embodiments, the first similarity may also be the similarity between each keyword in the candidate cluster and the benchmark cluster, and the second similarity may also be the similarity between each keyword in the candidate cluster and the benchmark cluster. For example, taking cluster 61 as an example, cluster 61 includes "skin". Based on the representation vector of "skin" and the representation vector of "skin quality" in cluster 82, a homogeneous match is performed to obtain the first similarity. Based on the representation vector of "skin" and the representation vector of cluster 82 as a whole (that is, the average value of the representation vectors of all keywords in cluster 82), a heterogeneous match is performed to obtain the second similarity. Similarly, the first similarity and the second similarity corresponding to "epidermis" in cluster 61 can also be calculated. If the first similarity and the second similarity corresponding to "skin" are respectively greater than the first similarity and the second similarity corresponding to "epidermis", then "skin" is merged into cluster 82, as shown in FIG. Figure 8 As shown. Similarly, when cluster 71 is a candidate cluster, the first similarity and the second similarity corresponding to each keyword in cluster 71 are calculated. For example, "dry skin" in cluster 71 and "oily" and "dry" in cluster 82 belong to the same category. Therefore, based on the representation vector of "dry skin" in cluster 71 and the representation vectors corresponding to "oily" and "dry" in cluster 82, a homogeneous matching is performed to obtain the first similarity. For example, the first similarity is calculated based on the representation vector of "dry skin" and the average value of the representation vectors of "oily" and "dry". Furthermore, based on the representation vector of "dry skin" and the representation vector of cluster 82 as a whole, a heterogeneous matching is performed to obtain the second similarity. If the first similarity corresponding to "dry skin" is greater than the first threshold, and / or the second similarity corresponding to "dry skin" is greater than the second threshold, then "dry skin" is merged into cluster 82, as shown in FIG. Figure 8 shown.
[0091] In some embodiments, to ensure accuracy, each time homogeneous matching and heterogeneous matching are performed, each reference cluster can be fused with at most one candidate cluster of the same category. For example, cluster 82 can be fused with one candidate cluster from multiple candidate clusters under the attribute name category, and one candidate cluster from multiple candidate clusters under the attribute value category. When the fusion process for cluster 82 is completed, since cluster 82 is fused with new keywords, an updated reference cluster is obtained, such as Figure 8 The cluster 91 shown in FIG. 1 is a cluster 92, and therefore, the representation vector of the entire cluster 82 will be updated to the representation vector of the entire cluster 91. Further, the reference cluster is replaced, for example, with cluster 81 or cluster 83, and the fusion process described above is performed. It can be understood that for the same reference cluster, after the fusion process, it may be fused to the candidate cluster or a keyword in the candidate cluster, or it may not be fused.
[0092] When cluster cluster 81, cluster cluster 82, and cluster cluster 83 have undergone the fusion process respectively, then further, the benchmark cluster clusters are selected one by one from the updated cluster cluster 81, the updated cluster cluster 82, and the updated cluster cluster 83, so as to continue the fusion process, that is, iterative matching, until no candidate clusters can be fused to the benchmark cluster cluster. It can be understood that when a candidate cluster is fused to the benchmark cluster cluster, the candidate cluster can continue to participate in the matching in the subsequent iterative matching, or it can no longer participate in the subsequent iterative matching. After multiple iterative matching, the final output attribute pair for each cluster cluster under this category includes all possible attribute names and attribute values in the same interactive attribute, and the final output attribute pair for each cluster cluster under this category can be used as a structured data for providing merchants with multiple rounds of shopping guide industry templates, thereby providing large, medium and small merchants with an out-of-the-box industry shopping guide template. For example, in the field of beauty, a specific interactive attribute cluster is composed of a set of attribute names (such as skin texture, skin type, oiliness, etc.) and a set of attribute values (such as dry skin, combination skin, oily skin, etc.).
[0093] The reason why this embodiment selects the two matching methods of homogeneous matching and heterogeneous matching is that attribute names and attribute values are semantically related, so it is beneficial to learn richer semantic information by integrating the matching calculations of the same dimension and the overall dimension at the same time. In addition, this proposal cleverly uses the natural co-occurrence of attribute names and attribute values in mixed questions to link keywords belonging to different categories (such as attribute names and attribute values) together through the data structure of attribute pairs. In other words, the relationship between different categories is associated and guaranteed by attribute pairs where attribute names and attribute values co-occur. For example, "skin quality problems" and "thick eyebrows" will not appear in the same sentence. Therefore, in the process of merging attribute pair clusters with clusters of other different categories, it is not easy to merge "skin quality problems" and "thick eyebrows" into the same attribute pair cluster (i.e., the cluster under the category of attribute pairs), thereby avoiding keywords such as "skin quality problems" and "thick eyebrows" that have no semantic relationship from being merged into the same attribute pair cluster, thereby improving the accuracy of structured data.
[0094] Fig. 9 The structure diagram of the dialogue processing device provided by the embodiment of the present disclosure is as follows. The dialogue processing device provided by the embodiment of the present disclosure can execute the processing flow provided by the dialogue processing method embodiment, such as Fig. 9 As shown, the dialogue processing device 90 includes:
[0095] An acquisition module 901 is used to acquire one or more shopping guide questions related to user demands in one or more conversations;
[0096] A first determining module 902 is configured to determine, for each of the one or more shopping guide questions, keywords in the shopping guide question corresponding to at least one of a plurality of preset categories according to a representation vector of each text unit in the shopping guide question, wherein the plurality of preset categories include a target category, and the target category is used to establish a semantic relationship between keywords of different preset categories in other preset categories among the plurality of preset categories except the target category;
[0097] A clustering module 903 is used to cluster the one or more keywords corresponding to each of the plurality of preset categories according to the representation vectors respectively corresponding to the one or more keywords in the one or more shopping guide questions corresponding to the preset category, to obtain one or more clusters;
[0098] A selection module 904 is used to select a reference cluster from one or more clusters corresponding to the target category;
[0099] A second determination module 905 is configured to determine one or more target clusters from the plurality of clusters corresponding to the other preset categories, wherein the similarity between the target cluster and the reference cluster satisfies a preset condition;
[0100] The fusion module 906 is used to fuse the keywords in the one or more target clusters into the reference cluster.
[0101] Optionally, the second determination module 905 determines one or more target clusters from the multiple clusters corresponding to the other preset categories, and when the similarity between the target cluster and the reference cluster meets a preset condition, is specifically used to:
[0102] Taking each of the plurality of clusters corresponding to the other preset categories as a candidate cluster, and determining a first similarity between the candidate cluster and the reference cluster according to the representation vectors of all keywords in the candidate cluster and the representation vectors of keywords in the reference cluster that have the same preset category as the candidate cluster;
[0103] Determining a second similarity between the candidate cluster and the reference cluster according to the representation vectors of all keywords in the candidate cluster and the representation vectors of all keywords in the reference cluster;
[0104] When the first similarity satisfies a first preset condition, and / or the second similarity satisfies a second preset condition, the candidate cluster is used as the target clustering cluster.
[0105] Optionally, when the second determination module 905 determines the first similarity between the candidate cluster and the reference cluster according to the representation vectors of all keywords in the candidate cluster and the representation vectors of keywords in the reference cluster that have the same preset category as the candidate cluster, it is specifically used to:
[0106] Calculate a first average value of the representation vectors of all keywords in the candidate cluster;
[0107] Calculating a second average value of representation vectors of keywords in the reference cluster that have the same preset category as the candidate cluster;
[0108] A first similarity between the candidate cluster and the reference cluster is determined according to the first average value and the second average value.
[0109] Optionally, when the second determination module 905 determines the second similarity between the candidate cluster and the reference cluster according to the representation vectors of all keywords in the candidate cluster and the representation vectors of all keywords in the reference cluster, it is specifically used to:
[0110] Calculate a first average value of the representation vectors of all keywords in the candidate cluster;
[0111] Calculating a third average value of the representation vectors of all keywords in the reference cluster;
[0112] A second similarity between the candidate cluster and the reference cluster is determined according to the first average value and the third average value.
[0113] Optionally, when the first determining module 902 determines the keywords in the shopping guide question corresponding to at least one of the plurality of preset categories respectively according to the representation vector of each text unit in the shopping guide question, it is specifically configured to:
[0114] Determining the sentence category of the shopping guide question;
[0115] Determining at least one preset category corresponding to the keyword in the shopping guide question according to the sentence category;
[0116] According to the representation vector of each text unit in the shopping guide question, keywords in the shopping guide question corresponding to the at least one preset category are determined.
[0117] Optionally, the multiple preset categories include attribute names, attribute values, and attribute pairs, and the attribute pairs include attribute names and attribute values; when the first determination module 902 determines at least one preset category corresponding to the keywords in the shopping guide question according to the sentence category, it is specifically used to: if the sentence category is a mixed question, determine that the shopping guide question contains keywords corresponding to the attribute names and the attribute values respectively.
[0118] Optionally, when the first determining module 902 determines the sentence category of the shopping guide question, it is specifically used to:
[0119] Adding a preset character at the beginning of the shopping guide question;
[0120] Inputting the preset character and the shopping guide question into an encoder, so that the encoder outputs a representation vector of the preset character in the context of the shopping guide question and a representation vector of each text unit in the shopping guide question in the context;
[0121] The sentence category of the shopping guide question is determined according to the representation vector of the preset character in the context of the shopping guide question.
[0122] Fig. 9 The dialogue processing device of the illustrated embodiment can be used to execute the technical solution of the above-mentioned method embodiment. Its implementation principle and technical effect are similar and will not be repeated here.
[0123] The above describes the internal functions and structure of the dialogue processing device, which can be implemented as an electronic device. Fig.10 This is a schematic diagram of the structure of an electronic device embodiment provided by the present disclosure. Fig.10 As shown, the electronic device includes a memory 1001 and a processor 1002 .
[0124] The memory 1001 is used to store programs. In addition to the above programs, the memory 1001 can also be configured to store various other data to support operations on the electronic device. Examples of these data include instructions for any application or method operating on the electronic device, contact data, phone book data, messages, pictures, videos, etc.
[0125] Memory 1001 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk.
[0126] The processor 1002 is coupled to the memory 1001 and executes the program stored in the memory 1001 to:
[0127] Obtain one or more shopping guide questions related to the user's demands in one or more conversations;
[0128] For each of the one or more shopping guide questions, determining keywords in the shopping guide question corresponding to at least one preset category among a plurality of preset categories according to a representation vector of each text unit in the shopping guide question, wherein the plurality of preset categories include a target category, and the target category is used to establish a semantic relationship between keywords of different preset categories in other preset categories among the plurality of preset categories except the target category;
[0129] For each preset category of the plurality of preset categories, clustering the one or more keywords corresponding to the preset category in the one or more shopping guide questions according to the representation vectors respectively corresponding to the one or more keywords corresponding to the preset category to obtain one or more cluster clusters;
[0130] Selecting a reference cluster from one or more clusters corresponding to the target category, determining one or more target clusters from the multiple clusters corresponding to the other preset categories, and the similarity between the target cluster and the reference cluster meets a preset condition;
[0131] The keywords in the one or more target clusters are merged into the reference clusters.
[0132] Further, if Fig.10 As shown, the electronic device may also include: a communication component 1003, a power component 1004, an audio component 1005, a display 1006 and other components. Fig.10Only some components are shown schematically, which does not mean that the electronic device only includes Fig.10 Components shown.
[0133] The communication component 1003 is configured to facilitate wired or wireless communication between the electronic device and other devices. The electronic device can access a wireless network based on a communication standard, such as WiFi, 2G or 3G, or a combination thereof. In an exemplary embodiment, the communication component 1003 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 1003 also includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology and other technologies.
[0134] The power supply component 1004 provides power to various components of the electronic device. The power supply component 1004 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to the electronic device.
[0135] The audio component 1005 is configured to output and / or input audio signals. For example, the audio component 1005 includes a microphone (MIC), and when the electronic device is in an operation mode, such as a call mode, a recording mode, and a voice recognition mode, the microphone is configured to receive an external audio signal. The received audio signal can be further stored in the memory 1001 or sent via the communication component 1003. In some embodiments, the audio component 1005 also includes a speaker for outputting audio signals.
[0136] The display 1006 includes a screen, which may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touch screen to receive input signals from a user. The touch panel includes one or more touch sensors to sense touch, slide, and gestures on the touch panel. The touch sensor may not only sense the boundaries of a touch or slide action, but also detect the duration and pressure associated with the touch or slide operation.
[0137] In addition, an embodiment of the present disclosure further provides a computer-readable storage medium on which a computer program is stored. The computer program is executed by a processor to implement the dialog processing method described in the above embodiment.
[0138] It should be noted that, in this article, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the existence of other identical elements in the process, method, article or device including the elements.
[0139] The above description is only a specific embodiment of the present disclosure, so that those skilled in the art can understand or implement the present disclosure. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present disclosure. Therefore, the present disclosure will not be limited to the embodiments described herein, but will conform to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method for processing a conversation, wherein: The method comprises: Obtain one or more shopping guide questions related to the user's demands in one or more conversations; For each of the one or more shopping guide questions, determining keywords in the shopping guide question corresponding to at least one preset category among a plurality of preset categories according to a representation vector of each text unit in the shopping guide question, wherein the plurality of preset categories include a target category, and the target category is used to establish a semantic relationship between keywords of different preset categories in other preset categories among the plurality of preset categories except the target category; For each preset category of the plurality of preset categories, clustering the one or more keywords corresponding to the preset category in the one or more shopping guide questions according to the representation vectors respectively corresponding to the one or more keywords corresponding to the preset category to obtain one or more cluster clusters; Selecting a reference cluster from one or more clusters corresponding to the target category, determining one or more target clusters from the multiple clusters corresponding to the other preset categories, and the similarity between the target cluster and the reference cluster meets a preset condition; The keywords in the one or more target clusters are merged into the reference clusters.
2. The method according to claim 1, wherein: Determining one or more target clusters from the multiple clusters corresponding to the other preset categories, wherein the similarity between the target cluster and the reference cluster satisfies a preset condition, includes: Taking each of the plurality of clusters corresponding to the other preset categories as a candidate cluster, and determining a first similarity between the candidate cluster and the reference cluster according to the representation vectors of all keywords in the candidate cluster and the representation vectors of keywords in the reference cluster that have the same preset category as the candidate cluster; Determining a second similarity between the candidate cluster and the reference cluster according to the representation vectors of all keywords in the candidate cluster and the representation vectors of all keywords in the reference cluster; When the first similarity satisfies a first preset condition, and / or the second similarity satisfies a second preset condition, the candidate cluster is used as the target clustering cluster.
3. The method according to claim 2, wherein: Determining a first similarity between the candidate cluster and the reference cluster according to the representation vectors of all keywords in the candidate cluster and the representation vectors of keywords in the reference cluster that have the same preset category as the candidate cluster, including: Calculate a first average value of the representation vectors of all keywords in the candidate cluster; Calculating a second average value of representation vectors of keywords in the reference cluster that have the same preset category as the candidate cluster; A first similarity between the candidate cluster and the reference cluster is determined according to the first average value and the second average value.
4. The method according to claim 2, wherein: Determining a second similarity between the candidate cluster and the reference cluster according to the representation vectors of all keywords in the candidate cluster and the representation vectors of all keywords in the reference cluster includes: Calculate a first average value of the representation vectors of all keywords in the candidate cluster; Calculating a third average value of the representation vectors of all keywords in the reference cluster; A second similarity between the candidate cluster and the reference cluster is determined according to the first average value and the third average value.
5. The method according to claim 1, wherein: Determining keywords in the shopping guide question corresponding to at least one of a plurality of preset categories according to the representation vector of each text unit in the shopping guide question, comprises: Determining the sentence category of the shopping guide question; Determine at least one preset category corresponding to the keyword in the shopping guide question according to the sentence category; According to the representation vector of each text unit in the shopping guide question, keywords in the shopping guide question corresponding to the at least one preset category are determined.
6. The method according to claim 5, wherein: The plurality of preset categories include attribute names, attribute values, and attribute pairs, wherein the attribute pairs include attribute names and attribute values; Determining at least one preset category corresponding to the keyword in the shopping guide question according to the sentence category includes: If the sentence category is a mixed question, it is determined that the shopping guide question contains keywords corresponding to the attribute name and the attribute value respectively.
7. The method according to claim 5, wherein: Determining the sentence category of the shopping guide question includes: Adding a preset character at the beginning of the shopping guide question; Inputting the preset character and the shopping guide question into an encoder, so that the encoder outputs a representation vector of the preset character in the context of the shopping guide question and a representation vector of each text unit in the shopping guide question in the context; The sentence category of the shopping guide question is determined according to the representation vector of the preset character in the context of the shopping guide question.
8. A conversation processing device, wherein: include: An acquisition module, used to acquire one or more shopping guide questions related to user demands in one or more conversations; A first determining module is configured to determine, for each of the one or more shopping guide questions, keywords in the shopping guide question corresponding to at least one preset category among a plurality of preset categories according to a representation vector of each text unit in the shopping guide question, wherein the plurality of preset categories include a target category, and the target category is used to establish a semantic relationship between keywords of different preset categories in other preset categories among the plurality of preset categories except the target category; A clustering module, configured to cluster the one or more keywords corresponding to the preset category in the one or more shopping guide questions according to the representation vectors respectively corresponding to the one or more keywords corresponding to the preset category, for each preset category in the multiple preset categories, to obtain one or more clusters; A selection module, used for selecting a reference cluster from one or more clusters corresponding to the target category; A second determination module, configured to determine one or more target clusters from the plurality of clusters corresponding to the other preset categories, wherein the similarity between the target cluster and the reference cluster meets a preset condition; A fusion module is used to fuse the keywords in the one or more target clusters into the reference cluster.
9. An electronic device, wherein: include: Memory; processor; as well as Computer programs; The computer program is stored in the memory and is configured to be executed by the processor to implement the method according to any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, wherein: When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Electronic device, text information detection method and storage medium
CN109614608A
Text classification method and device based on reinforcement learning, computer equipment and medium
CN114780727A