Dialogue data processing method, device, equipment and storage medium

By clustering and mining candidate subsequences of multiple rounds of speech content, the support degree and cohesion degree are calculated, and the key content in the dialogue data is determined, which solves the problem that the key content cannot be accurately determined in the prior art and achieves higher accuracy.

CN115129836BActive Publication Date: 2025-05-09ALIBABA (CHINA) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210643395.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-08
Publication Date
2025-05-09
Estimated Expiration
2042-06-08

AI Technical Summary

Technical Problem

The prior art cannot accurately determine the key content in conversation data.

Method used

By clustering multiple rounds of speech content, the cluster identification sequence is determined, and candidate subsequences are mined from it, their support and cohesion are calculated, and the score is determined, and the cluster clusters are sorted as the key content.

Benefits of technology

It realizes more accurate determination of key content in dialogue data, and is more accurate than directly using clustered clusters as key content without candidate subsequence mining.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115129836B_ABST
    Figure CN115129836B_ABST
Patent Text Reader

Abstract

The present disclosure relates to a method, device, equipment and storage medium for processing conversation data. The present disclosure clusters multiple rounds of speech content in a first set consisting of at least one round of speech content in each conversation data, thereby determining a cluster identification sequence corresponding to each conversation data, and then forming a second set by the cluster identification sequence corresponding to each conversation data. Candidate subsequences in each cluster identification sequence in the second set are mined, and the support of each candidate subsequence in the second set is calculated. According to the cohesion corresponding to at least one candidate subsequence whose support is greater than or equal to a threshold, the score of each cluster identification that has appeared in the at least one candidate subsequence is calculated. According to the score, multiple cluster clusters can be sorted, so that the cluster clusters with the highest sorting are used as key content. This embodiment can more accurately determine the key content in the conversation data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of information technology, and in particular to a conversation data processing method, device, equipment and storage medium. Background Art

[0002] With the continuous development of technology, intelligent customer service or manual customer service will talk or converse with users every day, thus generating a large amount of conversation data. Each conversation data includes multiple rounds of speech content, which are composed of the speeches of intelligent customer service or manual customer service and users in sequence. Some of the speeches in the multiple rounds of speech content are the key content of the conversation, which represents the main content and direction of the conversation.

[0003] However, the inventor of the present application found that in each conversation data, only a small amount of speech content is the key content of the conversation, and the prior art cannot accurately determine the key content in the conversation data. Summary of the invention

[0004] In order to solve the above technical problems or at least partially solve the above technical problems, the present disclosure provides a conversation data processing method, device, equipment and storage medium. Compared with directly using the clustering clusters obtained after clustering as key content without mining candidate subsequences, this embodiment can more accurately determine the key content in the conversation data.

[0005] In a first aspect, an embodiment of the present disclosure provides a method for processing conversation data, including:

[0006] Acquire at least one dialogue data, each dialogue data including multiple rounds of speech content;

[0007] Determine a cluster identification sequence corresponding to each of the dialogue data by clustering multiple rounds of speech content in a first set formed by at least one round of speech content in each of the dialogue data, wherein the cluster identification sequence includes cluster identifications of cluster clusters corresponding to the at least one round of speech content in the dialogue data;

[0008] For each cluster identification sequence corresponding to each conversation data, determine a candidate subsequence in the cluster identification sequence, and determine the support of the candidate subsequence in a second set, wherein the second set includes the cluster identification sequence corresponding to each conversation data;

[0009] Calculate the score of each cluster identifier that appears in at least one candidate subsequence according to the cohesion corresponding to each candidate subsequence whose support is greater than or equal to the threshold;

[0010] According to the score, the cluster corresponding to the cluster identifier that meets the preset condition is determined as the key content.

[0011] In a second aspect, an embodiment of the present disclosure provides a conversation data processing device, including:

[0012] An acquisition module, used to acquire at least one dialogue data, each dialogue data includes multiple rounds of speech content;

[0013] A first determination module is used to determine a cluster identification sequence corresponding to each of the dialogue data by clustering multiple rounds of speech content in a first set consisting of at least one round of speech content in each of the dialogue data, wherein the cluster identification sequence includes cluster identifications of cluster clusters corresponding to the at least one round of speech content in the dialogue data;

[0014] A second determination module is used to determine, for each cluster identification sequence corresponding to each conversation data, a candidate subsequence in the cluster identification sequence, and determine the support of the candidate subsequence in a second set, wherein the second set includes the cluster identification sequence corresponding to each conversation data;

[0015] a calculation module, configured to calculate the score of each cluster identifier that has appeared in at least one candidate subsequence according to the cohesion corresponding to each candidate subsequence whose support is greater than or equal to the threshold;

[0016] The third determination module is used to determine, according to the score, the cluster corresponding to the cluster identifier that meets the preset condition as the key content.

[0017] In a third aspect, an embodiment of the present disclosure provides an electronic device, including:

[0018] Memory;

[0019] Processor; and

[0020] Computer programs;

[0021] The computer program is stored in the memory and is configured to be executed by the processor to implement the method as described in the first aspect.

[0022] In a fourth aspect, an embodiment of the present disclosure provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program is executed by a processor to implement the method described in the first aspect.

[0023] The conversation data processing method, device, equipment and storage medium provided by the embodiments of the present disclosure cluster multiple rounds of speech content in the first set formed by at least one round of speech content in each conversation data, thereby determining the cluster identification sequence corresponding to each conversation data, and then forming the second set by the cluster identification sequence corresponding to each conversation data. Further, the candidate subsequences in each cluster identification sequence in the second set are mined, and the support of each candidate subsequence in the second set is calculated. The greater the support, the more times the candidate subsequence appears in the second set, and the cluster clusters corresponding to each cluster identification in the candidate subsequence are more likely to be key content. Further, at least one candidate subsequence with a support greater than or equal to a threshold is selected, and then the cohesion corresponding to the at least one candidate subsequence is calculated according to the support corresponding to the at least one candidate subsequence. Since the cohesion represents the stability of the candidate subsequence, the higher the stability, the more stable the combination formed by each cluster identification in the candidate subsequence. Therefore, when the score of each cluster identifier that appears in the at least one candidate subsequence is calculated based on the cohesion degree corresponding to each candidate subsequence, a higher score indicates that the cluster identifier appears more times in multiple combinations, or the cluster identifier is more important in multiple combinations. Therefore, multiple clusters can be sorted according to the score, so that the clusters with higher rankings are used as key content. Compared with directly using the clusters obtained after clustering as key content without mining candidate subsequences, this embodiment can more accurately determine the key content in the conversation data. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present disclosure and, together with the description, serve to explain the principles of the present disclosure.

[0025] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.

[0026] Figure 1 A flow chart of a method for processing conversation data provided by an embodiment of the present disclosure;

[0027] Figure 2 A schematic diagram of an application scenario provided by an embodiment of the present disclosure;

[0028] Figure 3 A flow chart of a method for processing conversation data provided by another embodiment of the present disclosure;

[0029] Figure 4A schematic diagram of the structure of a conversation data processing device provided in an embodiment of the present disclosure;

[0030] Figure 5 A schematic diagram of the structure of an electronic device embodiment provided by the present disclosure. DETAILED DESCRIPTION

[0031] In order to more clearly understand the above-mentioned objectives, features and advantages of the present disclosure, the scheme of the present disclosure will be further described below. It should be noted that the embodiments of the present disclosure and the features in the embodiments can be combined with each other without conflict.

[0032] In the following description, many specific details are set forth to facilitate a full understanding of the present disclosure, but the present disclosure may also be implemented in other ways different from those described herein; it is obvious that the embodiments in the specification are only part of the embodiments of the present disclosure, rather than all of the embodiments.

[0033] Normally, intelligent customer service or manual customer service will talk or converse with users every day, thereby generating a large amount of conversation data. Each conversation data includes multiple rounds of speech content, which are composed of the speeches of intelligent customer service or manual customer service, and users in sequence. Some of the speeches in the multiple rounds of speech content are key contents in the conversation, and the key contents represent the main content and direction of the conversation. However, in each conversation data, only a small amount of speech content is the key content of the conversation, and the prior art cannot accurately determine the key content in the conversation data. In response to this problem, the embodiments of the present disclosure provide a method for processing conversation data, which is described below in conjunction with specific embodiments.

[0034] Figure 1 This is a flow chart of the conversation data processing method provided by the embodiment of the present disclosure. The method can be executed by a conversation data processing device, which can be implemented in software and / or hardware. The device can be configured in an electronic device, such as a server or a terminal, wherein the terminal specifically includes a mobile phone, a computer or a tablet computer. In addition, the conversation data processing method described in this embodiment can be applied to Figure 2 In the application scenario shown, the application scenario includes a terminal 21 and a server 22, wherein the server 22 can provide a service similar to an intelligent customer service for the user of the terminal 21, so that the user and the intelligent customer service can have a conversation, and the conversation data generated during the conversation can be stored in the server 22, and the server 22 can process the conversation data using the method described in this embodiment. In some embodiments, the server 22 can also send the key content obtained after processing to the terminal 21. Figure 1 As shown, the specific steps of this method are as follows:

[0035] S101. Obtain at least one dialogue data, each dialogue data including multiple rounds of speech content.

[0036] For example, the server 22 may acquire at least one conversation data in advance, and the at least one conversation data may be generated during the conversation between the intelligent customer service and the user. Alternatively, the at least one conversation data may be acquired by the server 22 from other servers, and the other servers may provide services similar to the intelligent customer service to the user of the terminal 21. Specifically, each conversation data may be a conversation, and the conversation may be a conversation process between the intelligent customer service and the user. Each conversation data may be conversation data in voice form, or may be conversation data in text form converted from conversation data in voice form. In this embodiment, text-based conversation data is taken as an example, and the text-based conversation data may include multiple rounds of speech content, wherein each round of speech content may be a complete speech content of a role (such as an intelligent customer service or a user).

[0037] S102. Determine a cluster identification sequence corresponding to each of the conversation data by clustering multiple rounds of speech content in a first set consisting of at least one round of speech content in each of the conversation data, wherein the cluster identification sequence includes cluster identifications of cluster clusters corresponding to the at least one round of speech content in the conversation data.

[0038] For example, in this embodiment, the server 22 obtains N conversation data, each of which includes multiple rounds of speech content. Each round of speech content can be recorded as a speech, that is, a multi-round conversation can be composed of multiple speech sequences. For example, a multi-round conversation is recorded as d, d = {s 1 ,s 2 ,…,s |d|}, where s i The speech content of the i-th round is the speech of the i-th round. The speech content of the i-th round can be the speech text content of the i-th round. Among them, |d| represents the total number of rounds of speech content in a multi-round dialogue d. For example, a complete speech of a character corresponds to one round.

[0039] Furthermore, at least one round of speech content in each of the N dialogue data can be collected to form a speech set. The speech set can be recorded as the first set. Among them, at least one round of speech content in each dialogue data can be all speech content or part of speech content in each dialogue data. For example, all speech content in each of the N dialogue data is collected to form a total speech set S, Where D is a set consisting of N dialogue data. The speech set S can be recorded as the first set. Alternatively, collect part of the speech content that meets the preset conditions in each dialogue data in the N dialogue data, so as to form the first set S′. Taking the first set S′ as an example, the multiple rounds of speech content included in the first set S′ can be clustered. In the clustering process, for each round of speech content, that is, each speech in the first set S′, each speech can be converted into a representation vector through a sentence vector encoder, and the sentence vector encoder can be any sentence vector model. For example, the i-th speech s in the first set S′ i The corresponding representation vector is v i , v i =Encoder(s i ), where Encoder represents a sentence vector encoder. Further, the representation vector of each speech in the first set S′ forms a speech vector set V = {v i |v i =Encoder(s i )}. Clustering is performed on the speech vector set V. The clustering method can use any common clustering method, such as K-means clustering algorithm (K-Means), machine learning clustering algorithm (HDBSCAN), etc. After clustering, each speech in the first set S' will be assigned to a cluster cluster, and each cluster cluster can correspond to a cluster identifier, which can be a cluster ID. Therefore, after clustering, each speech in the first set S' will correspond to a cluster ID.

[0040] For any conversation d, if part of the speech content in the conversation d is collected in the first set S', then each round of speech content in the part of speech content will correspond to a cluster ID. At this time, the cluster IDs corresponding to each round of speech content in the part of speech content can constitute the cluster identification sequence corresponding to the conversation d. In other words, a conversation d can be converted into a sequence of cluster IDs d', d' = {c 1 ,c 2 ,…,c |d′|}, d′ represents the cluster identifier sequence corresponding to conversation d, where c 1 ,c 2 ,…,c |d′| denote cluster IDs respectively, and |d′| denotes the number of speeches in the conversation d collected into the first set S′.

[0041] Similarly, if all the speeches in the conversation d are collected into the first set S, after clustering all the speeches in the first set S, each round of speeches in the conversation d will correspond to a cluster ID. At this time, the cluster IDs corresponding to each round of speeches in the conversation d can constitute the cluster identification sequence corresponding to the conversation d.

[0042] S103: for the cluster identification sequence corresponding to each conversation data, determine a candidate subsequence in the cluster identification sequence, and determine the support of the candidate subsequence in a second set, where the second set includes the cluster identification sequence corresponding to each conversation data.

[0043] For example, the cluster identification sequences corresponding to each of the N conversation data are used to form a second set, and the second set is recorded as D′.

[0044] D′={{c 11 ,c 12 ,…,c |d1′|},{c 21 ,c 22 ,…,c |d2′|},…,{c N1 ,c N2 ,…,c |dN′|}}, where {c 11 ,c 12 ,…,c |d1′|} represents the cluster identification sequence corresponding to the first conversation data in the N conversation data, and |d1′| represents the number of speech contents in the first conversation data collected into the first set S′. 11 ,c 12 ,…,c |d1′| Respectively represent cluster IDs. 21 ,c 22 ,…,c |d2′|} represents the cluster identifier sequence corresponding to the second conversation data in the N conversation data, and |d2′| represents the number of speech contents in the second conversation data collected into the first set S′. 21 ,c 22 ,…,c |d2′| Respectively represent cluster IDs. N1 ,c N2 ,…,c |dN′|} represents the cluster identification sequence corresponding to the Nth dialogue data in the N dialogue data, and |dN′| represents the number of speech contents in the Nth dialogue data collected into the first set S′. N1 ,c N2 ,…,c |dN′| Respectively represent cluster IDs.

[0045] Optionally, determining the support of the candidate subsequence in the second set includes: taking the number of cluster identifier sequences in the second set that contain the candidate subsequence as the support of the candidate subsequence in the second set. The candidate subsequence in the cluster identifier sequence includes at least two cluster identifiers in the cluster identifier sequence, and the partial order relationship of the at least two cluster identifiers in the cluster identifier sequence is the same as the partial order relationship of the at least two cluster identifiers in the candidate subsequence.

[0046] For each cluster identifier sequence in the second set, determine the candidate subsequence in the cluster identifier sequence. For example, the cluster identifier sequence is abcde, that is, the cluster identifier sequence includes 5 cluster identifiers, each cluster identifier is recorded as an element, and the candidate subsequence is a sequence composed of at least two cluster identifiers randomly selected from the 5 cluster identifiers, and the partial order relationship of the at least two cluster identifiers selected in the cluster identifier sequence is the same as the partial order relationship of the at least two cluster identifiers selected in the candidate subsequence, that is, a new sequence composed of at least two cluster identifiers randomly selected from the 5 cluster identifiers according to the original partial order relationship is called a candidate subsequence or subsequence, for example, ace is a subsequence of abcde. Therefore, each cluster identifier sequence in the second set can correspond to multiple candidate subsequences, and for each candidate subsequence of each cluster identifier sequence, the support of the candidate subsequence in the second set can be counted first, and the support can be the number of cluster identifier sequences in the second set that contain the candidate subsequence. That is, for each candidate subsequence, first count how many cluster identifier sequences in the second set contain the candidate subsequence, and then the support of the candidate subsequence is how many. Further, at least one candidate subsequence whose support is greater than or equal to a threshold is determined. The threshold can be denoted as k, and the value of k can be 5-10. For example, the number of candidate subsequences whose support is greater than or equal to the threshold is denoted as |p|, and the |p| candidate subsequences constitute a subsequence set P, where P = {p = (c 1 ,c 2 ,…,c |p| )|sup(p)>k}, p=(c 1 ,c 2 ,…,c |p| ) represents the candidate subsequence whose support is greater than or equal to the threshold, and sup(p) represents the support of the candidate subsequence p.

[0047] S104 . Calculate the score of each cluster identifier that has appeared in at least one candidate subsequence according to the cohesion corresponding to each candidate subsequence whose support is greater than or equal to the threshold.

[0048] For example, for each candidate subsequence in the subsequence set P, the cohesion of the candidate subsequence is calculated. Further, according to the cohesion of each candidate subsequence in the subsequence set P, the score of each cluster identifier that has appeared in each candidate subsequence in the subsequence set P is calculated. Among them, each cluster identifier that has appeared in each candidate subsequence in the subsequence set P is each cluster identifier that has appeared in the subsequence set P.

[0049] S105: Determine, based on the score, that the cluster corresponding to the cluster identifier that meets a preset condition is the key content.

[0050] For example, each cluster identifier that has appeared in the subsequence set P can be sorted according to the score of each cluster identifier that has appeared in the subsequence set P, for example, sorted from large to small according to the score. The first n cluster identifiers are obtained from the sorting result, and the cluster clusters corresponding to the first n cluster identifiers are used as key content. It can be understood that in this embodiment, each cluster cluster includes multiple speech contents, and the similarity between the multiple speech contents is relatively high. Therefore, when a cluster cluster is used as the key content, it means that the multiple speech contents in the cluster cluster are respectively the key content. In some embodiments, the key content can also be called a key node, or the round corresponding to the key content is recorded as a key node.

[0051] The embodiment of the present disclosure clusters multiple rounds of speech content in the first set formed by at least one round of speech content in each dialogue data, thereby determining the cluster identification sequence corresponding to each dialogue data, and then forming the second set by the cluster identification sequence corresponding to each dialogue data. Further, the candidate subsequences in each cluster identification sequence in the second set are mined, and the support of each candidate subsequence in the second set is calculated. The greater the support, the more times the candidate subsequence appears in the second set, and the cluster clusters corresponding to each cluster identification in the candidate subsequence are more likely to be key content. Further, at least one candidate subsequence with a support greater than or equal to a threshold is selected, and then the cohesion corresponding to the at least one candidate subsequence is calculated according to the support corresponding to the at least one candidate subsequence. Since the cohesion represents the stability of the candidate subsequence, the higher the stability, the more stable the combination formed by each cluster identification in the candidate subsequence. Therefore, when the score of each cluster identification that has appeared in the at least one candidate subsequence is calculated according to the cohesion corresponding to the at least one candidate subsequence, the higher the score, the more times the cluster identification appears in multiple combinations, or the more important the cluster identification is in multiple combinations. Therefore, multiple clusters can be sorted according to the scores, so that the clusters with the highest sorting are used as key contents. Compared with directly using the clusters obtained after clustering as key contents without mining candidate subsequences, this embodiment can more accurately determine the key contents in the conversation data.

[0052] Figure 3 This is a flow chart of a method for processing conversation data provided by another embodiment of the present disclosure. In this embodiment, the specific steps of the method are as follows:

[0053] S301. Obtain at least one dialogue data, each dialogue data including multiple rounds of speech content.

[0054] For example, the server 22 obtains N conversation data, each of which includes multiple rounds of speech content. The sentence vector encoder as described above can convert each round of speech content in each conversation data into a representation vector. In some embodiments, the speech content can also be recorded as a speech text.

[0055] S302: Delete speech contents that are not key contents in each dialogue data, and obtain the at least one round of speech contents in each dialogue data.

[0056] For example, all speeches in each of the N conversation data are collected to form a total speech set S. Further, the total speech set S is recalled. The purpose of the speech recall is to preliminarily screen all speeches in the total speech set S, thereby excluding or deleting some speeches that are believed not to be key content, leaving speeches that may be key content. Among them, the speeches that may be key content are recorded as the recalled speech set, and clustering can be performed on the recalled speech set to obtain multiple clusters. The subsequence mining method described above is further used to determine the cluster of key content from the multiple clusters.

[0057] For example, in some conversation data, the speech content of some rounds is relatively simple, such as "hmm", "ok", "thank you", etc. This type of speech content appears frequently during the conversation, but it is usually not the key content. Therefore, this type of speech content can be recorded as invalid speech. Since the form of these invalid speech is relatively fixed, this embodiment can identify invalid speech through a pre-trained text classification model. The pre-trained text classification model can, for example, be a bidirectional encoder representation based on a transformer (Bidirectional Encoder Representation from Transformers, BERT). For example, each speech content in the total speech set S can be input into the pre-trained text classification model, and the model determines whether to recall the speech content. For example, a certain speech content is recorded as s i , will s i After being input into the pre-trained text classification model, the output of the model is recorded as y i ,y i=Recall(s i ), if y i =1, indicating s i Need to be recalled, that is, s i It is not an ineffective rhetoric. i =0, indicating s i No need to be recalled, i.e. s i Furthermore, the recalled speech content in the total speech set S can be collected to form a new speech set S′, S′={s i |y i =1}, the new speech set S′ here can be the first set S′ as described above.

[0058] S303: Determine a cluster identification sequence corresponding to each of the conversation data by clustering multiple rounds of speech content in a first set consisting of at least one round of speech content in each of the conversation data, wherein the cluster identification sequence includes cluster identifications of cluster clusters corresponding to the at least one round of speech content in the conversation data.

[0059] Specifically, the implementation method and specific principle of S303 and S102 are consistent, and will not be repeated here.

[0060] S304: for the cluster identification sequence corresponding to each conversation data, determine a candidate subsequence in the cluster identification sequence, and determine the support of the candidate subsequence in a second set, where the second set includes the cluster identification sequence corresponding to each conversation data.

[0061] Specifically, the implementation method and specific principle of S304 and S103 are consistent, and will not be repeated here.

[0062] S305 . Calculate the score of each cluster identifier that has appeared in at least one candidate subsequence according to the cohesion corresponding to each candidate subsequence whose support is greater than or equal to the threshold.

[0063] Optionally, the cohesion corresponding to any candidate subsequence in the at least one candidate subsequence is determined according to the support of the any candidate subsequence and the support corresponding to each cluster identifier in the any candidate subsequence.

[0064] For example, after determining the subsequence set P consisting of |p| candidate subsequences whose support is greater than or equal to the threshold, the cohesion of each candidate subsequence p in the subsequence set P whose support is greater than or equal to the threshold can be calculated. The cohesion can also be recorded as a cohesion score, which is recorded as score p (p), score p (p) can be calculated by the following formula (1):

[0065]

[0066] Among them, sup(c 1 ,c 2 ,…,c |p| ) represents the support of the candidate subsequence p, |D| represents the total number of conversation data, such as N. c represents each cluster identifier in the candidate subsequence p, and sup(c) represents the support of the cluster identifier c. Specifically, first count how many cluster identifier sequences in the second set as described above contain the cluster identifier c, then the support of the cluster identifier c is a. In other words, when calculating the support of the cluster identifier c, it is considered whether the cluster identifier c has appeared in a cluster identifier sequence in the second set. If it has appeared, the support of the cluster identifier c is increased by 1.

[0067] Optionally, the score of each cluster identifier that has appeared in the at least one candidate subsequence is calculated based on the cohesion corresponding to each candidate subsequence whose support is greater than or equal to a threshold, including: for each cluster identifier that has appeared in the at least one candidate subsequence, determining each candidate subsequence that includes the cluster identifier in the at least one candidate subsequence; and adding the cohesion of each candidate subsequence that includes the cluster identifier in the at least one candidate subsequence to obtain the score of the cluster identifier.

[0068] For example, after calculating the cohesion score of each candidate subsequence p in the subsequence set P whose support is greater than or equal to the threshold, the score of each cluster identifier that appears in the subsequence set P can also be calculated, where each cluster identifier that appears in the subsequence set P can be recorded as c, c∈C′, Where C′ is the set of all cluster identifiers that appear in the subsequence set P. The score of the cluster identifier c can be recorded as score c (c), score c (c) can be calculated by the following formula (2):

[0069] score c (c) = ∑ {p|c∈p} score p (p) (2)

[0070] Among them, ∑ {p|c∈p} score p (p) represents adding the cohesion scores of each candidate subsequence including the cluster identifier c in the subsequence set P to obtain the score of the cluster identifier c.

[0071] S306: Determine, based on the score, that the cluster corresponding to the cluster identifier that meets a preset condition is the key content.

[0072] Optionally, based on the score, determining the cluster corresponding to the cluster identifier that meets the preset condition as the key content includes: sorting each cluster identifier according to the score to obtain a sorting result; and taking the first preset number of cluster identifiers in the sorting result as the cluster identifier that meets the preset condition.

[0073] For example, after calculating the score of each cluster identifier that appears in the subsequence set P, all cluster identifiers in the subsequence set P are sorted according to the score of each cluster identifier. For example, the higher the cluster identifier is in the sorting, the greater the score of the cluster identifier is. The first n cluster identifiers are obtained from the sorting result, and the clusters corresponding to the first n cluster identifiers are used as key content. The value of n can be calculated by the following formula (3):

[0074] n=|C|·t (3)

[0075] Wherein, C represents the total number of clusters obtained after clustering processing, and t may be a node ratio parameter. For example, t may be adjusted according to business prior knowledge of different business scenarios. Typically, the value of t is 0.1-0.3.

[0076] This embodiment removes the speech contents that are not key contents in the dialogue data in advance, and then clusters the remaining speech contents that may be key contents, thereby eliminating the speech contents that are obviously not key contents. Even if some speech contents appear more times, they may be excluded because they are not key contents. This effectively prevents the speech contents that are obviously not key contents from entering the subsequent steps such as clustering processing and subsequence mining, thereby effectively reducing the impact of non-key contents on the subsequent steps and reducing error propagation, so that the key contents in the dialogue data can be further accurately determined.

[0077] Figure 4 The structure diagram of the conversation data processing device provided by the embodiment of the present disclosure is as follows. The conversation data processing device provided by the embodiment of the present disclosure can execute the processing flow provided by the conversation data processing method embodiment, such as Figure 4 As shown, the conversation data processing device 40 includes:

[0078] An acquisition module 41 is used to acquire at least one dialogue data, each dialogue data includes multiple rounds of speech content;

[0079] A first determining module 42 is used to determine a cluster identification sequence corresponding to each of the dialogue data by clustering multiple rounds of speech content in a first set consisting of at least one round of speech content in each of the dialogue data, wherein the cluster identification sequence includes cluster identifications of cluster clusters corresponding to the at least one round of speech content in the dialogue data;

[0080] A second determination module 43 is used to determine, for each cluster identification sequence corresponding to each conversation data, a candidate subsequence in the cluster identification sequence, and determine the support of the candidate subsequence in a second set, wherein the second set includes the cluster identification sequence corresponding to each conversation data;

[0081] A calculation module 44, configured to calculate the score of each cluster identifier that has appeared in at least one candidate subsequence according to the cohesion corresponding to each candidate subsequence whose support is greater than or equal to the threshold;

[0082] The third determination module 45 is configured to determine, based on the score, that the cluster corresponding to the cluster identifier that meets a preset condition is the key content.

[0083] Optionally, the conversation data processing device 40 further includes: a deleting module 46, which is used to delete the speech content that does not belong to the key content in each conversation data after the acquisition module 41 acquires at least one conversation data, so as to obtain the at least one round of speech content in each conversation data.

[0084] Optionally, when determining the support of the candidate subsequence in the second set, the second determination module 43 is specifically configured to:

[0085] The number of cluster identifier sequences in the second set that contain the candidate subsequence is used as the support of the candidate subsequence in the second set.

[0086] Optionally, the candidate subsequence in the cluster identifier sequence includes at least two cluster identifiers in the cluster identifier sequence, and a partial order relationship between the at least two cluster identifiers in the cluster identifier sequence is the same as a partial order relationship between the at least two cluster identifiers in the candidate subsequence.

[0087] Optionally, the cohesion corresponding to any candidate subsequence in the at least one candidate subsequence is determined according to the support of the any candidate subsequence and the support corresponding to each cluster identifier in the any candidate subsequence.

[0088] Optionally, when the calculation module 44 calculates the score of each cluster identifier that has appeared in at least one candidate subsequence according to the cohesion corresponding to the at least one candidate subsequence whose support is greater than or equal to the threshold, it is specifically used to:

[0089] For each cluster identifier that has appeared in the at least one candidate subsequence, determine each candidate subsequence in the at least one candidate subsequence that includes the cluster identifier;

[0090] The cohesion degree of each candidate subsequence including the cluster identifier in the at least one candidate subsequence is added together to obtain a score of the cluster identifier.

[0091] Optionally, when the third determination module 45 determines that the cluster corresponding to the cluster identifier that meets the preset condition is the key content, it is specifically used to:

[0092] Sort each cluster identifier according to the score to obtain a sorting result;

[0093] The first preset number of cluster identifiers in the sorting result are used as cluster identifiers that meet the preset conditions.

[0094] Figure 4 The conversation data processing device of the illustrated embodiment can be used to execute the technical solution of the above-mentioned method embodiment. Its implementation principle and technical effect are similar and will not be described in detail here.

[0095] The internal functions and structure of the conversation data processing device are described above. The device can be implemented as an electronic device. Figure 5 This is a schematic diagram of the structure of an electronic device embodiment provided by the present disclosure. Figure 5 As shown, the electronic device includes a memory 51 and a processor 52 .

[0096] The memory 51 is used to store programs. In addition to the above-mentioned programs, the memory 51 can also be configured to store various other data to support operations on the electronic device. Examples of these data include instructions for any application or method operating on the electronic device, contact data, phone book data, messages, pictures, videos, etc.

[0097] The memory 51 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk.

[0098] The processor 52 is coupled to the memory 51 and executes the program stored in the memory 51 to:

[0099] Acquire at least one dialogue data, each dialogue data including multiple rounds of speech content;

[0100] Determine a cluster identification sequence corresponding to each of the dialogue data by clustering multiple rounds of speech content in a first set formed by at least one round of speech content in each of the dialogue data, wherein the cluster identification sequence includes cluster identifications of cluster clusters corresponding to the at least one round of speech content in the dialogue data;

[0101] For each cluster identification sequence corresponding to each conversation data, determine a candidate subsequence in the cluster identification sequence, and determine the support of the candidate subsequence in a second set, wherein the second set includes the cluster identification sequence corresponding to each conversation data;

[0102] Calculate the score of each cluster identifier that appears in at least one candidate subsequence according to the cohesion corresponding to each candidate subsequence whose support is greater than or equal to the threshold;

[0103] According to the score, the cluster corresponding to the cluster identifier that meets the preset condition is determined as the key content.

[0104] Further, if Figure 5 As shown, the electronic device may also include: a communication component 53, a power component 54, an audio component 55, a display 56 and other components. Figure 5 Only some components are shown schematically, which does not mean that the electronic device only includes Figure 5 Components shown.

[0105] The communication component 53 is configured to facilitate wired or wireless communication between the electronic device and other devices. The electronic device can access a wireless network based on a communication standard, such as WiFi, 2G or 3G, or a combination thereof. In an exemplary embodiment, the communication component 53 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 53 also includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology and other technologies.

[0106] The power supply component 54 provides power to various components of the electronic device. The power supply component 54 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to the electronic device.

[0107] The audio component 55 is configured to output and / or input audio signals. For example, the audio component 55 includes a microphone (MIC), and when the electronic device is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode, the microphone is configured to receive an external audio signal. The received audio signal can be further stored in the memory 51 or sent via the communication component 53. In some embodiments, the audio component 55 also includes a speaker for outputting audio signals.

[0108] The display 56 includes a screen, which may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touch screen to receive input signals from a user. The touch panel includes one or more touch sensors to sense touch, slide, and gestures on the touch panel. The touch sensor may not only sense the boundary of a touch or slide action, but also detect the duration and pressure associated with the touch or slide operation.

[0109] In addition, an embodiment of the present disclosure further provides a computer-readable storage medium on which a computer program is stored. The computer program is executed by a processor to implement the conversation data processing method described in the above embodiment.

[0110] It should be noted that, in this article, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the existence of other identical elements in the process, method, article or device including the elements.

[0111] The above description is only a specific embodiment of the present disclosure, so that those skilled in the art can understand or implement the present disclosure. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present disclosure. Therefore, the present disclosure will not be limited to the embodiments described herein, but will conform to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for processing conversation data, wherein: The method comprises: Acquire at least one dialogue data, each dialogue data including multiple rounds of speech content; Determine a cluster identification sequence corresponding to each of the dialogue data by clustering multiple rounds of speech content in a first set formed by at least one round of speech content in each of the dialogue data, wherein the cluster identification sequence includes cluster identifications of cluster clusters corresponding to the at least one round of speech content in the dialogue data; For each cluster identification sequence corresponding to each conversation data, determine a candidate subsequence in the cluster identification sequence, and determine the support of the candidate subsequence in a second set, wherein the second set includes the cluster identification sequence corresponding to each conversation data; Calculate the score of each cluster identifier that has appeared in the at least one candidate subsequence according to the cohesion corresponding to each candidate subsequence whose support is greater than or equal to the threshold, wherein the cohesion corresponding to any candidate subsequence in the at least one candidate subsequence is determined according to the support of the any candidate subsequence and the support corresponding to each cluster identifier in the any candidate subsequence; According to the score, determining the cluster corresponding to the cluster identifier that meets the preset condition as the key content; The step of determining the support of the candidate subsequence in the second set includes: taking the number of cluster identifier sequences in the second set that contain the candidate subsequence as the support of the candidate subsequence in the second set.

2. The method according to claim 1, wherein: After obtaining at least one conversation data, the method further includes: The speech contents that are not key contents in each dialogue data are deleted to obtain the at least one round of speech contents in each dialogue data.

3. The method according to claim 1, wherein: The candidate subsequence in the cluster identifier sequence includes at least two cluster identifiers in the cluster identifier sequence, and the partial order relationship between the at least two cluster identifiers in the cluster identifier sequence is the same as the partial order relationship between the at least two cluster identifiers in the candidate subsequence.

4. The method according to claim 1, wherein: Calculating the score of each cluster identifier that appears in at least one candidate subsequence according to the cohesion corresponding to each candidate subsequence whose support is greater than or equal to the threshold, including: For each cluster identifier that has appeared in the at least one candidate subsequence, determine each candidate subsequence in the at least one candidate subsequence that includes the cluster identifier; The cohesion degree of each candidate subsequence including the cluster identifier in the at least one candidate subsequence is added together to obtain a score of the cluster identifier.

5. The method according to claim 1, wherein: According to the score, determining the cluster corresponding to the cluster identifier that meets the preset condition as the key content includes: Sort each cluster identifier according to the score to obtain a sorting result; The first preset number of cluster identifiers in the sorting result are used as cluster identifiers that meet the preset conditions.

6. A conversation data processing device, wherein: include: An acquisition module, used to acquire at least one dialogue data, each dialogue data includes multiple rounds of speech content; A first determination module is used to determine a cluster identification sequence corresponding to each of the dialogue data by clustering multiple rounds of speech content in a first set consisting of at least one round of speech content in each of the dialogue data, wherein the cluster identification sequence includes cluster identifications of cluster clusters corresponding to the at least one round of speech content in the dialogue data; A second determination module is used to determine, for each cluster identification sequence corresponding to each conversation data, a candidate subsequence in the cluster identification sequence, and determine the support of the candidate subsequence in a second set, wherein the second set includes the cluster identification sequence corresponding to each conversation data; a calculation module, configured to calculate the score of each cluster identifier that has appeared in the at least one candidate subsequence according to the cohesion corresponding to each candidate subsequence whose support is greater than or equal to a threshold, wherein the cohesion corresponding to any candidate subsequence in the at least one candidate subsequence is determined according to the support of the any candidate subsequence and the support corresponding to each cluster identifier in the any candidate subsequence; A third determination module is used to determine, based on the score, the cluster corresponding to the cluster identifier that meets the preset condition as the key content; The second determination module is further used to determine the support of the candidate subsequence in the second set by executing the following steps: taking the number of cluster identification sequences in the second set that contain the candidate subsequence as the support of the candidate subsequence in the second set.

7. An electronic device, wherein: include: Memory; processor; as well as Computer programs; The computer program is stored in the memory and is configured to be executed by the processor to implement the method according to any one of claims 1 to 5.

8. A computer-readable storage medium having a computer program stored thereon, wherein: When the computer program is executed by a processor, the method according to any one of claims 1 to 5 is implemented.

Citation Information

Patent Citations

  • Data mining method and device, server and readable storage medium

    CN111401388A

  • Data matching method and device, electronic equipment and storage medium

    CN111858869A