Topic label extraction method and related device
Through the splitting and word segmentation of dialogue text data, the topic tags are automatically determined, which solves the problem of inefficiency in the existing technology and realizes the establishment of an efficient topic tag system.
Patent Information
- Application Number
- CN202211661569.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-23
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2042-12-23
AI Technical Summary
In the field of telephone customer service, the existing technology needs to rely on manual experience to establish a topic tag system, which is inefficient and cannot efficiently process massive historical call text.
Through dialogue text data, split and word segmentation, calculate the frequency of univariate phrases, determine the hot topic groups, and determine whether the text segment contains these phrases, automatically determine the first-level topic tags, and further refine the tags in combination with binary phrases and clustering algorithms.
It realizes efficient topic tag extraction without manual participation, and improves the efficiency of establishing a topic tag system.
Smart Images

Figure CN115859959B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data processing, and particularly to a method for extracting topic tags and related devices. Background Art
[0002] In the field of telephone customer service, before designing the conversation process and conversation knowledge base of the manual outbound call or robot outbound call system, in order to effectively classify and identify the intent of the call process, it is first necessary to establish a topic tag system.
[0003] The traditional method for establishing a topic tag system needs to rely on manual experience. For a large amount of historical call texts, a large amount of manpower is invested in the extraction, merging, and hierarchical definition operations of topic tags, resulting in low efficiency. Summary of the Invention
[0004] In view of the above problems, the present invention provides a method for extracting topic tags and related devices that overcome the above problems or at least partially solve the above problems.
[0005] In a first aspect, a method for extracting topic tags includes:
[0006] Splitting the pre-obtained conversation text data to obtain multiple text segments;
[0007] Performing word segmentation on each of the text segments to obtain multiple single-word phrases;
[0008] Calculating the frequency of occurrence of each of the single-word phrases in the text segment;
[0009] Determining a hot topic group corresponding to the conversation text data according to the frequency of occurrence of each of the single-word phrases, where the hot topic group includes multiple single-word phrases;
[0010] For any one of the text segments, determining whether the text segment includes at least one of the single-word phrases in the hot topic group;
[0011] If it includes, determining the single-word phrases in the hot topic group included in the text segment as the first-level topic tags corresponding to the text segment.
[0012] In combination with the first aspect, in some optional embodiments, before the step of determining the hot topic group corresponding to the conversation text data according to the frequency of occurrence of each of the single-word phrases, the method further includes:
[0013] Freely combining each of the single-word phrases to obtain multiple two-word phrases, where the two-word phrases are composed of any two non-repeating single-word phrases;
[0014] For any of the above-mentioned binary phrases, the number of text segments that simultaneously include the two single phrases that form the binary phrase is calculated;
[0015] Determining the hot topic group corresponding to the dialogue text data according to the frequencies of occurrence of the single phrases includes:
[0016] Determining the hot topic group according to the frequencies of occurrence of the single phrases and the quantities corresponding to the binary phrases, wherein the hot topic group includes multiple single phrases and multiple binary phrases.
[0017] Combined with the previous embodiment, in some alternative embodiments, determining the hot topic group according to the frequencies of occurrence of the single phrases and the quantities corresponding to the binary phrases includes:
[0018] Jointly determining the top N single phrases with larger frequencies and the top M binary phrases with larger quantities as the hot topic group, where N and M are both integers greater than 0;
[0019] Alternatively, jointly determining the single phrases with frequencies greater than a preset frequency threshold and the binary phrases with quantities greater than a preset quantity threshold as the hot topic group.
[0020] Combined with the previous embodiment, in some alternative embodiments, determining whether any of the text segments includes at least one of the single phrases in the hot topic group includes:
[0021] For any of the text segments, it is determined whether the text segment includes at least one of the single phrases in the hot topic group and / or at least one of the binary phrases in the hot topic group;
[0022] If it includes, then determining the single phrases in the hot topic group included in the text segment as the corresponding first-level topic labels of the text segment includes:
[0023] If it includes, then determining the single phrases and / or binary phrases in the hot topic group included in the text segment as the corresponding first-level topic labels of the text segment.
[0024] Combined with the previous embodiment, in some alternative embodiments, after determining the single phrases and / or binary phrases in the hot topic group included in the text segment as the corresponding first-level topic labels of the text segment, the method further includes:
[0025] Setting the corresponding first-level topic labels on the corresponding text segments.
[0026] Combined with the previous embodiment, in some alternative embodiments, after setting the corresponding first-level topic tags on the corresponding text segments, the method further includes:
[0027] Extract G text segments with target tags from the multiple text segments, where the target tags are the first-level topic tags in the hot topic group, and G is an integer greater than 0;
[0028] After vectorizing the G text segments, input them into a pre-established clustering algorithm, so as to cluster the G text segments into K groups, where each group includes at least one of the G text segments, and K is an integer greater than 0 and less than G;
[0029] For any one of the K groups, extract the corresponding second-level topic tags through an algorithm for extracting keywords.
[0030] Combined with the previous embodiment, in some alternative embodiments, the step of extracting the corresponding second-level topic tags through an algorithm for extracting keywords for any one of the K groups includes:
[0031] For any one of the K groups, obtain a first word group according to the unigrams of the corresponding text segments, where the first word group includes multiple unigrams;
[0032] Calculate the first frequency of each unigram in the first word group appearing in the first word group respectively;
[0033] For any one of the unigrams in the first word group, calculate the corresponding second frequency, where the numerator of the second frequency is the number of groups in the K groups involving the corresponding unigram, and the denominator of the second frequency is K;
[0034] For any one of the unigrams in the first word group, calculate the quotient of the corresponding first frequency and the second frequency;
[0035] Determine the first J unigrams with larger quotients in the first word group as the second-level topic tags, where J is an integer greater than 0.
[0036] Optionally, in some alternative embodiments, after determining the unigrams and / or bigrams in the hot topic group included in the text segment as the corresponding first-level topic tags of the text segment, the method further includes:
[0037] Input the first-level topic label into a pre-established neural network model to obtain the extended topic label corresponding to the input first-level topic label output by the neural network model, where the similarity between the extended topic label and the input first-level topic label is greater than a preset similarity threshold.
[0038] In a second aspect, a topic label extraction device includes: a segmentation unit, a word segmentation unit, a word frequency unit, a hot phrase unit, a judgment unit, and a first-level label unit;
[0039] The segmentation unit is configured to split pre-obtained dialogue text data to obtain multiple text segments;
[0040] The word segmentation unit is configured to perform word segmentation on each of the text segments to obtain multiple unigrams;
[0041] The word frequency unit is configured to calculate the frequency of occurrence of each of the unigrams in each of the text segments;
[0042] The hot phrase unit is configured to determine a hot topic group corresponding to the dialogue text data according to the frequency of occurrence of each of the unigrams, where the hot topic group includes multiple unigrams;
[0043] The judgment unit is configured to, for any one of the text segments, judge whether at least one of the unigrams in the hot topic group is included in the text segment;
[0044] The first-level label unit is configured to, if included, determine the unigrams in the hot topic group included in the text segment as the first-level topic label of the corresponding text segment.
[0045] In a third aspect, a computer-readable storage medium stores a program, and when the program is executed by a processor, the topic label extraction method described in any one of the above is implemented.
[0046] In a fourth aspect, an electronic device includes at least one processor, at least one memory connected to the processor, and a bus; wherein, the processor and the memory communicate with each other through the bus; the processor is configured to call program instructions in the memory to execute the topic label extraction method described in any one of the above.
[0047] With the above technical solution, the topic label extraction method and related device provided by the present invention can split the pre-obtained conversation text data to obtain multiple text segments; perform word segmentation on each of the text segments to obtain multiple single-word phrases; calculate the frequency of occurrence of each of the single-word phrases in each of the text segments; determine the hot topic group corresponding to the conversation text data according to the frequency of occurrence of each of the single-word phrases, where the hot topic group includes multiple of the single-word phrases; for any one of the text segments, determine whether at least one of the single-word phrases in the hot topic group is included in the text segment; if so, determine the single-word phrases in the hot topic group included in the text segment as the first-level topic labels of the corresponding text segment. It can be seen from this that the present invention can process conversation text data to automatically determine topic labels without manual participation and has high efficiency.
[0048] The above description is only an overview of the technical solution of the present invention. In order to be able to understand the technical means of the present invention more clearly, it can be implemented according to the content of the description. And in order to make the above and other objects, features and advantages of the present invention more obvious and understandable, the following specifically illustrates the specific embodiments of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] By reading the following detailed description of the preferred embodiments, various other advantages and benefits will become clear to those of ordinary skill in the art. The drawings are only for the purpose of showing the preferred embodiments and are not considered to be a limitation of the present invention. And throughout the drawings, the same reference numerals are used to represent the same components. In the drawings:
[0050] Figure 1 The flowchart of the first topic label extraction method provided by the present invention is shown;
[0051] Figure 2 The flowchart of the second topic label extraction method provided by the present invention is shown;
[0052] Figure 3 The flowchart of the third topic label extraction method provided by the present invention is shown;
[0053] Figure 4 The flowchart of the fourth topic label extraction method provided by the present invention is shown;
[0054] Figure 5 The flowchart of the fifth topic label extraction method provided by the present invention is shown;
[0055] Figure 6 The flowchart of the sixth topic label extraction method provided by the present invention is shown;
[0056] Figure 7Shows the flowchart of the seventh topic tag extraction method provided by the present invention;
[0057] Figure 8 Shows the flowchart of the eighth topic tag extraction method provided by the present invention;
[0058] Figure 9 Shows the structural schematic diagram of a topic tag extraction device provided by the present invention;
[0059] Figure 10 Shows the structural schematic diagram of an electronic device provided by the present invention. Detailed implementation manners
[0060] Hereinafter, exemplary embodiments of the present disclosure will be described in more detail with reference to the accompanying drawings. Although the exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present disclosure can be more thoroughly understood and the scope of the present disclosure can be fully conveyed to those skilled in the art.
[0061] As Figure 1 shown, the present invention provides a topic tag extraction method, including: S100, S200, S300, S400, S500 and S600;
[0062] S100. Split the pre-obtained conversation text data to obtain multiple text segments;
[0063] Optionally, the execution subject of the present invention can pre-obtain the original conversation text data. For example, obtain the voice data between the customer service and the customer, and convert the voice data into conversation text data. The present invention does not limit this.
[0064] Optionally, taking the conversion of voice data into conversation text data as an example above, a complete call record includes the voice data of the customer and the customer. Therefore, in order to accurately extract the corresponding topic tags subsequently, the present invention can distinguish the voice data of the customer service and the customer, that is, split the conversation text data to obtain multiple text segments. For example, generally speaking, in the conversation text data, there is an AB flag bit, where "A" represents the data of the customer service, and "B" represents the data of the customer. For example, after the customer service finishes speaking a paragraph each time, the system can automatically set an "A" after converting the voice data spoken by the customer service this time into text. Then, after the customer finishes speaking a paragraph, the system can automatically set a "B" after converting the voice data spoken by the customer this time into text, and so on alternately. Therefore, the present invention can split the conversation text data into multiple text segments according to "A" and "B", and each text segment corresponds to a paragraph spoken by the customer service or a paragraph spoken by the customer. That is, the multiple text segments include multiple text segments carrying "A" and multiple text segments carrying "B", and the present invention does not limit this.
[0065] Optionally, after splitting to obtain multiple text segments, the present invention can subsequently perform word segmentation on each text segment separately to facilitate determining the corresponding topic tags for each text segment, and the present invention does not limit this.
[0066] Optionally, in order to improve the efficiency of the present invention, the present invention can filter the text segments that do not carry "A" or "B". For example, for the text segments of some outbound call records, the present invention can delete them to achieve filtering, and the present invention does not limit this.
[0067] S200. Perform word segmentation on each of the text segments to obtain a plurality of single-word phrases;
[0068] Optionally, taking the conversion of voice data into conversation text data as an example above, since a complete call record not only includes the voice data of the customer and the customer, but may also include voice data such as system prompt tones. Therefore, the conversation text data naturally also includes the text data corresponding to what the customer said, the text data corresponding to what the customer service said, and the text data corresponding to the system prompt tone, and even includes the text data corresponding to some tone voices of the customer and the customer service (for example, "um", "ah", and "ya", etc.).
[0069] For this reason, the present invention can obtain a plurality of single-word phrases through word segmentation, that is, obtain each single-word phrase involved in the entire conversation text data, so as to facilitate filtering some single-word phrases irrelevant to the topic tags subsequently.
[0070] Optionally, the present invention does not limit the word segmentation process, and any feasible method belongs to the protection scope of the present invention. For example, jieba (an excellent third-party Chinese word segmentation library for Python) can be used for word segmentation. At the same time, in order to improve the accuracy of word segmentation, the present invention can supplement some characteristic professional thesauruses in the field, that is, perform word segmentation based on jieba and the professional thesaurus, and the present invention does not limit this.
[0071] Optionally, in order to further improve the efficiency and accuracy of the subsequent steps of the present invention, after obtaining a plurality of single-word phrases, the present invention can perform corresponding stop word filtering, that is, filter some irrelevant words, such as some common words, modal words, and punctuation marks.
[0072] S300. Calculate the frequency of occurrence of each of the single-word phrases in each of the text segments;
[0073] Optionally, as described above, corresponding topic tags can be determined for each text segment subsequently. Therefore, the present invention can first determine the set of common topic tags (hot topic group) for all text segments, so as to facilitate determining the topic tags of each text segment from the set of topic tags subsequently. For this purpose, the present invention can calculate the frequency of occurrence of each single-word phrase in the entire dialogue text data (which can be the text segments of the filtered outbound call records and the data after filtering stop words).
[0074] Optionally, the present invention does not limit the method for calculating the frequency of occurrence of each single-word phrase. For example, the present invention can preset a fixed denominator, and then use the number of occurrences of each single-word phrase as the numerator respectively, so as to obtain the frequency of occurrence of the corresponding single-word phrase respectively. It should be noted that: the specific data of the denominator can be set according to actual needs, and generally a relatively large value can be selected to avoid the situation where the numerator is larger than the denominator, and the present invention does not limit this.
[0075] S400. Determine the hot topic group corresponding to the dialogue text data according to the frequency of occurrence of each of the single-word phrases;
[0076] Wherein, the hot topic group includes a plurality of the single-word phrases;
[0077] Optionally, the present invention can arrange each single-word phrase in sequence according to the size of the frequency, and obtain the hot topic group therefrom. For example, the present invention can select all single-word phrases with a frequency greater than a certain threshold (a specific value is set according to needs) to form the hot topic group, or can select the top N (a specific value is set according to needs) single-word phrases with the largest frequency to form the hot topic group, and the present invention does not limit this.
[0078] It should be noted that: the hot topic group includes a plurality of non-repeating single-word phrases, and the present invention does not limit this.
[0079] Optionally, as described above, the hot topic group includes a plurality of non-repeating single-word phrases, which has certain limitations. To further improve the efficiency of the present invention, some combined words can be added to the present invention, that is, two-word phrases obtained by combining two single-word phrases together.
[0080] For example, as Figure 2 shown, in combination with Figure 1 the embodiment shown, in some optional embodiments, before the S400, the method further includes: Step 1.1 and Step 1.2;
[0081] Step 1.1: Freely combine each of the single-word phrases to obtain a plurality of two-word phrases, where the two-word phrases are composed of any two non-repeating single-word phrases;
[0082] Step 1.2: For any one of the two-word phrases, calculate the number of text segments that simultaneously include the two single-word phrases that make up the two-word phrase;
[0083] Optionally, for the sake of clearly describing this embodiment, the present invention takes the example of freely combining to obtain 100 two-word phrases, and each two-word phrase is composed of two non-repeating single-word phrases. Subsequently, a part needs to be selected from the 100 two-word phrases to be used for constructing the hot topic group, so it is necessary to calculate the weight of each two-word phrase as a hot topic, that is, calculate the number of text segments that simultaneously include the two single-word phrases that make up the two-word phrase. Taking a total of 50 text segments obtained this time as an example, for any one two-word phrase, it is necessary to calculate how many of the 50 text segments simultaneously include the two single-word phrases of this two-word phrase, so as to obtain the weight of this two-word phrase as a hot topic. The larger the number, the greater the weight, and the present invention does not limit this.
[0084] The S400 includes: Step 1.3;
[0085] Step 1.3: Determine the hot topic group according to the frequencies of occurrence of each single-word phrase and the quantities corresponding to each two-word phrase;
[0086] Among them, the hot topic group includes a plurality of the single-word phrases and a plurality of the two-word phrases.
[0087] Optionally, as described above, a part of the single-word phrases and a part of the two-word phrases can be selected to construct the hot topic group. For example, as Figure 3 shown, in combination with the previous embodiment, in some optional embodiments, the Step 1.3 includes: Step 1.31;
[0088] Step 1.31: jointly determine the top N single-word phrases with relatively high frequencies and the top M two-word phrases with relatively large quantities as the hot topic group, or jointly determine the single-word phrases with frequencies greater than a preset frequency threshold and the two-word phrases with quantities greater than a preset quantity threshold as the hot topic group, where both N and M are integers greater than 0.
[0089] Optionally, the present invention does not limit N, M, the preset frequency threshold, and the preset quantity threshold, which can be set according to actual needs.
[0090] Optionally, referring to the processing process of two-word phrases, the present invention can also construct three-word phrases (formed by combining 3 single-word phrases), and select a part of the three-word phrases to jointly construct a hot topic group with single-word phrases and two-word phrases. The present invention does not limit this.
[0091] S500: For any one of the text segments, determine whether at least one of the single-word phrases in the hot topic group is included in the text segment;
[0092] If included, execute S600;
[0093] Optionally, as described above, the hot topic group is a set of common topic labels for all text segments. For any text segment, the corresponding topic label can be determined from the hot topic group. Taking the hot topic group including single-word phrase A, single-word phrase B, and single-word phrase C as an example, if only single-word phrase A is included in text segment 1, then single-word phrase A is used as the first-level topic label of text segment 1; if both single-word phrase A and single-word phrase B are included in text segment 2, then single-word phrase A and single-word phrase B are used as the two first-level topic labels of text segment 2. The present invention does not limit this.
[0094] Of course, in combination with the foregoing embodiments, if the hot topic group also includes two-word phrases, corresponding judgments also need to be made. For example, if the hot topic group also includes two-word phrase A (formed by combining single-word phrase A and single-word phrase 4) and two-word phrase B (formed by combining single-word phrase A and single-word phrase B), then there is no first-level topic label of two-word phrases in the above text segment 1, and two-word phrase B can be used as one of the first-level topic labels of the above text segment 2. The present invention does not limit this.
[0095] That is, as Figure 4 shown, in combination with the previous embodiment, in some optional embodiments, S500 includes: Step 2.1;
[0096] Step 2.1: For any one of the text segments, determine whether at least one of the single-word phrases in the hot topic group and / or at least one of the two-word phrases in the hot topic group is included in the text segment;
[0097] The S600 includes: step 2.2;
[0098] Step 2.2: Determine the single-word phrases and / or two-word phrases in the hot topic groups included in the text segment as the first-level topic tags of the corresponding text segment.
[0099] Optionally, if the text segment only includes the single-word phrases in the hot topic groups, determine the single-word phrases in the hot topic groups included in the text segment as the first-level topic tags of the corresponding text segment; if the text segment only includes the two-word phrases in the hot topic groups, determine the two-word phrases in the hot topic groups included in the text segment as the first-level topic tags of the corresponding text segment; if the text segment includes both the single-word phrases and two-word phrases in the hot topic groups (except for the case where the single-word phrase and the two-word phrase do not have an inclusion relationship, such as the single-word phrase being "credit card" and the two-word phrase being "credit card - cash withdrawal"), determine the single-word phrases and two-word phrases in the hot topic groups included in the text segment as the first-level topic tags of the corresponding text segment.
[0100] S600: Determine the single-word phrases in the hot topic groups included in the text segment as the first-level topic tags of the corresponding text segment.
[0101] Optionally, as Figure 5 shown, in combination with Figure 1 the embodiments shown, in some optional embodiments, after step 2.2, the method further includes: step 3.1;
[0102] Step 3.1: Set the corresponding first-level topic tags on the corresponding text segments.
[0103] Optionally, after determining the topic tags of each text segment, the topic tags of each text segment can be associated with the corresponding text segment, that is, set the corresponding topic tags on the text segment for subsequent viewing and use, and the present invention does not limit this.
[0104] Optionally, the foregoing obtained are first-level topic tags. Based on the foregoing process, the present invention can further obtain more refined second-level topic tags, third-level topic tags, fourth-level topic tags, etc. Among them, the second-level topic tags are tags obtained by further refining the first-level topic tags, the third-level topic tags are tags obtained by further refining the second-level topic tags, the fourth-level topic tags are tags obtained by further refining the third-level topic tags, and so on in a cycle until the number of topic tags meets the threshold condition. For example, set the upper limit of the number of topic tags to 1000, and the present invention does not limit this.
[0105] That is, as Figure 6As shown in combination with the previous embodiment, in some alternative embodiments, after step 3.1, the method further includes: step 4.1, step 4.2, and step 4.3;
[0106] Step 4.1: Extract G text segments with a target label from the multiple text segments, where the target label is a first-level topic label in the hot topic group, and G is an integer greater than 0;
[0107] Optionally, the target label mentioned in the present invention is the first-level topic label that needs to be refined this time. By performing subsequent processing on the text segments with the target label, the second-level topic labels of each text segment under this target label can be obtained, and the present invention does not limit this.
[0108] Step 4.2: After vectorizing the G text segments, input them into a pre-established clustering algorithm, so as to cluster the G text segments into K groups, where each group includes at least one of the G text segments, and K is an integer greater than 0 and less than G;
[0109] Optionally, text vectorization belongs to the well-known technology in the art, and the present invention does not describe it in detail. It should be noted that: text vectorization can first construct a dictionary library containing all words, and then assign "0" or "1" according to whether the words in the dictionary library exist in the text, and perform text vectorization accordingly. For example, there are 5 words in the dictionary library: "CCB", "credit card", "installment", "repayment", and "application", then the text vectorization result of a text segment "Credit card repayment requires real-name application by the applicant" is: 110001, and the present invention does not limit this.
[0110] Optionally, the clustering algorithm is an unsupervised learning algorithm, and the present invention can use the k-means algorithm. First, set the value of K, and then based on the distance between text vectors (the result after text vectorization) (the distance of text vectors can be calculated by multiplying and adding the corresponding bits of the vectors) and the iterative algorithm for finding the center point, cluster K classes with similar distances. The silhouette coefficient is a general index for evaluating the quality of clustering later. The larger the silhouette coefficient, the better the clustering effect, and the best value of K can be obtained through experimental comparison. Among them, the K classes with similar distances correspond to K groups, and the present invention does not limit this.
[0111] Step 4.3: For any one of the K groups, extract the corresponding second-level topic label through an algorithm for extracting keywords.
[0112] Optionally, the algorithm for extracting keywords belongs to the well-known technology in the art, and the present invention will not describe it in detail. It should be noted that in the algorithm for extracting keywords of the present invention, a weighting technique is adopted. Its main idea is that if a word or phrase appears frequently in an article and rarely appears in other articles, then this word or phrase is considered to have good category discrimination ability and is suitable for describing the article. Based on this assumption, two indicators are statistically calculated for all words. The A indicator is the frequency of the word in this article, and the B indicator is the frequency of the word in all articles. The A indicator or the B indicator is used as the keyword weight of the word, and the words with the top-ranked weights under each topic are selected as the keyword group representing the topic. For example, the representative keyword group for the secondary topic "Credit card cancellation - Points" is "Points, Clearing", and the present invention does not limit this.
[0113] Finally, if it is found that the keywords of two topics are exactly the same, the two topics are merged to obtain the deduplicated topics and the related tag system.
[0114] For example, as Figure 7 shown, in combination with the previous embodiment, in some optional embodiments, step 4.3 includes: step 4.31, step 4.32, step 4.33, step 4.34 and step 4.35;
[0115] Step 4.31: For any one of the K groups, a first word group is obtained according to the single-word groups of the corresponding text segments;
[0116] Wherein, the first word group includes multiple single-word groups;
[0117] Step 4.32: Calculate the first frequency of each single-word group in the first word group appearing in the first word group respectively;
[0118] Step 4.33: For any one of the single-word groups in the first word group, a corresponding second frequency is calculated;
[0119] Wherein, the numerator of the second frequency is the number of groups in the K groups involving the corresponding single-word group, and the denominator of the second frequency is the K;
[0120] Step 4.34: For any one of the single-word groups in the first word group, calculate the quotient of the corresponding first frequency and the second frequency;
[0121] Step 4.35: Determine the first J single-word groups with larger quotients in the first word group as the secondary topic tags;
[0122] Wherein, the J is an integer greater than 0.
[0123] Optionally, in order to further improve the efficiency of the present invention, the present invention may also determine extended topic tags such as first-level topic tags, second-level topic tags, and third-level topic tags, which can also be understood as determining synonyms or near-synonyms of each topic tag. The present invention does not limit this.
[0124] Optionally, as Figure 8 shown, in some alternative embodiments, after step 2.2, the method further includes: step 5.1;
[0125] Step 5.1: Input the first-level topic tag into a pre-established neural network model, so as to obtain the extended topic tag corresponding to the first-level topic tag output by the neural network model;
[0126] Wherein, the similarity between the extended topic tag and the input first-level topic tag is greater than a preset similarity threshold.
[0127] That is, based on a large amount of historical conversation text data (not limited to the customer service conversation field), the present invention can use a deep learning algorithm to train a word vector neural network model (in the present invention, the principle of Continuous Bag-of-Words can be adopted). That is, the training input is the word vector corresponding to the words related to the context of a certain feature word, and the output is the word vector of this specific word. Then, through the similarity between word vectors, various expressions of the feature word are found and added to the keyword group of the topic tag to realize the recognition of synonyms, near-synonyms, and misspelled words of the tag.
[0128] When calculating the similarity between word vectors, the cosine similarity algorithm is used. That is, the Cosine cosine similarity algorithm is used to determine the magnitude of the difference between two word vectors. Compared with distance metrics, cosine similarity pays more attention to the difference in the direction of two word vectors rather than the distance or length.
[0129] The calculation method of the cosine similarity algorithm is to transform the problem into a point in an n-dimensional coordinate system. By connecting this point with the origin of the coordinate system to form a straight line (vector), the similarity value between two problems is the cosine value of the angle between the two straight lines (vectors).
[0130] Since the straight lines connecting the points representing the problems to the origin will intersect at the origin, the smaller the angle, the more similar the two problem expressions are, and the larger the angle, the smaller the similarity of the two problem expressions.
[0131] Taking the credit card cancellation retention business scenario as an example, the similar expressions of the keyword "cancellation" in the present invention include "card cancellation", "cancellation", "account closure", "card suspension", and "service suspension", etc. Based on this, the keyword group containing "cancellation" is expanded.
[0132] Such asFigure 9 As shown in the figure, the present invention provides a topic tag extraction device, comprising: a segmentation unit 100, a word segmentation unit 200, a word frequency unit 300, a hot phrase unit 400, a judgment unit 500, and a first-level tag unit 600;
[0133] The segmentation unit 100 is configured to split the pre-obtained dialogue text data to obtain multiple text segments;
[0134] The word segmentation unit 200 is configured to perform word segmentation on each of the text segments to obtain multiple single-word phrases;
[0135] The word frequency unit 300 is configured to calculate the frequency of occurrence of each of the single-word phrases in each of the text segments;
[0136] The hot phrase unit 400 is configured to determine a hot topic group corresponding to the dialogue text data according to the frequency of occurrence of each of the single-word phrases;
[0137] Wherein, the hot topic group includes multiple of the single-word phrases;
[0138] The judgment unit 500 is configured to, for any one of the text segments, determine whether at least one of the single-word phrases in the hot topic group is included in the text segment. If so, the first-level tag unit 600 is triggered;
[0139] The first-level tag unit 600 is configured to determine the single-word phrases in the hot topic group included in the text segment as the first-level topic tags corresponding to the text segment.
[0140] Combined Figure 9 With the embodiments shown, in some alternative embodiments, the device further comprises: a two-word phrase unit and a quantity calculation unit;
[0141] The two-word phrase unit is configured to freely combine each of the single-word phrases to obtain multiple two-word phrases before determining the hot topic group corresponding to the dialogue text data according to the frequency of occurrence of each of the single-word phrases, wherein the two-word phrase is composed of any two non-repeating single-word phrases;
[0142] The quantity calculation unit is configured to calculate, for any one of the two-word phrases, the quantity of the text segments that simultaneously include the two single-word phrases that form the two-word phrase;
[0143] The hot phrase unit 400 includes: a first hot phrase sub-unit;
[0144] The first hot phrase sub-unit is configured to determine the hot topic group according to the frequency of occurrence of each of the single-word phrases and the quantity corresponding to each of the two-word phrases;
[0145] Among them, the hot topic group includes a plurality of the one-word phrases and a plurality of the two-word phrases.
[0146] Combined with the previous embodiment, in some alternative embodiments, the first hot phrase sub-unit includes: a second hot phrase sub-unit;
[0147] The second hot phrase sub-unit is configured to jointly determine the top N one-word phrases with larger frequencies and the top M two-word phrases with larger quantities as the hot topic group, where both N and M are integers greater than 0;
[0148] Alternatively, jointly determine the one-word phrases with frequencies greater than a preset frequency threshold and the two-word phrases with quantities greater than a preset quantity threshold as the hot topic group.
[0149] Combined with the previous embodiment, in some alternative embodiments, the determination unit 500 includes: a first determination sub-unit;
[0150] The first determination sub-unit is configured to, for any one of the text segments, determine whether at least one of the one-word phrases in the hot topic group and / or at least one of the two-word phrases in the hot topic group are included in the text segment. If included, trigger the first first-level label sub-unit;
[0151] The first-level label unit 600 includes: a first first-level label sub-unit;
[0152] The first first-level label sub-unit is configured to determine the one-word phrases and / or two-word phrases in the hot topic group included in the text segment as the first-level topic labels of the corresponding text segment.
[0153] Combined Figure 9 In the shown embodiment, in some alternative embodiments, the apparatus further includes: a label setting unit;
[0154] The label setting unit is configured to, after determining the one-word phrases in the hot topic group included in the text segment as the first-level topic labels of the corresponding text segment, set the corresponding first-level topic labels on the corresponding text segment.
[0155] Combined with the previous embodiment, in some alternative embodiments, the apparatus further includes: a text segment extraction unit, a clustering unit, and a second-level label unit;
[0156] A text segment extraction unit, configured to extract G text segments with a target tag from the multiple text segments after setting the corresponding first-level topic tag on the corresponding text segment, where the target tag is the first-level topic tag in the hot topic group, and G is an integer greater than 0;
[0157] A clustering unit, configured to vectorize the G text segments and then input them into a pre-established clustering algorithm, so as to cluster the G text segments into K groups, where each group includes at least one of the G text segments, and K is an integer greater than 0 and less than G;
[0158] A second-level label unit, configured to, for any one of the K groups, extract a corresponding second-level topic label through an algorithm for extracting keywords.
[0159] Combined with the previous embodiment, in some optional embodiments, the second-level label unit includes: a first word group subunit, a first frequency subunit, a second frequency subunit, a quotient calculation subunit, and a second-level label subunit;
[0160] The first word group subunit is configured to, for any one of the K groups, obtain a first word group according to the unigrams of the corresponding text segments, where the first word group includes multiple unigrams;
[0161] The first frequency subunit is configured to calculate the first frequency of each unigram in the first word group that appears in the first word group respectively;
[0162] The second frequency subunit is configured to, for any one of the unigrams in the first word group, calculate a corresponding second frequency, where the numerator of the second frequency is the number of groups among the K groups that involve the corresponding unigram, and the denominator of the second frequency is K;
[0163] The quotient calculation subunit is configured to, for any one of the unigrams in the first word group, calculate the quotient of the corresponding first frequency and the second frequency;
[0164] The second-level label subunit is configured to determine the first J unigrams with larger quotients in the first word group as the second-level topic label, where J is an integer greater than 0.
[0165] Combined Figure 9 In the embodiment shown, in some optional embodiments, the device further includes: a similar label unit;
[0166] A similar tag unit is configured to, after determining the one-word phrases in the hot topic group included in the text segment as the first-level topic tags of the corresponding text segment, input the first-level topic tags into a pre-established neural network model, so as to obtain extended topic tags corresponding to the first-level topic tags output by the neural network model, where the similarity between the extended topic tags and the input first-level topic tags is greater than a preset similarity threshold.
[0167] The present invention provides a computer-readable storage medium, on which a program is stored, and when the program is executed by a processor, the topic tag extraction method described in any one of the above is implemented.
[0168] As Figure 10 As shown, the present invention provides an electronic device 70, which includes at least one processor 701, and at least one memory 702 and a bus 703 connected to the processor 701; wherein, the processor 701 and the memory 702 complete communication with each other through the bus 703; the processor 701 is configured to call program instructions in the memory 702 to execute the topic tag extraction method described in any one of the above.
[0169] In this application, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprising", "including" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the phrase "including a..." does not exclude the existence of additional identical elements in the process, method, article or device including the element.
[0170] Each embodiment in this specification is described in a related manner. The same or similar parts between the embodiments can be referred to each other, and the differences between each embodiment and other embodiments are emphasized. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the related parts can be referred to the partial description of the method embodiment.
[0171] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but rather to the broadest scope consistent with the principles and novel features disclosed herein.
[0172] The above are only the preferred embodiments of the present invention, and are not intended to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention are included in the protection scope of the present invention.
Claims
1. A method for extracting topic tags, characterized in that, Including: Splitting the pre-obtained dialogue text data to obtain multiple text segments; Performing word segmentation on each of the text segments to obtain multiple single-word phrases; Calculating the frequency of occurrence of each of the single-word phrases in the text segment; Determining the hot topic group corresponding to the dialogue text data according to the frequency of occurrence of each of the single-word phrases, wherein the hot topic group includes multiple of the single-word phrases; For any one of the text segments, determining whether at least one of the single-word phrases in the hot topic group is included in the text segment; If included, determining the single-word phrases in the hot topic group included in the text segment as the first-level topic labels corresponding to the respective text segments; Extracting G text segments with target labels from the multiple text segments, wherein the target label is the first-level topic label in the hot topic group, and G is an integer greater than 0; After vectorizing the G text segments, inputting them into a pre-established clustering algorithm to cluster the G text segments into K groups, wherein each group includes at least one of the G text segments, and K is an integer greater than 0 and less than G; For any one of the K groups, extracting the corresponding second-level topic labels through an algorithm for extracting keywords, wherein for any one of the K groups, extracting the corresponding second-level topic labels through an algorithm for extracting keywords includes: for any one of the K groups, obtaining a first word group according to the single-word phrases of the respective text segments, and obtaining a second word group obtained by word segmentation of the respective text segments, wherein the first word group includes multiple non-repeating single-word phrases; respectively calculating the first frequency of occurrence of each of the single-word phrases in the first word group in the second word group; for any one of the single-word phrases in the first word group, calculating a corresponding second frequency, wherein the numerator of the second frequency is the number of groups in the K groups involving the corresponding single-word phrase, and the denominator of the second frequency is K; for any one of the single-word phrases in the first word group, calculating the quotient of the corresponding first frequency and the second frequency; determining the first J single-word phrases with larger quotients among the single-word phrases in the first word group as the second-level topic labels, wherein J is an integer greater than 0.
2. The method according to claim 1, characterized in that, Before determining the hot topic group corresponding to the dialogue text data according to the frequency of occurrence of each of the single-word phrases, the method further includes: Combining the single-word phrases freely to obtain multiple two-word phrases, wherein the two-word phrases are composed of any two non-repeating single-word phrases; For any one of the two-word phrases, calculating the number of text segments that simultaneously include the two single-word phrases that form the two-word phrase; The determining the hot topic group corresponding to the dialogue text data according to the frequency of occurrence of each of the single-word phrases includes: Determine the hot topic group according to the frequency of each of the unary phrases and the quantity corresponding to each of the binary phrases, where the hot topic group includes a plurality of the unary phrases and a plurality of the binary phrases.
3. The method according to claim 2, wherein The determining the hot topic group according to the frequency of each of the unary phrases and the quantity corresponding to each of the binary phrases includes: Jointly determine the top N unary phrases with larger frequencies and the top M binary phrases with larger quantities as the hot topic group, where both N and M are integers greater than 0; Alternatively, jointly determine the unary phrases with frequencies greater than a preset frequency threshold and the binary phrases with quantities greater than a preset quantity threshold as the hot topic group.
4. The method according to claim 3, wherein The determining whether any of the text segments includes at least one of the unary phrases in the hot topic group includes: For any of the text segments, determine whether the text segment includes at least one of the unary phrases in the hot topic group and / or at least one of the binary phrases in the hot topic group; If it includes, then determine the unary phrases in the hot topic group included in the text segment as the first-level topic labels corresponding to the respective text segments, including: If it includes, then determine the unary phrases and / or binary phrases in the hot topic group included in the text segment as the first-level topic labels corresponding to the respective text segments.
5. The method according to claim 4, characterized in that, After determining the unary phrases and / or binary phrases in the hot topic group included in the text segment as the first-level topic labels corresponding to the respective text segments, the method further includes: Set the corresponding first-level topic label on the corresponding text segment.
6. The method according to claim 4, characterized in that After determining the unary phrases and / or binary phrases in the hot topic group included in the text segment as the first-level topic labels corresponding to the respective text segments, the method further includes: Input the first-level topic label into a pre-established neural network model to obtain an extended topic label corresponding to the first-level topic label output by the neural network model, where the similarity between the extended topic label and the input first-level topic label is greater than a preset similarity threshold.
7. A topic tag extraction device, characterized in that Includes: A segmentation unit, a word segmentation unit, a word frequency unit, a hot phrase unit, a judgment unit, a first-level label unit, a text segment extraction unit, a clustering unit, and a second-level label unit; The segmentation unit is configured to split pre-obtained dialogue text data to obtain multiple text segments; The word segmentation unit is configured to perform word segmentation on each of the text segments to obtain a plurality of unary phrases; The word frequency unit is configured to calculate the frequency of each of the unary phrases appearing in each of the text segments; The hot phrase unit is configured to determine the hot topic group corresponding to the dialogue text data according to the frequency of each of the unary phrases, where the hot topic group includes a plurality of the unary phrases; The judgment unit is configured to, for any of the text segments, determine whether the text segment includes at least one of the unary phrases in the hot topic group; The first-level tag unit is configured to, if included, determine the single-word phrases in the hot topic group included in the text segment as the first-level topic tags of the corresponding text segment; The text segment extraction unit is configured to extract G text segments with target tags from the multiple text segments, where the target tag is the first-level topic tag in the hot topic group, and G is an integer greater than 0; The clustering unit is configured to vectorize the G text segments and then input them into a pre-established clustering algorithm, so as to cluster the G text segments into K groups, where each group includes at least one of the G text segments, and K is an integer greater than 0 and less than G; The second-level tag unit is configured to, for any one of the K groups, extract the corresponding second-level topic tags through an algorithm for extracting keywords, where, for any one of the K groups, extracting the corresponding second-level topic tags through an algorithm for extracting keywords includes: for any one of the K groups, obtaining a first word group according to the single-word phrases of the corresponding text segments, and obtaining a second word group obtained by segmenting the corresponding text segments, where the first word group includes multiple non-repeating single-word phrases; calculating the first frequency of each of the single-word phrases in the first word group in the second word group respectively; for any one of the single-word phrases in the first word group, calculating a corresponding second frequency, where the numerator of the second frequency is the number of groups in the K groups involving the corresponding single-word phrase, and the denominator of the second frequency is K; for any one of the single-word phrases in the first word group, calculating the quotient of the corresponding first frequency and the second frequency; determining the first J single-word phrases with larger quotients in the first word group as the second-level topic tags according to the quotients of the single-word phrases in the first word group, where J is an integer greater than 0.
8. An electronic device, characterized in that, The electronic device includes at least one processor, at least one memory connected to the processor, and a bus; wherein, the processor and the memory complete communication with each other through the bus; the processor is configured to call program instructions in the memory to execute the topic tag extraction method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Method and device for automatic detection of microblog hot topics
CN104615593A
Method and device for tracing hot topics and confirming keywords
CN104915447A