Method and device for generating word list for text classification

By clustering and word segmentation of texts, word lists are generated to determine text categories, the problem of low mutual exclusion of text categories in the prior art is solved, and the accuracy and reliability of text classification are improved.

CN115238681BActive Publication Date: 2025-05-09GUANGZHOU YOUMI INFORMATION TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210741184.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-28
Publication Date
2025-05-09
Estimated Expiration
2042-06-28

AI Technical Summary

Technical Problem

In the existing text classification technology, the single-label classification method leads to low mutual exclusion between text categories and unclear category characteristics, resulting in low classification accuracy.

Method used

By inputting the sample text into the pre-trained text analysis vector model, the sentence vector of the sample sentence is generated, and clustered based on these sentence vectors, the sample sentence set under the target class cluster is obtained. Then, word segmentation is performed on each class cluster, and a word list is generated to determine the text category.

Benefits of technology

It improves the mutual exclusion between text categories, enhances the reliability and accuracy of classification, and enables text to be classified more accurately.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115238681B_ABST
    Figure CN115238681B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and device for generating a word list for text classification, the method comprising: inputting a sample text into a trained text analysis vector model for analysis to obtain a sentence vector for each sample sentence in the sample text; based on the sentence vectors of all sample sentences, performing a clustering operation on all sample sentences to obtain a sample sentence set under at least one cluster; performing a word segmentation operation on the sample sentence set under each cluster to obtain a word list under each cluster. It can be seen that the implementation of the present invention can obtain a word list for text classification by clustering and word segmenting the sample sentences of the sample text, enriching the intelligent text category determination method of the text classification system, facilitating the mutual exclusivity between the determined text categories, making the category characteristics between the text categories clearer, and further facilitating the classification reliability and accuracy of the text to be analyzed, thereby facilitating the accurate classification of the text to be analyzed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of text processing, and in particular to a word list generation method and device for text classification. Background Art

[0002] In recent years, natural language processing technology has been widely used in all aspects of people's lives. For example, text classification technology has been widely used in search engines, question-answering systems, conversation systems and other systems that need to process important information in order to classify a large amount of information.

[0003] Currently, text classification technology often uses single-label classification to classify text. However, it is found in practice that when determining text categories through a single label, the mutual exclusivity between the determined text categories is low, that is, the category characteristics of the text categories are not clear enough, resulting in low text classification accuracy. It can be seen that it is particularly important to provide a method that can improve the accuracy of text classification. Summary of the invention

[0004] The technical problem to be solved by the present invention is to provide a method and device for generating a word list for text classification, which can help improve the mutual exclusivity between the determined text categories, and further help improve the classification reliability and classification accuracy of the text to be analyzed, thereby helping to accurately classify the text to be analyzed.

[0005] In order to solve the above technical problems, the first aspect of the present invention discloses a word list generation method for text classification, the method comprising:

[0006] Input the sample text into a pre-trained text analysis vector model for analysis to obtain a sentence vector for each sample sentence in the sample text;

[0007] Based on the sentence vectors of all the sample sentences, a clustering operation is performed on all the sample sentences to obtain a sample sentence set under at least one target cluster; each sample sentence set under the target cluster includes at least one sample sentence;

[0008] A word segmentation operation is performed on the sample sentence set under each of the target clusters to obtain a word list under each of the target clusters; the word list is used to determine the cluster to which each sentence in the to-be-determined cluster text belongs.

[0009] As an optional implementation, in the first aspect of the present invention, the clustering operation is performed on all the sample sentences based on the sentence vectors of all the sample sentences to obtain a sample sentence set under at least one target cluster, including:

[0010] For each of the sample sentences, according to the sentence vector of the sample sentence and the sentence vectors of all other sample sentences except the sample sentence, calculate the first vector cosine similarity between the sample sentence and each of the other sample sentences; determine the first sample sentence with the smallest first vector cosine similarity from all the other sample sentences, and merge the sample sentence with the first sample sentence to obtain a first sample sentence set;

[0011] After all the first sample sentence sets are obtained, sentence clustering operations are performed on all the first sample sentence sets according to the sentence vectors of all the first sample sentence sets to obtain sample sentence sets under at least one target cluster.

[0012] As an optional implementation manner, in the first aspect of the present invention, performing a sentence clustering operation on all the first sample sentence sets according to the sentence vectors of all the first sample sentence sets to obtain a sample sentence set under at least one target cluster includes:

[0013] For each of the first sample sentence sets, calculating the second vector cosine similarity between the first sample sentence set and each of the other first sample sentence sets according to the sentence vector of the first sample sentence set and the sentence vectors of all other first sample sentence sets except the first sample sentence set;

[0014] Determine whether the cosine similarities of all the second vectors corresponding to any of the first sample sentence sets are less than a preset similarity threshold, and when the judgment result is yes, calculate the sentence set center vector of each of the first sample sentence sets and determine the number of sample sentences of each of the first sample sentence sets, and determine all second sample sentence sets whose number of sample sentences is less than or equal to a preset sample sentence number threshold and all third sample sentence sets whose number of sample sentences is greater than the preset sample sentence number threshold from all the first sample sentence sets;

[0015] For each of the second sample sentences, according to the sentence center vector of the second sample sentence and the sentence center vector of each of the third sample sentences, calculate the center vector cosine similarity between the second sample sentence and each of the third sample sentences; determine a fourth sample sentence with the largest center vector cosine similarity from all the third sample sentences, and merge the fourth sample sentence with the second sample sentence to obtain the merged fourth sample sentence;

[0016] According to all the fourth sample sentence sets obtained after the merger, all the third sample sentence sets are updated, and according to all the sample sentences contained in each of the third sample sentence sets, the target cluster to which each of the third sample sentence sets belongs is determined as the sample sentence set under each of the target clusters.

[0017] As an optional embodiment, in the first aspect of the present invention, the method further comprises:

[0018] When it is determined that all the second vector cosine similarities corresponding to any one of the first sample sentence sets are not all less than the preset similarity threshold, all fifth sample sentence sets whose second vector cosine similarities are greater than the preset similarity threshold are determined from all the other first sample sentence sets corresponding to the first sample sentence set, and the first sample sentence set is merged with each of the fifth sample sentence sets to update all the fifth sample sentence sets;

[0019] After all the fifth sample sentence sets are updated, all the updated fifth sample sentence sets are determined as all the first sample sentence sets, and the operations of calculating, for each of the first sample sentence sets, the second vector cosine similarity between the first sample sentence set and each of the other first sample sentence sets according to the sentence vector of the first sample sentence set and the sentence vectors of all other first sample sentence sets except the first sample sentence set, and determining whether all the second vector cosine similarities corresponding to any of the first sample sentence sets are less than a preset similarity threshold are triggered.

[0020] As an optional implementation, in the first aspect of the present invention, performing a word segmentation operation on the sample sentence set under each of the target clusters to obtain a word list under each of the target clusters includes:

[0021] By using a preset word segmenter, a word segmentation operation is performed on the sample sentence set under each of the target clusters to obtain a word list under each of the target clusters; the word list under each of the target clusters includes a plurality of words;

[0022] For each vocabulary under the target cluster, determine the vocabulary difference set between the target cluster and all other target clusters, and determine the vocabulary difference set between the target cluster and all other target clusters as the first vocabulary under the target cluster;

[0023] Remove the first word list under the target cluster from the word list under the target cluster to obtain the second word list under the target cluster, and determine whether the number of words contained in the second word list under the target cluster is greater than or equal to a preset word number threshold. If not, determine the first word list under the target cluster as the word list under the target cluster.

[0024] As an optional embodiment, in the first aspect of the present invention, the method further comprises:

[0025] When it is determined that the number of all the words contained in the second vocabulary under the target cluster is greater than or equal to the preset word number threshold, determining the number of word combination connections under the current word combination round of the target cluster, and performing a word combination operation on the second vocabulary under the target cluster according to the number of word combination connections under the current word combination round to obtain a third vocabulary under the target cluster;

[0026] Acquire the fourth vocabulary under each of the other target clusters, and determine the vocabulary difference set between the third vocabulary under the target cluster and the fourth vocabulary under all the other target clusters; the number of word combination connections corresponding to the fourth vocabulary under each of the other target clusters matches the number of word combination connections corresponding to the third vocabulary under the target cluster;

[0027] The vocabulary difference set between the third vocabulary under the target cluster and the fourth vocabulary under all other target clusters is updated to the third vocabulary under the target cluster, and it is determined whether the current word combination round of the target cluster is greater than or equal to a preset word combination round threshold. When the judgment result is yes, the first vocabulary under the target cluster and the third vocabulary under the target cluster are determined as the word list under the target cluster.

[0028] As an optional embodiment, in the first aspect of the present invention, the method further comprises:

[0029] When it is determined that the current word combination round of the target cluster is less than the preset word combination round threshold, the current word combination round of the target cluster is increased by 1 to update the current word combination round of the target cluster, and the third word list under the target cluster is removed from the word list under the target cluster to obtain the fifth word list under the target cluster; the fifth word list under the target cluster is determined as the second word list under the target cluster, and the determination of the number of word combination connections under the current word combination round of the target cluster is triggered, and the word combination connection number under the current word combination round is determined according to the word combination under the current word combination round. Combine the number of connections, perform a word combination operation on the second word list under the target cluster to obtain the third word list under the target cluster; obtain the fourth word list under each of the other target clusters, and determine the word list difference set between the third word list under the target cluster and the fourth word lists under all the other target clusters; update the word list difference set between the third word list under the target cluster and the fourth word lists under all the other target clusters to the third word list under the target cluster, and determine whether the current word combination round of the target cluster is greater than or equal to the preset word combination round threshold.

[0030] A second aspect of the present invention discloses a word list generating device for text classification, the device comprising:

[0031] An input module, used to input the sample text into a pre-trained text analysis vector model for analysis, and obtain a sentence vector for each sample sentence in the sample text;

[0032] A clustering module, configured to perform a clustering operation on all the sample sentences based on the sentence vectors of all the sample sentences to obtain a sample sentence set under at least one target cluster; each sample sentence set under the target cluster includes at least one sample sentence;

[0033] The word segmentation module is used to perform word segmentation operations on the sample sentence set under each of the target clusters to obtain a word list under each of the target clusters; the word list is used to determine the cluster to which each sentence in the cluster text to be determined belongs.

[0034] As an optional implementation, in the second aspect of the present invention, the clustering module includes:

[0035] A calculation submodule, configured to calculate, for each of the sample sentences, a first vector cosine similarity between the sample sentence and each of the other sample sentences according to the sentence vector of the sample sentence and the sentence vectors of all other sample sentences except the sample sentence;

[0036] A first determination submodule, configured to determine a first sample sentence having the smallest first vector cosine similarity from all the other sample sentences;

[0037] a merging submodule, configured to merge the sample sentence with the first sample sentence to obtain a first sample sentence set;

[0038] The clustering submodule is used for performing sentence clustering operation on all the first sample sentence sets according to the sentence vectors of all the first sample sentence sets after obtaining all the first sample sentence sets, so as to obtain sample sentence sets under at least one target cluster.

[0039] As an optional implementation, in the second aspect of the present invention, the clustering submodule includes:

[0040] a calculating unit, configured to calculate, for each of the first sample sentence sets, a second vector cosine similarity between the first sample sentence set and each of the other first sample sentence sets according to the sentence vector of the first sample sentence set and the sentence vectors of all other first sample sentence sets except the first sample sentence set;

[0041] A judging unit, configured to judge whether the cosine similarities of all the second vectors corresponding to any of the first sample sentence sets are less than a preset similarity threshold;

[0042] The calculation unit is further configured to calculate a sentence set center vector of each of the first sample sentence sets when the judgment result of the judgment unit is yes;

[0043] a determining unit, configured to determine the number of sample sentences in each of the first sample sentence sets, and determine from all the first sample sentence sets all second sample sentence sets whose number of sample sentences is less than or equal to a preset sample sentence number threshold and all third sample sentence sets whose number of sample sentences is greater than the preset sample sentence number threshold;

[0044] The calculation unit is further configured to calculate, for each of the second sample sentences, a center vector cosine similarity between the second sample sentence set and each of the third sample sentence sets according to the sentence center vector of the second sample sentence set and the sentence center vector of each of the third sample sentence sets;

[0045] The determining unit is further configured to determine a fourth sample sentence set having the largest center vector cosine similarity from all the third sample sentence sets;

[0046] a merging unit, configured to merge the fourth sample sentence set with the second sample sentence set to obtain the merged fourth sample sentence set;

[0047] An updating unit, configured to update all the third sample sentence sets according to all the fourth sample sentence sets obtained after merging;

[0048] The determining unit is further configured to determine, based on all the sample sentences included in each of the third sample sentence sets, a target cluster to which each of the third sample sentence sets belongs, as a sample sentence set under each of the target clusters.

[0049] As an optional implementation manner, in the second aspect of the present invention, the determining unit is further configured to:

[0050] When the judging unit judges that all the second vector cosine similarities corresponding to any one of the first sample sentence sets are not all less than the preset similarity threshold, determining all fifth sample sentence sets whose second vector cosine similarities are greater than the preset similarity threshold from all the other first sample sentence sets corresponding to the first sample sentence set;

[0051] The merging unit is further configured to merge the first sample sentence set with each of the fifth sample sentence sets to update all of the fifth sample sentence sets;

[0052] The determining unit is further configured to, after all the fifth sample sentence sets are updated, determine all the updated fifth sample sentence sets as all the first sample sentence sets, and trigger the calculating unit to perform the operation of calculating, for each of the first sample sentence sets, the cosine similarity of the second vector between the first sample sentence set and each of the other first sample sentence sets according to the sentence vector of the first sample sentence set and the sentence vectors of all the other first sample sentence sets except the first sample sentence set, and trigger the judging unit to perform the operation of judging whether all the cosine similarities of the second vectors corresponding to any of the first sample sentence sets are less than a preset similarity threshold.

[0053] As an optional implementation, in the second aspect of the present invention, the word segmentation module includes:

[0054] A word segmentation submodule is used to perform a word segmentation operation on the sample sentence set under each target cluster through a preset word segmenter to obtain a word list under each target cluster; the word list under each target cluster includes a plurality of words;

[0055] A second determination submodule is used to determine, for each vocabulary under the target cluster, a vocabulary difference set between the target cluster and all other target clusters, and determine the vocabulary difference set between the target cluster and all other target clusters as the first vocabulary under the target cluster;

[0056] A vocabulary removal submodule, used to remove the first vocabulary under the target cluster from the vocabulary under the target cluster to obtain a second vocabulary under the target cluster;

[0057] A judgment submodule, used to judge whether the number of all the words contained in the second vocabulary under the target cluster is greater than or equal to a preset word number threshold;

[0058] The second determination submodule is further configured to determine the first word list under the target cluster as the word list under the target cluster when the determination result of the determination submodule is no.

[0059] As an optional implementation manner, in the second aspect of the present invention, the second determining submodule is further used to:

[0060] When the judging submodule judges that the number of all the words contained in the second word list under the target cluster is greater than or equal to the preset word number threshold, determining the number of word combination connections under the current word combination round of the target cluster;

[0061] A word combination submodule, configured to perform a word combination operation on the second word list under the target cluster according to the number of word combination connections under the current word combination round, to obtain a third word list under the target cluster;

[0062] An acquisition submodule, used for acquiring a fourth vocabulary under each of the other target clusters;

[0063] The second determination submodule is further used to determine a vocabulary difference set between the third vocabulary under the target cluster and the fourth vocabulary under all the other target clusters; the number of word combination connections corresponding to the fourth vocabulary under each of the other target clusters matches the number of word combination connections corresponding to the third vocabulary under the target cluster;

[0064] An updating submodule, configured to update a vocabulary difference set between the third vocabulary under the target cluster and the fourth vocabulary under all other target clusters to the third vocabulary under the target cluster;

[0065] The judgment submodule is further used to judge whether the current word combination round of the target cluster is greater than or equal to a preset word combination round threshold;

[0066] The second determination submodule is further configured to determine the first vocabulary list under the target cluster and the third vocabulary list under the target cluster as the vocabulary list under the target cluster when the determination result of the determination submodule is yes.

[0067] As an optional implementation, in the second aspect of the present invention, the updating submodule is further used for:

[0068] When the judging submodule judges that the current word combination round of the target cluster is less than the preset word combination round threshold, the current word combination round of the target cluster is increased by 1 to update the current word combination round of the target cluster;

[0069] The vocabulary removal submodule is further used to remove the third vocabulary under the target cluster from the vocabulary under the target cluster to obtain the fifth vocabulary under the target cluster;

[0070] The second determination submodule is also used to determine the fifth vocabulary under the target cluster as the second vocabulary under the target cluster, and trigger the execution of the operation of determining the number of word combination connections under the current word combination round of the target cluster, and trigger the word combination submodule to execute the operation of performing word combination operation on the second vocabulary under the target cluster according to the number of word combination connections under the current word combination round to obtain the third vocabulary under the target cluster, and trigger the acquisition submodule to execute the operation of acquiring the fourth vocabulary under each of the other target clusters, and trigger the execution of the operation of determining the vocabulary difference set between the third vocabulary under the target cluster and the fourth vocabulary under all the other target clusters, and trigger the update submodule to execute the operation of updating the vocabulary difference set between the third vocabulary under the target cluster and the fourth vocabulary under all the other target clusters to the third vocabulary under the target cluster, and trigger the judgment submodule to execute the operation of judging whether the current word combination round of the target cluster is greater than or equal to the preset word combination round threshold.

[0071] The third aspect of the present invention discloses another device for generating a word list for text classification, the device comprising:

[0072] A memory storing executable program code;

[0073] a processor coupled to the memory;

[0074] The processor calls the executable program code stored in the memory to execute the word list generation method for text classification disclosed in the first aspect of the present invention.

[0075] The fourth aspect of the present invention discloses a computer storage medium, wherein the computer storage medium stores computer instructions, and when the computer instructions are called, they are used to execute the word list generation method for text classification disclosed in the first aspect of the present invention.

[0076] Compared with the prior art, the embodiments of the present invention have the following beneficial effects:

[0077] In the embodiment of the present invention, the sample text is input into the pre-trained text analysis vector model for analysis to obtain the sentence vector of each sample sentence in the sample text; based on the sentence vectors of all sample sentences, a clustering operation is performed on all sample sentences to obtain a sample sentence set under at least one target cluster; a word segmentation operation is performed on the sample sentence set under each target cluster to obtain a word list under each target cluster. It can be seen that the implementation of the present invention can obtain a word list for text classification by clustering and word segmenting the sample sentences of the sample text, which can effectively solve the problem of fuzzy category characteristics of text categories in traditional single-label classification methods, enrich the intelligent determination of text categories in the text classification system, and is conducive to improving the mutual exclusivity between the determined text categories, making the category characteristics between text categories clearer, and thus helping to improve the classification reliability and classification accuracy of the text to be analyzed, thereby facilitating the accurate classification of the text to be analyzed. BRIEF DESCRIPTION OF THE DRAWINGS

[0078] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0079] Figure 1 It is a flowchart of a method for generating a word list for text classification disclosed in an embodiment of the present invention;

[0080] Figure 2 It is a flowchart of another method for generating a word list for text classification disclosed in an embodiment of the present invention;

[0081] Figure 3 It is a structural schematic diagram of a word list generating device for text classification disclosed in an embodiment of the present invention;

[0082] Figure 4 It is a structural schematic diagram of another device for generating a word list for text classification disclosed in an embodiment of the present invention;

[0083] Figure 5 It is a structural schematic diagram of another device for generating a word list for text classification disclosed in an embodiment of the present invention. DETAILED DESCRIPTION

[0084] In order to enable those skilled in the art to better understand the scheme of the present invention, the technical scheme in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0085] The terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish different objects rather than to describe a specific order. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, device, product or end including a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units that are not listed, or may optionally include other steps or units inherent to these processes, methods, products or ends.

[0086] Reference to "embodiments" herein means that a particular feature, structure, or characteristic described in conjunction with the embodiments may be included in at least one embodiment of the present invention. The appearance of the phrase in various places in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment that is mutually exclusive with other embodiments. It is explicitly and implicitly understood by those skilled in the art that the embodiments described herein may be combined with other embodiments.

[0087] The present invention discloses a word list generation method and device for text classification, which can help improve the mutual exclusivity between determined text categories, thereby helping to improve the classification reliability and classification accuracy of the text to be analyzed, thereby helping to accurately classify the text to be analyzed. The following are detailed descriptions.

[0088] Embodiment 1

[0089] See also Figure 1 , Figure 1 1 is a flow chart of a method for generating a word list for text classification disclosed in an embodiment of the present invention. Figure 1 The described method for generating a word list for text classification can be applied to the classification and detection of emails (such as filtering spam), and can also be applied to the classification and detection of online comments (such as sentiment analysis and classification of online comments), and the embodiments of the present invention do not limit this. Optionally, the method can be implemented by a text classification system, and the text classification system can be integrated into a text classification device, or can be a local server or cloud server for managing the text classification process, and the embodiments of the present invention do not limit this. Figure 1As shown, the word list generation method applied to text classification may include the following operations:

[0090] 101. Input the sample text into a pre-trained text analysis vector model for analysis to obtain a sentence vector for each sample sentence in the sample text.

[0091] In an embodiment of the present invention, specifically, the text analysis vector model may be a word vector model, wherein the sample text is input into a pre-trained text analysis vector model for analysis, and all word vectors of each sample sentence in the sample text may be obtained, and based on all word vectors of each sample sentence in the sample text, the sentence vector of each sample sentence in the sample text may be obtained. Optionally, the word vector carries at least one of semantic information, grammatical information, and word order information.

[0092] 102. Based on the sentence vectors of all the sample sentences, a clustering operation is performed on all the sample sentences to obtain a sample sentence set under at least one target cluster.

[0093] In the embodiment of the present invention, the sample sentence set under each target cluster includes at least one sample sentence. Specifically, before performing the clustering operation on all sample sentences, the number of clustering categories may not be pre-defined, and each sample sentence may be regarded as a sample point, that is, each sample sentence may be regarded as an independent cluster, wherein performing the clustering operation on all sample sentences may be understood as classifying all sample sentences until the sample sentences contained in each category are mutually exclusive.

[0094] 103. Perform word segmentation operation on the sample sentence set under each target cluster to obtain a word list under each target cluster.

[0095] In the embodiment of the present invention, the word list under each target cluster is used to determine the cluster to which each sentence in the to-be-determined cluster text belongs. For example, if the word list under the target cluster A includes words such as "card", "attribute", "mana", and "magic value", and the content of sentence B in a to-be-determined cluster text describes the content of interactive activities, then the cluster to which sentence B belongs can be determined to be the target cluster A.

[0096] It can be seen that the implementation of the embodiment of the present invention can obtain a word list applied to text classification by clustering and segmenting sample sentences of the sample text, which can effectively solve the problem of fuzzy category characteristics of text categories in traditional single-label classification methods, enrich the intelligent text category determination method of the text classification system, and is conducive to improving the mutual exclusivity between the determined text categories, making the category characteristics between text categories clearer, and further conducive to improving the classification reliability and classification accuracy of the text to be analyzed, thereby facilitating the accurate classification of the text to be analyzed.

[0097] In an optional embodiment, in the above step 102, based on the sentence vectors of all sample sentences, a clustering operation is performed on all sample sentences to obtain a sample sentence set under at least one target cluster, including:

[0098] For each sample sentence, the first vector cosine similarity between the sample sentence and each other sample sentence is calculated according to the sentence vector of the sample sentence and the sentence vectors of all other sample sentences except the sample sentence; a first sample sentence with the smallest first vector cosine similarity is determined from all other sample sentences, and the sample sentence is merged with the first sample sentence to obtain a first sample sentence set;

[0099] After all first sample sentence sets are obtained, sentence set clustering operations are performed on all first sample sentence sets according to sentence vectors of all first sample sentence sets to obtain sample sentence sets under at least one target cluster.

[0100] In this optional embodiment, specifically, calculating the first vector cosine similarity between the sample sentence and each other sample sentence can be understood as treating each sample sentence as an independent cluster, and calculating the vector cosine similarity (i.e., vector dot product) between the cluster and other clusters in each cluster; and merging the sample sentence with the first sample sentence to obtain the first sample sentence set can be understood as for each cluster, it is necessary to find other clusters with the smallest vector cosine similarity with the cluster, and merge the cluster with the other clusters. For example, for cluster A, in cluster BCD, the vector cosine similarity between cluster B and it is the smallest, in which case cluster A and cluster B are merged, and for cluster B, in cluster ACD, the vector distance between cluster D and it is the smallest, in which case cluster B and cluster D are merged at the same time.

[0101] It can be seen that this optional embodiment can preliminarily merge each sample sentence by calculating the vector cosine similarity between the sample sentences, which is beneficial to improving the merging reliability and merging accuracy between the sample sentences, and further beneficial to improving the subsequent clustering reliability and clustering accuracy of the sample sentences, thereby facilitating the accurate generation of the subsequent word list.

[0102] In another optional embodiment, the above step of performing sentence clustering operations on all first sample sentence sets according to the sentence vectors of all first sample sentence sets to obtain sample sentence sets under at least one target cluster includes:

[0103] For each first sample sentence set, calculating the second vector cosine similarity between the first sample sentence set and each other first sample sentence set according to the sentence vector of the first sample sentence set and the sentence vectors of all other first sample sentence sets except the first sample sentence set;

[0104] Determine whether the cosine similarities of all second vectors corresponding to any first sample sentence set are less than a preset similarity threshold, and when the judgment result is yes, calculate the sentence set center vector of each first sample sentence set and determine the number of sample sentences of each first sample sentence set, and determine all second sample sentence sets whose number of sample sentences is less than or equal to the preset sample sentence number threshold and all third sample sentence sets whose number of sample sentences is greater than the preset sample sentence number threshold from all first sample sentence sets;

[0105] For each second sample sentence set, according to the sentence set center vector of the second sample sentence set and the sentence set center vector of each third sample sentence set, the center vector cosine similarity between the second sample sentence set and each third sample sentence set is calculated; a fourth sample sentence set with the largest center vector cosine similarity is determined from all the third sample sentence sets, and the fourth sample sentence set is merged with the second sample sentence set to obtain a merged fourth sample sentence set;

[0106] According to all the merged fourth sample sentence sets, all the third sample sentence sets are updated, and according to all the sample sentences contained in each third sample sentence set, the target cluster to which each third sample sentence set belongs is determined as the sample sentence set under each target cluster.

[0107] In this optional embodiment, determining whether all second vector cosine similarities corresponding to any first sample sentence set are less than a preset similarity threshold can be understood as determining whether all first sample sentence sets have completed a preliminary clustering operation by determining the size of the vector cosine similarity between any first sample sentence set and each other first sample sentence set, that is, determining whether the vector cosine similarity between any two merged clusters is less than a preset similarity threshold. If so, it means that all merged clusters have completed a preliminary clustering operation, and then the next sentence set merging operation (that is, the next clustering operation) is performed. Specifically, the sentence center vector of each first sample sentence set is calculated, that is, the sentence vectors of each first sample sentence set are summed and averaged; and the fourth sample sentence set is merged with the second sample sentence set, which can be understood as merging the sample sentence set with less than N sample sentences (that is, the clusters that have completed preliminary clustering) with the sample sentence set with greater than or equal to N sample sentences and the largest cosine similarity of the center vector, so as to ensure that the number of sample sentences in each sample sentence set is greater than or equal to N and that the mutual exclusivity between the sample sentence sets is large enough.

[0108] Further, in this optional embodiment, according to all sample sentences included in each third sample sentence set, the target cluster to which each third sample sentence set belongs is determined, and the sample sentence set under each target cluster may include:

[0109] For each third sample sentence set, at least one sample sentence to be analyzed is screened out from all sample sentences contained in the third sample sentence set, and based on all sample sentences to be analyzed, a target cluster to which the third sample sentence set belongs is determined as a sample sentence set under each target cluster.

[0110] For example, if the contents of all sample sentences to be analyzed are interactive activity type contents, then the target cluster to which the third sample sentence set belongs can be determined as "interactive activity", that is, the third sample sentence set is a sample sentence set under "interactive activity".

[0111] It can be seen that this optional embodiment can obtain sample sentence sets under each target cluster by performing secondary clustering on the sample sentence sets. In this way, the mutual exclusivity between the sample sentence sets under each target cluster can be improved, and then the category characteristics between the word lists under each target cluster obtained subsequently can be improved, thereby solving the problem that the category characteristics of text categories in traditional text classification methods are too vague.

[0112] In yet another optional embodiment, the method may further include:

[0113] When it is determined that all second vector cosine similarities corresponding to any first sample sentence set are not less than a preset similarity threshold, all fifth sample sentence sets whose second vector cosine similarities are greater than the preset similarity threshold are determined from all other first sample sentence sets corresponding to the first sample sentence set, and the first sample sentence set is merged with each fifth sample sentence set to update all fifth sample sentence sets;

[0114] After all the fifth sample sentence sets are updated, all the updated fifth sample sentence sets are determined as all the first sample sentence sets, and the above steps are triggered to calculate, for each first sample sentence set, the second vector cosine similarity between the first sample sentence set and each other first sample sentence set according to the sentence vector of the first sample sentence set and the sentence vectors of all other first sample sentence sets except the first sample sentence set, and to determine whether all the second vector cosine similarities corresponding to any first sample sentence set are less than a preset similarity threshold.

[0115] In this optional embodiment, specifically, when it is determined that all second vector cosine similarities corresponding to any first sample sentence set are not all less than the preset similarity threshold, it means that all merged clusters have not completed the preliminary clustering operation. At this time, it is necessary to continue to iteratively merge the two sample sentence sets whose corresponding vector cosine similarities are greater than the preset similarity threshold until the vector cosine similarity between any two sample sentence sets is less than the preset similarity threshold, so as to complete the preliminary clustering operation of all merged clusters. For example, for the merged clusters AB, CD and EF, if the vector cosine similarity between cluster AB and cluster CD is greater than the preset similarity threshold, cluster AB needs to be further merged with cluster CD to obtain cluster ABCD. At the same time, if the vector cosine similarity between cluster CD and cluster EF is also greater than the preset similarity threshold, cluster CD needs to be further merged with cluster EF to obtain cluster CDEF. Then, continue to determine whether the vector cosine similarity between cluster ABCD and cluster CDEF is less than the preset similarity threshold. If not, continue to merge cluster ABCD with cluster CDEF. Otherwise, it means that the preliminary clustering operation has been completed and the next clustering operation can be performed.

[0116] It can be seen that this optional embodiment can automatically determine whether the preliminary clustering operation of the sample sentence set has been completed, which shows the intelligent determination method of the text classification system for the clustering of the sample sentence set, which is conducive to improving the clustering reliability and clustering accuracy of the sample sentence set, and further helps to ensure the mutual exclusivity of the sentence sets between the obtained sample sentence sets, thereby helping to clarify the sentence set characteristics between the sample sentence sets.

[0117] Embodiment 2

[0118] See also Figure 2 , Figure 2 1 is a flow chart of a method for generating a word list for text classification disclosed in an embodiment of the present invention. Figure 2 The described method for generating a word list for text classification can be applied to the classification and detection of emails (such as filtering spam), and can also be applied to the classification and detection of online comments (such as sentiment analysis and classification of online comments), and the embodiments of the present invention do not limit this. Optionally, the method can be implemented by a text classification system, and the text classification system can be integrated into a text classification device, or can be a local server or cloud server for managing the text classification process, and the embodiments of the present invention do not limit this. Figure 2 As shown, the word list generation method applied to text classification may include the following operations:

[0119] 201. Input the sample text into a pre-trained text analysis vector model for analysis to obtain a sentence vector for each sample sentence in the sample text.

[0120] 202. Based on the sentence vectors of all the sample sentences, perform a clustering operation on all the sample sentences to obtain a sample sentence set under at least one target cluster.

[0121] In the embodiment of the present invention, for other descriptions of step 201 and step 202, please refer to the detailed description of step 101 and step 102 in the first embodiment, and the embodiment of the present invention will not be repeated.

[0122] 203. Perform a word segmentation operation on the sample sentence set under each target cluster through a preset word segmenter to obtain a word list under each target cluster.

[0123] In the embodiment of the present invention, the vocabulary under each target cluster includes multiple words.

[0124] Specifically, as an optional implementation, the preset word segmenter can be determined by:

[0125] Determine the segmentation length of the sample text in the current segmentation round, and perform a word segmentation operation on the sample text according to the segmentation length in the current segmentation round to obtain multiple sample words of the sample text;

[0126] For each sample word, calculate the mutual information value between the sample word and the right adjacent word, and determine whether the mutual information value is greater than or equal to a preset mutual information threshold. If so, connect the sample word with the right adjacent word to update the sample word, otherwise, retain the sample word; calculate the target adjacent entropy between the sample word and the target adjacent word, and determine whether the target adjacent entropy is less than a preset adjacent entropy threshold. If so, delete the sample word from all sample words to update all sample words, otherwise, retain the sample word;

[0127] Determine the word frequency of each sample word among all the updated sample words; determine, from all the sample words according to the word frequencies of all the sample words, all the sample words whose word frequencies are greater than or equal to a preset word frequency threshold, to obtain a sample word set of the sample text, and delete the first and last stop words and / or historical words from the sample word set of the sample text to update the sample word set of the sample text;

[0128] Increasing the current segmentation round of the sample text by 1 to update the current segmentation round of the sample text, and determining whether the current segmentation round of the sample text is greater than or equal to a preset segmentation round threshold;

[0129] When the judgment result is no, all operations from determining the segmentation length of the sample text under the current segmentation round to determining whether the current segmentation round of the sample text is greater than or equal to the preset segmentation round threshold are triggered to execute; when the judgment result is yes, all sample word sets of the sample text are determined as word segmenters.

[0130] It can be seen that this optional embodiment can identify new words, uncommon words and proprietary words in the text to be segmented, which is beneficial to improving the segmentation effect of the word segmenter, thereby helping to improve the reliability and accuracy of the subsequent word list.

[0131] 204. For each vocabulary under the target cluster, determine the vocabulary difference set between the target cluster and all other target clusters, and determine the vocabulary difference set between the target cluster and all other target clusters as the first vocabulary under the target cluster.

[0132] In an embodiment of the present invention, specifically, before determining the vocabulary difference set between the target cluster and all other target clusters for the vocabulary under each target cluster, a word deduplication operation can be performed on the vocabulary under each target cluster, that is, to ensure that each word in the vocabulary under each target cluster is unique, so as to update the vocabulary under each target cluster.

[0133] 205. Remove the first vocabulary under the target cluster from the vocabulary under the target cluster to obtain the second vocabulary under the target cluster, and determine whether the number of words contained in the second vocabulary under the target cluster is greater than or equal to a preset word number threshold. If not, determine the first vocabulary under the target cluster as the word list under the target cluster.

[0134] In an embodiment of the present invention, specifically, the first vocabulary under the target cluster can be understood as a vocabulary with a word connection number of 1 under the target cluster, that is, when the number of words in the first vocabulary under the target cluster is too small, the vocabulary under the target cluster after word segmentation can be directly used as the word list under the target cluster, without setting the word combination rounds required for the target cluster and the word combination connection number under each word combination round, that is, there is no need to perform subsequent word combination operations on the first vocabulary of the target cluster.

[0135] It can be seen that the implementation of the embodiment of the present invention can segment the sample sentence set under each target cluster through a preset word segmenter, so that the reliability and accuracy of the word list obtained under each target cluster can be guaranteed, and then the mutual exclusivity between the word lists under each target cluster can be guaranteed, thereby ensuring the reliability and accuracy of the cluster to which each sentence in the analyzed cluster text to be determined belongs.

[0136] In an optional embodiment, the method may further include:

[0137] When it is determined that the number of words contained in the second vocabulary under the target cluster is greater than or equal to the preset word number threshold, the number of word combination connections under the current word combination round of the target cluster is determined, and according to the number of word combination connections under the current word combination round, the second vocabulary under the target cluster is subjected to a word combination operation to obtain a third vocabulary under the target cluster;

[0138] Obtaining the fourth vocabulary under each other target cluster, and determining the vocabulary difference between the third vocabulary under the target cluster and the fourth vocabulary under all other target clusters;

[0139] The vocabulary difference set between the third vocabulary under the target cluster and the fourth vocabulary under all other target clusters is updated to the third vocabulary under the target cluster, and it is determined whether the current word combination round of the target cluster is greater than or equal to a preset word combination round threshold. When the judgment result is yes, the first vocabulary under the target cluster and the third vocabulary under the target cluster are determined as the word list under the target cluster.

[0140] In this optional embodiment, the number of word combination connections corresponding to the fourth vocabulary under each other target cluster matches the number of word combination connections corresponding to the third vocabulary under the target cluster. Specifically, performing a word combination operation on the second vocabulary under the target cluster can be understood as performing a non-repetitive word combination operation with a word combination connection number of N on the words in the second vocabulary under the target cluster. Further, the number of word combination connections under the current word combination round of the target cluster can be determined based on the number of words of all words contained in the second vocabulary under the target cluster. If the number of words is less than 50, two word combination rounds are set, and the number of word combination connections of the first word combination round is 2 and the number of word combination connections of the second word combination round is 3.

[0141] For example, when it is determined that the number of words in the second vocabulary under the target cluster is greater than or equal to the preset word number threshold, the word combination rounds required for the target cluster and the number of word combination connections under each word combination round can be set. If the number of words in the second vocabulary is small, only one word combination round is set, and the corresponding word combination connection number is 2. Then, in the first word combination round, the second vocabulary under the target cluster is subjected to a word combination operation with a word combination connection number of 2 and no repetition, to obtain the third vocabulary under the target cluster (i.e., a 2-word combination table, such as Apple & mobile phone, Apple & protective film, etc.), and then the vocabulary difference set between it and the fourth vocabulary under all other target clusters (i.e., a vocabulary that is also a 2-word combination table) is determined, and this is updated as the third vocabulary under the target cluster. Finally, since the preset word combination round threshold is 1, the first vocabulary (1-word combination table) and the third vocabulary (2-word combination table) of the target cluster can be directly determined as the word list under the target cluster.

[0142] It can be seen that this optional embodiment can intelligently generate a corresponding word list from the word list under each target cluster by setting the word combination round, which shows the intelligent processing method of the word list by the text classification system, which is beneficial to improve the reliability and accuracy of the word list obtained under each target cluster, and further helps to improve the mutual exclusivity between the word lists under each target cluster, so as to improve the classification reliability and classification accuracy of the text of the determined cluster.

[0143] In another optional embodiment, the method may further include:

[0144] When it is determined that the current word combination round of the target cluster is less than the preset word combination round threshold, the current word combination round of the target cluster is increased by 1 to update the current word combination round of the target cluster, and the third word list under the target cluster is removed from the word list under the target cluster to obtain the fifth word list under the target cluster; the fifth word list under the target cluster is determined as the second word list under the target cluster, and the execution is triggered to determine the number of word combination connections under the current word combination round of the target cluster, and according to the number of word combination connections under the current word combination round, the word combination operation is performed on the second word list under the target cluster to obtain the third word list under the target cluster; the fourth word list under each other target cluster is obtained, and the word list difference set between the third word list under the target cluster and the fourth word lists under all other target clusters is determined; the word list difference set between the third word list under the target cluster and the fourth word lists under all other target clusters is updated to the third word list under the target cluster, and it is determined whether the current word combination round of the target cluster is greater than or equal to the preset word combination round threshold.

[0145] In this optional embodiment, for example, when it is determined that the number of words in the second vocabulary under the target cluster is greater than or equal to the preset word number threshold, the word combination rounds required for the target cluster and the number of word combination connections under each word combination round can be set. If the number of words in the second vocabulary is large, two word combination rounds are set, and the number of word combination connections in the first word combination round is 2 and the number of word combination connections in the second word combination round is 3. Then, in the first word combination round, a word combination operation with a word combination connection number of 2 and no repetition is performed on the second vocabulary under the target cluster to obtain the third vocabulary under the target cluster (i.e., a 2-word combination table, such as Apple & mobile phone, Apple & protective film, etc.), and then the vocabulary difference set between it and the fourth vocabulary under all other target clusters (i.e., a vocabulary that is also a 2-word combination table) is determined, and this is updated as the third vocabulary under the target cluster. Then, in the second word combination round, the fourth vocabulary under the target cluster in the previous round is removed from the vocabulary under the target cluster. Three word lists (word lists of 2 word combination lists), and the removed word list, that is, the fifth word list, that is, the second word list in the new word combination round, is subjected to a word combination operation with a word combination connection number of 3 and no repetition, to obtain the third word list in the new word combination round (that is, a 3-word combination list, such as Apple & mobile phone & protective film, Apple & protective film & mobile phone case, etc.) and determine the word list difference set between it and the fourth word list under all other target clusters (that is, the word list that is also a 3-word combination list), and update it as the third word list under the target cluster in the new word combination round. Finally, the first word list (1 word combination list) of the target cluster, the third word list (2 word combination list) under the target cluster in the first word combination round, and the third word list (3 word combination list) under the target cluster in the second word combination round are determined as the word list under the target cluster.

[0146] It can be seen that this optional embodiment can intelligently generate corresponding word lists with different word combination connection numbers from the word list under each target cluster, further demonstrating the intelligent processing method of the word list by the text classification system, which is conducive to further improving the reliability and accuracy of the word list obtained under each target cluster, and further conducive to further improving the mutual exclusivity between the word lists under each target cluster, thereby facilitating further improving the classification reliability and classification accuracy of the text to be determined in the cluster.

[0147] Embodiment 3

[0148] See also Figure 3 , Figure 3 Schematic diagram of a word list generating device for text classification disclosed in an embodiment of the present invention. Figure 3 As shown, the word list generating device applied to text classification may include:

[0149] An input module 301 is used to input the sample text into a pre-trained text analysis vector model for analysis to obtain a sentence vector for each sample sentence in the sample text;

[0150] A clustering module 302 is used to perform a clustering operation on all sample sentences based on the sentence vectors of all sample sentences to obtain a sample sentence set under at least one target cluster;

[0151] The word segmentation module 303 is used to perform word segmentation operations on the sample sentence set under each target cluster to obtain a word list under each target cluster.

[0152] In the embodiment of the present invention, the sample sentence set under each target cluster includes at least one sample sentence; the word list is used to determine the cluster to which each sentence in the to-be-determined cluster text belongs.

[0153] It can be seen that the implementation Figure 3 The described word list generation device for text classification can obtain a word list for text classification by clustering and segmenting sample sentences of sample text, which can effectively solve the problem of fuzzy category characteristics of text categories in traditional single-label classification methods, enrich the intelligent text category determination method of the text classification system, and is conducive to improving the mutual exclusivity between the determined text categories, making the category characteristics between text categories clearer, and further conducive to improving the classification reliability and classification accuracy of the text to be analyzed, thereby facilitating the accurate classification of the text to be analyzed.

[0154] In an optional embodiment, the clustering module 302 includes:

[0155] A calculation submodule 3021 is used to calculate, for each sample sentence, a first vector cosine similarity between the sample sentence and each other sample sentence according to the sentence vector of the sample sentence and the sentence vectors of all other sample sentences except the sample sentence;

[0156] A first determination submodule 3022, configured to determine a first sample sentence having the smallest first vector cosine similarity from all other sample sentences;

[0157] A merging submodule 3023 is used to merge the sample sentence with the first sample sentence to obtain a first sample sentence set;

[0158] The clustering submodule 3024 is used to perform sentence clustering operations on all the first sample sentence sets according to the sentence vectors of all the first sample sentence sets after obtaining all the first sample sentence sets, so as to obtain sample sentence sets under at least one target cluster.

[0159] It can be seen that the implementation Figure 4The described word list generation device applied to text classification can preliminarily merge each sample sentence by calculating the vector cosine similarity between the sample sentences, which is beneficial to improving the merging reliability and merging accuracy between the sample sentences, and further beneficial to improving the subsequent clustering reliability and clustering accuracy of the sample sentences, thereby facilitating the accurate generation of the subsequent word list.

[0160] In another optional embodiment, the clustering submodule 3024 includes:

[0161] The calculating unit 30241 is used to calculate, for each first sample sentence set, a second vector cosine similarity between the first sample sentence set and each other first sample sentence set according to the sentence vector of the first sample sentence set and the sentence vectors of all other first sample sentence sets except the first sample sentence set;

[0162] A judging unit 30242 is used to judge whether the cosine similarities of all second vectors corresponding to any first sample sentence set are less than a preset similarity threshold;

[0163] The calculation unit 30241 is further configured to calculate the sentence set center vector of each first sample sentence set when the judgment result of the judgment unit 30242 is yes;

[0164] A determining unit 30243 is used to determine the number of sample sentences in each first sample sentence set, and determine all second sample sentence sets whose number of sample sentences is less than or equal to a preset sample sentence number threshold and all third sample sentence sets whose number of sample sentences is greater than the preset sample sentence number threshold from all first sample sentence sets;

[0165] The calculation unit 30241 is further used to calculate, for each second sample sentence set, the center vector cosine similarity between the second sample sentence set and each third sample sentence set according to the sentence set center vector of the second sample sentence set and the sentence set center vector of each third sample sentence set;

[0166] The determining unit 30243 is further configured to determine a fourth sample sentence set having the largest center vector cosine similarity from all third sample sentence sets;

[0167] A merging unit 30244, configured to merge the fourth sample sentence set with the second sample sentence set to obtain a merged fourth sample sentence set;

[0168] An updating unit 30245, configured to update all third sample sentence sets according to all fourth sample sentence sets obtained after merging;

[0169] The determining unit 30243 is further configured to determine, based on all sample sentences included in each third sample sentence set, the target cluster to which each third sample sentence set belongs, as the sample sentence set under each target cluster.

[0170] It can be seen that the implementation Figure 4 The described word list generation device for text classification can obtain sample sentence sets under each target cluster by performing secondary clustering on sample sentence sets. In this way, the mutual exclusivity between sample sentence sets under each target cluster can be improved, and then the category characteristics between the word lists under each target cluster obtained subsequently can be improved, thereby solving the problem that the category characteristics of text categories in traditional text classification methods are too vague.

[0171] In yet another optional embodiment, the determining unit 30243 is further configured to:

[0172] When the judging unit 30242 judges that all the second vector cosine similarities corresponding to any first sample sentence set are not less than the preset similarity threshold, all fifth sample sentence sets whose second vector cosine similarities are greater than the preset similarity threshold are determined from all other first sample sentence sets corresponding to the first sample sentence set;

[0173] The merging unit 30244 is further configured to merge the first sample sentence set with each fifth sample sentence set to update all fifth sample sentence sets;

[0174] The determining unit 30243 is further used to determine all the updated fifth sample sentence sets as all the first sample sentence sets after updating all the fifth sample sentence sets, and trigger the calculation unit 30241 to perform an operation of calculating, for each first sample sentence set, the cosine similarity of the second vector between the first sample sentence set and each other first sample sentence set according to the sentence vector of the first sample sentence set and the sentence vectors of all other first sample sentence sets except the first sample sentence set, and trigger the judgment unit 30242 to perform an operation of judging whether all the cosine similarities of the second vectors corresponding to any first sample sentence set are less than a preset similarity threshold.

[0175] It can be seen that the implementation Figure 4 The described word list generation device applied to text classification can automatically determine whether the preliminary clustering operation of the sample sentence set has been completed, showing the intelligent determination method of the text classification system for the clustering of the sample sentence set, which is conducive to improving the clustering reliability and clustering accuracy of the sample sentence set, and further helps to ensure the mutual exclusivity between the obtained sample sentence sets, thereby helping to clarify the sentence characteristics between the sample sentence sets.

[0176] In yet another optional embodiment, the word segmentation module 303 includes:

[0177] The word segmentation submodule 3031 is used to perform word segmentation operations on the sample sentence set under each target cluster through a preset word segmenter to obtain a word list under each target cluster;

[0178] The second determination submodule 3032 is used to determine, for each vocabulary under the target cluster, a vocabulary difference set between the target cluster and all other target clusters, and determine the vocabulary difference set between the target cluster and all other target clusters as the first vocabulary under the target cluster;

[0179] The vocabulary removal submodule 3033 is used to remove the first vocabulary under the target cluster from the vocabulary under the target cluster to obtain the second vocabulary under the target cluster;

[0180] A judgment submodule 3034 is used to judge whether the number of all words contained in the second vocabulary under the target cluster is greater than or equal to a preset word number threshold;

[0181] The second determining submodule 3032 is further configured to determine the first word list under the target cluster as the word list under the target cluster when the determination result of the determining submodule 3034 is no.

[0182] In this optional embodiment, the vocabulary under each target cluster includes multiple words.

[0183] It can be seen that the implementation Figure 4 The described word list generation device for text classification can segment the sample sentence set under each target cluster through a preset word segmenter. In this way, the reliability and accuracy of the word list obtained under each target cluster can be guaranteed, and the mutual exclusivity between the word lists under each target cluster can be guaranteed, thereby ensuring the reliability and accuracy of the cluster to which each sentence in the analyzed cluster text to be determined belongs.

[0184] In yet another optional embodiment, the second determining submodule 3032 is further configured to:

[0185] When the judging submodule 3034 judges that the number of all words contained in the second vocabulary under the target cluster is greater than or equal to the preset word number threshold, the number of word combination connections under the current word combination round of the target cluster is determined;

[0186] The word combination submodule 3035 is used to perform a word combination operation on the second word list under the target cluster according to the word combination connection number under the current word combination round to obtain a third word list under the target cluster;

[0187] An acquisition submodule 3036 is used to acquire a fourth vocabulary under each other target cluster;

[0188] The second determining submodule 3032 is further used to determine a vocabulary difference set between the third vocabulary under the target cluster and the fourth vocabulary under all other target clusters;

[0189] An updating submodule 3037, configured to update a vocabulary difference set between the third vocabulary under the target cluster and the fourth vocabulary under all other target clusters to the third vocabulary under the target cluster;

[0190] The judging submodule 3034 is further used to judge whether the current word combination round of the target cluster is greater than or equal to a preset word combination round threshold;

[0191] The second determining submodule 3032 is further configured to determine the first vocabulary list under the target cluster and the third vocabulary list under the target cluster as the word list under the target cluster when the determination result of the determining submodule 3034 is yes.

[0192] In this optional embodiment, the number of word combination connections corresponding to the fourth vocabulary under each other target cluster matches the number of word combination connections corresponding to the third vocabulary under the target cluster.

[0193] It can be seen that the implementation Figure 4 The described word list generation device applied to text classification can intelligently generate a corresponding word list from the word list under each target cluster by setting the word combination rounds, which shows the intelligent processing method of the word list by the text classification system, which is beneficial to improve the reliability and accuracy of the word list obtained under each target cluster, and further helps to improve the mutual exclusivity between the word lists under each target cluster, so as to improve the classification reliability and classification accuracy of the text to be determined in the cluster.

[0194] In yet another optional embodiment, the updating submodule 3037 is further used to:

[0195] When the judging submodule 3034 judges that the current word combination round of the target cluster is less than the preset word combination round threshold, the current word combination round of the target cluster is increased by 1 to update the current word combination round of the target cluster;

[0196] The vocabulary removal submodule 3033 is further used to remove the third vocabulary under the target cluster from the vocabulary under the target cluster to obtain the fifth vocabulary under the target cluster;

[0197] The second determination submodule 3032 is also used to determine the fifth vocabulary under the target cluster as the second vocabulary under the target cluster, and trigger the operation of determining the number of word combination connections under the current word combination round of the target cluster, and trigger the word combination submodule 3035 to perform a word combination operation on the second vocabulary under the target cluster according to the number of word combination connections under the current word combination round to obtain the third vocabulary under the target cluster, and trigger the acquisition submodule 3036 to acquire the fourth vocabulary under each other target cluster, and trigger the operation of determining the vocabulary difference set between the third vocabulary under the target cluster and the fourth vocabulary under all other target clusters, and trigger the update submodule 3037 to update the vocabulary difference set between the third vocabulary under the target cluster and the fourth vocabulary under all other target clusters to the third vocabulary under the target cluster, and trigger the judgment submodule 3034 to judge whether the current word combination round of the target cluster is greater than or equal to the preset word combination round threshold.

[0198] It can be seen that the implementation Figure 4 The described word list generation device applied to text classification can intelligently generate corresponding word lists with different word combination connection numbers from the word list under each target cluster, further demonstrating the intelligent processing method of the word list by the text classification system, which is conducive to further improving the reliability and accuracy of the word list obtained under each target cluster, and further conducive to further improving the mutual exclusivity between the word lists under each target cluster, thereby facilitating further improving the classification reliability and classification accuracy of the text to be determined in the cluster.

[0199] Embodiment 4

[0200] See also Figure 5 , Figure 5 FIG. 1 is a schematic diagram of the structure of another device for generating a word list for text classification disclosed in an embodiment of the present invention. Figure 5 As shown, the word list generating device applied to text classification may include:

[0201] A memory 401 storing executable program codes;

[0202] a processor 402 coupled to the memory 401;

[0203] The processor 402 calls the executable program code stored in the memory 401 to execute the steps of the method for generating a word list for text classification described in the first embodiment of the present invention or the second embodiment of the present invention.

[0204] Embodiment 5

[0205] An embodiment of the present invention discloses a computer storage medium storing computer instructions. When the computer instructions are called, they are used to execute the steps of the word list generation method for text classification described in the first embodiment of the present invention or the second embodiment of the present invention.

[0206] Embodiment 6

[0207] An embodiment of the present invention discloses a computer program product, which includes a non-transitory computer-readable storage medium storing a computer program, and the computer program is operable to cause a computer to execute the steps of the word list generation method for text classification described in Embodiment 1 or Embodiment 2.

[0208] The device embodiments described above are only illustrative, wherein the modules described as separate components may or may not be physically separated, and the components displayed as modules may or may not be physical modules, i.e., they may be located in one place, or they may be distributed on multiple network modules. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Those of ordinary skill in the art may understand and implement it without creative work.

[0209] Through the specific description of the above embodiments, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the above technical solution can be essentially or partly contributed to the prior art in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, and the storage medium includes a read-only memory (ROM), a random access memory (RAM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), a one-time programmable read-only memory (OTPROM), an electronically erasable rewritable read-only memory (EEPROM), a compact disc (CD-ROM) or other optical disc storage, magnetic disk storage, magnetic tape storage, or any other computer-readable medium that can be used to carry or store data.

[0210] Finally, it should be noted that the method and device for generating a word list for text classification disclosed in the embodiment of the present invention only discloses a preferred embodiment of the present invention, which is only used to illustrate the technical solution of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, a person skilled in the art should understand that the technical solutions described in the aforementioned embodiments can still be modified, or some of the technical features therein can be replaced by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A word list generation method for text classification, characterized in that: The method comprises: Input the sample text into a pre-trained text analysis vector model for analysis to obtain a sentence vector for each sample sentence in the sample text; Based on the sentence vectors of all the sample sentences, a clustering operation is performed on all the sample sentences to obtain a sample sentence set under at least one target cluster; each sample sentence set under the target cluster includes at least one sample sentence; Performing a word segmentation operation on the sample sentence set under each of the target clusters to obtain a word list under each of the target clusters; the word list is used to determine the cluster to which each sentence in the to-be-determined cluster text belongs; The word segmentation operation is performed on the sample sentence set under each target cluster to obtain a word list under each target cluster, including: By using a preset word segmenter, a word segmentation operation is performed on the sample sentence set under each of the target clusters to obtain a word list under each of the target clusters; the word list under each of the target clusters includes a plurality of words; For each vocabulary under the target cluster, determine the vocabulary difference set between the target cluster and all other target clusters, and determine the vocabulary difference set between the target cluster and all other target clusters as the first vocabulary under the target cluster; Remove the first word list under the target cluster from the word list under the target cluster to obtain the second word list under the target cluster, and determine whether the number of words contained in the second word list under the target cluster is greater than or equal to a preset word number threshold. If not, determine the first word list under the target cluster as the word list under the target cluster.

2. The word list generation method for text classification according to claim 1, characterized in that: The step of performing a clustering operation on all the sample sentences based on the sentence vectors of all the sample sentences to obtain a sample sentence set under at least one target cluster includes: For each of the sample sentences, according to the sentence vector of the sample sentence and the sentence vectors of all other sample sentences except the sample sentence, calculate the first vector cosine similarity between the sample sentence and each of the other sample sentences; determine the first sample sentence with the smallest first vector cosine similarity from all the other sample sentences, and merge the sample sentence with the first sample sentence to obtain a first sample sentence set; After all the first sample sentence sets are obtained, sentence clustering operations are performed on all the first sample sentence sets according to the sentence vectors of all the first sample sentence sets to obtain sample sentence sets under at least one target cluster.

3. The word list generation method for text classification according to claim 2, characterized in that: The step of performing sentence clustering operations on all the first sample sentence sets according to the sentence vectors of all the first sample sentence sets to obtain a sample sentence set under at least one target cluster includes: For each of the first sample sentence sets, calculating the second vector cosine similarity between the first sample sentence set and each of the other first sample sentence sets according to the sentence vector of the first sample sentence set and the sentence vectors of all other first sample sentence sets except the first sample sentence set; Determine whether the cosine similarities of all the second vectors corresponding to any of the first sample sentence sets are less than a preset similarity threshold, and when the judgment result is yes, calculate the sentence set center vector of each of the first sample sentence sets and determine the number of sample sentences of each of the first sample sentence sets, and determine all second sample sentence sets whose number of sample sentences is less than or equal to a preset sample sentence number threshold and all third sample sentence sets whose number of sample sentences is greater than the preset sample sentence number threshold from all the first sample sentence sets; For each of the second sample sentences, according to the sentence center vector of the second sample sentence and the sentence center vector of each of the third sample sentences, calculate the center vector cosine similarity between the second sample sentence and each of the third sample sentences; determine a fourth sample sentence with the largest center vector cosine similarity from all the third sample sentences, and merge the fourth sample sentence with the second sample sentence to obtain the merged fourth sample sentence; According to all the fourth sample sentence sets obtained after the merger, all the third sample sentence sets are updated, and according to all the sample sentences contained in each of the third sample sentence sets, the target cluster to which each of the third sample sentence sets belongs is determined as the sample sentence set under each of the target clusters.

4. The word list generation method for text classification according to claim 3 is characterized in that: The method further comprises: When it is determined that all the second vector cosine similarities corresponding to any one of the first sample sentence sets are not all less than the preset similarity threshold, all fifth sample sentence sets whose second vector cosine similarities are greater than the preset similarity threshold are determined from all the other first sample sentence sets corresponding to the first sample sentence set, and the first sample sentence set is merged with each of the fifth sample sentence sets to update all the fifth sample sentence sets; After all the fifth sample sentence sets are updated, all the updated fifth sample sentence sets are determined as all the first sample sentence sets, and the operations of calculating, for each of the first sample sentence sets, the second vector cosine similarity between the first sample sentence set and each of the other first sample sentence sets according to the sentence vector of the first sample sentence set and the sentence vectors of all other first sample sentence sets except the first sample sentence set, and determining whether all the second vector cosine similarities corresponding to any of the first sample sentence sets are less than a preset similarity threshold are triggered.

5. The method for generating a word list for text classification according to any one of claims 1 to 4, characterized in that: The method further comprises: When it is determined that the number of all the words contained in the second vocabulary under the target cluster is greater than or equal to the preset word number threshold, determining the number of word combination connections under the current word combination round of the target cluster, and performing a word combination operation on the second vocabulary under the target cluster according to the number of word combination connections under the current word combination round to obtain a third vocabulary under the target cluster; Acquire the fourth vocabulary under each of the other target clusters, and determine the vocabulary difference set between the third vocabulary under the target cluster and the fourth vocabulary under all the other target clusters; the number of word combination connections corresponding to the fourth vocabulary under each of the other target clusters matches the number of word combination connections corresponding to the third vocabulary under the target cluster; The vocabulary difference set between the third vocabulary under the target cluster and the fourth vocabulary under all other target clusters is updated to the third vocabulary under the target cluster, and it is determined whether the current word combination round of the target cluster is greater than or equal to a preset word combination round threshold. When the judgment result is yes, the first vocabulary under the target cluster and the third vocabulary under the target cluster are determined as the word list under the target cluster.

6. The word list generation method for text classification according to claim 5, characterized in that: The method further comprises: When it is determined that the current word combination round of the target cluster is less than the preset word combination round threshold, the current word combination round of the target cluster is increased by 1 to update the current word combination round of the target cluster, and the third word list under the target cluster is removed from the word list under the target cluster to obtain the fifth word list under the target cluster; the fifth word list under the target cluster is determined as the second word list under the target cluster, and the determination of the number of word combination connections under the current word combination round of the target cluster is triggered, and the word combination connection number under the current word combination round is determined according to the word combination under the current word combination round. Combine the number of connections, perform a word combination operation on the second word list under the target cluster to obtain the third word list under the target cluster; obtain the fourth word list under each of the other target clusters, and determine the word list difference set between the third word list under the target cluster and the fourth word lists under all the other target clusters; update the word list difference set between the third word list under the target cluster and the fourth word lists under all the other target clusters to the third word list under the target cluster, and determine whether the current word combination round of the target cluster is greater than or equal to the preset word combination round threshold.

7. A word list generation device for text classification, characterized in that: The device comprises: An input module, used to input the sample text into a pre-trained text analysis vector model for analysis, and obtain a sentence vector for each sample sentence in the sample text; A clustering module, configured to perform a clustering operation on all the sample sentences based on the sentence vectors of all the sample sentences to obtain a sample sentence set under at least one target cluster; each sample sentence set under the target cluster includes at least one sample sentence; A word segmentation module is used to perform word segmentation operations on the sample sentence set under each of the target clusters to obtain a word list under each of the target clusters; the word list is used to determine the cluster to which each sentence in the to-be-determined cluster text belongs; Wherein, the word segmentation module includes: A word segmentation submodule is used to perform a word segmentation operation on the sample sentence set under each target cluster through a preset word segmenter to obtain a word list under each target cluster; the word list under each target cluster includes a plurality of words; A second determination submodule is used to determine, for each vocabulary under the target cluster, a vocabulary difference set between the target cluster and all other target clusters, and determine the vocabulary difference set between the target cluster and all other target clusters as the first vocabulary under the target cluster; A vocabulary removal submodule, used to remove the first vocabulary under the target cluster from the vocabulary under the target cluster to obtain a second vocabulary under the target cluster; A judgment submodule, used to judge whether the number of all the words contained in the second vocabulary under the target cluster is greater than or equal to a preset word number threshold; The second determination submodule is further configured to determine the first word list under the target cluster as the word list under the target cluster when the determination result of the determination submodule is no.

8. A word list generation device for text classification, characterized in that: The device comprises: A memory storing executable program code; a processor coupled to the memory; The processor calls the executable program code stored in the memory to execute the word list generation method for text classification according to any one of claims 1 to 6.

9. A computer storage medium, characterized in that The computer storage medium stores computer instructions, which, when called, are used to execute the word list generation method for text classification as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Text sentiment analysis method and device, storage medium and electronic equipment

    CN110413780A

  • Label extraction method based on short text clustering technology

    CN111414479A