Multi-layer text clustering method, system, device and storage medium

By performing word segmentation and IDF score dictionary processing on multi-layer text, combined with hierarchical TFIDF values, the problem of low correlation between lower text and upper text is solved, achieving a more accurate clustering effect and a better reading experience.

CN114036305BActive Publication Date: 2025-07-25CHINA PING AN LIFE INSURANCE CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202111434730.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-29
Publication Date
2025-07-25
Estimated Expiration
2041-11-29

AI Technical Summary

Technical Problem

When clustering multi-layer texts in the prior art, the correlation between the lower text and the upper text is low, resulting in poor clustering effect and affecting the reading experience.

Method used

By performing word segmentation on multi-layer text, the target words and their word frequency values are obtained, and the IDF score dictionary and hierarchical TFIDF values are clustered. Combining the intersection relationship between first-level text, second-level text and third-level text, the relevance of text is improved.

Benefits of technology

It improves the accuracy and relevance of multi-layer text clustering, improves the clustering effect, enhances the correlation between lower text and upper text, and improves the user's reading experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114036305B_ABST
    Figure CN114036305B_ABST
Patent Text Reader

Abstract

The present invention relates to artificial intelligence and provides a multi-layer text clustering method, which includes: obtaining first-level texts with first-level attributes and second-level texts with second-level attributes, where the second-level texts include multiple third-level texts; performing word segmentation on the third-level texts to obtain target words and third-level word frequency values, and the target words correspond to the third-level word frequency values; performing word and character statistics processing on the first-level texts and the second-level texts respectively according to the target words to obtain a first-level IDF score dictionary and a second-level IDF score dictionary; for each target word, obtaining a third-level word-level IDF score according to the first-level IDF score dictionary and the second-level IDF score dictionary; obtaining multiple hierarchical TFIDF values according to the third-level word-level IDF scores and the third-level word frequency values; performing clustering processing on the third-level texts according to the multiple hierarchical TFIDF values to obtain a clustering result. The present invention performs clustering on multi-layer texts, improves the relevance between lower-layer texts and upper-layer texts, and improves the clustering effect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and particularly relates to a multi-layer text clustering method, system, device and storage medium. Background Art

[0002] Currently, in order to help readers select their favorite types for reading, the types of articles are screened and classified, such as news reports, scientific research, military consultations, etc. Due to the rapid development of the Internet, the number of articles appearing on the Internet every day is huge, and the types of articles are also diverse. In the related art, clustering is used to classify the types of articles, but the clustering effect for multi-layer texts is poor. Even if the articles are clustered into multiple types, there are a large number of articles in one type at the lower layer, the scope of the type is large, and the relevance between the lower-layer articles and the type they belong to is low, resulting in scattered content of the lower-layer articles and affecting the reading experience of readers. Summary of the Invention

[0003] The following is an overview of the subject matter described in detail in this article. This overview is not intended to limit the scope of protection of the claims.

[0004] Embodiments of the present invention provide a multi-layer text clustering method, system, device and storage medium, which can cluster multi-layer texts, improve the relevance between the lower-layer texts and the upper-layer texts, and improve the clustering effect.

[0005] In a first aspect, an embodiment of the present invention provides a multi-layer text clustering method, which includes:

[0006] Obtain first-level texts with first-level attributes and second-level texts with second-level attributes, where the first-level attributes include the second-level attributes, and the second-level texts include multiple third-level texts;

[0007] Perform word segmentation processing on the third-level texts in the second-level texts to obtain target words and third-level word frequency values, and the target words correspond to the third-level word frequency values;

[0008] Perform word and character statistics processing on the first-level texts according to the target words to obtain a first-level IDF score dictionary, and perform word and character statistics processing on the second-level texts according to the target words to obtain a second-level IDF score dictionary;

[0009] For each of the target words, obtain a third-level word-level IDF score according to the first-level IDF score dictionary and the second-level IDF score dictionary;

[0010] Obtain multiple hierarchical TFIDF values according to the third-level word-level IDF score and the third-level word frequency value;

[0011] Cluster the third-level text based on the TFIDF values of multiple levels to obtain a clustering result.

[0012] According to some embodiments of the present invention, in the above multi-level text clustering method, the process of performing word segmentation on the third-level text in the second-level text to obtain target words and third-level word frequency values includes:

[0013] Perform word segmentation on the third-level text to obtain multiple target words;

[0014] For each of the target words, perform statistical processing on the target word according to the third-level text to obtain the number of occurrences of the word;

[0015] According to all the number of occurrences of the words and a preset upper limit truncation value, obtain the third-level word frequency values.

[0016] Extract the target words from the third-level text and perform statistical processing on the target words to obtain the number of occurrences of each target word. Use the preset upper limit truncation value to perform upper limit truncation processing on all the number of occurrences of the words to avoid too high an upper limit of the third-level word frequency value, which affects the discrimination ability of the target word and reduces the clustering effect.

[0017] According to some embodiments of the present invention, in the above multi-level text clustering method, the preset upper limit truncation value includes a truncation coefficient, a correction coefficient, a first compensation value, and a second compensation value. The process of obtaining the third-level word frequency values according to all the number of occurrences of the words and the preset upper limit truncation value includes:

[0018] For each of the target words, process the number of occurrences of the word according to the truncation coefficient and the first compensation value to obtain a first word frequency value;

[0019] Process the sum of all the number of occurrences of the words according to the correction coefficient and the second compensation value to obtain a second word frequency value;

[0020] Obtain multiple third-level word frequency values according to the ratio of each of the first word frequency values to the second word frequency value.

[0021] Use the truncation coefficient and the correction coefficient to perform upper limit truncation processing on the third-level word frequency values to limit their maximum values, avoid too high a number of occurrences of the words, which affects the third-level word frequency values, and perform correction compensation on the third-level word frequency values through the compensation values to avoid deviation of the third-level word frequency values when the number of occurrences of the words is low, thereby improving the accuracy of the third-level word frequency values and improving the clustering effect.

[0022] According to some embodiments of the present invention, in the above multi-level text clustering method, for each of the target words, obtaining a three-level word-level IDF score according to the first-level IDF score dictionary and the second-level IDF score dictionary includes:

[0023] Obtaining a hierarchical weight value according to the cross-relationship between the first-level text, the second-level text, and the third-level text;

[0024] Processing the first-level IDF score dictionary according to the hierarchical weight value and each of the target words to obtain a plurality of first-level hierarchical scores;

[0025] Processing the second-level IDF score dictionary according to the hierarchical weight value and each of the target words to obtain a plurality of second-level hierarchical scores;

[0026] For each of the target words, processing the first-level hierarchical score and the second-level hierarchical score to obtain a three-level word IDF score.

[0027] Different first-level texts include different second-level texts, and different second-level texts include different third-level texts. By combining the cross-relationship between the first-level text, the second-level text, and the third-level text, and integrating the discrimination ability of each target word in different texts, the IDF score of the target word in the third-level text is obtained, improving the relevance with the upper-level text and improving the clustering effect.

[0028] According to some embodiments of the present invention, in the above multi-level text clustering method, there are a plurality of first-level texts and a plurality of second-level texts. Obtaining a hierarchical weight value according to the cross-relationship between the first-level text, the second-level text, and the third-level text includes:

[0029] Counting the number of the first-level texts according to the cross-relationship between the third-level text and the first-level text to obtain a first text number;

[0030] Counting the number of the second-level texts according to the cross-relationship between the third-level text and the second-level text to obtain a second text number;

[0031] Obtaining a hierarchical weight value according to the ratio of the first text number to the second text number.

[0032] Since the primary attributes of the primary text include the secondary attributes of the secondary text, and the secondary text includes multiple tertiary texts, different primary texts include different secondary texts, and different secondary texts include different tertiary texts. By combining the cross-relationships among the primary text, secondary text, and tertiary text, calculate the number of primary texts and secondary texts containing the tertiary text to obtain a hierarchical weight value, thereby improving the relevance to the upper-layer text and enhancing the clustering accuracy.

[0033] According to some embodiments of the present invention, in the above multi-layer text clustering method, obtaining a plurality of hierarchical TFIDF values according to the tertiary word-level IDF score and the tertiary word frequency value includes:

[0034] Calculate a plurality of hierarchical TFIDF values according to the ratio between the tertiary word frequency value and the tertiary word-level IDF score.

[0035] For target words with a high occurrence frequency in the tertiary text but a low document frequency in the upper-layer text, increasing their hierarchical TFIDF values can filter out common words, retain important words related to the upper-layer text, improve the relevance to the upper-layer text, and enhance the clustering effect.

[0036] According to some embodiments of the present invention, in the above multi-layer text clustering method, performing clustering processing according to the hierarchical TFIDF values to obtain a clustering result includes:

[0037] Import the hierarchical TFIDF values into a similar hashing neural model for processing to obtain a plurality of text representation vectors, where the text representation vectors correspond to the tertiary text;

[0038] Cluster the tertiary text according to the distance between any two of the text representation vectors to obtain a clustering result.

[0039] Using a similar hashing neural network to process target words with their corresponding hierarchical TFIDF values to obtain the representation vectors of the corresponding tertiary text. Clustering the tertiary text based on the distance can improve the clustering effect.

[0040] In a second aspect, an embodiment of the present invention provides a multi-layer text clustering system, including:

[0041] A text acquisition module, configured to acquire a primary text with primary attributes and a secondary text with secondary attributes, where the primary attributes include the secondary attributes, and the secondary text includes a plurality of tertiary texts;

[0042] A text word segmentation module, configured to perform word segmentation processing on the third-level text in the second-level text to obtain target words and third-level word frequency values, where the target words correspond to the third-level word frequency values;

[0043] A word and character statistics module, configured to perform word and character statistics processing on the first-level text according to the target words to obtain a first-level IDF score dictionary, and perform word and character statistics processing on the second-level text according to the target words to obtain a second-level IDF score dictionary;

[0044] A word-level score calculation module, configured to obtain a third-level word-level IDF score corresponding to each target word according to the first-level IDF score dictionary and the second-level IDF score dictionary;

[0045] A level value calculation module, configured to obtain a plurality of level TFIDF values according to the third-level word-level IDF scores and the third-level word frequency values;

[0046] A multi-level clustering module, configured to perform clustering processing according to the plurality of level TFIDF values to obtain a clustering result.

[0047] In a third aspect, an embodiment of the present invention provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the multi-level text clustering method as described in the first aspect above is implemented.

[0048] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium, storing a computer program, where when the computer program is executed by a processor, the multi-level text clustering method as described in the first aspect above is implemented.

[0049] The multi-layer text clustering method according to the embodiments provided by the present invention has at least the following beneficial effects: The first-level text has first-level attributes, and the second-level text has second-level attributes. The first-level attributes include the second-level attributes. Different first-level attributes include different second-level attributes. The first-level text and the second-level text are in an intersection relationship, and the second-level text further includes multiple third-level texts. By performing word segmentation on the third-level texts, target words and their corresponding third-level word frequency values are extracted. Moreover, the target words obtained by extraction are used to perform word statistics on the first-level text and the second-level text respectively, obtaining corresponding first-level IDF score dictionaries and second-level IDF score dictionaries. Combining the discrimination ability of the target words in the first-level text and the second-level text, the discrimination ability of each target word in the third-level text is calculated, so that target words with strong discrimination ability can be selected as the representation vectors of the third-level text for clustering. By using the intersection relationship among the first-level text, the second-level text, and the third-level text and performing clustering according to the attributes of the upper-level text, the relevance between the third-level text and the upper-level text can be improved, thereby improving the clustering accuracy and the clustering effect.

[0050] Other features and advantages of the present invention will be described in the following specification. Moreover, some of them will become obvious from the specification or can be understood by implementing the present invention. The objectives and other advantages of the present invention can be achieved and obtained through the structures specifically pointed out in the specification, the claims, and the drawings. Brief Description of the Drawings

[0051] The drawings are used to provide a further understanding of the technical solutions of the present invention and constitute a part of the specification. They are used together with the embodiments of the present invention to explain the technical solutions of the present invention and do not constitute a limitation to the technical solutions of the present invention.

[0052] Figure 1 is a flowchart of the multi-layer text clustering method provided by the embodiments of the present invention;

[0053] Figure 2 is Figure 1 a schematic diagram of the specific implementation process of step S200 in

[0054] Figure 3 is Figure 2 a schematic diagram of the specific implementation process of step S230 in

[0055] Figure 4 is Figure 1 a schematic diagram of the specific implementation process of step S400 in

[0056] Figure 5 is Figure 4 a schematic diagram of the specific implementation process of step S410 in

[0057] Figure 6It is a flowchart of a multi - layer text clustering method provided by another embodiment of the present invention;

[0058] Figure 7 is Figure 1 a schematic diagram of the specific implementation process of step S600 in

[0059] Figure 8 It is a schematic structural diagram of a multi - layer text clustering system provided by an embodiment of the present invention;

[0060] Figure 9 It is a schematic structural diagram of an electronic device provided by an embodiment of the present invention. Detailed implementation manners

[0061] In order to make the objectives, technical solutions and advantages of the present invention more clear and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0062] It should be noted that although functional module division is carried out in the module schematic diagram and the logical sequence is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order from the module division in the module or the sequence in the flowchart. Terms such as "first", "second", etc. in the specification, claims and the above - mentioned drawings are used to distinguish similar objects and do not necessarily need to describe a specific order or sequence.

[0063] The present invention relates to artificial intelligence and provides a multi - layer text clustering method. It obtains first - level texts with first - level attributes and second - level texts with second - level attributes. The first - level attributes include second - level attributes, and the second - level texts include multiple third - level texts; perform word segmentation processing on the third - level texts to obtain target words and third - level word frequency values, and the target words correspond to the third - level word frequency values; perform word and character statistics processing on the first - level texts and the second - level texts respectively according to the target words to obtain a first - level IDF score dictionary and a second - level IDF score dictionary; for each target word, obtain the third - level word - level IDF score according to the first - level IDF score dictionary and the second - level IDF score dictionary; obtain multiple level TFIDF values according to the third - level word - level IDF score and the third - level word frequency values; perform clustering processing on the third - level texts according to the multiple level TFIDF values to obtain a clustering result. The present invention performs clustering on multi - layer texts, utilizes the cross - relationship between the first - level texts, the second - level texts and the third - level texts, and performs clustering according to the attributes of the upper - level texts, improves the relevance between the lower - level texts and the upper - level texts, thereby improving the clustering accuracy and the clustering effect.

[0064] Embodiments of the present invention can acquire and process relevant data based on artificial intelligence technology. Among them, artificial intelligence (AI) is a theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results.

[0065] It should be noted that artificial intelligence technology also includes dividing a dataset into different classes or clusters according to a specific standard, such as the distance criterion, so that the similarity of data objects within the same cluster is as large as possible, while the difference of data objects not in the same cluster is also as large as possible. That is, after clustering, data of the same class is gathered together as much as possible, and different data is separated as much as possible, namely clustering.

[0066] Cluster analysis is a statistical analysis method for studying classification problems and is also an important algorithm in data mining. Cluster analysis consists of several patterns. Usually, a pattern is a vector of metrics or a point in a multi-dimensional space. Cluster analysis is based on similarity, and there is more similarity between patterns in a cluster than between patterns not in the same cluster.

[0067] Clustering has a wide range of applications. For example, in business, clustering can help market analysts distinguish different consumer groups from a consumer database and summarize the consumption patterns or habits of each group of consumers. As a module in data mining, it can be used as a separate tool to discover some deep information distributed in the database, summarize the characteristics of each class, or focus on a specific class for further analysis; moreover, cluster analysis can also be used as a preprocessing step for other analysis algorithms in data mining algorithms.

[0068] Term Frequency–Inverse Document Frequency (TF-IDF) is a commonly used weighting technique for information retrieval and data mining. For example, the clustering data can be text data. In order to cluster the text data, it is necessary to mine the topics of the text data. Therefore, it is necessary to segment the text data. If a word or phrase has a high frequency TF in an article, that is, the clustering data, and rarely appears in other articles, then it is considered that this word or phrase has good category discrimination ability and is suitable for classifying the clustering data. Term frequency TF represents the frequency of a term in the first document. If the number of documents containing the first term is less, then the inverse document frequency IDF is larger, indicating that the first term has good category discrimination ability.

[0069] Refer to Figure 1 , Figure 1The flowchart of the multi - layer text clustering method provided by an embodiment of the present invention is shown. The multi - layer text clustering method includes but is not limited to the following steps:

[0070] Step S100, obtain first - level texts with first - level attributes and second - level texts with second - level attributes, where the first - level attributes include second - level attributes, and the second - level texts include multiple third - level texts;

[0071] Step S200, perform word segmentation on the third - level texts in the second - level texts to obtain target words and third - level word frequency values, and the target words correspond to the third - level word frequency values;

[0072] Step S300, perform word and character statistics processing on the first - level texts according to the target words to obtain a first - level IDF score dictionary, and perform word and character statistics processing on the second - level texts according to the target words to obtain a second - level IDF score dictionary;

[0073] Step S400, for each target word, obtain the third - level word - level IDF score according to the first - level IDF score dictionary and the second - level IDF score dictionary;

[0074] Step S500, obtain multiple - level TFIDF values according to the third - level word - level IDF scores and the third - level word frequency values;

[0075] Step S600, perform clustering processing on the third - level texts according to the multiple - level TFIDF values to obtain a clustering result.

[0076] It can be understood that the first - level texts have first - level attributes, the second - level texts have second - level attributes, the first - level attributes include multiple second - level attributes, and the second - level texts include multiple third - level texts. That is, the first - level texts can contain multiple second - level texts and also multiple third - level texts. Currently, with the rapid development of the network, the number of articles on the network is huge, and the types of articles are diverse. The articles are classified into multiple first - level attributes through clustering, such as sports, entertainment, economy, and technology, etc. And each first - level attribute includes multiple second - level attributes. For example, in the sports first - level attribute, there are also multiple second - level attributes such as basketball, football, and badminton. However, the number of texts with the same second - level attribute is still huge and the scope is wide. The second - level texts also include multiple third - level texts. For example, in the basketball articles, there are also articles about American professional basketball games, Chinese professional basketball games, articles about each team, and articles about each player, etc. The results obtained by directly performing clustering analysis on the lower - level attributes are poor, and the relevance between the lower - level classification and the higher - level attributes is low, resulting in the dispersion of the article content of the lower - level attributes and affecting the user's reading experience.

[0077] Therefore, in order to improve the content concentration of each article in the secondary text and avoid the low association between the lower-level classification results obtained by direct clustering and the attributes of the upper-level text, the cross-relationship between the primary text, the secondary text, and the tertiary text in the secondary text is combined to cluster the tertiary text, so as to improve the relevance between the tertiary text and the primary and secondary attributes. The tertiary text included in the secondary text is segmented to obtain multiple target words, and each target word is statistically analyzed to obtain the tertiary word frequency value of each target word. The word statistics of each target word extracted from the tertiary text are respectively performed in the primary text and the secondary text, and the inverse document frequency (IDF) score dictionaries of each target word in the primary text (i.e., the primary IDF score dictionary) and the secondary IDF score dictionary are obtained. Among them, the primary IDF score dictionary is obtained by statistically analyzing each target word, recording the number of occurrences of the target word, obtaining the IDF score corresponding to each target word in the primary text, and organizing the IDF scores of all target words. Therefore, for each target word, the primary category discrimination ability of the target word in the primary text can be obtained through the primary IDF score dictionary, and the secondary category discrimination ability of the target word in the secondary text can be obtained through the secondary IDF score dictionary. By combining the primary category discrimination ability and the secondary category discrimination ability, the document occurrence frequency of the target word in the tertiary text is obtained, that is, the IDF score at the tertiary word level. Since the secondary text includes multiple tertiary texts, and the primary attribute of the primary text includes the secondary attribute of the secondary text, the category discrimination ability of the target word in the tertiary text can be reflected by combining the IDF score of the target word in the primary text and its IDF score in the secondary text. Since the words contained in the text of the same category are close or similar, using words with similar meanings for division is avoided to affect the clustering effect. Therefore, obtaining the IDF value of the target word in the upper-level text, that is, obtaining the discrimination ability and importance degree of the word in its category, helps to distinguish the lower-level text.

[0078] Through the IDF scores at the tertiary word level corresponding to each target word and the tertiary word frequency values, the TFIDF values at the hierarchical level of each target word, that is, the hierarchical term frequency-inverse document frequency, are obtained. Different hierarchical TFIDF values are assigned according to the occurrence frequencies of different documents and word frequencies, and common words are filtered out, and the target words with important category discrimination ability are retained as the representation values of the tertiary text to improve the clustering accuracy. Therefore, according to the characteristics of the target word in the upper-level text, the corresponding hierarchical TFIDF value is obtained, and the hierarchical TFIDF value is used to cluster the tertiary text, so as to improve the relevance between the clustering result and the upper-level text, improve the clustering accuracy, and improve the clustering effect.

[0079] It should be noted that there can be multiple first-level texts, and correspondingly, there can also be multiple first-level attributes. The second-level attributes included in different first-level attributes are also different. Therefore, the second-level texts included in the first-level texts of different first-level attributes are also different, and the third-level texts included in different second-level texts are also different.

[0080] Referring to Figure 2 , Figure 1 the steps in the illustrated embodiment of S200 include but are not limited to the following steps:

[0081] Step S210, perform word segmentation on the third-level text to obtain multiple target words;

[0082] Step S220, for each target word, perform statistical processing on the target word according to the third-level text to obtain the word occurrence times;

[0083] Step S230, according to all the word occurrence times and a preset upper limit cut-off value, obtain the third-level word frequency value.

[0084] It can be understood that word segmentation is performed on each third-level text respectively to obtain multiple target words, and the target words are selected as the representation values of the text. Each target word is statistically processed in each third-level text to obtain the word occurrence times of each target word. The higher the number of times a certain word appears in the text, the higher the importance of the word in the text can be considered. Since the target words included in the same category are similar or the same, therefore, the target words that frequently appear in one text will also often appear in other texts. However, the higher the number of times this word appears in other texts, the more it can be considered that this word is a common term in a category. Using this word as the representation value of a text is difficult to distinguish from other texts, that is, the category discrimination ability of target words with high repetition rate and high occurrence frequency is weak. In order to avoid the influence of target words with high repetition rate and high occurrence frequency on subsequent text category discrimination, therefore, an upper limit cut-off process is performed on all the word occurrence times through the preset upper limit cut-off value to obtain the third-level word frequency value, thereby reducing the interference of high word occurrence frequency and high repetition rate and improving the clustering effect.

[0085] Referring to Figure 3 , Figure 2 the steps in the illustrated embodiment of S230 include but are not limited to the following steps:

[0086] Step S231, for each target word, process the word occurrence times according to the cut-off coefficient and the first compensation value to obtain the first word frequency value;

[0087] Step S232, process the sum of all the word occurrence times according to the correction coefficient and the second compensation value to obtain the second word frequency value;

[0088] Step S233: Obtain a plurality of third-level word frequency values according to the ratio of each first word frequency value to the second word frequency value.

[0089] Among them, the preset upper limit truncation value includes a truncation coefficient, a correction coefficient, a first compensation value, and a second compensation value.

[0090] It can be understood that the third-level word frequency value is obtained according to the ratio of the number of occurrences of the target word corresponding to the target word in the third-level text to the sum of the number of occurrences of all target words in the text. Therefore, the third-level word frequency value will increase as the number of occurrences of the target word increases. Even if the target word with a high occurrence frequency has a high repetition rate, its third-level word frequency value will also be relatively high, and it is impossible to accurately distinguish categories. The number of occurrences of each target word is processed with a truncation coefficient and a first compensation value, that is, the first compensation value is added to the number of occurrences of the word to obtain a first intermediate value, and the product of the truncation coefficient and the first intermediate value is used as the first word frequency value of the target word. At the same time, the sum of the number of occurrences of all target words is processed with a correction coefficient and a second compensation value, that is, the product of the correction coefficient and the sum of the number of occurrences of all words is used as the second intermediate value, and the second compensation value is added to the second intermediate value to obtain the second word frequency value. The third-level word frequency value is obtained according to the ratio of each first word frequency value to the second word frequency value. The upper limit truncation of the third-level word frequency value is completed through the truncation coefficient and the correction coefficient, reducing the third-level word frequency value of the target word with a high occurrence frequency, while the third-level word frequency value of the target word with a low occurrence frequency is compensated by the first compensation value and the second compensation value, avoiding the reduction of the third-level word frequency value of the target word with a low occurrence frequency, and being unable to suppress the third-level word frequency value of the high-frequency word, affecting the subsequent calculation and reducing the clustering effect.

[0091] It should be noted that the calculation formula of the third-level word frequency value can be:

[0092]

[0093] Among them, is the third-level word frequency value, is the label of the target word, is the label of the third-level text, is for the label of the target word, in the third-level text labeled is the number of occurrences of the word, is the truncation coefficient, is the first compensation value, is the correction coefficient, is the second compensation value. It can be seen from the above formula that is a monotonically increasing function of . When tends to positive infinity, The value approaches . And when the value is 0, then the value is the minimum value . For example, can take the value of 2, can take the value of 1, can take the value of 1.5, can take the value of 0.5, then when approaches positive infinity, the value approaches , so as to achieve the upper limit truncation of the three-level word frequency value of the target words with high occurrence frequency, avoid the influence of too high values on subsequent clustering analysis, and improve the accuracy of clustering analysis.

[0094] Refer to Figure 4 , Figure 1 The steps in the embodiment shown in

[0095] Step S410, obtaining a hierarchical weight value according to the cross relationship between the first-level text, the second-level text and the third-level text;

[0096] Step S420, processing the first-level IDF score dictionary according to the hierarchical weight value and each target word to obtain a plurality of first-level hierarchical scores;

[0097] Step S430, processing the second-level IDF score dictionary according to the hierarchical weight value and each target word to obtain a plurality of second-level hierarchical scores;

[0098] Step S440, for each target word, processing the first-level hierarchical score and the second-level hierarchical score to obtain a third-level word IDF score.

[0099] It can be understood that since the secondary texts included in different primary texts are different, and the tertiary texts included in different secondary texts can also be different, not all tertiary texts belong to the same primary text or the same secondary text. According to the cross-relationship among the primary text, secondary text, and tertiary text, a hierarchical weight value is obtained. Based on the hierarchical weight value, the primary IDF score dictionary of each target word is processed, that is, by multiplying the hierarchical weight value by the primary IDF score dictionary corresponding to the target word, the primary hierarchical score corresponding to the target word is obtained. And the tertiary text must belong to a secondary text with a second attribute and also belong to a primary text with a first attribute, and this secondary text must also belong to this primary text. Therefore, the hierarchical weight value of the target word in the primary text is complementary to the hierarchical weight value of the target word in the secondary text. According to the hierarchical weight value, the secondary IDF score dictionary of each target word is processed, that is, by multiplying the complementary value of the hierarchical weight value by the secondary IDF score dictionary corresponding to the target word, the secondary hierarchical score corresponding to the target word is obtained. By adding the primary hierarchical scores and secondary hierarchical scores corresponding to each target word, the tertiary word IDF score is obtained. Therefore, by combining the cross-relationship between the primary text, secondary text, and tertiary text, calculating the category discrimination ability of the target word in the tertiary text helps subsequent clustering analysis and improves the clustering effect.

[0100] It should be noted that the calculation formula for the tertiary word IDF score can be:

[0101]

[0102] Wherein, is the tertiary word IDF score, is the label of the target word, is the hierarchical weight value, is the primary IDF score dictionary corresponding to the target word with the label ; is the secondary IDF score dictionary corresponding to the target word with the label . For example, if the secondary attribute of the secondary text is "wealth planning" and the primary attribute of the corresponding primary text is "insurance sales conversation", when calculating the tertiary word IDF score of the tertiary text in the "wealth planning" secondary text, the primary IDF score dictionary of the target word in the "insurance sales conversation" primary text and the secondary IDF score dictionary of the target word in the "wealth planning" secondary text are selected for calculation. Clustering in combination with the characteristics of the target word in the upper-level text can improve the relevance to the attributes of the upper-level text and improve the clustering effect.

[0103] Refer to Figure 5 , Figure 4Step S410 in the illustrated embodiment includes but is not limited to the following steps:

[0104] Step S411, according to the cross-relationship between the tertiary text and the primary text, count the number of primary texts to obtain the first text quantity;

[0105] Step S412, according to the cross-relationship between the tertiary text and the secondary text, count the number of secondary texts to obtain the second text quantity;

[0106] Step S413, obtain the hierarchical weight value according to the ratio of the first text quantity to the second text quantity.

[0107] It can be understood that since there are multiple primary texts, secondary texts, and tertiary texts, not every primary text contains all secondary texts or all tertiary texts, and not every secondary text contains all tertiary texts. Therefore, according to the cross-relationship between the tertiary text and the primary text, count the number of primary texts that contain this tertiary text to obtain the first text quantity. According to the cross-relationship between the tertiary text and the secondary text, count the number of secondary texts that contain this tertiary text to obtain the second text quantity. That is, count the number of primary texts of the primary attribute to which the tertiary text belongs, and count the number of secondary texts of the secondary attribute to which the tertiary text belongs, so as to obtain the hierarchical weight value according to the ratio of the first text quantity to the second text quantity.

[0108] It should be noted that the calculation formula of the hierarchical weight value can be:

[0109]

[0110] Wherein, is the hierarchical weight value, is the number of primary texts of the primary attribute to which the tertiary text belongs, is the number of secondary texts of the secondary attribute to which the tertiary text belongs. For example, if the primary attribute to which the tertiary text belongs is "insurance sales conversation" and the secondary attribute is "wealth planning", then count the number of texts in the relevant category, and calculate the hierarchical weight value through the ratio of the number of primary texts and secondary texts of the corresponding attribute, so as to improve the relevance between clustering and the category characteristics of the upper-layer text and improve the clustering effect. In addition, by adding a compensation value in the calculation process of the hierarchical weight value, the hierarchical weight value can be reduced, avoiding errors in the hierarchical weight value when the text quantity is low and affecting the clustering effect.

[0111] Refer to Figure 6 , Figure 1 Step S500 in the illustrated embodiment includes but is not limited to the following steps:

[0112] Step S510, calculate multiple hierarchical TFIDF values according to the ratio between the three-level word frequency value and the three-level word hierarchical IDF score.

[0113] It can be understood that since the importance of the target word increases in direct proportion to the number of times it appears in the text, but at the same time decreases in inverse proportion to the frequency of its appearance in all texts. Therefore, according to the ratio between the three-level word frequency value and the three-level word hierarchical IDF score, the corresponding hierarchical TFIDF values of each target word are obtained, so as to determine the importance of the target word, that is, the category discrimination ability of the target word.

[0114] It should be noted that the calculation formula of the hierarchical TFIDF value can be:

[0115]

[0116] where, is the hierarchical TFIDF value, is the three-level word frequency value, is the three-level word IDF score, is the label of the target word, is the label of the three-level text.

[0117] Refer to Figure 7 , Figure 1 The steps in the embodiment shown in

[0118] Step S610, import the hierarchical TFIDF values into a similarity hashing neural model for processing to obtain multiple text representation vectors, and the text representation vectors correspond to the three-level texts;

[0119] Step S620, cluster the three-level texts according to the distance between any two text representation vectors to obtain a clustering result.

[0120] It can be understood that a similar hashing neural model is used to process the hierarchical TFIDF values of each target word to obtain the corresponding hash code of the target word, which is used as the text representation vector of the corresponding text. Based on the distance between any two text representation vectors, the corresponding three-level texts are clustered. Among them, the Hamming distance can be used to calculate the distance of each text representation vector, the Jaccard distance can also be used for calculation, the cosine distance can also be calculated, and the Euclidean distance can also be calculated. When clustering the three-level texts, multiple three-level texts can be selected as the initial clustering centers, calculate the distance between the text representation vector corresponding to each three-level text and each clustering center, and assign each three-level text to the clustering center with the closest distance. The clustering center and the three-level texts assigned to the clustering center represent a clustering cluster. When all the three-level texts are assigned, the clustering center of each clustering cluster will be recalculated according to the existing three-level texts in the clustering cluster. The calculation is continuously repeated until no three-level text is reassigned to a different clustering cluster, or no clustering center changes, or the sum of squared errors is locally minimized, so as to complete the clustering of the three-level texts according to the hierarchical TFIDF values corresponding to the target words, realize clustering by combining the attributes of the upper-level texts, improve the relevance with the upper-level texts, improve the clustering accuracy, and improve the clustering effect.

[0121] Refer to Figure 8 , Figure 8 FIG. shows a schematic structural diagram of a multi-level text clustering system 800 provided by an embodiment of the present invention.

[0122] It can be understood that the multi-level text clustering system 800 includes:

[0123] A text acquisition module 810, configured to acquire a first-level text with a first-level attribute and a second-level text with a second-level attribute, where the first-level attribute includes the second-level attribute, and the second-level text includes multiple third-level texts.

[0124] A text word segmentation module 820, configured to perform word segmentation on the third-level texts in the second-level text to obtain target words and third-level word frequency values, and the target words correspond to the third-level word frequency values.

[0125] A word and character statistics module 830, configured to perform word and character statistics processing on the first-level text according to the target words to obtain a first-level IDF score dictionary, and perform word and character statistics processing on the second-level text according to the target words to obtain a second-level IDF score dictionary.

[0126] A word hierarchical score calculation module 840, configured to obtain a third-level word hierarchical IDF score corresponding to each target word according to the first-level IDF score dictionary and the second-level IDF score dictionary.

[0127] The hierarchical numerical calculation module 850 is used to obtain multiple hierarchical TFIDF values based on the three-level word hierarchical IDF scores and the three-level word frequency values.

[0128] The multi-layer clustering module 860 is used to perform clustering processing based on multiple hierarchical TFIDF values to obtain a clustering result.

[0129] In addition, the text word segmentation module 820 includes:

[0130] The target word segmentation module 821 is used to perform word segmentation on the three-level text to obtain multiple target words.

[0131] The target word statistics module 822 is used to perform statistical processing on each target word according to the three-level text to obtain the number of times the word appears.

[0132] The word frequency value calculation module 823 is used to obtain the three-level word frequency value according to the number of times all words appear and the preset upper limit truncation value.

[0133] In addition, the preset upper limit truncation value includes a truncation coefficient, a correction coefficient, a first compensation value, and a second compensation value, and the word frequency value calculation module 823 includes:

[0134] The first word frequency value calculation module 824 is used to process the number of times the word appears for each target word according to the truncation coefficient and the first compensation value to obtain the first word frequency value.

[0135] The second word frequency value calculation module 825 is used to process the sum of the number of times all words appear according to the correction coefficient and the second compensation value to obtain the second word frequency value.

[0136] The three-level word frequency value calculation module 826 is used to obtain multiple three-level word frequency values according to the ratio of each first word frequency value to the second word frequency value.

[0137] In addition, the word hierarchical score calculation module 840 includes:

[0138] The hierarchical weight value acquisition module 841 is used to obtain the hierarchical weight value according to the cross relationship between the first-level text, the second-level text, and the third-level text.

[0139] The first-level hierarchical score calculation module 842 is used to process the first-level IDF score dictionary according to the hierarchical weight value and each target word to obtain multiple first-level hierarchical scores.

[0140] The second-level hierarchical score calculation module 843 is used to process the second-level IDF score dictionary according to the hierarchical weight value and each target word to obtain multiple second-level hierarchical scores.

[0141] The three - level word IDF score calculation module 844 is used to process the first - level hierarchical score and the second - level hierarchical score for each target word to obtain the three - level word IDF score.

[0142] In addition, there are multiple first - level texts, multiple second - level texts, and the hierarchical weight value acquisition module 841 includes:

[0143] The first text statistics module 845 is used to count the quantity of the first - level texts according to the cross - relationship between the third - level texts and the first - level texts to obtain the first text quantity.

[0144] The second text statistics module 846 is used to count the quantity of the second - level texts according to the cross - relationship between the third - level texts and the second - level texts to obtain the second text quantity.

[0145] The hierarchical weight value calculation module 847 is used to obtain the hierarchical weight value according to the ratio of the first text quantity to the second text quantity.

[0146] In addition, the word hierarchical score calculation module is also used to calculate a plurality of hierarchical TFIDF values according to the ratio between the third - level word frequency value and the third - level word hierarchical IDF score.

[0147] In addition, the multi - layer clustering module 860 includes:

[0148] The text representation vector calculation module 861 is used to import the hierarchical TFIDF values into the similar hashing neural model for processing to obtain a plurality of text representation vectors, and the text representation vectors correspond to the third - level texts.

[0149] The distance clustering module 862 is used to cluster the third - level texts according to the distance between any two text representation vectors to obtain a clustering result.

[0150] Refer to Figure 9 , Figure 9 shows the electronic device 900 provided by the embodiment of the present invention. The electronic device 900 includes a memory 910, a processor 920, and a computer program stored on the memory 910 and executable on the processor 920. When the processor 920 executes the computer program, the multi - layer text clustering method in the above - mentioned embodiment is implemented.

[0151] The memory 910, as a non - transitory computer - readable storage medium, can be used to store non - transitory software programs and non - transitory computer - executable programs, such as the multi - layer text clustering method in the above - mentioned embodiment of the present invention. The processor 920 realizes the multi - layer text clustering method in the above - mentioned embodiment of the present invention by running the non - transitory software programs and instructions stored in the memory 910.

[0152] The memory 910 may include a program storage area and a data storage area. The program storage area may store an operating system and application programs required for at least one function. The data storage area may store data required to execute the density-radius-based clustering method in the above embodiments, etc. In addition, the memory 910 may include high-speed random access memory and may also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. It should be noted that the memory 910 may optionally include a memory remotely provided with respect to the processor 920, and these remote memories may be connected to the terminal through a network. Examples of the above network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.

[0153] The non-transitory software programs and instructions required to implement the multi-layer text clustering method in the above embodiments are stored in the memory. When executed by one or more processors, the multi-layer text clustering method in the above embodiments is executed. For example, the method steps S100 to step S600 described above are executed Figure 1 in the method steps S100 to step S600, Figure 2 in the method steps S210 to step S230, Figure 3 in the method steps S231 to step S233, Figure 4 in the method steps S410 to step S440, Figure 5 in the method steps S411 to step S413, Figure 6 in the method step S510, Figure 7 in the method steps S610 to step S620.

[0154] The present invention also provides a computer-readable storage medium. The computer-readable storage medium stores computer-executable instructions, and the computer-executable instructions are used to cause a computer to execute the multi-layer text clustering method in the above embodiments. For example, the method steps S100 to step S600 described above are executed Figure 1 in the method steps S100 to step S600, Figure 2 in the method steps S210 to step S230, Figure 3 in the method steps S231 to step S233, Figure 4 in the method steps S410 to step S440, Figure 5 in the method steps S411 to step S413, Figure 6 in the method step S510, Figure 7 in the method steps S610 to step S620.

[0155] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0156] Those of ordinary skill in the art can understand that all or some of the steps and systems disclosed in the above methods can be implemented as software, firmware, hardware, and their appropriate combinations. Some physical components or all physical components can be implemented as software executed by a processor, such as a central processing unit, a digital signal processor, or a microprocessor, or implemented as hardware, or implemented as an integrated circuit, such as an application-specific integrated circuit. Such software can be distributed on a computer-readable medium, which can include a computer storage medium (or non-transitory medium) and a communication medium (or transitory medium). As is well known to those of ordinary skill in the art, the term computer storage medium includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information, such as computer-readable instructions, data structures, program modules, or other data. Computer storage media includes, but is not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disk (DVD) or other optical disk storage, magnetic cassette, tape, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired information and can be accessed by a computer. In addition, as is well known to those of ordinary skill in the art, a communication medium typically includes computer-readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transmission mechanism, and can include any information delivery medium.

[0157] It should be noted that the server can be an independent server or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms.

[0158] The embodiments of the present invention have been described in detail above in conjunction with the accompanying drawings. However, the present invention is not limited to the above embodiments. Within the scope of knowledge possessed by those of ordinary skill in the art, various changes can be made without departing from the spirit of the present invention.

Claims

1. A multi - layer text clustering method, the method comprising: Obtaining first - level texts with first - level attributes and second - level texts with second - level attributes, wherein the first - level attributes include the second - level attributes, and the second - level texts include multiple third - level texts; Performing word segmentation on the third - level texts in the second - level texts to obtain target words and third - level word frequency values, where the target words correspond to the third - level word frequency values; Performing word and character statistics processing on the first - level texts according to the target words to obtain a first - level IDF score dictionary, and performing word and character statistics processing on the second - level texts according to the target words to obtain a second - level IDF score dictionary; For each of the target words, obtaining third - level word - level IDF scores according to the first - level IDF score dictionary and the second - level IDF score dictionary; Obtaining multiple hierarchical TFIDF values according to the third - level word - level IDF scores and the third - level word frequency values; Performing clustering processing on the third - level texts according to the multiple hierarchical TFIDF values to obtain a clustering result; Wherein, a preset upper - limit truncation value includes a truncation coefficient, a correction coefficient, a first compensation value, and a second compensation value. The performing word segmentation on the third - level texts in the second - level texts to obtain target words and third - level word frequency values includes: Performing word segmentation on the third - level texts to obtain multiple target words; For each of the target words, performing statistical processing on the target word according to the third - level texts to obtain the number of times the word appears; For each of the target words, processing the number of times the word appears according to the truncation coefficient and the first compensation value to obtain a first word frequency value; Processing the sum of the number of times all the words appear according to the correction coefficient and the second compensation value to obtain a second word frequency value; Obtaining multiple third - level word frequency values according to the ratio of each of the first word frequency values to the second word frequency value.

2. The multi-layer text clustering method according to claim 1, wherein The for each of the target words, obtaining third - level word - level IDF scores according to the first - level IDF score dictionary and the second - level IDF score dictionary includes: Obtaining a hierarchical weight value according to the cross - relationship between the first - level texts, the second - level texts, and the third - level texts; Processing the first - level IDF score dictionary according to the hierarchical weight value and each of the target words to obtain multiple first - level hierarchical scores; Processing the second - level IDF score dictionary according to the hierarchical weight value and each of the target words to obtain multiple second - level hierarchical scores; For each of the target words, processing the first - level hierarchical score and the second - level hierarchical score to obtain a third - level word IDF score.

3. The multi-layer text clustering method according to claim 2, wherein, There are multiple first - level texts and multiple second - level texts; The obtaining a hierarchical weight value according to the cross - relationship between the first - level texts, the second - level texts, and the third - level texts includes: Counting the number of first - level texts according to the cross - relationship between the third - level texts and the first - level texts to obtain a first text quantity; Counting the number of second - level texts according to the cross - relationship between the third - level texts and the second - level texts to obtain a second text quantity; Obtain a hierarchical weight value according to the ratio of the first text quantity and the second text quantity.

4. The multi-layer text clustering method according to claim 1, wherein Based on the IDF scores of the three-level word hierarchies and the three-level word frequency values, obtain multiple hierarchical TFIDF values, including: Calculate multiple hierarchical TFIDF values according to the ratio between the three-level word frequency values and the IDF scores of the three-level word hierarchies.

5. The multi-layer text clustering method according to claim 1, characterized in that Based on the hierarchical TFIDF values, perform clustering processing to obtain a clustering result, including: Import the hierarchical TFIDF values into a similarity hashing neural model for processing to obtain multiple text representation vectors, where the text representation vectors correspond to the three-level texts; Cluster the three-level texts according to the distance between any two of the text representation vectors to obtain a clustering result.

6. A multi-layer text clustering system, characterized in that, Include: A text acquisition module, configured to acquire a first-level text with a first-level attribute and a second-level text with a second-level attribute, where the first-level attribute includes the second-level attribute, and the second-level text includes multiple three-level texts; A text word segmentation module, configured to perform word segmentation processing on the three-level texts in the second-level text to obtain target words and three-level word frequency values, where the target words correspond to the three-level word frequency values; A word and character statistics module, configured to perform word and character statistics processing on the first-level text according to the target words to obtain a first-level IDF score dictionary, and perform word and character statistics processing on the second-level text according to the target words to obtain a second-level IDF score dictionary; A word hierarchy score calculation module, configured to, for each of the target words, obtain an IDF score of the three-level word hierarchy according to the first-level IDF score dictionary and the second-level IDF score dictionary, where the IDF score of the three-level word hierarchy corresponds to the target word; A hierarchical value calculation module, configured to obtain multiple hierarchical TFIDF values according to the IDF scores of the three-level word hierarchies and the three-level word frequency values; A multi-level clustering module, configured to perform clustering processing according to multiple hierarchical TFIDF values to obtain a clustering result; Among them, the preset upper limit truncation value includes a truncation coefficient, a correction coefficient, a first compensation value, and a second compensation value, and the text word segmentation module includes: A target word segmentation module, configured to perform word segmentation processing on the three-level text to obtain multiple target words; A target word statistics module, configured to, for each target word, perform statistics processing on the target word according to the three-level text to obtain the number of times the word appears; A word frequency value calculation module, configured to obtain three-level word frequency values according to the number of times all words appear and the preset upper limit truncation value; The word frequency value calculation module includes: A first word frequency value calculation module, configured to, for each target word, process the number of times the word appears according to the truncation coefficient and the first compensation value to obtain a first word frequency value; A second word frequency value calculation module, configured to process the sum of the number of times all words appear according to the correction coefficient and the second compensation value to obtain a second word frequency value; A three-level word frequency value calculation module, configured to obtain multiple three-level word frequency values according to the ratio of each first word frequency value to the second word frequency value.

7. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the multi-layer text clustering method according to any one of claims 1 to 5.

8. A computer-readable storage medium, characterized in that, There is a computer program stored, and when the computer program is executed by a processor, it implements the multi-layer text clustering method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Text clustering method, electronic equipment and storage medium

    CN112256842A