Text category determination method, device, computer device, and storage medium
By calculating the minimum value of word segmentation combination and selecting the target word segmentation combination, the problem of inefficient traditional text classification is solved, and efficient text category determination and clustering is achieved.
Patent Information
- Application Number
- CN202210630294.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-06
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2042-06-06
AI Technical Summary
When traditional text classification methods face a large number of texts with determined categories, the increase in the number of times of calculating similarity results in inefficient text determination.
By obtaining the number of word segmentation in the target text, calculate the minimum number of word segmentation required for word segmentation combinations that meet the similarity condition, and selecting the number of word segmentation for combinations, determine that the cluster cluster represented by the target segmentation combination is text category, avoiding multiple similarity calculations.
The efficiency of text category determination is improved, especially in large-scale text clustering scenarios, and efficient clustering with near-linear time complexity is achieved.
Smart Images

Figure CN115129871B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular, to a method, apparatus, computer device, computer-readable storage medium, and computer program product for determining text categories. Background Art
[0002] With the development of computer technology, text classification and text clustering technologies have emerged. This technology can effectively organize, summarize, navigate, and recommend texts, enabling users to quickly locate the information they need.
[0003] In traditional technologies, generally, the similarity between an undetermined text and a determined text is continuously calculated, and the category to which the text belongs is determined based on the similarity. However, as the number of categories of the determined texts increases, the number of times of calculating the similarity will increase accordingly, resulting in a slower and slower speed of determining text categories and a relatively low efficiency of determining text categories. Summary of the Invention
[0004] Based on this, it is necessary to provide a method, apparatus, computer device, computer-readable storage medium, and computer program product for determining text categories that can improve the efficiency of determining text categories for the above technical problems.
[0005] In a first aspect, the present application provides a method for determining text categories. The method includes:
[0006] Obtain the word segments included in the target text;
[0007] According to the number of word segments included in the target text, calculate the minimum number of word segments required for a word segment combination whose clustering similarity to the target text meets the similarity condition;
[0008] Select a target number of word segments from the target text for combination to obtain a target word segment combination; the target number is greater than or equal to the minimum number of word segments; there is at least one different word segment in different target word segment combinations;
[0009] Determine the clustering cluster represented by the target word segment combination as the text category to which the target text belongs.
[0010] In a second aspect, the present application further provides a device for determining text categories. The device includes:
[0011] A word segment acquisition module, configured to obtain the word segments included in the target text;
[0012] A quantity calculation module, configured to calculate the minimum number of word segments required for a word segment combination whose clustering similarity to the target text meets the similarity condition according to the number of word segments included in the target text;
[0013] A combination determination module, configured to select a target number of word segments from the target text for combination to obtain a target word segment combination; the target number is greater than or equal to the minimum value of the number of word segments; there is at least one different word segment in different target word segment combinations;
[0014] A category determination module, configured to determine the clustering cluster represented by the target word segment combination as the text category to which the target text belongs.
[0015] In a third aspect, the present application further provides a computer device. The computer device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of the above method are implemented.
[0016] In a fourth aspect, the present application further provides a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the above method are implemented.
[0017] In a fifth aspect, the present application further provides a computer program product. The computer program product includes a computer program, and when the computer program is executed by a processor, the steps of the above method are implemented.
[0018] After obtaining the word segments included in the target text, the above text category determination method, device, computer device, computer-readable storage medium, and computer program product determine the minimum value of the number of word segments required for a word segment combination whose clustering similarity to the target text meets the similarity condition through a preset similarity condition, and then determine a target word segment combination composed of word segments greater than or equal to the minimum value of the number of word segments, so as to ensure that the target word segment combination can also meet the similarity condition, and there is no need to calculate the clustering similarity between the text corresponding to each word segment combination and the target text, avoiding multiple calculations of similarity. By determining the clustering cluster represented by the target word segment combination as the text category to which the target text belongs, the text category to which the target text belongs can be quickly determined, and further clustering of the text can be realized, thereby effectively improving the determination efficiency of the text type, and further improving the clustering efficiency of clustering text on the scale of hundreds of millions. Description of the Drawings
[0019] Figure 1 It is an application environment diagram of the text category determination method in an embodiment;
[0020] Figure 2 It is a flowchart of the text category determination method in an embodiment;
[0021] Figure 3 It is a schematic diagram of the clustering relationship diagram between the clustering clusters of a text set in an embodiment;
[0022] Figure 4 It is a flowchart showing the process of the method for determining text categories in a specific embodiment;
[0023] Figure 5 It is a schematic diagram showing the process of the method for determining text categories in a specific embodiment;
[0024] Figure 6 It is a schematic diagram showing the processing content of the method for determining text categories in a specific embodiment;
[0025] Figure 7 It is a schematic diagram of the interface of the application environment of the method for determining text categories in a specific embodiment;
[0026] Figure 8 It is a block diagram showing the structure of the text category determination device in an embodiment;
[0027] Figure 9 It is an internal structure diagram of a computer device in an embodiment;
[0028] Figure 10 It is an internal structure diagram of a computer device in another embodiment. Detailed implementation manners
[0029] In order to make the objectives, technical solutions and advantages of the present application clearer, the following further describes the present application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0030] It should be noted that the text data involved in the present application are data that have been fully authorized by all parties, and the collection, use and processing of the relevant data need to comply with the relevant laws, regulations and standards of the relevant countries and regions.
[0031] In one embodiment, the method for determining text categories provided by the present application can be applied to an application environment as shown in Figure 1 In the application environment, it involves a terminal 102 and a server 104. In some embodiments, it may also involve a terminal 106 at the same time. Among them, the terminal 102 and the terminal 106 communicate with the server 104 through a network. The data storage system can store the data that the server 104 needs to process. The data storage system can be integrated on the server 104, or can be placed in the cloud or on other servers.
[0032] A text set composed of multiple texts is stored in the terminal 102. The texts can be obtained from a publicly available dataset, input by a user, or obtained by converting the user's speech into text. The terminal 102 can determine the target text from the text set and send it to the server 104, or the server 104 can directly obtain the target text from the text set.
[0033] Then, the server 104 obtains the word segmentation included in the target text; calculates the minimum number of word segments required for a word segment combination whose clustering similarity to the target text meets the similarity condition according to the number of word segments included in the target text; selects a target number of word segments from the target text for combination to obtain a target word segment combination; the target number is greater than or equal to the minimum number of word segments; there is at least one different word segment in different target word segment combinations; determines the clustering cluster represented by the target word segment combination as the text category to which the target text belongs. Then, the server 104 can store the text category to which the target text belongs, or return the text category to which the target text belongs to the terminal 102 for display or storage on the terminal 102, or send the text category to which the target text belongs to the terminal 106 for display or storage on the terminal 106. By continuously determining the text categories to which different target texts belong, text clustering of texts is ultimately achieved.
[0034] Among them, the terminal 102 and the terminal 106 can be, but are not limited to, various desktop computers, laptop computers, smart phones, tablet computers, Internet of Things devices, and portable wearable devices. The Internet of Things devices can be smart TVs, in-vehicle intelligent devices, etc. The portable wearable devices can be smart watches, smart bracelets, etc. The server 104 can be implemented by an independent server or a server cluster composed of multiple servers.
[0035] Natural Language Processing (NLP) is an important direction in the fields of computer science and artificial intelligence. It studies various theories and methods that can achieve effective communication between humans and computers in natural language. Natural language processing is a science that integrates linguistics, computer science, and mathematics. Therefore, the research in this field will involve natural language, that is, the language people use in daily life, so it has a close connection with the research of linguistics. Natural language processing technologies usually include text processing, semantic understanding, machine translation, robot question answering, knowledge graph and other technologies. The method involved in the embodiments of this application is mainly related to text processing.
[0036] In one embodiment, as Figure 2 shown, a method for determining a text category is provided. Taking the method applied to the Figure 1 server 104 as an example for illustration, the method includes the following steps:
[0037] Step S202, obtain the word segments included in the target text.
[0038] A text refers to data that uses written language for semantic representation. A target text refers to a text for which the text category to which it belongs needs to be determined. A word segment is the basic unit that makes up a text. Specifically, a text set is preset, the target text is obtained from the text set, and the target text is segmented to obtain the word segments included in the target text.
[0039] A text set is a set composed of multiple texts, and the target text is the text determined from the text set. The segmentation process refers to the processing method of extracting the word segments included in the target text. The specific method of the segmentation process can be selected according to actual technical needs. For example, any one of the forward maximum matching method, reverse maximum matching method, bidirectional matching segmentation method, or neural network model can be used to segment the target text to obtain the word segments included in the target text.
[0040] In one embodiment, in order to improve the accuracy of the word segments of the obtained target text, after segmenting the target text to obtain the word segments included in the target text, it may further include: performing noise removal processing on the word segments included in the target text to obtain the processed word segments included in the target text.
[0041] The noise removal processing of word segments refers to the processing method of further removing invalid words after segmenting the target text. Invalid words may include punctuation marks, invalid letters, etc. The word segments obtained after the noise removal processing of word segments are called processed word segments. The noise removal processing of word segments can be selected according to actual scenario needs. When the text contains punctuation marks or letters, the noise removal processing of word segments can be selected. It can be understood that if the noise removal processing of word segments is performed, the word segments included in the target text in the subsequent steps refer to the processed word segments after the noise removal processing of word segments.
[0042] It should be noted that the texts involved in this embodiment include but are not limited to short texts. A short text is a text whose number of included word segments does not exceed a predetermined number. The predetermined number can be determined according to the usage scenario or text type of the short text. For example, if the text type of the short text is a company name, the predetermined number can be set to 20, that is, the number of word segments included in the short text does not exceed 20 word segments. When the number of word segments included in the text exceeds the predetermined number, the text can be split to obtain the corresponding short text, and then the processing is performed.
[0043] Step S204, according to the number of word segments included in the target text, calculate the minimum number of word segments required for a word segment combination whose clustering similarity to the target text meets the similarity condition.
[0044] The number of word segments refers to the total number of word segments contained in the text. A word segment combination is formed by combining multiple word segments, and a word segment combination can also correspond to a text. Text clustering refers to a method of grouping multiple texts with relatively high similarity into one category according to the similarity between texts, so that the similarity between texts in the same category is relatively high, and the similarity between texts in different categories is relatively low. The clustering similarity refers to the similarity between the target text and the text corresponding to the word segment combination when the target text and the text corresponding to the word segment combination are grouped into one category. The similarity condition refers to the condition that the similarity between the target text and the word segment combination needs to meet when the target text and the text corresponding to the word segment combination are grouped into one category.
[0045] The minimum number of word segments refers to the minimum number of identical word segments that exist between the word segments contained in the word segment combination that meets the similarity condition and the word segments contained in the target text. The minimum number of word segments required for the word segment combination that meets the similarity condition is less than or equal to the number of word segments contained in the target text. That is, there are some identical word segments between the word segments in the word segment combination and the word segments contained in the target text. For a word segment combination, the word segment combination can be formed only by combining one or more word segments contained in the target text, or can be formed by combining one or more word segments contained in the target text and other word segments.
[0046] The similarity between texts can be specifically divided into lexical similarity and semantic similarity, and the calculation method of the similarity between texts can be selected according to actual technical needs. For example, it can be determined by methods such as the Jaccard similarity algorithm and the cosine similarity algorithm. The calculation formulas corresponding to different calculation methods are different. The similarity condition can be set according to actual technical needs, and specifically, it can be set that the similarity between the target text and the word segment combination is greater than or equal to the set similarity threshold. The similarity threshold can be set according to actual technical needs. For example, the similarity threshold is set according to the accuracy requirements of the application scenario of the text. In application scenarios with relatively high matching accuracy requirements, for example, by a company name, all company names associated with the company name are determined, then the set similarity threshold can be 0.8.
[0047] When the clustering similarity between a word segmentation combination and the target text meets the similarity condition, the text corresponding to the word segmentation combination and the target text can be classified into the same category. Specifically, the similarity condition can be set first, and according to the number of word segments contained in the target text, the minimum number of word segments required for the word segmentation combination that meets the similarity condition can be determined, and then the specific content of the word segmentation combination can be determined. Thus, subsequently, the target text and the text corresponding to the word segmentation combination can be classified into the same category. Accordingly, it is not necessary to first determine the specific content of all the word segmentation combinations corresponding to the target text and then calculate whether the similarity between the text corresponding to each word segmentation combination and the target text can meet the similarity condition, thereby avoiding calculating the similarity multiple times and effectively improving the data processing efficiency.
[0048] For example, the number of word segments contained in the target text is 4, and the minimum number of word segments required for the word segmentation combination whose clustering similarity with the target text meets the similarity condition is calculated to be 3, that is, there need to be at least 3 identical word segments between the word segments contained in the word segmentation combination that meets the similarity condition and the word segments contained in the target text. When there are 3 or 4 word segments in the word segmentation combination that are the same as the word segments contained in the target text, the word segmentation combination can meet the similarity condition.
[0049] Step S206: Select a target number of word segments from the target text for combination to obtain a target word segmentation combination; the target number is greater than or equal to the minimum number of word segments; there is at least one different word segment in different target word segmentation combinations.
[0050] The target number refers to the number of word segments selected from the word segments contained in the target text, and the target number is greater than or equal to the minimum number of word segments. For example, the number of word segments contained in the target text is 4. If the minimum number of word segments required for the word segmentation combination whose clustering similarity with the target text meets the similarity condition calculated according to the number of word segments contained in the target text is 3, then the target number can be 3, that is, any 3 word segments are selected from the word segments contained in the target text for combination, and the target number can also be 4, that is, all the word segments contained in the target text are selected.
[0051] The target word segmentation combination refers to the word segmentation combination determined after combining and screening the target number of word segments, and the target word segmentation combination is also the word segmentation combination whose clustering similarity with the target text meets the similarity condition. Since the word segments contained in the target text include one or more, generally multiple, after selecting the target number of word segments for combination and screening, the determined target word segmentation combination can also include one or more, generally multiple.
[0052] For example, if the target text contains 4 word segments, and the minimum number of word segments required for a word segment combination whose clustering similarity with the target text meets the similarity condition is 3, that is, there are at least 3 identical word segments between the word segments included in the word segment combination and those included in the target text. At this time, 3 word segments can be arbitrarily selected from the target text for combination, or 4 word segments can be selected from the target text for combination to obtain the target word segment combination.
[0053] Step S208: Determine the clustering cluster represented by the target word segment combination as the text category to which the target text belongs.
[0054] The clustering cluster of texts is a set of texts generated by text clustering. Multiple texts can be clustered according to the similarity between texts to obtain multiple clustering clusters. The similarity between texts in the same clustering cluster is relatively large, while the similarity between texts in different clustering clusters is relatively small. A clustering cluster can correspond to a unique identifier, which can be represented by numbers, letters, etc. The identifier corresponding to the target word segment combination corresponds to the identifier corresponding to the clustering cluster, that is, the target word segment combination can be used to represent the clustering cluster.
[0055] Specifically, match the identifier corresponding to the target word segment combination with the identifier corresponding to the clustering cluster, so as to determine the clustering cluster represented by the target word segment combination. Since the target word segment combination is a word segment combination whose clustering similarity with the target text meets the similarity condition, the target text and the target word segment combination can be classified into the same category, that is, the clustering cluster represented by the target word segment combination is the text category to which the target text belongs.
[0056] It can be understood that there can be one or more target word segment combinations. In the case where there is one target word segment combination, one target text corresponds to one clustering cluster, that is, one target text corresponds to one text category, realizing non-overlapping clustering of texts. In the case where there are multiple target word segment combinations, one target text can belong to multiple clustering clusters at the same time, that is, one target text corresponds to multiple text categories, realizing overlapping clustering.
[0057] In the above method for determining the text category, the word segments included in the target text are obtained; according to the number of word segments included in the target text, the minimum number of word segments required for a word segment combination whose clustering similarity with the target text meets the similarity condition is calculated; a target number of word segments are selected from the target text for combination to obtain a target word segment combination; the target number is greater than or equal to the minimum number of word segments; there is at least one different word segment in different target word segment combinations; accordingly, by presetting the similarity condition, the minimum number of word segments required for a word segment combination whose clustering similarity with the target text meets the similarity condition is determined, and then the target word segment combination composed of word segments greater than or equal to the minimum number of word segments is determined, so as to ensure that the target word segment combination can also meet the similarity condition, and there is no need to calculate the clustering similarity between each word segment combination and the target text, avoiding multiple calculations of the similarity; by determining the clustering cluster represented by the target word segment combination as the text category to which the target text belongs, the text category to which the target text belongs can be quickly determined, and further the text can be clustered, so that the efficiency of determining the text type can be effectively improved, and further the clustering efficiency of clustering texts on the scale of hundreds of millions can be improved.
[0058] In one embodiment, a text set is preset, the target text refers to the text determined from the text set, and when determining the minimum number of word segments, it is necessary to jointly determine in combination with other texts in the text set. Specifically, according to the number of word segments included in the target text, calculating the minimum number of word segments required for a word segment combination whose clustering similarity with the target text meets the similarity condition includes:
[0059] From the text set where the target text is located, the classified texts other than the target text are determined; according to the number of word segments included in each of the classified texts, the clustering word segment number corresponding to the text set is determined; based on the number of word segments included in the target text and the clustering word segment number, the minimum number of word segments required for a word segment combination whose clustering similarity with the target text meets the similarity condition is calculated.
[0060] The text set is a set composed of multiple texts, and the target text is the text determined from the text set. The classified text refers to the text whose text category has been determined among the texts in the text set other than the target text, that is, the clustering cluster where the classified text is located has been determined, or it can be understood that the currently determined clustering cluster is determined based on the classified text. The clustering word segment number corresponds to the text set and can be determined based on the word segment numbers of the classified texts in the text set.
[0061] Specifically, from the text set where the target text is located, determine the classified texts other than the target text, and based on the number of word segments contained in each classified text, determine the number of clustering word segments corresponding to the text set, and represent the number of clustering word segments as n. Based on the number of word segments contained in the target text and the number of clustering word segments, the minimum number of word segments required for a word segment combination whose clustering similarity to the target text meets the similarity condition can be calculated.
[0062] The number of clustering word segments can be determined based on the number of word segments of the classified texts in the text set. The number of clustering word segments can be the minimum value among the numbers of word segments contained in each classified text, represented as n1, or the average value of the numbers of word segments contained in each classified text, represented as n2, and can be specifically set according to actual technical needs. For example, it can be set according to the accuracy requirements of the application scenario.
[0063] It can be understood that when the number of clustering word segments is the minimum value among the numbers of word segments contained in each classified text, then when calculating the minimum number of word segments required for a word segment combination whose clustering similarity to the target text meets the similarity condition and then determining the text category to which the target text belongs, the inter-text similarity of all texts in the same category after clustering can meet the similarity condition. When the number of clustering word segments is the average value of the numbers of word segments contained in each classified text, then when calculating the minimum number of word segments required for a word segment combination whose clustering similarity to the target text meets the similarity condition and then determining the text category to which the target text belongs, the inter-text similarity of most texts in the same category after clustering can meet the similarity condition.
[0064] In this embodiment, by determining the number of clustering word segments corresponding to the text set according to the number of word segments contained in each classified text in the text set where the target text is located, and then calculating the minimum number of word segments required for a word segment combination whose clustering similarity to the target text meets the similarity condition, it can be ensured that in the clustering result of the subsequent obtained text set, the similarity between any two texts in each clustering cluster meets the similarity condition, or the similarity between most texts in each clustering cluster meets the similarity condition, improving the accuracy of determining the text category to which the target text belongs and further improving the accuracy of the clustering result of the text set.
[0065] In one embodiment, the minimum number of word segments also needs to be determined in combination with the specific calculation method of similarity. Specifically, based on the number of word segments contained in the target text and the number of clustering word segments, calculating the minimum number of word segments required for a word segment combination whose clustering similarity to the target text meets the similarity condition includes:
[0066] Obtain the clustering similarity threshold that matches the target text; Based on the similarity conditions that the word segmentation combination and the clustering similarity of the target text need to satisfy and the clustering similarity calculation method, calculate the maximum value of the number of different words between the word segmentation combination that meets the similarity conditions and the target text according to the number of clustering word segments and the clustering similarity threshold; Based on the number of word segments contained in the target word segmentation and the maximum value of the number of different words, determine the minimum value of the number of word segments required for the word segmentation combination that meets the similarity conditions.
[0067] The clustering similarity threshold refers to the similarity threshold preset between the target text and the word segmentation combination when classifying the target text and the text corresponding to the word segmentation combination into the same category. The clustering similarity threshold can be determined in advance according to the accuracy requirements of the application scenario, and the value range can be between 0 and 1. Denote the clustering similarity threshold as f.
[0068] The clustering similarity calculation method refers to the calculation method of the similarity between texts. In this embodiment, taking the Jaccard similarity algorithm adopted by the clustering similarity calculation method as an example, the clustering similarity threshold is the Jaccard similarity threshold. The greater the Jaccard similarity between texts, the higher the similarity between texts. The similarity condition is set that the Jaccard similarity between the target text and the word segmentation combination is greater than or equal to the clustering similarity threshold. The maximum value of the number of different words refers to the number of different words between the target text and the word segmentation combination that meets the similarity conditions. The maximum value of the number of different words and the minimum value of the number of word segments can be understood as corresponding concepts. Denote the number of different words as m.
[0069] The principle of the Jaccard similarity algorithm can be understood as that the more common word segments there are between two texts, the more similar the two texts are. When given two text sets A and B, the Jaccard similarity between texts is defined as the ratio of the size of the intersection of A and B to the size of the union of A and B, expressed as:
[0070]
[0071] Specifically, assume there is text A and text B. The number of word segments contained in text B is n1, and the number of different words between text A and text B is m. It means that there are n1 - m identical word segments between text A and text B. After merging and removing duplicates of all word segments of text A and text B, n1 + m word segments can be obtained. If it is required that the Jaccard similarity between text A and text B is greater than or equal to the clustering similarity threshold f, the formula is expressed as:
[0072]
[0073] Through the transformation of the inequality, determine that the number of different words m is expressed as:
[0074]
[0075] Specifically, obtain the clustering similarity threshold f that matches the target text. Based on the similarity condition that the clustering similarity of the word segmentation combination and the target text needs to satisfy and the clustering similarity calculation method, according to the number of clustering word segments and the clustering similarity threshold, through the above formula for the number of different word segments, the maximum value of the number of different word segments between the word segmentation combination that meets the similarity condition and the target text can be calculated.
[0076] Then, based on the number of word segments included in the target word segmentation and the maximum value of the number of different word segments, subtract the number of word segments from the maximum value of the number of different word segments, and the minimum value of the number of word segments required for the word segmentation combination that meets the similarity condition can be determined.
[0077] It can be understood that in this embodiment, the Jaccard similarity algorithm is used as an example for the clustering similarity calculation method. If the clustering similarity calculation method is other methods, the formula for the number of different word segments can be determined by substituting the corresponding known and unknown numbers into the formula of the specific clustering similarity calculation method for inequality transformation, and subsequent processing can be carried out to determine the minimum value of the number of word segments.
[0078] In this embodiment, by presetting the clustering similarity threshold and determining the calculation formula for the number of different word segments according to the clustering similarity calculation method, substituting for calculation and finally determining the minimum value of the number of word segments required for the word segmentation combination that meets the similarity condition can improve the calculation efficiency and accuracy.
[0079] In one embodiment, since there are identical word segments between the word segmentation combination that meets the similarity condition and the target text, therefore, the target word segmentation combination can be obtained by selecting word segments from the target text and combining them.
[0080] Specifically, selecting a target number of word segments from the target text for combination to obtain the target word segmentation combination includes:
[0081] Selecting a target number of word segments from the word segments included in the target text for permutation and combination to obtain multiple candidate word segmentation combinations; screening out the target word segmentation combinations whose included word segments and word segment order simultaneously meet the combination conditions from the multiple candidate word segmentation combinations.
[0082] A candidate word segmentation combination refers to a word segmentation combination obtained by arranging and combining a target number of word segmentations. The combination conditions are the conditions that need to be satisfied to screen out the target word segmentation combination from the candidate word segmentation combinations, and can be determined according to the specific word segmentations and the word order included in the candidate word segmentation combinations. Since in the process of calculating the clustering similarity, it is only necessary to know the number of identical word segmentations shared between two texts and the number of word segmentations after merging, without distinguishing the word order of the word segmentations in the texts, it is necessary to determine the target word segmentation combination that meets the combination conditions from multiple candidate word segmentation combinations obtained after arrangement and combination. So as to determine the identifier corresponding to the target word segmentation combination, and further determine the clustering cluster represented by the target word segmentation combination.
[0083] For example, the target text contains 4 word segmentations, namely word segmentation 1, word segmentation 2, word segmentation 3, and word segmentation 4. Assuming the target number is 3, and the selected word segmentations are word segmentation 1, word segmentation 2, and word segmentation 4. In this case, 6 candidate word segmentation combinations can be obtained after arrangement and combination: [word segmentation 1, word segmentation 2, word segmentation 4], [word segmentation 1, word segmentation 4, word segmentation 2], [word segmentation 2, word segmentation 1, word segmentation 4], [word segmentation 2, word segmentation 4, word segmentation 1], [word segmentation 4, word segmentation 1, word segmentation 2], [word segmentation 4, word segmentation 2, word segmentation 1]. In actual text clustering, as long as the word segmentation combination contains the above 3 word segmentations, the target text and the text corresponding to the word segmentation combination can be classified into the same category. Therefore, it is necessary to screen out the target word segmentation combination whose included word segmentations and word order simultaneously meet the combination conditions from the candidate word segmentation combinations, and then determine the clustering cluster represented by the target word segmentation combination.
[0084] In this embodiment, by determining the candidate word segmentation combinations in the way of arrangement and combination, it is possible to avoid missing the candidate word segmentation combinations composed of word segmentations. By screening out the target word segmentation combinations whose included word segmentations and word order simultaneously meet the combination conditions from multiple candidate word segmentation combinations, and then processing the target word segmentation combinations subsequently, it is possible to avoid repeated processing of candidate word segmentation combinations with the same word segmentation content, and improve the data processing efficiency.
[0085] In one embodiment, the target word segmentation combination is obtained by screening from multiple candidate word segmentation combinations, and the target word segmentation combination can be one or more. Specifically, screening out the target word segmentation combinations whose included word segmentations and word order simultaneously meet the combination conditions from multiple candidate word segmentation combinations includes:
[0086] Performing sorting processing on the word segmentations included in multiple candidate word segmentation combinations to determine the respective word orders of the candidate word segmentation combinations; screening out the target word segmentation combinations with at least one different word segmentation from each word order of the candidate word segmentation combinations.
[0087] The sorting process refers to sorting the word segments. The sorting process can be based on the number of strokes corresponding to the word segments, the order of the pinyin letters corresponding to the first characters in the word segments, etc., and arranging them in a descending or ascending manner, as long as the word segments are in order. Thus, it is ensured that multiple word segment combinations with the same specific content of the word segments will not result in different represented clustering clusters due to different word orders within the word segment combinations.
[0088] Specifically, sort the word segments included in multiple candidate word segment combinations to determine the respective word segment orders of the candidate word segment combinations. Thus, from each word segment order of the candidate word segment combinations, filter out the candidate word segment combinations with at least one different word segment as the target word segment combinations. Different word segments refer to the different word segments existing between the target word segment combinations.
[0089] For example, the target text contains 4 word segments, namely word segment 1, word segment 2, word segment 3, and word segment 4. Assuming the target quantity is 3, when the selected word segments are word segment 1, word segment 2, and word segment 3, 6 candidate word segment combinations can be obtained after permutation and combination: [word segment 1, word segment 2, word segment 3], [word segment 1, word segment 3, word segment 2], [word segment 2, word segment 1, word segment 3], [word segment 2, word segment 3, word segment 1], [word segment 3, word segment 1, word segment 2], [word segment 3, word segment 2, word segment 1]. When the selected word segments are word segment 1, word segment 2, and word segment 4, 6 candidate word segment combinations can be obtained after permutation and combination: [word segment 1, word segment 2, word segment 4], [word segment 1, word segment 4, word segment 2], [word segment 2, word segment 1, word segment 4], [word segment 2, word segment 4, word segment 1], [word segment 4, word segment 1, word segment 2], [word segment 4, word segment 2, word segment 1]. Sorting them in ascending order, the sorted candidate word segment combinations are represented as [word segment 1, word segment 2, word segment 3] and [word segment 1, word segment 2, word segment 4]. Then the target word segment combinations are [word segment 1, word segment 2, word segment 3], [word segment 1, word segment 2, word segment 4], and there are 2 identical word segments and 1 different word segment between the target word segment combinations.
[0090] It should be noted that when selecting the target quantity of word segments from the word segments included in the target text for permutation and combination to obtain multiple candidate word segment combinations, if it is preset that the permutation and combination method is in an ordered manner, and at this time the word segments in the obtained multiple candidate word segment combinations are already in order, then there is no need to perform the step of sorting the word segments included in the multiple candidate word segment combinations to determine the respective word segment orders of the candidate word segment combinations, and directly filter out the target word segment combinations with at least one different word segment.
[0091] In this embodiment, by sorting the words included in multiple candidate word segmentation combinations and then screening out the target word segmentation combination from the multiple candidate word segmentation combinations, it can be ensured that the candidate word segmentation combinations composed of the same words will not result in different encoded strings and represented clustering clusters due to different word orders, which can improve the accuracy of text category determination.
[0092] In one embodiment, taking the clustering cluster represented by the target word segmentation combination as the text category to which the target text belongs is convenient for the subsequent use of the clustering result. Specifically, determining the clustering cluster represented by the target word segmentation combination as the text category to which the target text belongs includes:
[0093] Encoding the words in the target word segmentation combination to respectively obtain the encodings corresponding to the words in the target word segmentation combination; concatenating the encodings to obtain the encoded string matched by the target word segmentation combination; the encoded string corresponds one-to-one with the identifier of the clustering cluster; determining the clustering cluster represented by the encoded string matched by the target word segmentation combination as the text category to which the target text belongs.
[0094] Encoding processing means representing the words in the target combination in another form, specifically, it can be represented by numbers, letters, symbols, etc. The word and the encoding correspond to each other, and the encoding corresponding to the word is unique. Concatenating processing means concatenating the encodings corresponding to the words so that multiple encodings are concatenated into an encoded string. Since the encodings corresponding to the words in the target word segmentation combination are unique, the encoded string matched by the target word segmentation combination is also unique, and the encoded string can also be directly used as the identifier of the target word segmentation combination, which is convenient for the subsequent use of the clustering result. Since the identifier of the target word segmentation combination corresponds to the identifier of the clustering cluster, the encoded string can correspond one-to-one with the identifier of the clustering cluster.
[0095] Specifically, the association relationship between the word and the encoding can be pre-stored, so that, according to the association relationship between the word and the encoding, the encoding corresponding to the word in the target word segmentation combination can be determined, and then the encodings are concatenated subsequently. The hash algorithm can also be used to perform hash encoding processing on the words in the target word segmentation combination to respectively obtain the hash encodings corresponding to the words in the target word segmentation combination, and then the hash encodings are concatenated subsequently to form a hash string. Concatenating the encodings can be to concatenate the encodings in the order of the words using a fixed delimiter, and the fixed delimiter can be "#". It can also be to directly concatenate the encodings to obtain the encoded string matched by the target word segmentation combination, so that the clustering cluster represented by the encoded string matched by the target word segmentation combination is determined as the text category to which the target text belongs.
[0096] It should be noted that when performing hash encoding processing on the words in the target word segmentation combination using the hash algorithm, it can also be to first concatenate the words in the target word segmentation combination using a fixed delimiter to obtain a word segmentation hash string, and then perform hash encoding processing on the word segmentation hash string to obtain the hash string encoding corresponding to the target word segmentation combination, and use the hash string encoding as the identifier of the clustering cluster represented by the target word segmentation combination.
[0097] For example, the hash string encoding corresponding to an initial word segmentation hash string is set to 10001. When encountering a new word segmentation hash string, 10001 can be assigned to the new word segmentation hash string and an increment operation is performed. At this time, the hash string encoding corresponding to the initial word segmentation hash string is 10002, and 10002 is assigned when encountering the next new word segmentation hash string.
[0098] In this embodiment, by determining the encoding string corresponding to the target word segmentation combination and determining the clustering cluster represented by the encoding string matched by the target word segmentation combination as the text category to which the target text belongs, the efficiency of determining the text category can be improved. Determining the clustering clusters corresponding to the text set can also be determined through the encoding string, which is convenient for using the clustering results of the text set. Moreover, the storage space occupied by the encoding string is small, which can save storage space.
[0099] In one embodiment, after determining the text category to which the target text belongs, that is, determining the clustering cluster where the target text is located. When determining the text categories to which all texts in the text set where the target text is located belong, that is, determining the clustering clusters corresponding to the text set, the text clustering of the texts in the text set is realized. It can be understood that in most texts, there is a certain same word. This same word can be called a high-frequency word. In text clustering, the information volume of high-frequency words is relatively low. In order to facilitate the subsequent use of the clustering results, clustering denoising processing can be performed after clustering to avoid false alarms caused by high-frequency words.
[0100] Specifically, the method further includes: obtaining the clustering clusters corresponding to the text set where the target text is located; determining the words included in the texts in each clustering cluster and the respective word numbers of each word; screening out the target clustering clusters whose word numbers meet the clustering cluster optimization conditions from each clustering cluster according to the respective word numbers of each word; and performing data cleaning processing on the words included in the texts in the target clustering clusters to obtain the optimized clustering clusters of the text set.
[0101] The clustering cluster optimization condition refers to the condition that needs to be met for optimizing the clustering cluster by performing clustering denoising processing on the clustering cluster. The target clustering cluster refers to the clustering cluster that meets the clustering cluster optimization condition. Data cleaning processing refers to the processing of verifying data to delete duplicate data and correct data errors. The optimized clustering cluster refers to the clustering cluster obtained after performing data cleaning processing on each clustering cluster of the text set.
[0102] Specifically, obtain each cluster corresponding to the text set where the target text is located. Then, determine the word segments included in the text in each cluster and the respective word segment quantities of each word segment. The cluster optimization condition can be set such that the word segment quantity of a certain word segment in the cluster exceeds a quantity threshold. The quantity threshold can be set according to actual technical needs. For example, it can be set according to the accuracy requirements and experience of the application scenario, and a word segment scale quantity that is unlikely to be reached under normal circumstances can be set in the cluster. When the word segment quantity of one or more word segments in the cluster exceeds the quantity threshold, it is determined that the cluster optimization condition is met, and the cluster is determined as the target cluster. Thus, one or more word segments whose word segment quantities in the text included in the target cluster exceed the quantity threshold can be deleted to obtain each optimized cluster of the text set. Or, the target cluster can be directly deleted to obtain each optimized cluster of the text set.
[0103] For example, in the clustering scenario of company names, the text set is composed of company names. For this text set, high-frequency words can be "Limited", "Company", "Technology", etc. The quantity threshold can be set to 1000, which can be used to remove the clusters formed by high-frequency word combinations to avoid false clustering reports. When the quantity of a certain word segment exceeds the quantity threshold, the word segment in the cluster is deleted to obtain the optimized cluster. Or, the cluster can be directly deleted to obtain the optimized cluster.
[0104] In this embodiment, by setting the cluster optimization condition, the clusters corresponding to the text set are optimized to obtain the optimized clusters. Accordingly, high-frequency words with less information content included in the clusters can be removed to avoid invalid clustering, thereby optimizing the clustering result and effectively improving the accuracy of the clustering result.
[0105] In one embodiment, a community can reflect the local characteristics of individual behaviors in a network and the correlation relationships between them. Studying the communities in the network plays a crucial role in understanding the structure and function of the entire network and can help analyze and predict the interaction relationships between elements of the entire network. In this embodiment, text clustering is combined with a community discovery algorithm to optimize the clustering result.
[0106] Specifically, the method further includes: obtaining each cluster corresponding to the text set where the target text is located; converting each cluster into a clustering relationship graph including overlapping parts according to the text in each cluster; the overlapping part indicates that there is at least one same text in the overlapping clusters; based on the clustering relationship graph and the community discovery algorithm, optimizing the overlapping part in the clustering relationship graph to obtain an optimized clustering relationship graph; and obtaining each optimized cluster of the text set according to the optimized clustering relationship graph.
[0107] The clustering relationship graph is a relationship graph transformed based on the relationships between texts in each clustering cluster. Since a target text in this embodiment can belong to multiple clustering clusters simultaneously, that is, a target text corresponds to multiple text categories, and overlapping clustering is achieved. Therefore, when transforming each clustering cluster into a clustering relationship graph with overlapping parts based on the texts in each clustering cluster, the clustering relationship graph contains overlapping parts, and the overlapping parts represent that there is at least one identical text in the overlapping clustering clusters. Please refer to Figure 3 , taking the example that the clustering clusters corresponding to the text set where the target text is located include 2, namely clustering cluster 1 and clustering cluster 2. If text A can belong to both clustering cluster 1 and clustering cluster 2 simultaneously, then when transformed into a clustering relationship graph, there is an overlapping part formed based on text A between clustering cluster 1 and clustering cluster 2.
[0108] The community discovery algorithm can partition the clustering relationship graph into closely related groups, obtain the group numbers of the nodes in the clustering relationship graph, and achieve the optimization of the clustering relationship graph. The community discovery algorithm can include non - overlapping community discovery algorithms and overlapping community discovery algorithms. In non - overlapping community discovery algorithms, a graph node corresponds to only one group number, and the non - overlapping community discovery algorithm can be the community detection (FastUnfolding) algorithm, label propagation (LPA) algorithm, etc. In overlapping community discovery algorithms, a graph node can have multiple group numbers, and the overlapping community discovery algorithm can be the overlapping community discovery based on label propagation (COPRA) algorithm, etc., and can be specifically selected according to actual optimization technology needs. Specifically, based on the clustering relationship graph and the community discovery algorithm, the clustering relationship graph can be input into the community discovery algorithm, and the overlapping parts in the clustering relationship graph can be optimized through the community discovery algorithm to obtain an optimized clustering relationship graph. Thus, according to the optimized clustering relationship graph, each optimized clustering cluster of the text set can be obtained.
[0109] In this embodiment, since a target text can belong to multiple text categories simultaneously, overlapping clustering has been achieved. Further, by combining the community discovery algorithm, the clustering result of the text set can be optimized. When combining the non - overlapping community discovery algorithm, non - overlapping clustering can be achieved. When combining the overlapping community discovery algorithm, the clustering effect can be adjusted, that is, the clustering effect can be flexibly adjusted.
[0110] The following further details the present application in conjunction with the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0111] In a specific embodiment, please refer to Figure 4, through the text category determination method, text clustering is then achieved. The main steps include word segmentation, word segmentation denoising, word segmentation hashing, hash string encoding, and clustering denoising. Please refer to Figure 5 , taking the text as a short text as an example, the specific steps of the text category determination method are as follows:
[0112] Step S501, obtain the word segmentation contained in the target text.
[0113] Specifically, a short text dataset X = [x1, x2,..., x n composed of n short texts x is preset in advance, and the target text x i is obtained from the short text dataset, where 1 ≤ i ≤ n. Use a word segmentation tool to perform word segmentation on the target text x i to obtain the word segmentation contained in the target text.
[0114] The word segmentation tool can be any one of the Jieba word segmentation tool, the natural language processing tool HanLP, etc. After word segmentation, the word segmentation result set word_list_i = [w1, w2,..., w k1 can be obtained, indicating that x i is divided into k1 words. Remove meaningless characters, such as punctuation marks, from the word segmentation result set word_list_i, or keywords can also be extracted from word_list_i using a pre-extracted keyword dictionary. After word segmentation denoising, f_word_list_i = [w1, w2,..., w k2 is obtained, where k2 is the number of word segments after denoising, and k2 ≤ k1.
[0115] Step S502, determine the classified texts other than the target text from the text set where the target text is located; determine the clustering word segmentation quantity corresponding to the text set according to the number of word segments contained in each of the classified texts.
[0116] Specifically, the clustering word segmentation quantity can be the minimum value of the number of word segments contained in each of the classified texts, denoted as n1, or it can be the average value of the number of word segments contained in each of the classified texts, denoted as n2.
[0117] Step S503, obtain the clustering similarity threshold that matches the target text.
[0118] Specifically, the clustering similarity threshold that matches the target text is denoted as f.
[0119] Step S504, based on the similarity condition that the word segmentation combination needs to satisfy with the target text for clustering similarity and the clustering similarity calculation method, calculate the maximum value of the number of different word segments between the word segmentation combination that meets the similarity condition and the target text according to the clustering word segmentation quantity and the clustering similarity threshold.
[0120] Specifically, the clustering similarity calculation method is the Jaccard similarity algorithm, and the number of different segmented words is represented as m. When the number of clustered segmented words is the minimum value n1 among the numbers of segmented words contained in the classified texts respectively, the maximum value of the number of different segmented words is expressed as:
[0121]
[0122] When the number of clustered segmented words is the average value n2 among the numbers of segmented words contained in the classified texts respectively, the maximum value of the number of different segmented words is expressed as:
[0123]
[0124] Step S505: Based on the number of segmented words contained in the target segmented words and the maximum value of the number of different segmented words, determine the minimum number of segmented words required for the segmented word combination that meets the similarity condition.
[0125] Specifically, subtract the maximum value of the number of different segmented words from the number of segmented words contained in the target segmented words to determine the minimum number of segmented words required for the segmented word combination that meets the similarity condition, which is expressed as k2 - m.
[0126] Step S506: Select a target number of segmented words from the segmented words contained in the target text for permutation and combination to obtain multiple candidate segmented word combinations; the target number is greater than or equal to the minimum number of segmented words.
[0127] Specifically, at least k2 - m segmented words and no more than k2 segmented words can be selected from the target text for permutation and combination to obtain multiple candidate segmented word combinations.
[0128] Step S507: Sort the segmented words contained in the multiple candidate segmented word combinations to determine the respective segmented word orders of the candidate segmented word combinations; from each segmented word order of the candidate segmented word combinations, filter out the target segmented word combinations with at least one different segmented word.
[0129] Specifically, sort the segmented words contained in the multiple candidate segmented word combinations in ascending order of the number of strokes to determine the respective segmented word orders of the candidate segmented word combinations; from each segmented word order of the candidate segmented word combinations, filter out the target segmented word combinations with at least one different segmented word.
[0130] Step S508: Encode the segmented words in the target segmented word combination to obtain the corresponding codes of the segmented words in the target segmented word combination respectively; concatenate the codes to obtain the code string matched by the target segmented word combination; the code string corresponds one-to-one with the identifier of the clustering cluster; determine the clustering cluster represented by the code string matched by the target segmented word combination as the text category to which the target text belongs.
[0131] Specifically, using the hash coding method, the words in the target word segmentation combination are hash-coded to obtain the hash codes corresponding to the words in the target word segmentation combination respectively, and the hash codes are concatenated into a hash string. The hash string is used as the id of the clustering cluster, that is, the identifier. The clustering cluster represented by the hash string matched by the target word segmentation combination is determined as the text category to which the target text belongs. Further, if the target word segmentation combination has not been encoded, the hash code of the target text and the target word segmentation combination is stored correspondingly. If the target word segmentation combination has been encoded, the previous hash code is retrieved and stored correspondingly.
[0132] Step S509: Obtain each clustering cluster corresponding to the text set where the target text is located; determine the words included in the text in each clustering cluster and the number of words for each word respectively.
[0133] Specifically, count the words and their word numbers in each clustering cluster id, expressed as {cluster_id1:n1,cluster_id2:n2,…,cluster_id n :n n}, where cluster_id n :n n means that there are n n words in the nth clustering cluster.
[0134] Step S510: According to the number of words for each word respectively, select the target clustering clusters whose word numbers meet the clustering cluster optimization conditions from each clustering cluster; perform data cleaning processing on the words included in the text in the target clustering clusters to obtain each optimized clustering cluster of the text set.
[0135] Specifically, set the quantity threshold mx of the word number, and set the clustering cluster optimization condition to that the word number is greater than or equal to the quantity threshold. According to the number of words for each word respectively, select the target clustering clusters whose word numbers meet the clustering cluster optimization conditions from each clustering cluster; delete the words in the target clustering clusters that are greater than or equal to the quantity threshold to obtain each optimized clustering cluster of the text set.
[0136] In this embodiment, the above steps S501 to S508 are referred to as the word segmentation hash algorithm, which can be specifically divided into a method for selecting the word segmentation combination for hashing and a method for forming the word segmentation hash string. Specifically, the word segmentation combination for hashing can be selected according to the requirements of clustering similarity.
[0137] Assume that after word segmentation and word noise removal of the target text, the number of words included is w. If it is required that at least w - m words of other texts are the same as this text in order to perform similarity aggregation of the target text and other texts, then the calculation formula for the number of extractable word segmentation combinations corresponding to the target text is expressed as:
[0138]
[0139] Among them, the calculation formula for the value of the number of differential words m is:
[0140]
[0141] Among them, f is the Jaccard similarity threshold required for clustering. Only when the similarity between texts is greater than or equal to f can they be clustered into one category. If n1 is the minimum value of the number of words segmented for each text input into the word segmentation hashing algorithm in the dataset where the target word segmentation is located, then the calculated value of m can satisfy that the Jaccard similarity between all texts in the same category after clustering is greater than or equal to the threshold. If n1 is the average value of the number of words segmented for each text input into the word segmentation hashing in the dataset where the target word segmentation is located, then the calculated value of m can satisfy that the Jaccard similarity between most texts in the same category after clustering is greater than or equal to the threshold.
[0142] In a specific embodiment, please refer to Figure 6 , and the method for determining text categories will be described in detail in combination with two specific target texts, specifically including:
[0143] Obtain the target text, which is "Male Emotional Counseling Studio". After performing word segmentation processing and word segmentation denoising processing on the target text, 4 words segmented from the target text are obtained, namely "male", "emotion", "counseling", and "studio". If n1 is the minimum value of the number of words segmented for each text input into the word segmentation hashing algorithm in the dataset where the target word segmentation is located, and the Jaccard similarity threshold required for clustering is 0.6, then the maximum value of the number of differential words m can be determined to be 1. Then there are 5 combinations of words segmented by hashing that the target text "Male Emotional Counseling Studio" can select, namely [male, emotion, counseling, studio], [male, emotion, studio], [male, emotion, counseling], [male, counseling, studio], [emotion, counseling, studio]. Similarly, for the target text "Male Psychological Counseling Studio", the corresponding combinations of words segmented by hashing are [male, psychology, counseling, studio], [male, psychology, studio], [male, psychology, counseling], [male, counseling, studio], [psychology, counseling, studio].
[0144] Then, the words in each word combination are sorted, and the sorted words in the word combination are concatenated using a fixed separator "#" to form a word hash string corresponding to the word combination. Specifically, [male, emotion, consultation, studio] is processed as “#consulting#studio#emotion#male”, [male, emotion, studio] is processed as “#studio#emotion#male”, [male, emotion, consultation] is processed as “#consulting#emotion#male”, [male, consultation, studio] is processed as “#consulting#studio#male”, [emotion, consultation, studio] is processed as “#consulting#studio#emotion”, [male, psychology, consultation, studio] is processed as “#consulting#studio#psychology#male”, [male, psychology, studio] is processed as “#studio#psychology#male”, [male, psychology, consultation] is processed as “#consulting#psychology#male”, [male, consultation, studio] is processed as “#consulting#studio#male”, [psychology, consultation, studio] is processed as “#consulting#studio#psychology”.
[0145] Thus, the segmentation hash string corresponding to the segmentation combination is encoded to obtain the encoded hash string, and the encoded hash string is used as the category id encoding to facilitate the use of clustering results. Specifically, "#consulting#studio#emotion#male" is encoded as 00001, "#studio#emotion#male" is encoded as 00002, "#consulting#emotion#male" is encoded as 00003, "#consulting#studio#male" is encoded as 00004, "#consulting#studio#emotion" is encoded as 00005, "#consulting#studio#psychology#male" is encoded as 00006, "#studio#psychology#male" is encoded as 00007, "#consulting#psychology#male" is encoded as 00008, "#consulting#studio#male" has been encoded as 00004, and "#consulting#studio#psychology" is encoded as 00009.
[0146] Finally, clustering denoising processing can be performed on the clustering results of the text, the segmented words and their number under each category id can be counted, and invalid clusters whose number of objects under each category id is greater than the threshold can be removed.
[0147] In a specific embodiment, the process of using the text clustering results is explained in combination with an actual specific application scenario. Take the scenario of using it in the identification of associations between enterprises as an example, such as identifying similar enterprises based on short texts such as address, company name, product name, and product content. In the backend, the short texts corresponding to the address, company name, product name, and product content of the enterprise have been clustered in advance, and the clustering results of each dimension are collected in the same table. In the frontend, please refer to Figure 7The interface diagram allows users to enter the company name for which similar companies need to be identified at the company name location. By retrieving the backend library table, companies that are similar in terms of company name, attribute association, product name similarity, registration time proximity, same business industry, address proximity, etc. can be searched and displayed on the interface to facilitate subsequent operations by the user.
[0148] The word segmentation hash clustering algorithm of the embodiment of the present application is designed starting from the similarity conditions required for clustering similarity, thereby avoiding the steps of calculating and comparing similarity during the clustering process in the traditional method, and thus achieving efficient clustering with a nearly linear time complexity.
[0149] To verify the beneficial effects produced by the method of the embodiment of the present application compared with the traditional method, the processing efficiency of the method of the embodiment of the present application and the traditional method was compared. Table 1 shows the comparison table of the theoretical time complexity corresponding to each method.
[0150] Table 1 Comparison table of theoretical time complexity
[0151]
[0152] Among them, n is the data scale, w is the number of word segments of the target text input into the word segmentation hash clustering algorithm, m is the number of different word segments that meet the similarity conditions calculated in the embodiment of the present application. k is the number of categories formed after clustering, and t is the number of K-means iterations.
[0153] Taking a specific large-scale short text clustering scenario as an example, assuming that the text data scale n = 100 million, the average value of the number of word segments w of the word segmentation hash input into the word segmentation hash clustering algorithm is 7, and the clustering similarity threshold f = 0.75, then the maximum value of the number of different word segments m can be calculated as 1. Assuming that every 20 text data form a cluster, that is, k = 100 million / 20 = 5 million, that is, the number of categories formed after text clustering is 5 million, and the number of iterations t of the K-means algorithm is 20. Substituting the data into the formula of the theoretical time complexity in Table 1, the time complexity of each method in this specific scenario can be calculated. Table 2 shows the comparison table of the theoretical time complexity corresponding to each method in the specific scenario.
[0154] Table 2 Comparison table of theoretical time complexity in the specific scenario
[0155] Algorithm Name Theoretical Time Complexity in Specific Scenarios The Word Segmentation Hash Clustering Algorithm of the Embodiment of the Present Application O(8 * 100 million) A Text Clustering Method in a Traditional Way O(10 million * 100 million) Distance-Based Clustering Algorithm (K-means) O(20 * 10 million * 100 million) Density-Based Clustering Algorithm (DB-SCAN) O(100 million * 100 million)
[0156] In the above scenario, the clustering time overhead ratio of the word segmentation hash clustering algorithm of the embodiment of the present application and a traditional text clustering method is:
[0157]
[0158] The clustering time overhead ratio between the word segmentation hash clustering algorithm and the K-means algorithm in the embodiments of this application is as follows:
[0159]
[0160] It can be seen that the clustering efficiency of the word segmentation hash clustering algorithm in the embodiments of this application is nearly linear time complexity, and it has an advantage in clustering efficiency in the scenario of large-scale short text clustering.
[0161] It should be understood that although the steps in the flowcharts involved in the above-mentioned embodiments are displayed in sequence according to the indication of the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear description in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, at least some of the steps in the flowcharts involved in the above-mentioned embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or alternately with at least some of the steps or stages in other steps or other steps.
[0162] Based on the same inventive concept, this application also provides a text category determination device for implementing the text category determination method involved above. The implementation solutions for solving problems provided by this device are similar to the implementation solutions described in the above text category determination method. Therefore, the specific limitations in one or more embodiments of the text category determination device provided below can refer to the limitations on the text category determination method in the above text, and will not be repeated here.
[0163] In one embodiment, as Figure 8 shown, a text category determination device is provided, including: a word segmentation acquisition module 10, a quantity calculation module 20, a combination determination module 30, and a category determination module 40, where:
[0164] The word segmentation acquisition module 10 is used to acquire the word segmentation included in the target text.
[0165] The quantity calculation module 20 is used to calculate the minimum value of the number of word segments required for a word segment combination that satisfies the similarity condition with the clustering similarity of the target text according to the number of word segments included in the target text.
[0166] The combination determination module 30 is used to select a target number of word segments from the target text for combination to obtain a target word segment combination; the target number is greater than or equal to the minimum value of the number of word segments; there is at least one different word segment in different target word segment combinations.
[0167] A category determination module 40, configured to determine the clustering cluster represented by the target word segmentation combination as the text category to which the target text belongs.
[0168] In one embodiment, the quantity calculation module 20 includes:
[0169] A classified text determination unit, configured to determine, from the text set where the target text is located, the classified texts other than the target text.
[0170] A clustering word segmentation quantity determination unit, configured to determine the clustering word segmentation quantity corresponding to the text set according to the word segmentation quantities included in the respective classified texts.
[0171] A minimum word segmentation quantity determination unit, configured to calculate the minimum value of the word segmentation quantity required for a word segmentation combination whose clustering similarity with the target text meets the similarity condition, based on the word segmentation quantity included in the target text and the clustering word segmentation quantity.
[0172] In one embodiment, the minimum word segmentation quantity determination unit includes:
[0173] A clustering similarity threshold acquisition unit, configured to acquire a clustering similarity threshold that matches the target text.
[0174] A maximum difference word segmentation quantity calculation unit, configured to calculate, according to the clustering word segmentation quantity and the clustering similarity threshold, the maximum value of the difference word segmentation quantity between a word segmentation combination that meets the similarity condition and the target text, based on the similarity condition that the clustering similarity between the word segmentation combination and the target text needs to meet and the clustering similarity calculation method.
[0175] A minimum word segmentation quantity calculation unit, configured to determine the minimum value of the word segmentation quantity required for a word segmentation combination that meets the similarity condition, based on the word segmentation quantity included in the target word segmentation and the maximum value of the difference word segmentation quantity.
[0176] In one embodiment, the combination determination module 30 includes:
[0177] A candidate word segmentation combination determination unit, configured to select a target number of word segmentations from the word segmentations included in the target text for permutation and combination to obtain a plurality of candidate word segmentation combinations.
[0178] A target word segmentation combination determination unit, configured to screen out a target word segmentation combination whose included word segmentations and word segmentation order simultaneously meet the combination condition from the plurality of candidate word segmentation combinations.
[0179] In one embodiment, the target word segmentation combination determination unit includes:
[0180] A word segmentation sorting processing unit, configured to sort the word segments included in the multiple candidate word segmentation combinations to determine the word segment order of each candidate word segmentation combination.
[0181] A word segmentation combination screening unit, configured to screen out target word segmentation combinations with at least one different word segment from each of the word segment orders of the candidate word segmentation combinations.
[0182] In one embodiment, the category determination module 40 includes:
[0183] A word segmentation encoding processing unit, configured to perform encoding processing on the word segments in the target word segmentation combination to respectively obtain the encodings corresponding to the word segments in the target word segmentation combination.
[0184] A word segmentation encoding concatenation unit, configured to concatenate the encodings to obtain an encoding string matched by the target word segmentation combination; the encoding string corresponds one-to-one to the identifier of the clustering cluster.
[0185] A text category determination unit, configured to determine the clustering cluster represented by the encoding string matched by the target word segmentation combination as the text category to which the target text belongs.
[0186] In one embodiment, the device further includes: a first optimization processing module. The first optimization processing module includes:
[0187] A clustering cluster acquisition unit, configured to acquire each clustering cluster corresponding to the text set where the target text is located.
[0188] A word segment determination unit, configured to determine the word segments included in the text in each clustering cluster and the word segment quantity of each word segment.
[0189] A target clustering cluster determination unit, configured to screen out target clustering clusters whose word segment quantity meets the clustering cluster optimization condition from each clustering cluster according to the word segment quantity of each word segment.
[0190] An optimization processing unit, configured to perform data cleaning processing on the word segments included in the text in the target clustering cluster to obtain each optimized clustering cluster of the text set.
[0191] In one embodiment, the device further includes: a second optimization processing module. The second optimization processing module includes:
[0192] A clustering cluster obtaining unit, configured to acquire each clustering cluster corresponding to the text set where the target text is located.
[0193] A relationship graph transformation unit, configured to transform each of the clustering clusters into a clustering relationship graph including overlapping parts according to the texts in each of the clustering clusters; the overlapping parts indicate that there is at least one identical text in the overlapping clustering clusters.
[0194] A relationship graph optimization unit, configured to perform an optimization process on the overlapping parts in the clustering relationship graph based on the clustering relationship graph and a community discovery algorithm to obtain an optimized clustering relationship graph.
[0195] An optimization result acquisition unit, configured to obtain each optimized clustering cluster of the text set according to the optimized clustering relationship graph.
[0196] Each module in the above text category determination device can be implemented in whole or in part by software, hardware, and their combination. Each of the above modules can be embedded in the processor in the computer device in the form of hardware or be independent of the processor, or can be stored in the memory in the computer device in the form of software, so that the processor can call and execute the operations corresponding to each of the above modules.
[0197] In one embodiment, a computer device is provided. The computer device can be a server, and its internal structure diagram can be as Figure 9 shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O), and a communication interface. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store text data. The input / output interface of the computer device is used for the processor to exchange information with external devices. The communication interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it implements a text category determination method.
[0198] In one embodiment, a computer device is provided. The computer device can be a terminal, and its internal structure diagram can be as Figure 10As shown in the figure. The computer device includes a processor, a memory, an input / output interface, a communication interface, a display unit, and an input device. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface, the display unit, and the input device are connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals in a wired or wireless manner, and the wireless manner can be implemented through WIFI, a mobile cellular network, NFC (Near Field Communication), or other technologies. When the computer program is executed by the processor, it implements a method for determining text categories. The display unit of the computer device is used to form a visually visible picture, which can be a display screen, a projection device, or a virtual reality imaging device. The display screen can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer covering the display screen, or a button, a trackball, or a touchpad provided on the computer device housing, or an external keyboard, touchpad, or mouse, etc.
[0199] Those skilled in the art can understand that Figure 9 and Figure 10 the structure shown in the figure is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have a different component layout.
[0200] In one embodiment, a computer device is provided, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the steps of the above method are implemented.
[0201] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by the processor, the steps of the above method are implemented.
[0202] In one embodiment, a computer program product is provided, including a computer program. When the computer program is executed by the processor, the steps of the above method are implemented.
[0203] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memories. Non-volatile memories can include read-only memory (ROM), magnetic tapes, floppy disks, flash memories, optical memories, high-density embedded non-volatile memories, resistive random access memories (ReRAM), magnetoresistive random access memories (MRAM), ferroelectric random access memories (FRAM), phase change memories (PCM), graphene memories, etc. Volatile memories can include random access memory (RAM) or external cache memories, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in the present application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in the present application can be general-purpose processors, central processors, graphics processors, digital signal processors, programmable logics, data processing logics based on quantum computing, etc., without limitation.
[0204] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered to be within the scope described in this specification.
[0205] The above embodiments only represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the patent of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the appended claims.
Claims
1. A method for determining text categories, characterized in that The method includes: Obtaining the word segments included in the target text; Based on the similarity condition that the combined word segments need to satisfy with the clustering similarity of the target text and the calculation method of the clustering similarity, according to the number of clustering word segments and the clustering similarity threshold matching the target text, calculating the maximum value of the number of different word segments between the combined word segments that satisfy the similarity condition and the target text; Based on the number of word segments included in the target text and the maximum value of the number of different word segments, determining the minimum number of word segments required for the combined word segments that satisfy the similarity condition; Selecting a target number of word segments from the target text for combination to obtain a target combined word segment; the target number is greater than or equal to the minimum number of word segments; there is at least one different word segment in different target combined word segments; Determining the clustering cluster represented by the target combined word segment as the text category to which the target text belongs.
2. The method according to claim 1, characterized in that The method further includes: Determining the classified texts other than the target text from the text set where the target text is located; According to the number of word segments included in each of the classified texts, determining the number of clustering word segments corresponding to the text set.
3. The method according to claim 1, wherein The step of selecting a target number of word segments from the target text for combination to obtain a target combined word segment includes: Selecting a target number of word segments from the word segments included in the target text for permutation and combination to obtain a plurality of candidate combined word segments; Filtering out the target combined word segments whose included word segments and word segment order simultaneously satisfy the combination conditions from the plurality of candidate combined word segments.
4. The method according to claim 3, wherein The step of filtering out the target combined word segments whose included word segments and word segment order simultaneously satisfy the combination conditions from the plurality of candidate combined word segments includes: Performing a sorting process on the word segments included in the plurality of candidate combined word segments to determine the word segment order of each candidate combined word segment; Filtering out the target combined word segments with at least one different word segment from each word segment order of the candidate combined word segments.
5. The method according to claim 4, wherein The step of determining the clustering cluster represented by the target combined word segment as the text category to which the target text belongs includes: Performing an encoding process on the word segments in the target combined word segment to respectively obtain the encodings corresponding to the word segments in the target combined word segment; Performing a concatenation process on the encodings to obtain an encoding string matched by the target combined word segment; the encoding string corresponds one-to-one with the identifier of the clustering cluster; Determining the clustering cluster represented by the encoding string matched by the target combined word segment as the text category to which the target text belongs.
6. The method according to claim 1, characterized in that The method further includes: Obtaining each clustering cluster corresponding to the text set where the target text is located; Determining the word segments included in the texts in each clustering cluster and the number of word segments of each word segment respectively; According to the number of word segments of each word segment respectively, filtering out the target clustering clusters whose word segment numbers satisfy the clustering cluster optimization condition from each clustering cluster; Performing a data cleaning process on the word segments included in the texts in the target clustering clusters to obtain each optimized clustering cluster of the text set.
7. The method according to claim 1, wherein The method further includes: Obtaining each clustering cluster corresponding to the text set where the target text is located; According to the text in each of the clusters, transform each of the clusters into a cluster relationship graph with overlapping parts; the overlapping parts represent that there is at least one same text in the overlapping clusters. Based on the cluster relationship graph and the community discovery algorithm, perform optimization processing on the overlapping parts in the cluster relationship graph to obtain an optimized cluster relationship graph. According to the optimized cluster relationship graph, obtain each optimized cluster of the text set.
8. A text category determination device, characterized in that The device includes: A word segmentation acquisition module, configured to acquire the word segmentation included in the target text. A calculation module, configured to calculate the maximum value of the number of different word segments between the word segment combination that meets the similarity condition and the target text based on the similarity condition required for the cluster similarity between the word segment combination and the target text and the cluster similarity calculation method, and according to the number of word segments included in the target text and the maximum value of the number of different word segments; determine the minimum number of word segments required for the word segment combination that meets the similarity condition. A combination determination module, configured to select a target number of word segments from the target text for combination to obtain a target word segment combination; the target number is greater than or equal to the minimum number of word segments; there is at least one different word segment in different target word segment combinations. A category determination module, configured to determine the cluster represented by the target word segment combination as the text category to which the target text belongs.
9. The text category determination device according to claim 8, characterized in that, The device further includes: A classified text determination unit, configured to determine the classified texts other than the target text from the text set where the target text is located. A cluster word segment number determination unit, configured to determine the cluster word segment number corresponding to the text set according to the number of word segments included in each of the classified texts.
10. The text category determination device according to claim 8, characterized in that The combination determination module includes: A candidate word segment combination determination unit, configured to select a target number of word segments from the word segments included in the target text for permutation and combination to obtain a plurality of candidate word segment combinations. A target word segment combination determination unit, configured to screen out the target word segment combinations whose included word segments and word segment order simultaneously meet the combination conditions from the plurality of candidate word segment combinations.
11. The text category determination device according to claim 10, characterized in that, The target word segment combination determination unit includes: A word segment sorting processing unit, configured to perform sorting processing on the word segments included in the plurality of candidate word segment combinations to determine the word segment order of each candidate word segment combination. A word segment combination screening unit, configured to screen out the target word segment combinations with at least one different word segment from each word segment order of the candidate word segment combinations.
12. The text category determination device according to claim 11, characterized in that, The category determination module includes: A word segment encoding processing unit, configured to perform encoding processing on the word segments in the target word segment combination to obtain the encodings corresponding to the word segments in the target word segment combination respectively. A word segment encoding concatenation unit, configured to perform concatenation processing on each of the encodings to obtain the encoding string matched by the target word segment combination; the encoding string corresponds one-to-one to the identifier of the cluster. A text category determination unit, configured to determine the cluster represented by the encoding string matched by the target word segment combination as the text category to which the target text belongs.
13. The text category determination device according to claim 8, characterized in that, The device further includes: The first optimization processing module, and the first optimization processing module includes: A clustering cluster acquisition unit, configured to acquire each clustering cluster corresponding to the text set where the target text is located; A word segmentation determination unit, configured to determine the word segmentation included in the text in each of the clustering clusters, and the word segmentation quantity of each word segmentation; A target clustering cluster determination unit, configured to screen out target clustering clusters whose word segmentation quantity meets the clustering cluster optimization condition from each of the clustering clusters according to the word segmentation quantity of each word segmentation; An optimization processing unit, configured to perform data cleaning processing on the word segmentation included in the text in the target clustering cluster to obtain each optimized clustering cluster of the text set.
14. The text category determination device according to claim 8, characterized in that, The apparatus further includes: A second optimization processing module, and the second optimization processing module includes: A clustering cluster obtaining unit, configured to acquire each clustering cluster corresponding to the text set where the target text is located; A relationship graph conversion unit, configured to convert each of the clustering clusters into a clustering relationship graph including overlapping parts according to the text in each of the clustering clusters; the overlapping parts indicate that there is at least one same text in the overlapping clustering clusters; A relationship graph optimization unit, configured to perform optimization processing on the overlapping parts in the clustering relationship graph based on the clustering relationship graph and a community discovery algorithm to obtain an optimized clustering relationship graph; An optimization result acquisition unit, configured to obtain each optimized clustering cluster of the text set according to the optimized clustering relationship graph.
15. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, the method according to any one of claims 1 to 7 is implemented.
16. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, the method according to any one of claims 1 to 7 is implemented.
17. A computer program product, comprising computer instructions, characterized in that, When the computer instruction is executed by the processor, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Setting method and device of content label and storage medium
CN108009228A
Text processing method and device
CN113919344A