A method and device for text data clustering
Through the methods of feature weight sorting and topic clustering, the problem of low text clustering accuracy in the prior art is solved, and more efficient and accurate text data clustering is achieved.
Patent Information
- Application Number
- CN201910221823.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2019-03-22
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2039-03-22
AI Technical Summary
In the prior art, in the process of text clustering, the clustering results are easily inaccurate because the cluster class created first affects the accuracy of subsequent clustering, and the sequential processing based on the time dimension.
The text data is sorted through feature weights, and priority is given to clustering text data containing rich information to form topic categories, and then text clustering is carried out based on the clustered topic categories.
Improve the accuracy of text clustering, especially when processing streaming data, it can greatly improve the accuracy and efficiency of clustering.
Smart Images

Figure CN111723201B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and in particular to a method and device for text data clustering. Background Art
[0002] Generally, in the process of text clustering, the previously created clusters are used to determine whether the subsequent text data belongs to the created clusters. Therefore, if clusters are first created based on text data that does not contain much information, the accuracy of subsequent clustering will be affected. In the prior art, when calculating text clusters, the texts are clustered based on the order of the time dimension, which can easily lead to inaccurate clustering results. Summary of the invention
[0003] In view of this, an embodiment of the present invention provides a method and device for text data clustering, which can rearrange the order of text clustering by using feature weights, and can preferentially cluster text data containing rich information to form topic clusters, and then perform text clustering based on the clustered topic clusters, thereby improving the accuracy of clustering.
[0004] To achieve the above objective, according to one aspect of an embodiment of the present invention, a method for text data clustering is provided.
[0005] The method for text data clustering of an embodiment of the present invention includes: acquiring batch text data, and determining a feature word set for each text data in the batch text data; for the feature word set of each text data, determining the weight of each feature word in the feature word set; sorting the batch text data according to the weight of the feature words in the feature word set; and performing clustering calculation on the batch text data based on the sorting result.
[0006] Optionally, the step of sorting the batch text data according to the weights of the feature words in the feature word set includes: determining the sum of weights of all feature words in each feature word set; and sorting the batch text data according to the sum of weights.
[0007] Optionally, based on the sorting result, the step of performing clustering calculation on the batch text data includes: based on the sorting result, traversing the batch text data, and determining the similarity value between the text data and the cluster center of the created topic class in turn; judging whether the similarity value meets a preset threshold; if so, adding the text data to the corresponding created topic class and updating the cluster center of the created topic class; otherwise, creating a new topic class.
[0008] Optionally, the step of determining the similarity value between each text data and the cluster center of the created topic class includes: determining the hash value of the feature word of each text data through a hash algorithm; weighting the hash value of each feature word according to the determined weight to obtain the weighted hash value of the feature word; accumulating the weighted hash value of the feature word of each text data to obtain a sequence string representation of the text data; determining the sim-hash value of the text data according to the sequence string representation of the text data; and determining the similarity value of the text data and the cluster center of the created topic class according to the sim-hash value of the text data.
[0009] Optionally, the similarity value includes a title similarity value and a content similarity value;
[0010] The step of judging whether the similarity value meets a preset threshold comprises: when it is determined that the title similarity value does not meet a first preset threshold, judging whether the content similarity value meets a second preset threshold.
[0011] Optionally, for each feature word set of text data, the step of determining the weight of each feature word in the feature word set includes: determining the TF-IDF weight value of the feature word in each feature word set by a term frequency-inverse document frequency algorithm.
[0012] To achieve the above objective, according to another aspect of an embodiment of the present invention, a device for text data clustering is provided.
[0013] The device for text data clustering according to the embodiment of the present invention comprises:
[0014] A feature word determination module, used to obtain batch text data and determine a feature word set for each text data in the batch text data;
[0015] A weight determination module is used for a feature word set of each text data to determine the weight of each feature word in the feature word set;
[0016] A sorting module, used for sorting the batch text data according to the weights of the feature words in the feature word set;
[0017] A clustering module is used to perform clustering calculation on the batch text data based on the sorting result.
[0018] Optionally, the sorting module is further used to determine the sum of weights of all feature words in each feature word set; and sort the batch text data according to the sum of weights.
[0019] Optionally, the clustering module is also used to traverse the batch text data based on the sorting results, and determine the similarity values between the text data and the cluster centers of the created topic classes in turn; and, determine whether the similarity values meet a preset threshold; if so, add the text data to the corresponding created topic class and update the cluster center of the created topic class; otherwise, create a new topic class.
[0020] Optionally, the clustering module is also used to determine the hash value of the feature word of each text data through a hash algorithm; perform weighted processing on the hash value of each feature word according to the determined weight to obtain the weighted hash value of the feature word; accumulate the weighted hash value of the feature word of each text data to obtain a sequence string representation of the text data; determine the sim-hash value of the text data according to the sequence string representation of the text data; and determine the similarity value between the text data and the cluster center of the created topic class according to the sim-hash value of the text data.
[0021] Optionally, the clustering module is further configured to: the step of determining whether the similarity value meets a preset threshold comprises: determining whether the content similarity value meets a second preset threshold when determining that the title similarity value does not meet a first preset threshold;
[0022] The similarity values include a title similarity value and a content similarity value.
[0023] Optionally, the weight determination module is further configured to determine the TF-IDF weight value of the feature word in each feature word set by using a term frequency-inverse document frequency algorithm.
[0024] To achieve the above objective, according to another aspect of an embodiment of the present invention, an electronic device is provided.
[0025] The electronic device of an embodiment of the present invention includes: one or more processors; a storage device for storing one or more programs, when the one or more programs are executed by the one or more processors, the one or more processors implement any of the above methods for text data clustering.
[0026] To achieve the above objective, according to another aspect of an embodiment of the present invention, a computer-readable medium is provided, on which a computer program is stored, characterized in that when the program is executed by a processor, any of the above methods for text data clustering is implemented.
[0027] One embodiment of the above invention has the following advantages or beneficial effects: by rearranging the order of text clustering through feature weights, text data containing rich information can be clustered into topic clusters first, and then text clustering can be performed based on the clustered topic clusters to improve the accuracy of clustering.
[0028] The further effects of the above-mentioned non-conventional optional manner will be described below in conjunction with the specific implementation manner. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] The accompanying drawings are used to better understand the present invention and do not constitute an improper limitation of the present invention.
[0030] Figure 1 is a schematic diagram of the main process of a method for text data clustering according to an embodiment of the present invention;
[0031] Figure 2 is a schematic diagram of a method for discovering hot topics according to an embodiment of the present invention;
[0032] Figure 3 is a schematic diagram of performing clustering calculation on text data according to an embodiment of the present invention;
[0033] Figure 4 is a schematic diagram of main modules of an apparatus for text data clustering according to an embodiment of the present invention;
[0034] Figure 5 is an exemplary system architecture diagram to which embodiments of the present invention may be applied;
[0035] Figure 6 It is a schematic diagram of the structure of a computer system of a terminal device or a server suitable for implementing an embodiment of the present invention. DETAILED DESCRIPTION
[0036] The following is a description of exemplary embodiments of the present invention in conjunction with the accompanying drawings, including various details of the embodiments of the present invention to facilitate understanding, which should be considered as merely exemplary. Therefore, it should be recognized by those of ordinary skill in the art that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present invention. Similarly, for clarity and conciseness, the description of well-known functions and structures is omitted in the following description.
[0037] Figure 1 is a schematic diagram of the main process of the method for text data clustering according to an embodiment of the present invention, such as Figure 1 As shown, the method for text data clustering in the embodiment of the present invention mainly includes:
[0038] Step S101: Acquire batch text data, and determine the feature word set of each text data in the batch text data. Through this process, words with relatively strong representative meanings can be screened out from the text data as feature words. And, in order to reduce the complexity of calculation and improve the speed and efficiency of text processing, try to make the number of words to be processed relatively small when selecting. Therefore, further, the initial feature words screened out can be scored, sorted according to the size of the score, and the initial feature words with low scores can be deleted (the size of the screening score can be set according to business needs), and several items with high scores can be selected as feature words extracted from the text. For a text data, the extracted feature words are generally multiple, and the multiple feature words constitute the feature word set of the text data.
[0039] Step S102: for each feature word set of text data, determine the weight of each feature word in the feature word set. The weight of a feature word refers to the role played by the feature word in a certain text data. In the embodiment of the present invention, the weight can be calculated by data statistics or by scoring.
[0040] Step S103: Sort the batch text data according to the weights of the feature words in the feature word set. In an embodiment of the present invention, the sum of the weights of all the feature words in each feature word set is determined; and the batch text data is sorted according to the sum of the weights. Specifically, the feature weights are sorted from high to low, and the text data at the front contains a relatively rich amount of information, which can be preferentially clustered to form topic clusters (topic classes), and then the text is clustered according to the clustered topic clusters, which is more accurate. Among them, as long as the technical point of sorting by feature word weights to reflect the amount of information contained in the text data, such as sorting batch text data according to the weight product of the feature words in each feature word set, are all within the protection scope of the present invention.
[0041] Step S104: Based on the sorting results, clustering calculation is performed on the batch text data. In this process, based on the sorting results, the batch text data is traversed, and the similarity values between the text data and the cluster centers of the created topic classes are determined in turn. Also, it is determined whether the similarity value meets the preset threshold; if it does, the text data is added to the corresponding created topic class, and the cluster center of the created topic class is updated; otherwise, a new topic class is created.
[0042] In the process of determining the similarity value between each text data and the cluster center of the created topic class, the hash value of the feature word of each text data is determined by the hash algorithm, and each feature word is represented by a string of numbers, thereby achieving the purpose of dimensionality reduction. Then, according to the weight of each feature word determined, its hash value is weighted to obtain the weighted hash value of the feature word. After that, the weighted hash value of the feature word of each text data is accumulated to obtain the sequence string representation of the text data. According to the sequence string representation of the text data, the sim-hash value of the text data is determined. And, according to the sim-hash value of the text data, the similarity value of the text data and the cluster center of the created topic class is determined. Simhash is a local sensitive hash. For local sensitivity, it means that if two strings have a certain similarity, and this similarity can still be maintained after hashing, it is called a local sensitive hash. Ordinary hash does not have this property. When determining the similarity value between text data and the cluster center of the created topic class, the Hamming distance between the sim-hash value of the text data and the sim-hash value of the cluster center of the created topic class can be calculated to determine the similarity value between the Hamming distance text data and the cluster center of the created topic class.
[0043] The similarity value includes a title similarity value and a content similarity value; the step of judging whether the similarity value meets a preset threshold value includes: if it is determined that the title similarity value does not meet a first preset threshold value, judging whether the content similarity value meets a second preset threshold value.
[0044] In the embodiment of the present invention, the order of text clustering is rearranged by using feature weights, for example, the text data at the front is sorted from high to low according to the feature weights, and the information contained in the text data is relatively rich. The text data containing rich information can be clustered to form topic classes first, and then the text clustering can be performed according to the clustered topic classes, which can improve the accuracy of clustering. In particular, the clustering processing for streaming data can greatly improve the accuracy of clustering. And, the cluster center of the topic class is set, and the cluster center represents the topic of the text data in this topic class. The new text data only needs to be compared with the cluster center in the topic class, which improves the speed of clustering calculation. In addition, when calculating the similarity of text data, the Hamming distance between the text data is determined based on the sim-hash value of the text data, so that the calculation complexity is greatly reduced. If the calculated Hamming distance is less than the preset threshold, it means that the similarity value between the texts meets the preset threshold, and the text can be clustered into a topic class.
[0045] Hot topic discovery refers to the technology of automatically discovering topics that have attracted great attention in a certain period of time through various information resources and integrating the content related to the topic together. Hot topic discovery relies on incremental text clustering technology. Information flows are clustered into limited topic clusters. The topics within the cluster are highly similar, and the similarity between different clusters is low, so as to integrate massive data. Among them, the identification of hot topics aims to obtain corresponding topics from semi-structured massive network data and aggregate them into new hot events for analysis. Hot topic analysis has great practical significance for public opinion analysis. It can not only provide netizens' focus in real time to content marketing platforms and assist in network public opinion analysis, but also help the public relations teams of major companies to handle crisis events in a timely manner.
[0046] Figure 2 FIG. 1 is a schematic diagram of a method for discovering hot topics according to an embodiment of the present invention. Figure 2 As shown, the method for discovering hot topics in an embodiment of the present invention includes:
[0047] Step 201: Get batch text data regularly based on keywords. In order to determine real-time topics, you can crawl data from web news, Weibo, WeChat public accounts, Zhihu, etc. regularly (for example, every 15 minutes) based on keywords to obtain batch text data to be clustered. By clustering the acquired text data, you can identify hot topics. The keyword can be one or more, and the keyword can be set according to business needs.
[0048] Step 202: Extracting feature words of text data. The purpose of extracting text feature words is to select words with relatively strong representative meanings from the text content or title. Under the premise of not affecting the theme of the text, the number of words to be processed should be relatively small, which can greatly reduce the complexity of calculation and improve the speed and efficiency of text processing. The less a word appears in a document, the lower the influence of the word on this document. Therefore, in the embodiment of the invention, feature words are extracted by word frequency method.
[0049] Specifically, the text data can be first segmented to obtain the initial feature words. Since the part of speech has a great influence on the effect of topic clustering, all words except nouns and verbs in the initial feature words are filtered out. Then, according to the preset feature evaluation index, the filtered initial feature words are scored, sorted according to the size of the score, and finally several items with the highest scores are selected as the feature words of the text (forming a feature word set).
[0050] Step 203: Calculate the weight of the feature word. In the embodiment of the present invention, the TF-IDF weight value of the feature word in each feature word set is determined by the term frequency-inverse document frequency (TF-IDF) algorithm. The weight value determined by this method is more accurate and has higher calculation efficiency. Among them, TF represents the number of times a feature word appears in the text d; IDF is the inverse document frequency (obtained by dividing the total number of files by the number of files containing the word, and then taking the logarithm of the quotient), which is inversely proportional to the number of times a word appears in the document set. The starting formula for calculating the TF-IDF weight value is:
[0051] w ij =tf ij ×idf ij
[0052] w ij is the weight of feature word i in text j, tf ij is the TF value of feature word i in text j, idf ij is the IDF value of feature word i in text j.
[0053] Step 204: Sort the batch text data according to the weights of the feature words. Since the single-pass algorithm is sensitive to the order of document clustering, the embodiment of the present invention adjusts the clustering order of the text data. Specifically, for each feature word in the feature word set of each text data, calculate its TF-IDF value. After normalization, the kth text data can be represented as D k =(w k1 d k1 , w k2 d k2 , ...w kn d kn ), k = 1, 2, ..., m, i = 1, 2, ..., n, m is the number of batch text data, n is the number of feature words of all text data, w kn is the weight of the normalized feature word, d ki For text data d k Whether the i-th word is contained in the : 1 if it contains the word, 0 if it does not. kn , w kn =
[0054]
[0055] Among them, weight(i, d k ) is represented as text data d k The TF-IDF weight value of the i-th feature word in .
[0056] The weight sum of the feature words of the kth text data is D ki =w ki d ki , indicating that the i-th feature word in the text data d k According to the weight of the feature word, the m text data are sorted. The text data at the front contains richer topic information and is preferentially clustered to form a topic cluster.
[0057] Step 205: Calculate the sim-hash value of the text data. Each feature word in the text data is compiled into a hash value through a hash algorithm, and each character string is converted into a string of numbers to achieve the purpose of dimensionality reduction. Then, the weights of the feature words obtained according to the above steps need to form a weighted digital string. For example, the hash value of the word "United States" is 100101, and its weight is 5, so its weighted calculation is "5 -5 -5 5 -5 5". In addition, each bit of the weighted digital string calculated by each feature word of the text data is accumulated, and the accumulated digital string is converted into a 0 1 string, that is, greater than 0 is recorded as 1, and less than 0 is recorded as 0, so as to obtain the sim-hash value of each text data, which can improve the accuracy of subsequent text data similarity value calculation.
[0058] Step 206: Calculate the Hamming distance between the text data and the cluster center of the created topic class. After obtaining the sim-hash value of each text data, the number of differences in the sim-hash values of the two text data is compared, which is called the Hamming distance. For example, if the third, fourth and fifth digits of '10101' and '10010' are different, the Hamming distance between the two is 3. When clustering batch text data, a threshold can be set in advance. When the Hamming distance is lower than this threshold (indicating that the similarity value meets the preset threshold), they are considered similar and classified as a topic cluster. For the determination of the distance center, the centroid method or the center method can be used. The center method refers to that the topic class is represented by a certain document as the topic center. This method requires a lot of time to search for a suitable center. For the centroid method, the average value of all document vectors in the class is used for similarity comparison, which can improve the efficiency of the single-pass clustering algorithm.
[0059] Step 207: Determine the topic category to which the text data belongs based on the calculated Hamming distance.
[0060] Figure 3 is a schematic diagram of clustering calculation of text data according to an embodiment of the present invention, such as Figure 3As shown, in the embodiment of the present invention, based on the sorting of batch texts, text data is obtained in sequence. After a text data is obtained, it is first determined whether it is the first text data. If it is, it means that there is no topic class created at this time, so the first topic class is established. According to the sorting result, the text data is obtained in sequence and clustering calculation is performed. It is determined that the obtained text data is not the first text data, indicating that there is a topic class that has been created. Then the title in the text data is first judged for similarity, for example, the first threshold T e1 Set to 3, if the Hamming distance D between the title of the text data and the title of the cluster center of the created topic class 标题 If the distance between titles is greater than the threshold 3, the Hamming distance between contents will be calculated. e2 Set to 10, if the Hamming distance D between the content of the text data and the content of the cluster center of the created topic class 内容 If the distance between the text data content and the content of the cluster center of all created topic classes is greater than this threshold of 10, a new topic cluster is created.
[0061] Step 208: Output all created topic classes.
[0062] Figure 4 is a schematic diagram of the main modules of the apparatus for text data clustering according to an embodiment of the present invention, such as Figure 4 As shown, the apparatus 400 for text data clustering according to the embodiment of the present invention includes a feature word determination module 401 , a weight determination module 402 , a sorting module 403 , and a clustering module 404 .
[0063] The feature word determination module 401 is used to obtain batch text data and determine a feature word set of each text data in the batch text data.
[0064] The weight determination module 402 is used to determine the weight of each feature word in each feature word set of each text data. The weight determination module is also used to determine the TF-IDF weight value of the feature word in each feature word set by using the word frequency-inverse document frequency algorithm.
[0065] The sorting module 403 is used to sort the batch text data according to the weights of the feature words in the feature word set. The sorting module is also used to determine the sum of the weights of all the feature words in each feature word set; and sort the batch text data according to the sum of the weights.
[0066] The clustering module 404 is used to perform clustering calculations on the batch text data based on the sorting results. The clustering module is also used to traverse the batch text data based on the sorting results, and determine the similarity values between the text data and the cluster centers of the created topic classes in turn; and determine whether the similarity values meet a preset threshold; if so, add the text data to the corresponding created topic class and update the cluster center of the created topic class; otherwise, create a new topic class.
[0067] The clustering module is also used to determine the hash value of the feature word of each text data through a hash algorithm; perform weighted processing on the hash value of each feature word according to the determined weight to obtain the weighted hash value of the feature word; accumulate the weighted hash value of the feature word of each text data to obtain a sequence string representation of the text data; determine the sim-hash value of the text data according to the sequence string representation of the text data; and determine the similarity value between the text data and the cluster center of the created topic class according to the sim-hash value of the text data. When determining the similarity value between the text data and the cluster center of the created topic class, the Hamming distance between the sim-hash value of the text data and the sim-hash value of the cluster center of the created topic class can be calculated to determine the similarity value between the Hamming distance text data and the cluster center of the created topic class.
[0068] The clustering module is also used for determining whether the similarity value meets the preset threshold, including: determining whether the content similarity value meets the second preset threshold when determining that the title similarity value does not meet the first preset threshold. The similarity value includes the title similarity value and the content similarity value.
[0069] In the embodiment of the present invention, the order of text clustering is rearranged by using feature weights, for example, the text data at the front is sorted from high to low according to the feature weights, and the information contained in the text data is relatively rich. The text data containing rich information can be clustered to form topic classes first, and then the text clustering can be performed according to the clustered topic classes, which can improve the accuracy of clustering. In particular, the clustering processing for streaming data can greatly improve the accuracy of clustering. And, the cluster center of the topic class is set, and the cluster center represents the topic of the text data in this topic class. The new text data only needs to be compared with the cluster center in the topic class, which improves the speed of clustering calculation. In addition, when calculating the similarity of text data, the Hamming distance between the text data is determined based on the sim-hash value of the text data, so that the calculation complexity is greatly reduced. If the calculated Hamming distance is less than the preset threshold, it means that the similarity value between the texts meets the preset threshold, and the text can be clustered into a topic class.
[0070] Figure 5An exemplary system architecture 500 is shown in which a method for text data clustering or an apparatus for text data clustering according to an embodiment of the present invention can be applied.
[0071] like Figure 5 As shown, system architecture 500 may include terminal devices 501, 502, 503, a network 504 and a server 505. Network 504 is used to provide a medium for communication links between terminal devices 501, 502, 503 and server 505. Network 504 may include various connection types, such as wired, wireless communication links or optical fiber cables, etc.
[0072] Users can use terminal devices 501, 502, 503 to interact with server 505 through network 504 to receive or send messages, etc. Various communication client applications can be installed on terminal devices 501, 502, 503, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social platform software, etc. (only as examples).
[0073] The terminal devices 501 , 502 , and 503 may be various electronic devices having a display screen and supporting web browsing, including but not limited to smart phones, tablet computers, laptop computers, desktop computers, and the like.
[0074] Server 505 may be a server that provides various services, such as a backend management server (only an example) that provides support for shopping websites browsed by users using terminal devices 501, 502, and 503. The backend management server may analyze and process the received data such as product information query requests, and feed back the processing results to the terminal device.
[0075] It should be noted that the method for text data clustering provided in the embodiment of the present invention is generally executed by the server 505 , and accordingly, the device for text data clustering is generally disposed in the server 505 .
[0076] It should be understood that Figure 5 The number of terminal devices, networks and servers in the embodiment is only for illustration. Any number of terminal devices, networks and servers may be provided according to implementation requirements.
[0077] Reference below Figure 6 , which shows a schematic diagram of the structure of a computer system 600 of a terminal device suitable for implementing an embodiment of the present invention. Figure 6 The terminal device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present invention.
[0078] like Figure 6As shown, the computer system 600 includes a central processing unit (CPU) 601, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 602 or a program loaded from a storage part 608 into a random access memory (RAM) 603. In the RAM 603, various programs and data required for the operation of the system 600 are also stored. The CPU 601, the ROM 602, and the RAM 603 are connected to each other via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.
[0079] The following components are connected to the I / O interface 605: an input section 606 including a keyboard, a mouse, etc.; an output section 607 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 608 including a hard disk, etc.; and a communication section 609 including a network interface card such as a LAN card, a modem, etc. The communication section 609 performs communication processing via a network such as the Internet. A drive 610 is also connected to the I / O interface 605 as needed. A removable medium 611, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 610 as needed, so that a computer program read therefrom is installed into the storage section 608 as needed.
[0080] In particular, according to the embodiments disclosed in the present invention, the process described above with reference to the flowchart can be implemented as a computer software program. For example, the embodiments disclosed in the present invention include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication part 609, and / or installed from the removable medium 611. When the computer program is executed by the central processing unit (CPU) 601, the above-mentioned functions defined in the system of the present invention are executed.
[0081] It should be noted that the computer-readable medium shown in the present invention may be a computer-readable signal medium or a computer-readable storage medium or any combination of the above two. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present invention, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in combination with an instruction execution system, device or device. In the present invention, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries a computer-readable program code. This propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. Computer-readable signal media may also be any computer-readable medium other than computer-readable storage media, which may send, propagate or transmit a program for use by or in conjunction with an instruction execution system, apparatus or device. The program code contained on the computer-readable medium may be transmitted using any appropriate medium, including but not limited to: wireless, wire, optical cable, RF, etc., or any suitable combination of the above.
[0082] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, function and operation of the system, method and computer program product according to various embodiments of the present invention. In this regard, each box in the flow chart or block diagram can represent a module, a program segment, or a part of a code, and the above-mentioned module, program segment, or a part of a code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order from the order marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or flow chart, and the combination of the boxes in the block diagram or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0083] The modules involved in the embodiments of the present invention may be implemented in software or hardware. The modules described may also be set in a processor, for example, may be described as: a processor includes a feature word determination module, a weight determination module, a sorting module, and a clustering module. The names of these modules do not constitute limitations on the modules themselves in certain circumstances, for example, the feature word determination module may also be described as a "module for obtaining batch text data, and determining a feature word set for each text data in the batch text data".
[0084] As another aspect, the present invention also provides a computer-readable medium, which may be included in the device described in the above embodiment; or may exist independently without being assembled into the device. The above computer-readable medium carries one or more programs, and when the above one or more programs are executed by a device, the device includes: obtaining batch text data, and determining a feature word set of each text data in the batch text data; for each feature word set of text data, determining the weight of each feature word in the feature word set; sorting the batch text data according to the weight of the feature words in the feature word set; and performing clustering calculation on the batch text data based on the sorting result.
[0085] In the embodiment of the present invention, the order of text clustering is rearranged by using feature weights, for example, the text data at the front contains more information according to the feature weights. The text data containing rich information can be clustered into topic groups first, and then the text clustering can be performed according to the clustered topic groups, which can improve the accuracy of clustering.
[0086] The above specific implementations do not constitute a limitation on the protection scope of the present invention. It should be understood by those skilled in the art that various modifications, combinations, sub-combinations and substitutions may occur depending on design requirements and other factors. Any modification, equivalent substitution and improvement made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A method for clustering text data, It is characterized in that include: Acquire batch text data, and determine a feature word set of each text data in the batch text data; For each feature word set of text data, determine the weight of each feature word in the feature word set; Determine the sum of weights of all feature words in each feature word set; sort the batch text data according to the sum of weights; Based on the sorting result, clustering calculation is performed on the batch text data; wherein, the weights and are sorted from high to low, and the text data with the highest weights and the highest weights contain relatively rich information. The text data are first clustered to form topic classes based on the weights and the highest weights, and then text clustering is performed based on the clustered topic classes.
2. The method according to claim 1, It is characterized in that Based on the sorting result, the step of performing clustering calculation on the batch text data includes: Based on the sorting result, traverse the batch text data, and determine the similarity values between the text data and the cluster centers of the created topic classes in turn; Determine whether the similarity value meets a preset threshold; if so, add the text data to the corresponding created topic class and update the cluster center of the created topic class; otherwise, create a new topic class.
3. The method according to claim 2, It is characterized in that The steps of determining the similarity value between each text data and the cluster center of the created topic class include: Determine the hash value of each feature word of text data through a hash algorithm; According to the determined weight of each feature word, the hash value thereof is weighted to obtain a weighted hash value of the feature word; Accumulate the weighted hash values of the feature words of each text data to obtain a sequence string representation of the text data; Determine the sim-hash value of the text data according to the sequence string representation of the text data; According to the sim-hash value of the text data, a similarity value between the text data and the cluster center of the created topic class is determined.
4. The method according to claim 2, It is characterized in that The similarity value includes a title similarity value and a content similarity value; The step of judging whether the similarity value meets a preset threshold comprises: when it is determined that the title similarity value does not meet a first preset threshold, judging whether the content similarity value meets a second preset threshold.
5. The method according to claim 1, It is characterized in that For each feature word set of text data, the step of determining the weight of each feature word in the feature word set includes: The TF-IDF weight value of the feature word in each feature word set is determined by the word frequency-inverse document frequency algorithm.
6. A device for clustering text data, It is characterized in that include: A feature word determination module, used to obtain batch text data and determine a feature word set for each text data in the batch text data; A weight determination module is used for a feature word set of each text data to determine the weight of each feature word in the feature word set; A sorting module, used to determine the sum of weights of all feature words in each feature word set; and sort the batch text data according to the sum of weights; A clustering module is used to perform clustering calculation on the batch text data based on the sorting results; wherein the weights and are sorted from high to low, and the text data with the highest weights and the highest weights contain relatively rich information. The text data are first clustered to form topic classes based on the weights and the highest weights, and then text clustering is performed based on the clustered topic classes.
7. The device according to claim 6, It is characterized in that The clustering module is also used to traverse the batch text data based on the sorting results, and determine the similarity values between the text data and the cluster centers of the created topic classes in turn; and determine whether the similarity values meet a preset threshold; if so, add the text data to the corresponding created topic class and update the cluster center of the created topic class; otherwise, create a new topic class.
8. The device according to claim 7, It is characterized in that The clustering module is also used to determine the hash value of the feature word of each text data through a hash algorithm; and perform weighted processing on the hash value of each feature word according to the determined weight to obtain a weighted hash value of the feature word; The weighted hash values of the feature words of each text data are accumulated to obtain a sequence string representation of the text data; the sim-hash value of the text data is determined according to the sequence string representation of the text data; and the similarity value between the text data and the cluster center of the created topic class is determined according to the sim-hash value of the text data.
9. The device according to claim 7, It is characterized in that The clustering module is further configured to: the step of determining whether the similarity value meets a preset threshold comprises: determining whether the content similarity value meets a second preset threshold when determining that the title similarity value does not meet a first preset threshold; The similarity values include a title similarity value and a content similarity value.
10. The device according to claim 6, It is characterized in that The weight determination module is also used to determine the TF-IDF weight value of the feature word in each feature word set by using a word frequency-inverse document frequency algorithm.
11. An electronic device, It is characterized in that include: one or more processors; a storage device for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 5.
12. A computer readable medium having a computer program stored thereon, It is characterized in that When the program is executed by a processor, the method according to any one of claims 1 to 5 is implemented.
Citation Information
Patent Citations
Microblog topic detection method and system based on incremental clustering algorithm
CN107291886A