Text processing method, device, computer equipment and computer-readable storage medium

By calculating the mixing distance and average similarity between the to be processed text and the pre-constructed text set, the problem of the single-channel clustering method dependence on input order is solved, and the accuracy of text clustering is improved.

CN113918719BActive Publication Date: 2025-05-30PING AN BANK CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111270485.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-10-29
Publication Date
2025-05-30
Estimated Expiration
2041-10-29

AI Technical Summary

Technical Problem

The single-channel clustering method in the prior art depends on the order of data input, resulting in low accuracy of text clustering.

Method used

By calculating the mixing distance between the Mahayana distance and Pearson distance between the to-be-processed text and each text in the pre-constructed first text set, the average similarity between each first text set and the to-be-processed text is further calculated and the to-be-processed text is categorized based on the average similarity.

Benefits of technology

This method can reduce the dependence on the order of input to be processed, improve the accuracy of text clustering, and provide more accurate similarity characterization by eliminating the impact of uneven distribution of components in each dimension of text on similarity calculation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113918719B_ABST
    Figure CN113918719B_ABST
Patent Text Reader

Abstract

This application is applicable to the field of data processing technology, and provides a text processing method, a text processing device, a computer device and a storage medium. The text processing method includes: obtaining a text to be processed; calculating a first mixed distance between each text in a pre-constructed first text set and the text to be processed respectively based on the Mahalanobis distance and the Pearson distance, wherein the number of the first text sets is greater than or equal to 2; calculating the average similarity between the text to be processed and each first text set respectively based on the first mixed distance; and classifying the text to be processed based on the average similarity. By the method of this application, the dependence on the input order of the text to be processed can be reduced, and the accuracy of text clustering can be improved. In addition, this application also relates to blockchain technology.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of data processing, and in particular, to a text processing method, a text processing device, a computer device, and a computer-readable storage medium. Background Art

[0002] Public opinion monitoring relies on search engine technology and text mining technology. It mainly realizes the need for supervision and management of relevant network public opinion through automatic collection and processing of web page content, sensitive word filtering, intelligent aggregation and classification, topic detection, special topic focus, and statistical analysis. Adding a public opinion monitoring function to the risk control system of users' financial product positions can effectively monitor the risk factors existing in the market, timely remind users and initiate early warnings, and avoid risks such as losses caused by untimely responses.

[0003] Currently, common public opinion clustering methods include K-means, single-channel classification algorithms, and distributed network classification methods. Among them, the single-channel clustering method is a classic method for classifying streaming data. For the data stream that arrives sequentially, this method processes one data each time according to the input order, and determines whether the data belongs to an existing class or creates a new data class based on the matching degree between the current data and the existing classes, realizing incremental and dynamic clustering of streaming data. It is suitable for mining streaming data and has high clustering efficiency. However, for the same clustering object, due to differences in the input order, different clustering results will appear, that is, this clustering method is dependent on the input order of data, resulting in a problem of low clustering accuracy when using this clustering method to cluster some data. Summary of the Invention

[0004] In view of this, the embodiments of this application provide a text processing method, a text processing device, a computer device, and a computer-readable storage medium, which can reduce the dependence on the input order of the text to be processed and improve the accuracy of text clustering.

[0005] The first aspect of the embodiments of this application provides a text processing method, including:

[0006] Obtain the text to be processed;

[0007] Based on the Mahalanobis distance and the Pearson distance, calculate the first mixed distance between each text in the pre-constructed first text set and the text to be processed respectively, where the number of the first text set is greater than or equal to 2;

[0008] Based on the first mixed distance, calculate the average similarity between the text to be processed and each first text set respectively;

[0009] Classify the text to be processed based on the average similarity.

[0010] The second aspect of the embodiments of the present application provides a text processing device, including:

[0011] A first acquisition module, configured to acquire a text to be processed;

[0012] A first calculation module, configured to calculate a first mixed distance between each text in a pre-constructed first text set and the text to be processed respectively based on the Mahalanobis distance and the Pearson distance, where the number of the first text set is greater than or equal to 2;

[0013] A second calculation module, configured to calculate an average similarity between the text to be processed and each of the first text sets respectively based on the first mixed distance;

[0014] A classification module, configured to classify the text to be processed based on the average similarity.

[0015] The third aspect of the embodiments of the present application provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the computer device. When the processor executes the computer program, each step of the text processing method provided in the first aspect is implemented.

[0016] The fourth aspect of the embodiments of the present application provides a computer-readable storage medium. The computer-readable storage medium stores a computer program. When the computer program is executed by a processor, each step of the text processing method provided in the first aspect is implemented.

[0017] In a fifth aspect, the present application provides a computer program product. The computer program product includes a computer program. When the computer program is executed by one or more processors, the steps of the text processing method provided in the first aspect are implemented.

[0018] Implementing a text processing method, a text processing device, a computer device, and a computer-readable storage medium provided by the embodiments of the present application has the following beneficial effects:

[0019] After calculating the mixed distance between each piece of text in the first text set and the text to be processed, calculate the average similarity between each first text set and the text to be processed based on the mixed distance, and finally classify the text to be processed based on the average similarity. This method characterizes the similarity between two texts through the mixed distance of the two texts, and can eliminate the influence of the uneven distribution of each dimensional component between the two texts on the similarity calculation. Specifically, the Mahalanobis distance can overcome the interference of the correlation between variables, and the Pearson distance is suitable for scenarios with a large amount of data. The mixed distance calculated based on these two distances can improve the accuracy of characterizing the similarity between two texts. By calculating the average similarity, the influence of the input order of the text to be processed on the text classification result can be further reduced, thereby improving the accuracy of text classification. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0021] Figure 1 is a flowchart of the implementation of a text processing method provided by an embodiment of the present application;

[0022] Figure 2 is a block diagram of the structure of a text processing device provided by an embodiment of the present application;

[0023] Figure 3 is a block diagram of the structure of a computer device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0024] In order to make the objectives, technical solutions, and advantages of the present application clearer, the following further describes the present application in detail with reference to the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0025] The text processing method involved in the embodiments of the present application can be executed by a computer device, such as a notebook computer, an ultra-mobile personal computer (UMPC), a netbook, or a personal digital assistant (PDA), or a server.

[0026] In order to illustrate the technical solutions proposed by the present application, the following will be described through specific embodiments.

[0027] Please refer to Figure 1 , Figure 1 which shows the implementation flowchart of a text processing method provided by an embodiment of the present application. The text processing method includes:

[0028] Step 101, obtain the text to be processed.

[0029] The text to be processed refers to the text data to be classified. The text to be processed can be text data crawled from the Internet through relevant keywords, or text data obtained after recognizing relevant voice data based on a speech recognition model. Among them, for the acquisition method of crawling through keywords on the Internet, the selected keywords can be hot words, entity names, or a combination of both. For example, for the public opinion that Company A acquired Company B, which is relatively well-known in the industry, acquisition belongs to the hot word in this public opinion, and the names of the two companies are entity names. When determining the keywords, the acquisition, Company A, or Company B can be used alone as keywords, or the three words can be combined as keywords. After obtaining the keywords, relevant text data can be crawled based on the keywords as the text to be processed to monitor the public opinion.

[0030] Step 102, calculate the first mixed distance between each text in the pre-constructed first text set and the text to be processed based on the Mahalanobis distance and the Pearson distance respectively.

[0031] In order to reduce the influence of the input order of the text to be processed on the text classification result, the Mahalanobis distance and the Pearson distance can be used to calculate the mixed distance between each text in the first text and the text to be processed respectively. This mixed distance can be denoted as the first mixed distance. It can be considered that in the embodiment of the present application, the first mixed distance is actually used to represent the similarity between the text in the first text set and the text to be processed. This method of representing similarity can eliminate the influence of the uneven distribution of each dimensional component of the two texts to be calculated on the similarity calculation. Specifically, the Mahalanobis distance can overcome the interference of the correlation between variables, and the Pearson distance has the characteristic of being applicable to scenarios with a large amount of data. The first mixed distance obtained by combining these two distance calculation methods can more accurately represent the similarity between the text to be processed and each text in the first text set, thereby reducing the dependence of the classification result on the input order of the text to be processed.

[0032] Among them, the first text set is pre-constructed, and its quantity is greater than or equal to 2. For each first text set, the texts it contains all point to the same public opinion or topic. That is to say, one first text set corresponds to one public opinion or topic. Specifically, to construct the first text set, the similarity between every two texts among the existing multiple texts can be calculated, and the existing multiple texts can be classified based on the calculated similarity to obtain the first text set. For example, if there are 8 existing texts, these 8 texts can be combined in pairs to obtain 28 combination methods, that is, 28 similarities can be calculated. Then, based on these 28 similarities, these 8 texts can be classified to obtain the first text set. If the existing multiple texts correspond to n topics or public opinions, then n first text sets can be obtained after classification. For each first text set, the first mixed distance between each text in the first text set and the text to be processed can be calculated respectively. If there are n first text sets as described above, then n groups of first mixed distances can be obtained.

[0033] For the convenience of understanding, an example is given: Suppose there are three first text sets, namely a, b, and c. Among them, there are 46 texts in a, 25 texts in b, and 88 texts in c. When calculating the first mixed distance, for a, the first mixed distance between the text to be processed and each of the 46 texts in a can be calculated respectively to obtain a set of data including 46 first mixed distances; for b, the first mixed distance between the text to be processed and each of the 25 texts in b can be calculated respectively to obtain a set of data including 25 first mixed distances; for c, and so on. That is, after calculation, the first mixed distances corresponding to the number of first text sets can be obtained, and the number of first mixed distances included in each group of data is the same as the number of texts included in the corresponding first text set.

[0034] Step 103: Calculate the average similarity between the text to be processed and each first text set based on the first mixed distance.

[0035] After obtaining the first mixing distance, calculate the average similarity between each first text set and the text to be processed based on this first mixing distance. That is, for each first text set, after calculating the first mixing distance between the text to be processed and each text in this first text set, directly sum all the obtained first mixing distances, or perform a weighted sum operation on all the first mixing distances according to the labels of each text, and then take the average value after obtaining the sum as the average similarity between the text to be processed and this first text set. For example, if there are a total of 5 texts in the first text set d, the first mixing distance between each text and the text to be processed can be calculated respectively, and 5 first mixing distances are obtained. It can be understood that these 5 first mixing distances are the similarities between the text to be processed and the 5 texts in d respectively. By adding these 5 first mixing distances and then taking the average value, the average similarity between the text to be processed and d can be calculated. The calculation of the average similarity can further reduce the influence of the input order of the text to be processed on the accuracy of text classification.

[0036] Step 104: Classify the text to be processed based on the average similarity.

[0037] After obtaining the average similarity, the first text set to which the text to be processed belongs can be found among at least two first text sets based on the average similarity, so as to realize the classification of the text to be processed, further reduce the influence of the input order of the text to be processed on the accuracy of text classification, and improve the accuracy of text classification.

[0038] The reason for saying that using the average similarity can further reduce the influence of the input order of the text to be processed on the text classification accuracy is that when classifying the text to be processed using the average similarity, what is considered is the similarity between the text to be processed and all texts in each first text set. Therefore, it can reduce the influence of the input order of the text to be processed on the text classification result. For ease of understanding, an example is given: Suppose there are two first text sets, namely a and b. Among them, a contains texts a1, a2, and a3, and b contains texts b1 and b2. Assume a text e to be processed, and calculate the first mixed distance between each text in the two first text sets and e in turn. After calculation, it can be known that the first mixed distance between a1 and e is 2, the first mixed distance between a2 and e is 1, the first mixed distance between a3 and e is 3, the first mixed distance between b1 and e is 2.5, and the first mixed distance between b2 and e is 2.5. If the existing single-channel algorithm is used, the maximum value of the first mixed distance will be determined first, and then e will be classified based on this maximum value. That is to say, since the first mixed distance between a3 and e is the largest, e will be classified into the first text set a. However, if e is input before a3, then e will be classified into the first text set b. The reason for the above two results is the input order of a3 and e. That is to say, if the input order of some texts is different, different classification results will be obtained. In order to avoid this situation and reduce the text classification accuracy, the method of this application can be adopted, that is, using the average similarity between e and the two text sets as the classification criterion for e. After calculation, it can be known that when e is input before a3, the average similarity between e and a is 1.5, and the average similarity between e and b is 2.5. According to the result of the average similarity, e can be classified into the first text set b; when e is input before a3, the average similarity between e and a is 2, and the average similarity between e and b is 2.5. According to the calculation result of the average similarity, e is still classified into the first text set b. Obviously, using the method of this application, regardless of the input order of e and a3, the classification result of e is the same. That is to say, the method of this application can be not affected by the order between a3 and e, effectively reducing the dependence of the classification method on the input order of the text to be processed, thereby improving the accuracy of text classification.

[0039] After calculating the hybrid distance between each piece of text in the first text set and the text to be processed, the average similarity between each first text set and the text to be processed is calculated based on the hybrid distance, and finally the text to be processed is classified based on the average similarity. In this method, by calculating the hybrid distance between two pieces of text and using this hybrid distance to represent the similarity between the two pieces of text, the influence caused by the uneven distribution of each dimensional component between the two pieces of text on the similarity calculation can be eliminated. Specifically, the Mahalanobis distance can overcome the interference of the correlation between variables, and the Pearson distance has the characteristic of being applicable to scenarios with a large amount of data. The hybrid distance calculated based on these two distances can improve the accuracy of representing the similarity between two pieces of text. By calculating the average similarity, the influence of the input order of the text to be processed on the text classification result can be further reduced, thereby improving the accuracy of text classification.

[0040] In some embodiments, step 102 above specifically includes:

[0041] 1021. Input the first text set into a pre-constructed word vector model to obtain a first feature vector corresponding to each piece of text.

[0042] 1022. Input the text to be processed into the word vector model to obtain a second feature vector of the text to be processed.

[0043] 1023. Calculate the first hybrid distance between each first feature vector and the second feature vector based on the Mahalanobis distance and the Pearson distance respectively.

[0044] To calculate the hybrid distance between two pieces of text, the text can be first represented in the form of a vector, that is, a vectorization operation is performed on the text. After the text is vectorized, it can be mapped into a digital representation and become a format that can be processed by a computer, thereby improving the rate of calculating the similarity between different texts. In the embodiments of the present application, after the word vector model is constructed, each piece of text in the first text set can be vectorized based on the word vector model to obtain a feature vector corresponding to each piece of text, and then the text to be processed can be vectorized based on the word vector model to obtain a feature vector of the text to be processed. To facilitate the distinction between these two feature vectors, the feature vector corresponding to each piece of text in the first text set can be denoted as the first feature vector, and the feature vector corresponding to the text to be processed can be denoted as the second feature vector. After obtaining the two types of feature vectors, the first hybrid distance between each first feature vector and the second feature vector can be calculated based on the Mahalanobis distance and the Pearson distance respectively.

[0045] In some embodiments, to improve the accuracy of classifying the text to be processed, for the word vector model, the Vector Space Model (VSM), word2vec / doc2vec distributed representation, or BBERT / ELMo / GPT deep pre-training model can be selected. Among them, the BERT deep pre-training model is preferably used for the word vector model. The BERT deep pre-training model was proposed in 2018 and is a milestone in the field of pre-training. It has reached the highest level in many Natural Language Processing (NLP) tasks and has also opened a new chapter for pre-training models. Literally, the BERT deep pre-training model is a pre-training model of deep bidirectional transformers for language understanding. The reason for choosing this model is as follows: First, it can learn the semantic information of the text, and tasks such as classification and semantic similarity calculation can be achieved through the vector-form output. Second, it is a pre-trained language model, that is to say, it has already trained the parameters on a large-scale corpus. When using it, only the parameters need to be trained and updated on this basis to achieve unsupervised learning, and the information related to the left and right of the text will be trained in each layer of the model. Therefore, the training efficiency of this model is higher.

[0046] In some embodiments, step 104 above specifically includes:

[0047] 1041. Determine whether the maximum value in the average similarity is greater than a preset similarity threshold.

[0048] 1042. If the maximum value is greater than the similarity threshold, classify the text to be processed into the target text set, where the target text set is the first text set corresponding to the maximum value.

[0049] After calculating the average similarity between the text to be processed and each first text set, the first text sets can be sorted based on the average similarity. For example, the first text sets can be sorted in descending or ascending order of the average similarity. In the case of sorting in descending order, the first text set corresponding to the largest average similarity value is ranked first, and the remaining first text sets are sorted in turn according to the corresponding average similarities. Among them, the first text set ranked first is the first text set most likely to receive the text to be processed. For the convenience of description, this first text set can be denoted as the target text set. To determine whether the text to be processed can be classified into this target text set, the average similarity corresponding to this target text set, that is, the maximum value of the average similarity, can be compared with a preset similarity threshold. If the maximum value is greater than the similarity threshold, it means that the text to be processed and the texts in this first text set belong to the same topic, and the text to be processed can be classified into this first text set.

[0050] In some embodiments, in order to improve the comprehensiveness of text classification, after the above step 1041, the following is further included: if the maximum value of the obtained average similarity is less than or equal to the similarity threshold, the text to be processed is classified into a pre-constructed second text set.

[0051] After determining whether the maximum value in the average similarity is greater than the preset similarity threshold, if the conclusion is that the maximum value is less than or equal to the similarity threshold, it indicates that the text to be processed and the texts in the target text set do not belong to the same topic. The text to be processed at this time may be the beginning of a new event or some news fragments that cannot form public opinion. For such text data, in order to improve the comprehensiveness of text classification, it can be classified into the second text set, where the second text set records the text set of the text to be processed that has not been classified into the first text set. That is, all texts to be processed that cannot be directly classified into the first text set will be classified into this second text set. For the construction of the first text set in the above step 103, the method of this step can also be used for construction to obtain at least two first text sets.

[0052] In some embodiments, in order to improve the timeliness of public opinion monitoring, the hybrid distance between the text to be processed and each text in the second text set can be calculated respectively, denoted as the third hybrid distance, and the maximum calculated third hybrid distance is compared with the newly set similarity threshold. The newly set similarity threshold is mainly used to judge whether a certain public opinion or topic can be formed. If the maximum third hybrid distance is greater than the newly set similarity threshold, a new first text set can be constructed based on the text corresponding to the maximum third hybrid distance and the text to be processed, that is, a new public opinion is formed.

[0053] In some embodiments, in order to improve the accuracy of classifying the text to be processed, the similarity threshold can be determined through the following steps:

[0054] A1. Obtain the popularity value of each text in the target text set.

[0055] A2. Determine the core text as the text with the highest popularity value.

[0056] A3. Calculate the second hybrid distance between each text in the target text set except the core text and the core text based on the Mahalanobis distance and the Pearson distance.

[0057] A4. Determine the similarity threshold based on the second hybrid distance.

[0058] After determining the target text, the popularity value of each piece of text in the target text can be obtained. Based on this popularity value, a more reasonable similarity threshold can be determined. Whether a piece of text is popular depends on the content quality of this piece of text. That is, the numerical expression of the content quality of the text can be approximately regarded as the popularity value of this text. Specifically, from which aspects to measure the popularity value of the text can be defined according to actual needs. For example, for communication texts with high timeliness, the popularity value of such texts can be determined from two aspects: the number of likes and the number of forwards. For texts with low timeliness requirements and relatively persistent public opinions or topics, in addition to determining the popularity value of such texts from the two aspects of the number of likes and the number of forwards, the popularity value of such texts can also be determined by combining the influence factor of the corresponding author of the text and the number of comments.

[0059] After obtaining the popularity value, the core text can be determined from the target text set based on the popularity value, that is, the text with the highest popularity value is determined. After obtaining the core text, the second mixed distance between the core text and each piece of text other than the core text in the target text set can be calculated using the Mahalanobis distance and the Pearson distance. Finally, the similarity threshold of the target text set can be determined according to the second mixed distance.

[0060] For ease of understanding, an example is given: Suppose the target text set includes 6 pieces of text, and the popularity values of these 6 pieces of text are obtained, which are 67, 53, 50, 77, 85, and 90 respectively. Based on the popularity value, the sixth piece of text can be determined as the core text. After determining the core text, the second mixed distance between the sixth piece of text and each of the first five pieces of text is calculated, and 5 second mixed distances are obtained. After obtaining these 5 second mixed distances, the similarity threshold of the target text can be further determined. Optionally, the maximum and minimum second mixed distances can be removed from the 5 second mixed distances, and then the average value of the remaining second mixed distances is calculated, and this average value is determined as the similarity threshold of the target text set. Or, each second mixed distance is compared with the average value of the 5 second mixed distances. If there is a second mixed distance whose difference from the average value exceeds a reasonable range, it is deleted. Finally, the average value of the remaining second mixed distances is calculated, and this average value is determined as the similarity threshold of the target text set.

[0061] In some embodiments, since each first text set may be determined as the target text set, the similarity threshold corresponding to each first text set can be pre-calculated using steps A1 to A4. When a certain first text set is determined as the similarity threshold, the corresponding similarity threshold of the first text and the first text can be directly called for text classification, thereby improving the efficiency of text classification.

[0062] In some embodiments, to improve the accuracy of determining the similarity threshold, the similarity threshold for each first text set can be dynamically updated. For example, after receiving new text, the similarity threshold can be updated based on the current texts in the first text set, or the texts in each first text set can be regularly checked. If there is an addition or deletion of a text in a certain first text set, the similarity threshold for that first text set can also be updated; or, when the number of added or deleted texts in a certain first text set reaches a set number of texts, the similarity threshold update operation is triggered.

[0063] In some embodiments, to improve the timeliness of the texts in the first text set, the above text processing method may further include:

[0064] For each first text set:

[0065] B1. Within a preset time period, count the click-through rate of each text in the first text set.

[0066] B2. If there is a target text, delete the target text from the first text set, where the target text is the text with a click-through rate less than the preset click-through rate threshold.

[0067] To ensure the timeliness of the texts in the first text set, for each first text set, the click-through rate of each text in the first text set by the user within a preset time period can be monitored, and a click-through rate threshold can be set. The text with a click-through rate lower than the click-through rate threshold is used as the target text, and the target text is deleted to improve the timeliness of the texts in the first text set. Optionally, texts that are relatively old in the first text set can also be regularly deleted to maintain the timeliness of the texts in the first text set.

[0068] In some embodiments, the calculation of the hybrid distance can refer to the following formula:

[0069]

[0070] where a and b represent two texts for which the hybrid distance is to be calculated. For example, a can represent text a in the first text set, and b can represent the text to be processed. I ab The hybrid distance between text a and text b is the Mahalanobis distance between text a and text b is the Pearson distance between text a and text b. That is, the hybrid distance between two texts is the sum of the Mahalanobis distance and the Pearson distance between these two texts. Optionally, corresponding weights can be assigned to the Mahalanobis distance and the Pearson distance respectively to further improve the accuracy of the hybrid distance calculation. Specifically, S is the variance matrix;

[0071] In some embodiments, after calculating the first hybrid distance between each text in a pre-constructed first text set and the text to be processed based on the Mahalanobis distance and the Pearson distance respectively, the above text processing method further includes:

[0072] Uploading the text to be processed, the first hybrid distance, and / or the first text set to a blockchain.

[0073] Among them, to ensure the security of data and fairness and transparency to users, the text to be processed, the first hybrid distance, and / or the first text set can be uploaded to the blockchain for evidence preservation. Subsequently, users can download the text to be processed, the first hybrid distance, and / or the first text set from the blockchain through their respective devices to verify whether these data have been tampered with. The blockchain referred to in this embodiment is a new application mode that uses computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanisms, and encryption algorithms. Essentially, a blockchain is a decentralized database, a string of data blocks generated by using cryptographic methods. Each data block contains information about a batch of network transactions, which is used to verify the validity of the information (anti-counterfeiting) and generate the next block. A blockchain can include a blockchain underlying platform, a platform product service layer, and an application service layer, etc.

[0074] In addition, an embodiment of the present application also provides a text processing device.

[0075] Please refer to Figure 2 , Figure 2 which is a structural block diagram of a text processing device provided by an embodiment of the present application. In this embodiment, each unit included in the computer device is used to execute Figure 2 the corresponding steps in the corresponding embodiment. Specifically, please refer to Figure 2 the relevant descriptions in the corresponding embodiment. For the sake of convenience of description, only the parts related to this embodiment are shown. Refer to Figure 2 , the text processing device 20 includes:

[0076] A first acquisition module 21, configured to acquire the text to be processed;

[0077] A first calculation module 22, configured to calculate the first hybrid distance between each text in a pre-constructed first text set and the text to be processed based on the Mahalanobis distance and the Pearson distance respectively, where the number of the first text sets is greater than or equal to 2;

[0078] A second calculation module 23, configured to calculate the average similarity between the text to be processed and each first text set based on the first hybrid distance;

[0079] A classification module 24, configured to classify the text to be processed based on the average similarity.

[0080] As an embodiment of the present application, the above-mentioned first calculation module 22 may include:

[0081] A first input unit, configured to input a first text set into a pre-constructed word vector model to obtain a first feature vector corresponding to each text;

[0082] A second input unit, configured to input a text to be processed into the word vector model to obtain a second feature vector of the text to be processed;

[0083] A first calculation unit, configured to calculate a first mixed distance between each first feature vector and the second feature vector respectively based on the Mahalanobis distance and the Pearson distance.

[0084] As an embodiment of the present application, the above-mentioned classification module 24 may include:

[0085] A judgment unit, configured to judge whether the maximum value in the average similarity is greater than a preset similarity threshold;

[0086] A first classification unit, configured to classify the text to be processed into a target text set if the maximum value is greater than the similarity threshold, where the target text set is the first text set corresponding to the maximum value.

[0087] As an embodiment of the present application, the above-mentioned classification module 24 may also be used for:

[0088] A second classification unit, configured to, after judging whether the maximum value in the average similarity is greater than a preset similarity threshold, if the maximum value is less than or equal to the similarity threshold, classify the text to be processed into a pre-constructed second text set, and the second text set records the text to be processed that has not been classified into the first text set.

[0089] As an embodiment of the present application, the above-mentioned text processing device 20 may also include:

[0090] A second acquisition module, configured to acquire the popularity value of each text in the target text set;

[0091] A first determination module, configured to determine the text with the highest popularity value as the core text;

[0092] A third calculation module, configured to calculate a second mixed distance between each text in the target text set except the core text and the core text based on the Mahalanobis distance and the Pearson distance;

[0093] A second determination module, configured to determine the similarity threshold based on the second mixed distance.

[0094] As an embodiment of the present application, the above-mentioned text processing device 20 may also include:

[0095] After classifying the text to be processed based on the average similarity, for each first text set:

[0096] Within a preset time period, count the click-through rate of each text in the first text set;

[0097] If there is a target text, delete the target text from the first text set, where the target text is a text with a click-through rate less than the preset click-through rate threshold.

[0098] As an embodiment of the present application, the above text processing device 20 may further include:

[0099] A data upload module, configured to upload the text to be processed, the first mixed distance, and / or the first text set to the blockchain after calculating the first mixed distance between each text in the pre-constructed first text set and the text to be processed based on the Mahalanobis distance and the Pearson distance respectively.

[0100] It should be understood that Figure 2 In the structural block diagram of the text processing device shown, each unit is used to execute Figure 1 the corresponding steps in the embodiments, and for Figure 1 the corresponding steps in the embodiments have been explained in detail in the above embodiments. For details, please refer to Figure 1 as well as Figure 1 the relevant descriptions in the corresponding embodiments, which will not be elaborated here.

[0101] Figure 3 This is the structural block diagram of a computer device provided by another embodiment of the present application. As Figure 3 shown, the computer device 30 of this embodiment includes: a processor 31, a memory 32, and a computer program 33 stored in the above memory 32 and executable on the above processor 31, such as a program for the text processing method. When the processor 31 executes the above computer program 33, it implements the steps in each embodiment of the above text processing method, such as Figure 1 shown as 101 to 104. Alternatively, when the processor 31 executes the computer program 33, it implements the functions of each unit in the above Figure 2 corresponding embodiment, for example, Figure 2 the functions of the units 21 to 24 shown. For details, please refer to Figure 2 the relevant descriptions in the corresponding embodiments, which will not be elaborated here.

[0102] Exemplarily, the computer program 33 can be divided into one or more units, which are stored in the memory 32 and executed by the processor 31 to complete this application. The one or more units can be a series of computer program instruction segments capable of performing specific functions, and these instruction segments are used to describe the execution process of the computer program 33 in the terminal 30. For example, the computer program 33 can be divided into a data acquisition module and a prediction module, and the specific functions of each module are as described above.

[0103] The turntable device may include, but is not limited to, a processor 31 and a memory 32. Those skilled in the art can understand that Figure 3 merely examples of the computer device 30, which do not constitute a limitation on the computer device 30, may include more or fewer components than shown in the figure, or combine certain components, or different components. For example, the turntable device may further include an input / output device, a network access device, a bus, etc.

[0104] The so-called processor 31 may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0105] The memory 32 may be an internal storage unit of the computer device 30, such as the hard disk or memory of the computer device 30. The memory 32 may also be an external storage device of the computer device 30, such as a plug-in hard disk equipped on the computer device 30, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. Further, the memory 32 may also include both the internal storage unit and the external storage device of the computer device 30. The memory 32 is used to store the computer program and other programs and data required by the turntable device. The memory 32 may also be used to temporarily store data that has been output or will be output.

[0106] It can be understood that the method provided by the embodiments of the present application can be applied not only to computer devices, but also to servers, such as cloud servers.

[0107] The embodiments of the present application also provide a computer-readable storage medium storing a computer program, which when executed by a processor can implement the steps in the above-mentioned various method embodiments. Exemplarily, the following steps can be implemented:

[0108] Obtain the text to be processed;

[0109] Based on the Mahalanobis distance and the Pearson distance, calculate the first mixed distance between each text in the pre-constructed first text set and the text to be processed respectively.

[0110] Based on the first mixed distance, calculate the average similarity between the text to be processed and each first text set respectively.

[0111] Classify the text to be processed based on the average similarity.

[0112] The embodiments of the present application provide a computer program product, which when running on a mobile terminal enables the mobile terminal to implement the steps in the above-mentioned various method embodiments when executed. Exemplarily, the following steps can be implemented:

[0113] Obtain the text to be processed;

[0114] Based on the Mahalanobis distance and the Pearson distance, calculate the first mixed distance between each text in the pre-constructed first text set and the text to be processed respectively.

[0115] Based on the first mixed distance, calculate the average similarity between the text to be processed and each first text set respectively.

[0116] Classify the text to be processed based on the average similarity.

[0117] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be included in the protection scope of the present application.

Claims

1. A text processing method, characterized in that, the text processing method includes: obtaining the text to be processed; calculating a first mixed distance between each text in a pre-constructed first text set and the text to be processed respectively based on the Mahalanobis distance and the Pearson distance, wherein the number of the first text set is greater than or equal to 2; the mixed distance is the sum of the Mahalanobis distance and the Pearson distance between two texts; calculating the average similarity between the text to be processed and each of the first text sets respectively based on the first mixed distance; classifying the text to be processed based on the average similarity; the calculating a first mixed distance between each text in a pre-constructed first text set and the text to be processed respectively based on the Mahalanobis distance and the Pearson distance includes: inputting the first text set into a pre-constructed word vector model to obtain a first feature vector corresponding to each text; inputting the text to be processed into the word vector model to obtain a second feature vector of the text to be processed; calculating a first mixed distance between each of the first feature vectors and the second feature vector respectively based on the Mahalanobis distance and the Pearson distance.

2. The text processing method according to claim 1, characterized in that, the classifying the text to be processed based on the average similarity includes: judging whether the maximum value in the average similarity is greater than a preset similarity threshold; if the maximum value is greater than the similarity threshold, classifying the text to be processed into a target text set, wherein the target text set is the first text set corresponding to the maximum value.

3. The text processing method according to claim 2, characterized in that, after the judging whether the maximum value in the average similarity is greater than a preset similarity threshold, the text processing method further includes: if the maximum value is less than or equal to the similarity threshold, classifying the text to be processed into a pre-constructed second text set, and the second text set records the texts to be processed that are not classified into the first text set.

4. The text processing method according to claim 2 or 3, characterized in that, the similarity threshold is determined by the following steps: obtaining the popularity value of each text in the target text set; determining the text with the highest popularity value as the core text; calculating a second mixed distance between each text in the target text set except the core text and the core text based on the Mahalanobis distance and the Pearson distance; determining the similarity threshold based on the second mixed distance.

5. The text processing method according to claim 1, characterized in that, after the classifying the text to be processed based on the average similarity, the text processing method further includes: for each first text set: counting the click-through rate of each text in the first text set within a preset time period; if there is a target text, deleting the target text from the first text set, and the target text is the text with a click-through rate less than a preset click-through rate threshold.

6. The text processing method according to claim 1, characterized in that, After calculating the first hybrid distance between each text in a pre-constructed first text set and the text to be processed based on the Mahalanobis distance and the Pearson distance respectively, the text processing method further includes: Uploading the text to be processed, the first hybrid distance, and / or the first text set to a blockchain.

7. A text processing device Characterized in that The text processing device includes: A first acquisition module for acquiring a text to be processed; A first calculation module for calculating the first hybrid distance between each text in a pre-constructed first text set and the text to be processed based on the Mahalanobis distance and the Pearson distance respectively, wherein the number of the first text sets is greater than or equal to 2; the hybrid distance is the sum of the Mahalanobis distance and the Pearson distance between two texts; A second calculation module for calculating the average similarity between the text to be processed and each of the first text sets based on the first hybrid distance; A classification module for classifying the text to be processed based on the average similarity; The first calculation module includes: A first input unit for inputting the first text set into a pre-constructed word vector model to obtain a first feature vector corresponding to each text; A second input unit for inputting the text to be processed into the word vector model to obtain a second feature vector of the text to be processed; A first calculation unit for calculating the first hybrid distance between each of the first feature vectors and the second feature vector based on the Mahalanobis distance and the Pearson distance respectively.

8. A computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, Characterized in that When the processor executes the computer program, the steps of the method according to any one of claims 1 to 6 are implemented.

9. A computer-readable storage medium storing a computer program, Characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.

Citation Information

Patent Citations

  • System and method for online to offline services

    CN111460248A

  • Text processing method and device, terminal equipment and storage medium

    CN112148843A