Text Classification Method, Apparatus, Storage Medium and Electronic Device

By forming a sample set and clustering processing, combining a classification model of multiple classification parameters, the problems of low accuracy and high computing resource consumption in text classification are solved, and more efficient text classification and business processing are achieved.

CN114020916BActive Publication Date: 2025-06-27泰康保险集团股份有限公司
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202111301572.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-04
Publication Date
2025-06-27
Estimated Expiration
2041-11-04

AI Technical Summary

Technical Problem

The prior art has problems such as low accuracy and high computing resource consumption in text classification, especially in large-scale text information annotation and automatic classification applications.

Method used

By forming the first sample set and the second sample set, the text is clustered using the clustering center, and a pre-selected classification model is used to classify the text to be classified in combination with multiple classification parameters, and the target subcategory of the text to be classified is finally determined.

Benefits of technology

It improves the accuracy of text classification, reduces the resource consumption of classification calculations, and improves the business processing efficiency based on text classification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114020916B_ABST
    Figure CN114020916B_ABST
Patent Text Reader

Abstract

The present disclosure relates to the field of computer technologies, and relates to a text classification method, an apparatus, a storage medium, and an electronic device. The text classification method includes: forming a first sample set according to a text to be classified and a text knowledge base, where the classification knowledge base includes texts with a target number of categories; determining clustering centers according to the target number and corresponding categories, and performing clustering processing on the first sample set based on the clustering centers to obtain a rough classification to which the sample to be classified belongs; performing fusion processing on the text to be classified and the texts in the rough classification in the text knowledge base to form a second sample set; obtaining a plurality of classification parameters, respectively classifying the second sample set by using a preselected classification model based on each classification parameter, and determining a target subcategory to which the sample to be classified belongs according to the obtained plurality of classification results. The present disclosure can adjust classification parameters according to the text categories and quantities in the existing text knowledge base, reduce the consumption of classification computing resources by narrowing the classification scope, and has high classification efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technologies, and more particularly, to a text classification method, a text classification device, a computer storage medium, and an electronic device. Background Art

[0002] Text classification refers to the process of classifying different texts into relevant categories according to predefined topic categories based on the text content. Classifying different texts to be classified not only facilitates browsing but also enables quick query of the required text by category, thereby improving text processing efficiency.

[0003] With the development of the field of computer technologies, text classification has changed from being completely dependent on manual classification by professionals in the past to automatic text classification implemented by machines. However, in actual production applications, the annotation of many text information often still relies on manual processing, and the quantity may reach the level of hundreds, thousands, or even millions as the business grows. In some application scenarios, automatic classification consumes a large amount of computing resources, and the text classification results are also limited by the classification samples, resulting in low classification accuracy.

[0004] It should be noted that the information disclosed in the above background art is only used to enhance the understanding of the background of the present disclosure, and thus may include information that does not constitute the prior art known to those of ordinary skill in the art. Summary of the Invention

[0005] The purpose of the present disclosure is to provide a text classification method and device, a computer storage medium, and an electronic device, thereby at least to a certain extent improving the accuracy of text classification and reducing the resource consumption of classification calculation.

[0006] Other features and advantages of the present disclosure will become apparent through the following detailed description, or will be partially learned through the practice of the present disclosure.

[0007] According to one aspect of the present disclosure, there is provided a text classification method, including: forming a first sample set according to the text to be classified and a text knowledge base, where the classification knowledge base includes texts with a target number of categories; determining cluster centers according to the target number and corresponding categories, and performing clustering processing on the first sample set based on the cluster centers to obtain a rough classification to which the sample to be classified belongs, where the rough classification includes a plurality of sub-classifications; performing fusion processing on the text to be classified and the texts in the rough classification in the text knowledge base to form a second sample set; obtaining a plurality of classification parameters, respectively classifying the second sample set by using a preselected classification model based on each classification parameter, and determining a target sub-category to which the sample to be classified belongs according to the obtained plurality of classification results.

[0008] In an exemplary embodiment of the present disclosure, determining the cluster centers according to the target quantity and the corresponding categories, and performing clustering processing on the first sample set based on the cluster centers to obtain the rough classification to which the sample to be classified belongs includes: determining the number of the cluster centers as the target quantity; randomly selecting samples with the target quantity from the first sample set as the first cluster centers, and performing clustering processing on the first sample set based on the first cluster centers to obtain the rough classification to which the sample to be classified belongs.

[0009] In an exemplary embodiment of the present disclosure, randomly selecting samples with the target quantity from the first sample set as the first cluster centers, and performing clustering processing on the first sample set based on the first cluster centers to obtain the rough classification to which the sample to be classified belongs includes: randomly selecting one sample from each category of texts in the text knowledge base included in the first sample set as the first cluster centers; based on the first cluster centers, using the K-means clustering algorithm to perform clustering processing on the first sample set to obtain the rough classification to which the sample to be classified belongs.

[0010] In an exemplary embodiment of the present disclosure, performing clustering processing on the first sample set based on the first cluster centers to obtain the rough classification to which the sample to be classified belongs includes: obtaining the distances between the samples in the first sample set and each of the first cluster centers, and allocating the samples in the first sample set to the categories corresponding to each of the first cluster centers according to the distances to obtain candidate sample sets corresponding to multiple categories; obtaining the number of samples in each candidate sample set; if the number of samples in a target candidate sample set is less than a first preset quantity threshold, then discarding the category corresponding to the target candidate sample set, and allocating the samples in the target candidate sample set to other candidate sample sets; re-determining the second cluster centers corresponding to each of the other candidate sample sets, and based on the second cluster centers, using the K-means clustering algorithm to perform clustering processing on the first sample set to obtain target sample sets corresponding to multiple categories, the number of the target sample sets being less than the number of the candidate sample sets; taking the category of the target sample set to which the sample to be classified belongs as the rough classification to which the sample to be classified belongs.

[0011] In an exemplary embodiment of the present disclosure, obtaining multiple classification parameters, and respectively classifying the second sample set using preselected classification models based on each of the classification parameters, and determining the target subcategory to which the sample to be classified belongs according to the obtained multiple classification results includes: classifying the second sample set using the nearest neighbor node algorithm based on each of the classification parameters to obtain multiple classification results; if there are more than a preset quantity of the same classification results, then determining the category to which the sample to be classified belongs among the same classification results as the target subcategory.

[0012] In an exemplary embodiment of the present disclosure, before classifying the second sample set using the nearest neighbor node algorithm based on each of the classification parameters to obtain a plurality of classification results, the method further includes: dividing the texts in the second sample set except the text to be classified into a first test set and a second test set; using the nearest neighbor node algorithm to classify each sample in the second test set by using the first test set; and deleting the samples with incorrect classification in the second test set from the second sample set.

[0013] In an exemplary embodiment of the present disclosure, before classifying the second sample set using the nearest neighbor node algorithm based on each of the classification parameters to obtain a plurality of classification results, the method further includes: if the number of samples in a sub-class corresponding to a rough classification is greater than a second preset quantity threshold, performing under-sampling on the samples in the sub-class corresponding to the rough classification to balance the number of samples in each sub-class of the rough classification.

[0014] In an exemplary embodiment of the present disclosure, after obtaining the target sub-category, the method further includes: establishing a correspondence relationship between the text to be classified and the target display information in the target sub-category.

[0015] According to one aspect of the present disclosure, there is provided a text classification device, including: a first sample set acquisition module, configured to form a first sample set according to a text to be classified and a text knowledge base, where the classification knowledge base includes texts with a target number of categories; a clustering processing module, configured to determine a clustering center according to the target number and the corresponding category, and perform clustering processing on the first sample set based on the clustering center to obtain a rough classification to which the sample to be classified belongs, where the rough classification includes a plurality of sub-classifications; a second sample set acquisition module, configured to perform fusion processing on the text to be classified and the texts in the rough classification in the text knowledge base to form a second sample set; and a classification processing module, configured to obtain a plurality of classification parameters, classify the second sample set respectively using a preselected classification model based on each of the classification parameters, and determine a target sub-category to which the sample to be classified belongs according to the obtained plurality of classification results.

[0016] According to one aspect of the present disclosure, there is provided a computer storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the text classification method described in any one of the above is implemented.

[0017] According to one aspect of the present disclosure, there is provided an electronic device, including: one or more processors; and a storage device, configured to store one or more programs, and when the one or more programs are executed by the one or more processors, the one or more processors are caused to implement the text classification method as described in any one of the above.

[0018] In the text classification method in the exemplary embodiments of the present disclosure, the first clustering center is determined according to the categories and quantities of texts in the text knowledge base, so as to fuse the text to be classified and the text knowledge base according to the first clustering center to obtain a first sample set for clustering processing, thereby obtaining multiple rough classifications; then, the texts in the rough classifications are fused with the text to be classified to obtain a second sample set, and the second sample set is classified respectively by using a preselected classification model based on multiple classification parameters to obtain the target sub-classification to which the text to be classified belongs. On the one hand, the present disclosure can adjust the clustering center according to the categories and quantities of texts in the existing text knowledge base, control the quantity of rough classifications, so as to control the quantity of selected texts, and further control the quantity of samples in the rough classification to which the text to be classified belongs, thereby classifying the clustered texts, reducing the consumption of computing resources for classification, and having great flexibility; on the other hand, by combining two classification methods, that is, clustering the texts first and then classifying them, the sensitivity to model parameters in the classification process is reduced, the accuracy of text classification is improved, and the efficiency of business processing based on text classification can also be improved.

[0019] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. Brief Description of the Drawings

[0020] By reading the following detailed description with reference to the drawings, the above and other objects, features and advantages of the exemplary embodiments of the present disclosure will become easy to understand. In the drawings, several embodiments of the present disclosure are shown in an exemplary rather than restrictive manner, wherein:

[0021] Figure 1 A flowchart of the text classification method according to the exemplary embodiment of the present disclosure is shown;

[0022] Figure 2 A flowchart of clustering the first sample set based on the first clustering center to obtain the rough classification to which the text to be classified belongs according to the exemplary embodiment of the present disclosure is shown;

[0023] Figure 3 A flowchart of classifying the second sample set respectively by using a preselected classification model based on each classification parameter and determining the target sub-classification to which the text to be classified belongs according to the obtained multiple classification results according to the exemplary embodiment of the present disclosure is shown;

[0024] Figure 4 A flowchart of deleting the overlapping parts between the detailed classifications corresponding to the rough classifications according to the exemplary embodiment of the present disclosure is shown;

[0025] Figure 5 A schematic structural diagram of the text classification device according to the exemplary embodiment of the present disclosure is shown;

[0026] Figure 6 shows a schematic diagram of a storage medium according to an exemplary embodiment of the present disclosure; and

[0027] Figure 7 shows a block diagram of an electronic device according to an exemplary embodiment of the present disclosure.

[0028] In the drawings, the same or corresponding reference numerals indicate the same or corresponding parts. Detailed implementation manners

[0029] Exemplary embodiments will now be described more fully with reference to the accompanying drawings. However, the exemplary embodiments can be implemented in various forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the concept of the exemplary embodiments to those skilled in the art. The same reference numerals in the figures denote the same or similar structures, and thus their detailed descriptions will be omitted.

[0030] In addition, the described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments. In the following description, numerous specific details are provided to give a thorough understanding of the embodiments of the present disclosure. However, those skilled in the art will realize that the technical solutions of the present disclosure can be practiced without one or more of the specific details, or other methods, components, devices, steps, etc. may be adopted. In other cases, well-known structures, methods, devices, implementations, or operations are not shown or described in detail to avoid obscuring aspects of the present disclosure.

[0031] The block diagrams shown in the drawings are only functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities can be implemented in software form, or in one or more software-hardened modules, or in different networks and / or processor devices and / or microcontroller devices.

[0032] In many industries, such as property management, banks, securities companies, insurance companies, trust investment companies, fund management companies, hotel management, etc., text classification is often involved in business processing. Accurately classifying texts can effectively improve business processing efficiency, such as disease analysis, financial analysis, and intelligent security analysis. However, in actual production applications, many text information annotation tasks often still rely on manual processing, and the volume may reach hundreds, thousands, or even millions as the business grows. Currently, text automatic classification includes supervised models and unsupervised models. For example, KNN (K-Nearest Neighbor) is a typical supervised model, and K-means (K-means clustering algorithm) is a typical unsupervised model. In actual production applications, when the value of K selected by the KNN algorithm is too small, the classification result is easily affected by abnormal data, and when the value of K is too large, it is easily affected by uneven samples. In the case of uneven samples, the K-means algorithm is prone to converge to a local optimal solution. That is, these two algorithms are relatively sensitive to model parameters, and the classification result is easily limited by the classification samples. Moreover, when using the KNN algorithm for classification, the computational resource consumption is large.

[0033] Based on this, in an exemplary embodiment of the present disclosure, a text classification method is first provided. Referring to Figure 1 as shown, the text classification method includes the following steps:

[0034] Step S110: Form a first sample set according to the text to be classified and a text knowledge base, where the classification knowledge base includes texts with a target number of categories;

[0035] Step S120: Determine the clustering centers according to the target number and the corresponding categories, and perform clustering processing on the first sample set based on the clustering centers to obtain the rough classification to which the sample to be classified belongs, and the rough classification includes multiple sub-classifications;

[0036] Step S130: Perform a fusion process on the text to be classified and the texts in the rough classification in the text knowledge base to form a second sample set;

[0037] Step S140: Obtain multiple classification parameters, and respectively classify the second sample set using a preselected classification model based on each classification parameter, and determine the target sub-category to which the sample to be classified belongs according to the obtained multiple classification results.

[0038] According to the text classification method in this exemplary embodiment, the clustering centers can be adjusted based on the categories and quantities of texts in the existing text knowledge base, the quantity of rough classifications can be controlled to control the quantity of selected texts, and further the quantity of samples in the rough classification to which the samples to be classified belong can be controlled, so as to classify the clustered texts, reduce the consumption of computing resources for classification, and have great flexibility; by combining two classification methods, the way of clustering the texts first and then classifying them can reduce the sensitivity to model parameters in the classification process, improve the accuracy of text classification, and also improve the business processing efficiency based on text classification.

[0039] The following will describe Figure 1 the text classification method in the exemplary embodiment of the present disclosure.

[0040] In step S110, a first sample set is formed according to the text to be classified and the text knowledge base.

[0041] In the exemplary embodiment of the present disclosure, the text to be classified and the texts in the text knowledge base are fused to form a first sample set. Among them, the text knowledge base includes a large number of texts with known categories, and each text belongs to a sub-classification corresponding to the rough classification in the case of belonging to a large-range rough classification. For example, a certain text A belongs to the rough classification "community", and the rough classification "community" includes multiple sub-classifications (such as community environment, community organization, and community culture, etc.), and text A can also belong to the fine classification "community culture" corresponding to the rough classification "community". A certain text B belongs to the rough classification "insurance", and the rough classification "insurance" includes multiple sub-classifications (such as vehicle insurance, accident insurance, health insurance, etc.), and text B also belongs to the fine classification "vehicle insurance" corresponding to the rough classification "insurance". In an exemplary embodiment of the present disclosure, the text knowledge base may include a large number of question texts, such as "What is the form of community organization?", "How to purchase accident insurance?", etc., and different questions have matching answers.

[0042] In some possible implementation manners, there are intersections among the texts in different categories in the text knowledge base, that is, the same text belongs to multiple rough classifications at the same time. For example, text C belongs to both the rough classifications "community text" and "insurance"; in other possible implementation manners, there are intersections among the texts in multiple sub-classifications corresponding to the same rough classification. For example, text D belongs to the rough classification "insurance" and, at the same time, belongs to the fine classifications "vehicle insurance" and "accident insurance" corresponding to the rough classification "insurance".

[0043] In step S120, the clustering centers are determined according to the target quantity and the corresponding categories, and the first sample set is clustered based on the clustering centers to obtain the rough classification to which the samples to be classified belong.

[0044] In an exemplary embodiment of the present disclosure, the clustering process is a process of classifying texts into different categories, and the texts in the same category have great similarity; the target number is the number of text categories in the existing classification knowledge base.

[0045] In some possible implementation manners, first, determine that the number of clustering centers is the target number of text categories in the classification knowledge base, then randomly select samples with the target number in the first sample set as the first clustering centers, and perform clustering processing on the first sample set based on the first clustering centers to obtain the rough classification to which the sample to be classified belongs.

[0046] Taking the K-means algorithm as an example, the process of performing clustering processing based on the first clustering centers is described in detail: First, randomly select samples with the target number in the first sample set as the first clustering centers; second, calculate the distance from each sample in the first sample set to each first clustering center; then, use the first clustering center corresponding to the minimum distance value among the multiple distance values obtained for each sample in the first sample set as the first clustering center to which the sample belongs, and divide the sample into the category corresponding to the first clustering center to which it belongs to obtain multiple rough classifications; finally, for each rough classification, calculate the mean value of the samples included as the first clustering center, and continue iterative calculation until the first clustering center no longer changes, so as to determine the rough classification to which the sample to be classified belongs from the multiple obtained rough classifications. Among them, calculating the distance from each sample in the first sample set to each first clustering center can be the Euclidean distance.

[0047] In some possible implementation manners, one sample can be randomly selected from each category of texts in the text knowledge base included in the first sample set as the first clustering center. For example, if the categories of the text knowledge base samples include "insurance", "community", and "finance", then one sample is randomly selected from the texts corresponding to these three categories as the first clustering center, and based on the obtained 3 first clustering centers, the K-means clustering algorithm is used to perform clustering processing on the first sample set to obtain the rough classification to which the sample to be classified belongs. The specific clustering processing process is as described above and will not be elaborated here.

[0048] Based on this embodiment, by respectively selecting the first clustering centers from the samples included in the known categories in the text knowledge base, so that samples belonging to different categories are selected as the first clustering centers before the clustering calculation process, to a certain extent, the number of iterations of the clustering process can be reduced and the clustering calculation amount can be reduced.

[0049] In some possible implementation manners, the number of clustering centers can be adjusted during the process of performing clustering processing on the first sample set based on the first clustering centers. Specifically, Figure 2The flowchart shows the clustering process of the first sample set based on the first clustering center according to an exemplary embodiment of the present disclosure to obtain the rough classification to which the sample to be classified belongs, as Figure 2 shown. This process includes:

[0050] In step S210, the distances between the samples in the first sample set and each first clustering center are obtained, and the samples in the first sample set are assigned to the categories corresponding to each first clustering center according to the distances, obtaining candidate sample sets corresponding to multiple categories.

[0051] In an exemplary embodiment of the present disclosure, the distance between the samples in the first sample set and each first clustering center can be the Euclidean distance. After obtaining the distances between the samples in the first sample set and each first clustering center, for any sample, the first clustering center corresponding to the minimum distance is obtained and the sample is assigned to the category corresponding to the first clustering center, obtaining multiple candidate sample sets corresponding to the categories.

[0052] In step S220, the number of samples in each candidate sample set is obtained, and the number of samples in each candidate sample set is compared with a first preset quantity threshold.

[0053] In an exemplary embodiment of the present disclosure, after obtaining multiple candidate sample sets corresponding to the categories, the number of samples in each candidate sample set is obtained, and the number of samples in each candidate sample set is respectively compared with a first preset quantity threshold.

[0054] In step S230, if the number of samples in the target candidate sample set is less than the first preset quantity threshold, the category corresponding to the target candidate sample set is discarded, and the samples in the target candidate sample set are assigned to other candidate sample sets.

[0055] In an exemplary embodiment of the present disclosure, if the number of samples in the target candidate sample set is less than the first preset quantity threshold, the category corresponding to the target candidate sample set is discarded, and the samples in the target candidate sample set are assigned to other candidate sample sets. Optionally, the samples in the target candidate sample set can be randomly assigned to other candidate sample sets; optionally, the center of each other candidate sample set (for example, it can be a clustering center) can be determined, then the distances between the samples in the target candidate sample set and the centers of each other candidate sample set are calculated, and the samples in the target candidate sample set are assigned to the other candidate sample corresponding to the minimum distance, etc. The present disclosure does not make special limitations on the assignment method of the samples in the target candidate sample set.

[0056] Based on this exemplary embodiment, discard the category corresponding to the target candidate sample set with the number of discarded samples less than the first preset number threshold. Since the number of samples included in this category is small enough, it indicates that the possibility of this category being the category to which the text to be classified belongs is small. Discarding this category not only reduces the probability of misallocation of the target to be classified but also reduces the clustering calculation amount.

[0057] In step S240, re-determine the second clustering center corresponding to each other candidate sample set. Based on the second clustering center, use the K-means clustering algorithm to perform clustering processing on the first sample set to obtain target sample sets corresponding to multiple categories.

[0058] In the exemplary embodiment of the present disclosure, since the number of second clustering centers is less than the number of first clustering centers, correspondingly, the number of target sample sets is less than the number of candidate sample sets. After discarding the target candidate sample set in step S230, re-determine the second clustering center corresponding to each other candidate sample set, that is, the number of second clustering centers is less than the number of first clustering centers. For example, the centroid of all samples in each other candidate sample set can be used as the second clustering center corresponding to each other candidate sample set. Of course, according to the actual text classification situation, other methods can also be used to determine the second clustering center, and the present disclosure does not make special limitations on this. Among them, the process of using the K-means clustering algorithm to perform clustering processing on the first sample set based on the second clustering center to obtain target sample sets corresponding to multiple categories is the same as the above-mentioned clustering processing method and will not be elaborated here.

[0059] Through this exemplary embodiment, in the process of performing clustering processing on the first sample set based on the first clustering center, by adjusting the number of clustering centers, the number of iterations in the clustering process is reduced, and the processing efficiency is improved.

[0060] In some possible implementation manners, it is also possible to, in the process of using the K-means clustering algorithm to perform clustering processing on the first sample set based on the second clustering center to obtain target sample sets corresponding to multiple categories, after each clustering is completed, obtain the number of samples in each candidate sample set in the clustering result, and compare this number with the preset number threshold, so as to discard the category according to the comparison result, and allocate the samples in the candidate sample set corresponding to the discarded category to other candidate sample sets, and continue to re-determine the clustering center of each other candidate sample set and continue to iterate until the preset number of iterations is reached, and use the final clustering result corresponding to the preset number of iterations as the target sample set. Based on this, the clustering center can be dynamically adjusted during the clustering processing to improve the clustering efficiency.

[0061] In step S130, fuse the text to be classified and the text in the rough classification in the text knowledge base to form a second sample set.

[0062] In an exemplary embodiment of the present disclosure, the text to be classified is fused with the text in the rough classification in the text knowledge base to form a second sample set.

[0063] In some possible implementation manners, the text to be classified can be directly merged with the text in the rough classification in the text knowledge base to be used as the second sample set.

[0064] In this exemplary embodiment, by fusing the text to be classified with the text in the rough classification in the text knowledge base and using it as the second sample set for subsequent classification processing, the memory overhead of the subsequent classification algorithm is reduced by reducing the classification range, and the classification calculation speed is improved.

[0065] In step S140, multiple classification parameters are obtained, and based on each classification parameter, a preselected classification model is respectively used to classify the second sample set, and the target subcategory to which the sample to be classified belongs is determined according to the obtained multiple classification results.

[0066] In an exemplary embodiment of the present disclosure, the classification parameters can be flexibly selected according to the actual classification situation. Optionally, the classification parameters can be randomly selected from different odd numbers, such as 3, 5, 7, 9, 11, etc.; optionally, a classification parameter benchmark can be determined by the method of cross-validation, and then different odd numbers are randomly selected near this classification parameter benchmark as the classification parameters.

[0067] Figure 3 The flowchart shows that based on each classification parameter, a preselected classification model is respectively used to classify the second sample set, and the target subcategory to which the sample to be classified belongs is determined according to the obtained multiple classification results, as Figure 3 shown, and the process includes:

[0068] In step S310, based on each classification parameter, the nearest neighbor node algorithm is used to classify the second sample set to obtain multiple classification results.

[0069] In an exemplary embodiment of the present disclosure, for each obtained classification parameter, the nearest neighbor point algorithm is respectively used to classify the second sample set. This process specifically includes: First, calculate the distance between the text to be classified in the second sample set and other samples in the second sample set, which can be, for example, the Euclidean distance; second, sort the obtained distances in ascending order to form a sequence; then, obtain the N other samples in the second sample set corresponding to the first N distances in the sequence, where N is the classification parameter; finally, respectively determine the frequencies of the categories to which the N other samples in the second sample set belong, and determine the category with the highest frequency as the classification result of the sample to be classified.

[0070] It should be noted that, based on each classification parameter, the nearest neighbor node algorithm is respectively used to classify the second sample set, and multiple classification results of the samples to be classified are obtained.

[0071] In some possible implementation manners, the texts in the second sample set other than the text to be classified are texts in a determined rough classification. When the number of fine classifications corresponding to the rough classification is multiple, before using the nearest neighbor node algorithm based on each classification parameter to classify the second sample set and obtain multiple classification results, the overlapping parts between the fine classifications can also be deleted by the following method. See Figure 4 As shown, this process includes: in step S410, the texts in the second sample set other than the text to be classified are divided into a first test set and a second test set; wherein, the division method of the first test set and the second test set can be random division, and the present disclosure does not make special limitations on this; in step S420, the nearest neighbor node algorithm is used to classify each sample in the second test set by using the first test set; in step S430, the samples in the second test set that are misclassified are deleted from the second sample set.

[0072] Through this implementation manner, since the overlapping parts are fuzzy and close in distance, they are easily misclassified. Therefore, by deleting the overlapping parts, the "misleading samples" are deleted, the influence of the overlapping parts on classification is reduced, and the classification accuracy is improved.

[0073] In some possible implementation manners, before using the nearest neighbor node algorithm based on each classification parameter to classify the second sample set and obtain multiple classification results, it is judged whether the number of samples in multiple sub-classifications corresponding to the rough classification is greater than a second preset quantity threshold. If there is a sub-classification corresponding to the rough classification in which the number of samples is greater than the second preset quantity threshold, then under-sampling is performed on the samples in the sub-classification corresponding to the rough classification to balance the number of samples in each sub-classification of the rough classification. For example, some samples can be randomly selected and deleted from the sub-classification corresponding to the rough classification, and the specific quantity can be adjusted according to the number of samples in each sub-classification of the rough classification. By deleting samples, the samples in each sub-classification are balanced, and to a certain extent, the classification calculation consumption can also be reduced.

[0074] In step S320, if there are more classification results that are the same than a preset quantity, then the category to which the sample to be classified belongs among the same classification results is determined as the target sub-category.

[0075] In an exemplary embodiment of the present disclosure, based on each classification parameter, a preselected classification model is respectively used to classify the second sample set to obtain multiple classification results. If there are more than a preset number of identical classification results, the category to which the sample to be classified belongs among the identical classification results is determined as the target subcategory. The preset number can be set according to actual classification requirements, such as 3 times, 5 times, etc., and the present disclosure does not make special limitations on this.

[0076] In addition, after obtaining the target subcategory to which the text to be classified belongs, a corresponding relationship can be established between the text to be classified and the target display information in the target subcategory. Among them, multiple target display information is pre-stored in the target subclassification, corresponding to different texts to be classified respectively. For example, if the text to be classified is a question Q, by classifying the question Q into the target subclassification M, a corresponding relationship is established between the question Q and the answer A in the question M under the target subclassification.

[0077] Optionally, an operator can select the answer A corresponding to the question Q in the target subclassification M and establish a corresponding relationship between the question Q and the answer A. Since the question Q has been classified from coarse to fine and has been accurately located in the target subclassification M, the operator can quickly select from this target subclassification M, greatly improving the work efficiency of the staff.

[0078] Optionally, after classifying the question Q into the target subclassification M, through text recognition and detection technologies, keywords in the question Q can be detected, and the keywords are matched with the text of the answer A. If they match, a corresponding relationship between the question Q and the answer A is automatically established, reducing manual intervention and improving work efficiency.

[0079] According to the text classification method in this exemplary embodiment, the clustering center can be adjusted according to the categories and quantities of texts in the existing text knowledge base, the quantity of rough classification can be controlled to control the quantity of selected texts, and further the quantity of samples in the rough classification to which the sample to be classified belongs can be controlled, so as to classify the clustered texts, reduce the consumption of computing resources for classification, and have great flexibility; by clustering the text first and then classifying it, the sensitivity of the model to parameters in the classification process is reduced, the accuracy of text classification is improved, and the business processing efficiency based on text classification can also be improved. It is a method applicable to various text classification application scenarios. For example, in an application scenario, the question text that cannot be answered by an intelligent question and answer robot can be classified into the corresponding category in the text knowledge base by the method of the present disclosure, and the corresponding answer is selected under this category, thereby improving the accuracy of intelligent questions and reducing the cost of manual annotation.

[0080] In addition, in an exemplary embodiment of the present disclosure, a text classification device is also provided. Refer to Figure 5As shown, the text classification device 500 may include a first sample set acquisition module 510, a clustering processing module 520, a second sample set acquisition module 530, and a classification processing module 540. Specifically,

[0081] The first sample set acquisition module 510 is configured to form a first sample set according to the text to be classified and a text knowledge base, where the classification knowledge base includes texts with a target number of categories;

[0082] The clustering processing module 520 is configured to determine clustering centers according to the target number and corresponding categories, and perform clustering processing on the first sample set based on the clustering centers to obtain a rough classification to which the sample to be classified belongs, where the rough classification includes multiple sub-classifications;

[0083] The second sample set acquisition module 530 is configured to fuse the text to be classified and the texts in the rough classification in the text knowledge base to form a second sample set;

[0084] The classification processing module 540 is configured to obtain multiple classification parameters, respectively classify the second sample set by using a preselected classification model based on each classification parameter, and determine the target sub-category to which the sample to be classified belongs according to the obtained multiple classification results.

[0085] In an exemplary embodiment of the present disclosure, the clustering processing module 520 may include:

[0086] A quantity determination unit configured to determine that the number of clustering centers is the target number;

[0087] A clustering processing unit configured to randomly select samples with the target number in the first sample set as the first clustering centers, and perform clustering processing on the first sample set based on the first clustering centers to obtain a rough classification to which the sample to be classified belongs.

[0088] In an exemplary embodiment of the present disclosure, the clustering processing module 520 may further include:

[0089] A clustering center selection unit configured to randomly select one sample from each category of texts in the text knowledge base included in the first sample set as the first clustering centers;

[0090] The clustering processing unit is further configured to perform clustering processing on the first sample set based on the first clustering centers by using the K-means clustering algorithm to obtain a rough classification to which the sample to be classified belongs.

[0091] In an exemplary embodiment of the present disclosure, the clustering processing module 520 may further include:

[0092] A sample quantity acquisition unit configured to acquire the quantity of samples in each candidate sample set;

[0093] A sample processing unit, configured to discard the category corresponding to the target candidate sample set if the number of samples in the target candidate sample set is less than the first preset number threshold, and allocate the samples in the target candidate sample set to other candidate sample sets

[0094] A clustering center determination unit, configured to re-determine the second clustering center corresponding to each other candidate sample set, and the clustering processing unit is further configured to perform clustering processing on the first sample set based on the second clustering center by using the K-means clustering algorithm to obtain target sample sets corresponding to multiple categories, and the number of the target sample sets is less than the number of the candidate sample sets;

[0095] A rough classification determination unit, configured to use the category of the target sample set to which the sample to be classified belongs as the rough classification to which the sample to be classified belongs.

[0096] In an exemplary embodiment of the present disclosure, the classification processing module 540 may include:

[0097] A classification processing unit, configured to classify the second sample set based on each classification parameter by using the nearest neighbor node algorithm to obtain multiple classification results;

[0098] A target sub-category determination unit, configured to determine the category to which the sample to be classified belongs in the same classification results as the target sub-category if there are more than a preset number of the same classification results.

[0099] In an exemplary embodiment of the present disclosure, the classification processing module 540 may further include:

[0100] A sample division unit, configured to divide the texts in the second sample set except the text to be classified into a first test set and a second test set;

[0101] The classification processing unit is further configured to classify each sample in the second test set by using the first test set by using the nearest neighbor node algorithm;

[0102] A sample deletion unit, configured to delete the samples with incorrect classification in the second test set from the second sample set.

[0103] In an exemplary embodiment of the present disclosure, the classification processing module 540 may further include:

[0104] A sample preprocessing unit, configured to perform undersampling on the samples in the sub-category corresponding to the rough classification if the number of samples in the sub-category corresponding to the rough classification is greater than the second preset number threshold, so as to balance the number of samples in each sub-category of the rough classification.

[0105] In an exemplary embodiment of the present disclosure, the text classification device of the present disclosure may further include:

[0106] A relationship establishment module, configured to establish a correspondence between the text to be classified and the target display information in the target subcategory.

[0107] Since each functional module of the text classification device according to the exemplary embodiments of the present disclosure is the same as that in the inventive embodiments of the above-described text classification method, it will not be described in detail herein.

[0108] It should be noted that although several modules or units of the text classification device are mentioned in the above detailed description, such a division is not mandatory. In fact, according to the embodiments of the present disclosure, the features and functions of the two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by a plurality of modules or units.

[0109] In addition, in the exemplary embodiments of the present disclosure, a computer storage medium capable of implementing the above method is also provided. A program product capable of implementing the above method of this specification is stored thereon. In some possible embodiments, various aspects of the present disclosure can also be implemented in the form of a program product, which includes program code. When the program product runs on a terminal device, the program code is used to cause the terminal device to execute the steps according to various exemplary embodiments of the present disclosure described in the above "Exemplary Method" section of this specification.

[0110] Reference Figure 6 As shown, a program product 600 for implementing the above method according to the exemplary embodiments of the present disclosure is described. It can be a portable compact disc read-only memory (CD-ROM) and includes program code, and can run on a terminal device, such as a personal computer. However, the program product of the present disclosure is not limited thereto. In this document, a readable storage medium can be any tangible medium that contains or stores a program, and the program can be used by or in combination with an instruction execution system, apparatus, or device.

[0111] The program product can adopt any combination of one or more readable media. The readable media can be a readable signal medium or a readable storage medium. The readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the readable storage medium include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.

[0112] A computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries readable program code. Such a propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the foregoing. The readable signal medium may also be any readable medium other than a readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device.

[0113] The program code contained on the readable medium may be transmitted using any appropriate medium, including but not limited to wireless, wired, optical fiber cable, RF, etc., or any suitable combination of the foregoing.

[0114] The program code for performing the operations of the present disclosure may be written in any combination of one or more programming languages, including object-oriented programming languages such as Java, C++, etc., and also including conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computing device, partially on the user's device, executed as a stand-alone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device may be connected to the user's computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computing device (e.g., by connecting through the Internet using an Internet service provider).

[0115] In addition, in an exemplary embodiment of the present disclosure, an electronic device capable of implementing the above method is also provided. Those skilled in the art can understand that various aspects of the present disclosure can be implemented as a system, method, or program product. Therefore, various aspects of the present disclosure can be specifically implemented in the following forms: a complete hardware embodiment, a complete software embodiment (including firmware, microcode, etc.), or an embodiment combining hardware and software aspects, which can be collectively referred to as "circuit", "module", or "system" here.

[0116] Reference will be made below Figure 7 to describe the electronic device 700 according to such an embodiment of the present disclosure. Figure 7 The illustrated electronic device 700 is merely an example and should not impose any limitation on the functions and scope of use of the embodiments of the present disclosure.

[0117] As Figure 7As shown, the electronic device 700 is presented in the form of a general-purpose computing device. The components of the electronic device 700 may include, but are not limited to: at least one of the above-mentioned processing units 710, at least one of the above-mentioned storage units 720, a bus 730 connecting different system components (including the storage unit 720 and the processing unit 710), and a display unit 740.

[0118] Among them, the storage unit stores program code, and the program code can be executed by the processing unit 710, so that the processing unit 710 executes the steps according to various exemplary embodiments of the present disclosure described in the "Exemplary Method" section of this specification.

[0119] The storage unit 720 may include a readable medium in the form of a volatile storage unit, such as a random access storage unit (RAM) 7201 and / or a cache storage unit 7202, and may further include a read-only storage unit (ROM) 7203.

[0120] The storage unit 720 may further include a program / utility 7204 having a set (at least one) of program modules 7205. Such program modules 7205 include, but are not limited to: an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include the implementation of a network environment.

[0121] The bus 730 may represent one or more of several types of bus structures, including a storage unit bus or a storage unit controller, a peripheral bus, a graphics acceleration port, a processing unit, or a local bus using any bus structure in a variety of bus structures.

[0122] The electronic device 700 can also communicate with one or more external devices 800 (such as a keyboard, a pointing device, a Bluetooth device, etc.), and can also communicate with one or more devices that enable a user to interact with the electronic device 700, and / or communicate with any device that enables the electronic device 700 to communicate with one or more other computing devices (such as a router, a modem, etc.). Such communication can be carried out through an input / output (I / O) interface 750. And, the electronic device 700 can also communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) through a network adapter 760. As shown in the figure, the network adapter 760 communicates with other modules of the electronic device 700 through the bus 730. It should be understood that although not shown in the figure, other hardware and / or software modules can be used in combination with the electronic device 700, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc.

[0123] Through the description of the above embodiments, those skilled in the art can easily understand that the exemplary embodiments described herein can be implemented by software or by a combination of software and necessary hardware. Therefore, the technical solution according to the embodiments of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, including several instructions to enable a computing device (such as a personal computer, a server, a terminal device, or a network device, etc.) to execute the method according to the embodiments of the present disclosure.

[0124] In addition, the above drawings are only schematic illustrations of the processes included in the method according to the exemplary embodiments of the present disclosure, rather than for limiting purposes. It is easy to understand that the processes shown in the above drawings do not indicate or limit the chronological order of these processes. Additionally, it is also easy to understand that these processes can be executed, for example, synchronously or asynchronously in multiple modules.

[0125] After considering the specification and practicing the invention disclosed herein, those skilled in the art will readily conceive of other embodiments of the present disclosure. The present disclosure is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include well-known knowledge or conventional technical means in the technical field not disclosed by the present disclosure. The specification and embodiments are only regarded as exemplary, and the true scope and spirit of the present disclosure are pointed out by the claims.

[0126] It should be understood that the present disclosure is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present disclosure is only limited by the appended claims.

Claims

1. A text classification method, characterized in that, Including: Forming a first sample set according to the text to be classified and a text knowledge base, where the text knowledge base includes texts with a target number of categories; Determining the number of cluster centers according to the target number, and randomly selecting samples with the target number in the first sample set as the first cluster centers, and performing clustering processing on the first sample set based on the first cluster centers to obtain the rough classification to which the text to be classified belongs, where the rough classification includes multiple sub-classifications; Performing fusion processing on the text to be classified and the texts in the rough classification in the text knowledge base to form a second sample set; Obtaining a plurality of classification parameters, respectively classifying the second sample set by using a preselected classification model based on each classification parameter, and determining the target sub-category to which the text to be classified belongs according to the obtained plurality of classification results; Wherein, before the step of obtaining a plurality of classification parameters and respectively classifying the second sample set by using a preselected classification model based on each classification parameter, the method further includes: Dividing the texts in the second sample set except the text to be classified into a first test set and a second test set; Using the nearest neighbor node algorithm to classify each sample in the second test set by using the first test set; Deleting the samples with incorrect classification in the second test set from the second sample set; The step of obtaining a plurality of classification parameters, respectively classifying the second sample set by using a preselected classification model based on each classification parameter, and determining the target sub-category to which the text to be classified belongs includes: Classifying the second sample set by using the nearest neighbor node algorithm based on each classification parameter to obtain a plurality of classification results; If there are more than a preset number of identical classification results, determining the category to which the text to be classified belongs among the identical classification results as the target sub-category.

2. The method according to claim 1, wherein The step of randomly selecting samples with the target number in the first sample set as the first cluster centers and performing clustering processing on the first sample set based on the first cluster centers to obtain the rough classification to which the text to be classified belongs includes: Randomly selecting one sample from each category of texts in the text knowledge base included in the first sample set as the first cluster center; Based on the first cluster centers, performing clustering processing on the first sample set by using the K-means clustering algorithm to obtain the rough classification to which the text to be classified belongs.

3. The method according to claim 1, wherein The step of performing clustering processing on the first sample set based on the first cluster centers to obtain the rough classification to which the text to be classified belongs includes: Obtaining the distances between the samples in the first sample set and each of the first cluster centers, and allocating the samples in the first sample set to the categories corresponding to each of the first cluster centers according to the distances to obtain candidate sample sets corresponding to multiple categories; Obtaining the number of samples in each candidate sample set; If the number of samples in a target candidate sample set is less than a first preset quantity threshold, discarding the category corresponding to the target candidate sample set and allocating the samples in the target candidate sample set to other candidate sample sets; Redetermine the second clustering center corresponding to each of the other candidate sample sets, and based on the second clustering center, use the K-means clustering algorithm to perform clustering processing on the first sample set to obtain target sample sets corresponding to multiple categories, where the number of the target sample sets is less than the number of the candidate sample sets; Use the category of the target sample set to which the text to be classified belongs as the rough classification to which the text to be classified belongs.

4. The method according to claim 1, characterized in that, Before classifying the second sample set using the nearest neighbor node algorithm based on each of the classification parameters to obtain multiple classification results, it further includes: If there are samples in the sub-classification corresponding to the rough classification whose number is greater than the second preset quantity threshold, perform undersampling on the samples in the sub-classification corresponding to the rough classification to balance the number of samples in each sub-classification of the rough classification.

5. The method according to any one of claims 1 to 4, characterized in that After obtaining the target sub-category, it further includes: Establish a correspondence relationship between the text to be classified and the target display information in the target sub-category.

6. A text classification device, characterized in that, It includes: A first sample set acquisition module, configured to form a first sample set according to the text to be classified and a text knowledge base, where the text knowledge base includes texts with a target number of categories; A clustering processing module, configured to determine the number of clustering centers according to the target number, and randomly select samples with the target number in the first sample set as the first clustering centers, so as to perform clustering processing on the first sample set based on the first clustering centers to obtain the rough classification to which the text to be classified belongs, where the rough classification includes multiple sub-classifications; A second sample set acquisition module, configured to fuse the text to be classified and the texts in the rough classification in the text knowledge base to form a second sample set; A classification processing module, configured to obtain multiple classification parameters, and respectively classify the second sample set using a preselected classification model based on each of the classification parameters, and determine the target sub-category to which the text to be classified belongs according to the multiple classification results obtained; The classification processing module is further configured to execute: Divide the texts in the second sample set except the text to be classified into a first test set and a second test set; Use the nearest neighbor node algorithm to classify each sample in the second test set using the first test set; Delete the samples in the second test set that are misclassified from the second sample set; Wherein, the classification processing module is configured to execute: Classify the second sample set using the nearest neighbor node algorithm based on each of the classification parameters to obtain multiple classification results; If there are more than a preset number of identical classification results, determine the category to which the text to be classified belongs among the identical classification results as the target sub-category.

7. A computer storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the text classification method according to any one of claims 1 to 5.

8. An electronic device, characterized in that, It includes: One or more processors; A storage device, configured to store one or more programs, and when the one or more programs are executed by the one or more processors, enable the one or more processors to implement the text classification method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Event mining method and device

    CN111767404A