Methods, devices and electronic equipment for detecting diversity in text data
By performing multiple clustering operations and processing seed point data, category clusters are generated, which solves the problem of accuracy in detecting text data diversity and improves the classification effect of the classification model.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- MASHANG CONSUMER FINANCE CO LTD
- Filing Date
- 2022-11-18
- Publication Date
- 2026-05-26
AI Technical Summary
In existing technologies, the lack of diversity in text data leads to poor classification performance of classification models. How to accurately detect the diversity of text data has become an urgent technical problem to be solved.
Seed point data is obtained and clustered through multiple clustering operations to generate category clusters. Diversity is detected based on the similarities and differences between the clusters and the original text data. This includes using methods such as BERT, SimCSE, and Word2vec to transform the text data vectors and using the K-means algorithm for clustering to generate category clusters that reflect the clustering results at each stage.
It achieves accurate detection of text data diversity, can identify categories with insufficient diversity, and improves the classification accuracy of the classification model.
Smart Images

Figure CN116226369B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data processing, and in particular to a method, apparatus, and electronic device for detecting the diversity of text data. Background Technology
[0002] Text data includes various types of data composed of Chinese, English, numbers, and other characters. Text data comes from diverse sources, such as chat software, agent recording software, and social media platforms.
[0003] In related technologies, to facilitate processing large amounts of text data, it is necessary to perform classification operations on the text data beforehand. During the classification process, a pre-trained classification model is typically used. Training this model requires acquiring a large amount of text data as training samples.
[0004] The greater the diversity of text data in the training samples, the higher the classification accuracy of the final trained classification model. However, in reality, the diversity of text data is often insufficient, resulting in poor classification performance. Therefore, how to detect the diversity of text data has become a pressing technical challenge. Summary of the Invention
[0005] This application provides a method, apparatus, and electronic device for detecting the diversity of text data, used to accurately detect whether the diversity of various categories of text data meets business requirements.
[0006] Firstly, this application provides a method for detecting diversity in text data, the method comprising:
[0007] Obtain the original text data of m categories corresponding to m category labels, wherein the original text data of the m categories belong to the same business scenario;
[0008] Extract j seed point data from the original text data of each category to obtain j sets of seed points; each set of seed points contains m seed point data, and the m seed point data correspond to m categories respectively; where m and j are natural numbers;
[0009] Based on the m seed point data contained in each seed point set, perform j clustering operations on the original text data of the m categories to obtain j clustering results; wherein, each clustering result contains m clusters, and the m clusters correspond to the m categories respectively;
[0010] For each category label, the j clusters corresponding to the j clustering results of the category are determined as the category clusters of the category;
[0011] For each category, the data diversity detection result of the category is obtained based on the clustered text data contained in the category cluster and the original text data corresponding to the category label of the category.
[0012] Secondly, this application provides a text data diversity detection device, comprising:
[0013] The acquisition module is adapted to acquire the original text data of m categories corresponding to m category labels, wherein the original text data of the m categories belong to the same business scenario.
[0014] The extraction module is adapted to extract j seed point data from the original text data of each category, resulting in j sets of seed points; each set of seed points contains m seed point data, and the m seed point data correspond to m categories respectively; where m and j are natural numbers;
[0015] The clustering module is adapted to perform j clustering processes on the original text data of the m categories based on the m seed point data contained in each seed point set, to obtain j clustering results; wherein each clustering result contains m clusters, and the m clusters correspond to the m categories respectively;
[0016] The determination module is adapted to determine the j clusters corresponding to the j clustering results of the category for each category label as the category clusters of the category;
[0017] The detection module is adapted to obtain the data diversity detection result of each category based on the clustered text data contained in the category clusters of the category and the original text data corresponding to the category label of the category.
[0018] Thirdly, this application provides an electronic device comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores one or more computer programs executable by the at least one processor, and the one or more computer programs are executed by the at least one processor to enable the at least one processor to perform the above-described method.
[0019] Fourthly, this application provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program implements the above-described method when executed by a processor / processor core.
[0020] The embodiments provided in this application can perform j clustering processes on original text data of m categories. Correspondingly, for each category, j clusters corresponding to the j clustering results of that category are determined as category clusters for that category. Based on the similarities and differences between the clustering results of each category's category clusters and the original text data corresponding to the category's category label, it is possible to detect whether the data diversity of that category meets business requirements. Typically, if the data diversity of a certain category is insufficient, the clustering results for that category will be identical in all iterations; that is, the clustering results of the category clusters for that category will be highly consistent with the original text data corresponding to the category's category label. Therefore, through multiple clustering processes, the diversity of text data in each category can be accurately detected.
[0021] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this application, nor is it intended to limit the scope of this application. Other features of this application will become readily apparent from the following description. Attached Figure Description
[0022] The accompanying drawings are provided to further illustrate the present application and form part of the specification. They are used together with the embodiments of the present application to explain the application and do not constitute a limitation thereof. The above and other features and advantages will become more apparent to those skilled in the art from the detailed example embodiments described with reference to the accompanying drawings, in which:
[0023] Figure 1 A flowchart illustrating a text data diversity detection method provided in one embodiment of this application;
[0024] Figure 2 A flowchart of a text data diversity detection method provided in another embodiment of this application;
[0025] Figure 3 This diagram illustrates a block diagram of a text data diversity detection device provided in an embodiment of this application;
[0026] Figure 4 This is a block diagram of an electronic device provided in an embodiment of this application. Detailed Implementation
[0027] To enable those skilled in the art to better understand the technical solutions of this application, exemplary embodiments of this application are described below in conjunction with the accompanying drawings, including various details of the embodiments of this application to aid understanding. These should be considered merely exemplary. Therefore, those skilled in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this application. Similarly, for clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0028] Where there is no conflict, the various embodiments of this application and the features thereof may be combined with each other.
[0029] As used herein, the term “and / or” includes any and all combinations of one or more related enumerated entries.
[0030] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the application. As used herein, the singular forms “a” and “the” are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will also be understood that when the terms “comprising” and / or “made of” are used in this specification, the presence of the stated feature, integral, step, operation, element, and / or component is specified, but the presence or addition of one or more other features, integrals, steps, operations, elements, components, and / or groups thereof is not excluded. Terms such as “connected” or “linked” are not limited to physical or mechanical connections but can include electrical connections, whether direct or indirect.
[0031] Unless otherwise specified, all terms used herein (including technical and scientific terms) have the same meaning as commonly understood by one of ordinary skill in the art. It will also be understood that terms such as those defined in commonly used dictionaries should be interpreted as having a meaning consistent with their meaning in the context of the relevant art and this application, and will not be interpreted as having an idealized or overly formal meaning, unless expressly so defined herein.
[0032] The text data diversity detection method according to embodiments of this application can be executed by electronic devices such as terminal devices or servers. Terminal devices can be in-vehicle devices, user equipment (UE), mobile devices, user terminals, terminals, cellular phones, cordless phones, personal digital assistants (PDAs), handheld devices, computing devices, in-vehicle devices, wearable devices, etc. The server can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services. Specifically, the method can be implemented by a processor calling a computer program stored in memory.
[0033] In related technologies, to facilitate processing large amounts of text data, it is necessary to perform classification operations on the text data beforehand. During classification, a pre-trained classification model is typically used. Training this model requires acquiring a large amount of text data as training samples. The greater the diversity of the text data in the training samples, the higher the classification accuracy of the final trained classification model. However, in practice, if the text data lacks diversity, the classification model will perform poorly. To address this issue, this application proposes a diversity detection scheme based on multiple clustering operations. By analyzing the differences between the results of multiple clustering operations, the diversity of various types of text data can be detected quickly and accurately.
[0034] Figure 1 A flowchart of a text data diversity detection method according to an embodiment of this application is provided. (Refer to...) Figure 1 The method includes:
[0035] Step S110: Obtain the original text data of m categories corresponding to m category labels. The original text data of the m categories belong to the same business scenario.
[0036] Here, m is a natural number representing the number of categories in the original text data, with each category corresponding to a category label. Accordingly, in this step, original text data for m categories is obtained, with each category corresponding to a category label. Specifically, the labeled original text data can be directly obtained, and the category of the original text data can be determined based on the labeled category labels corresponding to each piece of original text data. The labeling method can be manual or machine-generated.
[0037] The original text data in the m categories belong to the same business scenario. Within the same business scenario, the similarity between the original text data is high. Furthermore, the business scenario includes one or more business projects. Optionally, the original text data in the m categories all belong to the same business project within the same business scenario, thus resulting in even higher similarity between the data. Optionally, the text data corresponding to each category label has a different length, and the similarity between multiple text data corresponding to different category labels is greater than a preset similarity threshold.
[0038] Step S120: Extract j seed point data from the original text data of each category to obtain j sets of seed points; each set of seed points contains m seed point data, and the m seed point data correspond to m categories respectively.
[0039] Where m and j are both natural numbers. The number of j values is less than the amount of data in the original text data for each category. For example, if each category contains 20 data entries, the number of j values should be less than 20, specifically 5. Furthermore, if the amount of data in the original text data differs between categories, the value of j should be less than the amount of data in the category with the smallest amount of data.
[0040] In this clustering operation, each set of seed points contains m seed points, and each seed point corresponds to one of the m categories. Therefore, each set of seed points is used to perform one clustering operation, and the number of categories obtained after each clustering operation is always m.
[0041] Step S130: Based on the m seed point data contained in each seed point set, perform j clustering processes on the original text data of m categories to obtain j clustering results; wherein, each clustering result contains m clusters, and the m clusters correspond to m categories respectively.
[0042] Since each set of seed points contains m seed points, and each m seed point corresponds to one of the m categories, each clustering result produces m clusters, each corresponding to one of the m categories. Specifically, each cluster is identified by the category to which its seed points belong. Accordingly, based on the categories to which the seed points in a cluster belong, the mapping relationship between the m clusters and the m categories can be determined.
[0043] Step S140: For each category label, determine the j clusters corresponding to the j clustering results of the category as the category clusters of that category.
[0044] Since each clustering result generates a corresponding cluster for each category label, after j clustering processes, a total of j clusters are generated for each category label. Accordingly, the j clusters corresponding to the category label are defined as the category clusters for that category. Therefore, each category has j category clusters, and each category cluster corresponds to the clustering result of that category in one clustering process. These j category clusters comprehensively reflect the clustering situation of that category in each iteration.
[0045] Step S150: For each category, based on the clustered text data contained in the category clusters of that category and the original text data corresponding to the category label of that category, obtain the detection result of the data diversity of that category.
[0046] In this context, the clustered text data contained within a category cluster refers to the text data belonging to that category cluster. Therefore, the clustered text data contained within a category cluster is text data incorporated into that category cluster through clustering operations.
[0047] The detection results are used to evaluate the data diversity of the category. In one optional implementation, the detection results are used to characterize whether the data diversity of a certain category meets preset diversity conditions. When performing diversity detection based on the clustered text data contained in the category's clusters and the original text data corresponding to the category's category label, the similarities and differences between the clustered text data contained in the category's clusters and the original text data corresponding to the category's category label can be compared to determine whether the diversity meets the preset diversity conditions. For example, in one implementation, the difference between the number of data entries in the clustered text data contained in the category's clusters and the number of data entries in the original text data corresponding to the category's category label is compared. If the difference is larger, it indicates greater dissimilarity, i.e., better diversity (meets the preset diversity conditions); conversely, if the difference is smaller, it indicates less dissimilarity, i.e., worse diversity (does not meet the preset diversity conditions). For example, in another implementation, the number of identical text entries between the clustered text data contained in the category cluster and the original text data corresponding to the category label is compared. A higher number of identical text entries indicates less difference, i.e., poorer diversity (not meeting the preset diversity condition); conversely, a lower number of identical text entries indicates better diversity (meeting the preset diversity condition). In short, the preset diversity condition can be set in various ways, and this application does not limit the specific type of condition.
[0048] In the embodiments provided in this application, j clustering processes can be performed on m categories of original text data. Correspondingly, for each category, j clusters corresponding to the j clustering results of that category are determined as category clusters for that category. Based on the similarities and differences between the clustering results of each category's category clusters and the original text data corresponding to the category's category label, it is possible to detect whether the data diversity of that category meets business requirements. Typically, if the data diversity of a certain category is insufficient, the clustering results for that category will be identical, meaning the clustering results of the category clusters are highly consistent with the original text data corresponding to the category's category label. Therefore, through multiple clustering processes, the diversity of text data in each category can be accurately detected.
[0049] See Figure 2 This is a flowchart illustrating another text data diversity detection method provided in this application. (Refer to...) Figure 2 The method may include the following steps:
[0050] Step S210: Obtain the original text data of the m categories corresponding to the m category labels.
[0051] The original text data in the m categories belong to the same business scenario. Within the same business scenario, the similarity between the original text data is high. Furthermore, the business scenario includes one or more business projects. Optionally, the original text data in the m categories all belong to the same business project within the same business scenario, thus resulting in even higher similarity between the data. Optionally, the text data corresponding to each category label has a different length, and the similarity between multiple text data corresponding to different category labels is greater than a preset similarity threshold.
[0052] Here, m is a natural number representing the number of categories in the original text data, with each category corresponding to a category label. Accordingly, in this step, original text data for m categories is obtained, with each category corresponding to a category label. Specifically, the labeled original text data can be directly obtained, and the category of the original text data can be determined based on the labeled category labels corresponding to each piece of original text data. The labeling method can be manual or machine-generated.
[0053] The original text data is also called the original labeled data. Assume the total number of samples is N, and there are m categories, C1, C2, ..., C6. m The number of samples for each category are N1, N2, ..., N. m Where N = N1 + N2 + ... + N m .
[0054] Step S220: Extract j seed point data from the original text data of each category to obtain j sets of seed points; each set of seed points contains m seed point data, and the m seed point data correspond to m categories respectively.
[0055] Where m and j are both natural numbers. The number of j values is less than the amount of data in the original text data for each category. For example, if each category contains 20 data entries, the number of j values should be less than 20, specifically 5. Furthermore, if the amount of data in the original text data differs between categories, the value of j should be less than the amount of data in the category with the smallest amount of data.
[0056] In this clustering operation, each set of seed points contains m seed points, and each seed point corresponds to one of the m categories. Therefore, each set of seed points is used to perform one clustering operation, and the number of categories obtained after each clustering operation is always m.
[0057] Step S230: Based on the m seed point data contained in each seed point set, perform j clustering processes on the original text data of m categories to obtain j clustering results; wherein, each clustering result contains m clusters, and the m clusters correspond to m categories respectively.
[0058] Since each set of seed points contains m seed points, and each m seed point corresponds to one of the m categories, each clustering result produces m clusters, each corresponding to one of the m categories. Specifically, each cluster is identified by the category to which its seed points belong. Accordingly, based on the categories to which the seed points in a cluster belong, the mapping relationship between the m clusters and the m categories can be determined.
[0059] To facilitate clustering, each text data point is first converted into a corresponding text data vector (also called a sample vector). The text data vector representation can be implemented using methods such as BERT, SimCSE, and Word2vec. Specifically, the K-means algorithm can be used for clustering, where K equals the number of sample categories, i.e., K = m. First, from each category C1, C2, ..., C... m In the first step, j samples are randomly selected as the seed point set for each category, meaning there are j seed points for each category. Then, each time, one sample corresponding to a seed point is randomly selected without replacement from the seed point set of each category as the seed point for this clustering, resulting in m categories and m seed points selected each time. Finally, a clustering algorithm is used to cluster the N original text data in the labeled dataset. Because each category has j seed points, the above clustering process needs to be performed j times, and each clustering process yields m clusters.
[0060] Step S240: For each category label, determine the j clusters corresponding to the j clustering results of that category as the category clusters of that category.
[0061] Since each clustering result generates a corresponding cluster for each category label, after j clustering processes, a total of j clusters are generated for each category label. Accordingly, the j clusters corresponding to the category label are defined as the category clusters for that category. Therefore, each category has j category clusters, and each category cluster corresponds to the clustering result of that category in one clustering process. These j category clusters comprehensively reflect the clustering situation of that category in each iteration.
[0062] Step S250: For each category, based on the clustered text data contained in the category clusters of that category and the original text data corresponding to the category label of that category, detect whether the data diversity of that category meets the preset diversity conditions.
[0063] In this context, the clustered text data contained within a category cluster refers to the text data belonging to that category cluster. Therefore, the clustered text data contained within a category cluster is text data incorporated into that category cluster through clustering operations.
[0064] In the first implementation, data diversity detection for a category is achieved by comparing the similarity between the clustered text data in each category cluster and the original text data corresponding to the category label of that category.
[0065] First, for each category cluster, obtain the consistency count between the clustered text data in the category cluster and the original text data corresponding to the category label of that category. The consistency count represents the number of identical text entries in the clustered text data of the category cluster and the original text data corresponding to the category label. For example, perform data statistics on j clustering results: C 1i C 2i ..., C mi Let C1, C2, ..., Cn be the clusters after the i-th (1≤i≤j) clustering. m The number of samples in the cluster containing the i-th seed point in each category that are identical to samples in the corresponding category of that seed point (i.e., the actual category of the seed point, represented by the category label). Therefore, C 11 C 12 ;…;C 1j C 21 C 22 ;…;C 2j C m1 C m2 ;…;C mj Let C1, C2, ..., Cn be the clustering order after j-th clustering. m The number of samples in the cluster containing the j seed points in each category that are identical to the samples in the category corresponding to the seed point (i.e., the actual category of the seed point, represented by the category label). Let C1, C2, ..., Cn be the clusters after the i-th (1≤i≤j) clustering. m The total number of samples in the cluster containing the i-th seed point in each category. Therefore, Let C1, C2, ..., Cn be the clusters after j-th clustering. mThe total number of samples in the cluster containing the j seed points in each category (i.e., a cluster corresponding to that category).
[0066] Then, based on the number of consistent clusters in each category and the total number of categories in the original text data corresponding to the category label, the clustering bias value of the category is calculated.
[0067] The number of consistent clusters for each category corresponds to C mentioned above. 11 C 12 ;…;C 1j C 21 C 22 ;…;C 2j C m1 C m2 ;…;C mj The total number of categories in the original text data corresponding to the category labels corresponds to N1, N2, ..., N in the preceding text. m In one alternative implementation, the clustering bias of the categories is calculated using the following formula:
[0068]
[0069] Where, ρ θ Let θ be the clustering bias value for category θ. Therefore, the clustering bias value of a category is based on the total number of categories (N) in the original text data corresponding to that category's category label. θ The number of clusters consistent with each category of that category (C) θ1 …C θj The average of the differences between them is determined.
[0070] Finally, based on the comparison between the clustering deviation value of the category and the preset clustering deviation threshold, it is determined whether the data diversity of the category meets the preset diversity conditions.
[0071] The larger the clustering bias value, the greater the difference between the clustering result and the true category, and thus the better the diversity. Therefore, if the clustering bias value is greater than the preset clustering bias threshold, the data diversity of that category is determined to meet the preset diversity conditions; if the clustering bias value is not greater than the preset clustering bias threshold, the data diversity of that category is determined to not meet the preset diversity conditions. The preset diversity conditions can be set according to business needs, for example, based on the classification accuracy and classification type of the data classification model to be trained.
[0072] The preset clustering bias threshold can be set based on the clustering bias values of multiple categories.
[0073]
[0074] Where ρ is the mean clustering bias of multiple categories (i.e., the clustering bias threshold), 1 ≤ α ≤ m. Therefore, the preset clustering bias threshold is used to characterize the average clustering bias value of multiple categories. If ρ θ If the value is greater than ρ, it indicates that category θ meets the preset diversity condition; otherwise, it indicates that category θ does not meet the preset diversity condition. In practice, the clustering bias threshold can also be flexibly set in other ways, and this invention does not limit the specific details.
[0075] In the second implementation, data diversity detection for a category is achieved by comparing the number of data entries in the clustered text data within each category's clusters with the number of data entries in the original text data corresponding to the category's category label. Specifically, when detecting data diversity for each category based on the clustered text data contained in that category's clusters and the original text data corresponding to the category's category label, it is implemented in the following way:
[0076] First, for each category cluster, obtain the number of cluster texts in the clustered text data within that category cluster. Specifically, for a given category cluster, the number of cluster texts refers to the total number of cluster text data entries contained within that category cluster. As mentioned above, Let C1, C2, ..., Cn be the clusters after the i-th (1≤i≤j) clustering. m The total number of samples (i.e., the number of cluster texts) in the cluster containing the i-th seed point in each category. Therefore, Let C1, C2, ..., Cn be the clusters after j-th clustering. m The total number of samples (i.e., the number of cluster texts) in the cluster containing the j seed points in each category (i.e., a category cluster corresponding to that category).
[0077] Then, based on the number of cluster texts in each cluster of that category and the total number of original text data corresponding to the category's category label, the quantity deviation value for that category is calculated. Here, the total number of original text data corresponding to the category label refers to the actual number of original text data entries contained in that category. As mentioned above, the number of samples corresponding to each category are N1, N2, ..., N. m (That is, the total number of categories in the original text data corresponding to the category labels).
[0078] In one implementation, the category quantity deviation value is calculated using the following formula:
[0079]
[0080] Where, γ θN represents the quantity deviation value of category θ. θ The total number of categories in the original text data corresponding to the category label of category θ. The total number of samples (i.e., the number of cluster texts) in the cluster containing the first seed point in the category.
[0081] Finally, based on the comparison between the quantity deviation value and the preset quantity deviation threshold, the data diversity of that category is checked to see if it meets the preset diversity conditions. A larger quantity deviation value indicates a greater difference between the clustering result and the actual category, meaning there are greater differences among the data within that category, thus indicating better diversity; conversely, a smaller value indicates insufficient diversity. Therefore, when the quantity deviation value of a category is greater than the preset quantity deviation threshold, the data diversity of that category is determined to meet the preset diversity conditions. The preset quantity deviation threshold can be flexibly set according to business needs; it can be a fixed value or a variable value.
[0082] In one implementation, the preset quantity deviation threshold is set based on the quantity deviation values of multiple categories:
[0083]
[0084] Where γ is the mean of the overall quantity deviation across multiple categories (i.e., the quantity deviation threshold), 1 ≤ β ≤ m. Therefore, the preset quantity deviation threshold is used to characterize the average of the quantity deviation values across multiple categories. If γ θ If the value is greater than γ, it indicates that category θ meets the preset diversity conditions; conversely, if the value is less than γ, it indicates that category θ does not meet the preset diversity conditions. In practice, the quantity deviation threshold can also be flexibly set in other ways, and this invention does not limit the specific details.
[0085] The two implementation methods described above can be used individually or in combination. When used in combination, the condition is met only if category θ simultaneously satisfies ρ. θ >ρ, and γ θ Only when the value is greater than γ can it be determined that category θ meets the preset diversity conditions.
[0086] In this context, ρ and γ are used to measure the data diversity of the entire labeled dataset. Since ρ and γ are not fixed values but dynamically change with the dataset, they can flexibly adapt to various dataset formats, ensuring more accurate diversity assessments. It should also be noted that in the above formula, abs() represents taking the absolute value. ρ and γ are two different dimensions for measuring data diversity. ρ measures the average distance between samples across all classes in the entire labeled dataset; a larger value indicates better overall diversity, meaning a richer sample set. γ measures the distance between samples within the entire labeled dataset; a larger value indicates better overall diversity, meaning a richer sample set. Correspondingly, the clustering bias value ρ for class θ... θ and the quantity deviation threshold γ θ This is used to measure the data diversity of category θ. Where ρ θ This is used to measure the distance between samples from different classes. A higher value indicates better inter-class diversity, meaning a richer variety of samples within that class. θ It is used to measure the distance between samples of this category and samples of other categories. The larger the value, the better the inter-class diversity, that is, the richer the samples contained in this category.
[0087] Step S260: Identify categories whose data diversity does not meet the preset diversity conditions as categories to be enhanced; perform data enhancement category processing on the text data contained in the categories to be enhanced.
[0088] There may be one or more categories whose diversity does not meet the preset diversity conditions, and the total number of categories to be enhanced must be less than the total number of categories m. In fact, since the threshold for judging diversity is set based on the overall situation of m categories, it is usually possible to select at least one category as the category to be enhanced from the m categories.
[0089] The term "category to be augmented" refers to categories with insufficient diversity and high similarity of text data within those categories. If the original text data containing these categories is directly used as training samples to train a classification model, the accuracy of the model will be affected by the lack of sample diversity. Therefore, data augmentation category processing is required for the text data contained within the categories to be augmented. Specifically, data augmentation category processing is used to perform data augmentation on a specified category to be augmented. Since the augmentation method in this application aims to augment a specific category to be augmented, it is called data augmentation category processing.
[0090] Specifically, when performing data augmentation processing on text data contained in the category to be augmented, it can be achieved in a variety of ways, and the present invention does not limit the specific implementation details.
[0091] In one implementation, the intersection of j clusters of the category to be augmented is obtained, and the semantic similarity between every two text data points contained in this intersection is calculated. Text data with semantic similarity greater than a preset similarity threshold in the intersection is then deleted to reduce the number of text data points in the category to be augmented. Therefore, this method aims to achieve data augmentation by reducing the number of samples, primarily by reducing samples with high similarity. In a specific example, firstly, the clustering result C of the j seed points of class C1 is... 11 C 12 ;…;C 1j Find the intersection to obtain the intersection. Here, we assume the first category is the category to be enhanced. Then, we remove the intersections with the first category in class C1. Identical samples (i.e., if the intersection) (The sample appears multiple times in C1, but only one is retained). Finally, the intersection is calculated. Within the C1 class, if the semantic similarity between any two samples exceeds a certain preset threshold, the sample is removed from the C1 class. The data in the intersection of the j-category clusters of the class to be enhanced typically consists of highly similar data with poor diversity (grouping together in each cluster indicates high similarity). Therefore, removing the highly similar data from the intersection can improve the diversity of that class.
[0092] In another implementation, the union of the j clusters of the category to be enhanced is obtained; text data in the original text data corresponding to the category labels of the category to be enhanced that does not belong to this union is identified as data to be enhanced; data augmentation processing is performed on the data to be enhanced. This method aims to identify dissimilar data with good diversity within a category. Therefore, performing data augmentation processing on this dissimilar data helps to improve the diversity of the category. Specifically, the union of the j clusters of the category to be enhanced contains the text data that was clustered into the category to be enhanced in each clustering process. If one or more text data in the original text data does not belong to this union, it means that one or more text data that do not belong to this union are relatively unique data in the category and have a certain significance. Therefore, performing augmentation processing on one or more text data that do not belong to this union helps to strengthen the differences between the data in the category, thereby improving diversity. In a specific example, for the clustering result C1 of the j seed points of category C1... 11 C 12 ;…;C 1j Take the union of sets to get the union of sets. against (union) Remove intersection The samples in the sample, and (Remove union from data in category C1) The data augmentation is performed on samples from the dataset (in the sample list), increasing the number of samples. For the former, 2-3 samples can be augmented per sample; for the latter, 3-5 samples can be augmented per sample. Data augmentation methods include synonym replacement, back-translation, and random punctuation insertion. Therefore, the data to be augmented is determined based on the union of the original text data corresponding to the category labels of the category to be augmented. This means that the text data that does not belong to the union in the original text data corresponding to the category to be augmented can be identified as the data to be augmented, or the text data that does not belong to the intersection in the union can be identified as the data to be augmented.
[0093] In another implementation, multiple text data points are extracted from m categories excluding the category to be enhanced to form a reference text set. The average text distance between each text data point in the category to be enhanced and the text data in the reference text set is calculated. Based on the average text distance of each text data point, the multiple text data points in the category to be enhanced are divided into at least two text data sets, each corresponding to a different text distance interval. For each text data set, the number of texts to be added to the set is determined according to the corresponding text distance interval. Data augmentation is then performed on the text data sets to add texts corresponding to the number of texts to be added. Specifically, the larger the value of the text distance interval, the smaller the number of texts to be added to the corresponding text data set. Conversely, the smaller the value of the text distance interval, the larger the number of texts to be added to the corresponding text data set. Therefore, this method constructs a reference text set based on categories other than the category to be enhanced. It then calculates the average text distance between each text data point in the category to be enhanced and the text data in the reference text set. Based on this average text distance, it determines the processing method for each text data point in the category to be enhanced, allowing for focused enhancement of text data with large average text distances (i.e., better diversity). For example, this can be achieved as follows:
[0094] First, assuming C1 is the category to be enhanced, from C2, C3, ..., C m t texts are randomly selected from each category, for a total of t*(m-1) texts, forming a sample set Q (i.e., the control text set).
[0095] Then, the average text distance between each text in category C1 and the sample set Q is calculated using the following formula:
[0096] Remove the large summation symbols. Representing text With text Q δ The Jaccard distance; N1 represents the number of texts contained in category C1, and t*(m-1) represents the number of texts contained in sample set Q; Indicates the first in category C1 Sample, Q δ Let represent the δ-th sample in the sample set Q, where 1 ≤ δ ≤ t*(m-1); Indicates sample The average text distance to the sample set Q.
[0097] Next, The values are sorted in descending order (from largest to smallest). Select the largest The value will Corresponding text Divide into ε text combinations, with each text corresponding to... The values are arranged in order within the range For the first text combination, the text corresponding to The values are arranged in order within the range The text is the second text combination, and so on, corresponding to the text. The values are arranged in order within the range Let the text be the ε-th text combination. Assume that each combination contains μ1, μ2, ..., μ... texts respectively. ε .
[0098] Finally, data augmentation is achieved through text retrieval. Text retrieval is performed within the context of the labeled dataset to enrich the data diversity of this category. The number of additional texts required for each of the ε text combinations within this category varies. corresponding The smaller, the better The more text combinations that need to be added, the more text combinations that need to be added. Let's assume ε text combinations require μ text combinations. 1 μ 2 、…、μ ε Among them, μ 1 <μ 2 <…<μ ε The method to enrich the diversity of this category of data is to cluster the text data in the scene using the text contained in each of the above combinations as seed points. Taking the first text combination as an example, the first text combination contains μ1 texts, and the number of texts that need to be added is μ1. 1 Using μ1 texts from the first text combination as seed points, the text data in the scene is clustered, and each cluster is selected from the μ1 clusters of the clustering results. Samples are selected to enrich the data diversity of category C1. Similarly, other text combinations are processed using the same method, selecting samples from the scene data to enrich the data diversity of category C1. For ease of understanding, a specific example is used below to describe a specific application scenario of this embodiment: In related technologies, data diversity issues have received increasing attention in the field of machine learning, becoming a major challenge faced by machine learning models in practical deployment. When the data diversity distribution does not reach the coverage of the actual scene, the data diversity distribution to be predicted and the data diversity distribution used for training show a significant deviation, leading to poor model performance. The poor model performance caused by data diversity issues is difficult to solve by improving the model's generalization ability, because current machine learning methods are basically based on the premise of independent and identically distributed data. Under a real distribution, the observable training data is limited. When a well-trained model encounters samples that conform to the same distribution but are not observed during prediction, the accuracy decreases. In this case, the model's generalization ability can be effectively improved by selecting appropriate algorithms, cross-validation, regularization, etc. However, the essence of the data diversity problem is that the real data distribution differs greatly from the distribution of the actual scene; therefore, simply improving generalization ability cannot effectively improve the model's performance. Currently, industry and academia address these challenges through methods such as: manually constructing, screening, and labeling samples, which is effective but requires significant human and material resources and specialized expertise; and data augmentation, which is simple and easy to implement, solving the resource constraints, but can introduce noise and negatively impact model performance if used improperly. These methods all aim to improve the model's generalization ability by increasing the diversity of training data. However, these methods also have significant drawbacks: they lack scientific analysis of the original data. Some categories in the original data have sufficient diversity and do not require additional data, while others lack diversity and need enrichment. Blindly augmenting data globally wastes resources (time, machine costs, etc.) and can even be counterproductive. Therefore, this example proposes a data diversity detection method to identify which categories in the data lack diversity and then selectively enrich their diversity to improve model performance.
[0099] This example is primarily used in a voice quality inspection system. In this scenario, it's necessary to classify statements containing sensitive words to assess call quality. For instance, the texts "Sir, you owe money and aren't paying it back, and you're going to complain? That doesn't make sense.", "Sir, you can complain, but you still owe money and don't pay it back, so that needs to be resolved.", "Sir, you're the one who owes money and isn't paying it back, and you're going to complain about me? Will complaining about me make you pay it back?", and "Go ahead and complain, it's your right." All three sentences contain the sensitive word "complain," but the first and third sentences are not at fault, while the second and fourth are. In this scenario, the text in each category has the following characteristics: the text length varies; the texts are similar, but the semantics or labels may differ. In other words, in this scenario, the original text data corresponding to m category labels has the following characteristics: the length of each original text data in each category is different, and the semantic labels are not strongly correlated with the similarity between the texts. For example, m category labels belong to the same business scenario or the same business project; and the text data corresponding to each category label has a different length, and the similarity between multiple text data corresponding to different category labels is greater than a preset similarity threshold.
[0100] In this scenario, statistical methods based on words, sentences, and other dimensions cannot objectively reflect the diversity of the data. This invention measures data diversity from the perspective of text vector representation combined with clustering results. Before model training, the business system obtains the corresponding text data for each category from historical call data through sensitive word matching, and assigns corresponding labels to these text data to obtain the original labeled dataset (i.e., the original text data of m categories corresponding to m category labels). Then, statistics are performed on the overall labeled data and the text data of each category, including the total number of samples, the number of categories, and the number of samples in each category. Based on the labeled dataset, a seed point set for each category is constructed. One seed point is selected from the seed point set of each category at a time, and the labeled dataset is clustered. The K value is equal to the number of categories in the labeled dataset. Multiple clustering operations are performed until all seed points in the seed point set are selected. After clustering, the results of multiple clustering operations are statistically analyzed. The statistical data includes: the number of samples in each cluster of each seed point in each category that are identical to samples in the corresponding category of the seed point, and the total number of samples in each cluster of each seed point in each category. Based on the above statistical data and the calculation formula mentioned above, categories with insufficient diversity are identified. For categories with insufficient diversity, various solutions are used to improve or resolve inter-class diversity issues. Finally, the corrected labeled dataset is used to train a model for classifying sentences containing sensitive words.
[0101] In summary, this example achieves at least the following results: By detecting diversity across categories, it accurately identifies categories lacking diversity, avoiding blind model training and increased business cycle time. Furthermore, the diversity judgment rules can assess and determine categories lacking diversity. Moreover, it provides a scientific method for enriching inter-class diversity, improving the model's generalization ability and solving the problems of blindly augmenting data globally, which wastes resources (time costs, machine costs, etc.) and can even have adverse effects (easily introducing noise, causing data pollution; difficulty in controlling the semantics of augmented data; and failure to achieve global optimization). Additionally, this solution features high computational speed and a high degree of process automation.
[0102] In related technologies, methods such as synonym replacement, random insertion, and random deletion are used to modify the original text to achieve data increment, thereby improving the model's generalization ability. However, the above methods have at least the following drawbacks: they cannot accurately locate categories with insufficient inter-class diversity; they are prone to introducing noise, causing data pollution; the semantics of the enlarged sentences are not easy to control, affecting the model's generalization effect; and they cannot be used to specifically increase diversity, which can easily lead to the overall effect not reaching the optimal level.
[0103] The method in this invention detects sample diversity and improves the model's generalization ability by statistically analyzing sample information from labeled datasets and clustered sample information, designing evaluation methods for overall diversity and inter-class diversity, diversity judgment rules, and diversity enrichment methods. The novel methods for calculating overall diversity and inter-class diversity, as well as the diversity judgment method proposed in this invention, can complete sample diversity detection and accurately locate categories with insufficient inter-class diversity. The diversity enrichment scheme proposed in this invention can improve the model's generalization ability and solve the drawbacks of overall data augmentation. This invention features high computational speed and a high degree of automation.
[0104] Therefore, the diversity detection method provided in this application is particularly suitable for diversity detection processing of multiple subcategories within the same business scenario or project. In the same business scenario or project, the text data of multiple subcategories exhibits high data similarity, making it difficult to ensure consistency between the clustering results and the classification results of the subcategories using conventional clustering methods. In such cases, the diversity detection method in this application can fully explore the similarities and differences between the data of each subcategory through multiple clustering operations, thereby providing more accurate diversity detection results.
[0105] For example, in another specific example, the text data diversity detection method in this application is applicable to customer service scenarios, specifically for intent recognition based on user-inputted intent text. Assume there are two category labels: "processing conditions" and "processing procedure".
[0106] The "Application Requirements" category label contains the following text data:
[0107] How long after withdrawing cash should I wait before applying for installment payments?
[0108] I have a bachelor's degree and I want to apply for a certificate. Where can I find the supporting documents?
[0109] Do I need to provide proof of my academic qualifications?
[0110] When is it generally advisable to apply for installment payments?
[0111] How long after withdrawing cash can I apply for installment payments?
[0112] I withdrew cash and the bill has come out, can I still apply for installment payments...?
[0113] In addition, the "Processing Procedure" category label contains the following text data:
[0114] How to apply;
[0115] Hello, how do I apply for installment payments?
[0116] How do I purchase?
[0117] How to apply;
[0118] How do I apply for installment payments?
[0119] How do I apply...?
[0120] In this example, the text lengths in different category labels vary, and the similarity between the texts in different category labels is high, that is, there are a large number of identical words in the texts in different category labels. In this case, it is difficult to accurately assess the diversity of different categories. With the implementation method based on multiple clustering in this application, more targeted diversity detection processing can be carried out by analyzing the similarities and differences of each cluster.
[0121] It is understood that the various method embodiments mentioned above in this application can be combined with each other to form combined embodiments without violating the principle and logic. Due to space limitations, this application will not elaborate further. Those skilled in the art will understand that in the above methods of specific implementation, the specific execution order of each step should be determined by its function and possible internal logic.
[0122] In addition, this application also provides a text data diversity detection device and electronic device, and a computer-readable storage medium. All of the above can be used to implement any of the text data diversity detection methods provided in this application. The corresponding technical solutions and descriptions are described in the corresponding descriptions in the method section, and will not be repeated here.
[0123] Figure 3 This is a schematic diagram of the structure of a text data diversity detection device 30 provided in an embodiment of this application.
[0124] Reference Figure 3 This application provides a text data diversity detection device 30, which includes:
[0125] The acquisition module 31 is adapted to acquire the original text data of m categories corresponding to m category labels, wherein the original text data of the m categories belong to the same business scenario.
[0126] Extraction module 32 is adapted to extract j seed point data from the original text data of each category to obtain j sets of seed points; each set of seed points contains m seed point data, and the m seed point data correspond to m categories respectively; where m and j are natural numbers;
[0127] Clustering module 33 is adapted to perform j clustering processes on the original text data of the m categories based on the m seed point data contained in each seed point set, to obtain j clustering results; wherein each clustering result contains m clusters, and the m clusters correspond to the m categories respectively;
[0128] The determination module 34 is adapted to determine the j clusters corresponding to the j-th clustering results of the category for each category label as the category clusters of the category;
[0129] The detection module 35 is adapted to obtain the data diversity detection result of each category based on the clustered text data contained in the category clusters of the category and the original text data corresponding to the category label of the category.
[0130] Optionally, the detection module is specifically adapted to:
[0131] For each category cluster of the category, obtain the number of consistent data between the clustered text data in the category cluster and the original text data corresponding to the category label of the category;
[0132] The clustering bias value of the category is calculated based on the number of consistent clusters in each category and the total number of categories in the original text data corresponding to the category label.
[0133] Based on the comparison between the clustering deviation value and the preset clustering deviation threshold, it is determined whether the data diversity of the category meets the preset diversity conditions.
[0134] The consistency count is used to characterize the number of identical texts in the clustered text data of the category cluster and the original text data corresponding to the category label of the category.
[0135] Optionally, the detection module is specifically adapted to:
[0136] For each category cluster of the stated category, obtain the number of cluster texts in the clustered text data within that category cluster;
[0137] The quantity deviation value of the category is calculated based on the number of cluster texts in each category cluster and the total number of categories of the original text data corresponding to the category label.
[0138] Based on the comparison between the quantity deviation value and the preset quantity deviation threshold, it is determined whether the data diversity of the category meets the preset diversity conditions.
[0139] Optionally, the detection module is further adapted to:
[0140] Categories whose data diversity does not meet the preset diversity conditions are identified as categories that need to be enhanced;
[0141] Perform data augmentation category processing on the text data contained in the category to be augmented.
[0142] Optionally, the detection module is specifically adapted to:
[0143] Obtain the intersection of j category clusters of the category to be enhanced, and calculate the semantic similarity between every two text data contained in the intersection;
[0144] Delete text data in the intersection that have a semantic similarity greater than a preset similarity threshold to reduce the amount of text data in the category to be enhanced.
[0145] Optionally, the detection module is specifically adapted to:
[0146] Obtain the union of the j-category clusters of the category to be enhanced;
[0147] Based on the union, determine the data to be augmented, and perform data augmentation processing on the data to be augmented;
[0148] The data to be enhanced includes: text data in the original text data corresponding to the category label of the category to be enhanced that does not belong to the union, and / or text data in the union that does not belong to the intersection.
[0149] Optionally, the detection module is specifically adapted to:
[0150] Extract multiple text data from m categories other than the category to be enhanced to form a comparison text set;
[0151] Calculate the average text distance between each text data in the category to be enhanced and the text data in the reference text set;
[0152] Based on the average text distance of each text data item, multiple text data items in the category to be enhanced are divided into at least two text data sets; wherein each text data set corresponds to a different text distance interval;
[0153] For each text data set, the number of texts to be added to the text data set is determined based on the text distance interval corresponding to the text data set. Data augmentation processing is then performed on the text data set to add texts to the text data set corresponding to the number of texts to be added.
[0154] The larger the interval value of the text distance interval, the smaller the number of texts to be added to the corresponding text data set.
[0155] Figure 4 This is a block diagram of an electronic device provided in an embodiment of this application.
[0156] Reference Figure 4 This application provides an electronic device, which includes: at least one processor 501; at least one memory 502; and one or more I / O interfaces 503 connected between the processor 501 and the memory 502; wherein the memory 502 stores one or more computer programs that can be executed by the at least one processor 501, and the one or more computer programs are executed by the at least one processor 501 to perform the above-mentioned text data diversity detection method.
[0157] This application also provides a computer-readable storage medium storing a computer program thereon, wherein the computer program, when executed by a processor / processor core, implements the above-described method for detecting the diversity of text data. The computer-readable storage medium can be volatile or non-volatile.
[0158] This application also provides a computer program product, including computer-readable code, or a non-volatile computer-readable storage medium carrying computer-readable code. When the computer-readable code is run in the processor of an electronic device, the processor in the electronic device executes the above-described text data diversity detection method.
[0159] Those skilled in the art will understand that all or some of the steps, systems, or devices disclosed above, as well as their functional modules / units, can be implemented as software, firmware, hardware, and thereof.
[0160] Appropriate combinations. In hardware implementations, the division between functional modules / units mentioned above does not necessarily correspond to the division of physical components; for example, a physical component may have multiple functions, or a function or step may be performed collaboratively by several physical components. Some or all physical components may be implemented as software executed by a processor, such as a central processing unit, digital signal processor, or microprocessor, or as hardware, or as an integrated circuit, such as an application-specific integrated circuit. Such software may be distributed on a computer-readable storage medium, which may include computer storage media (or non-transitory media) and communication media (or transient media).
[0161] As is known to those skilled in the art, the term computer storage medium includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information (such as computer-readable program instructions, data structures, program modules, or other data). Computer storage media includes, but is not limited to, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), static random access memory (SRAM), flash memory or other memory technologies, portable compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical disc storage, magnetic cartridges, magnetic tape, disk storage or other magnetic storage devices, or any other medium that can be used to store desired information and is accessible to a computer. Furthermore, it is known to those skilled in the art that communication media typically contain computer-readable program instructions, data structures, program modules, or other data in modulated data signals such as carrier waves or other transmission mechanisms, and may include any information delivery medium.
[0162] The computer-readable program instructions described herein can be downloaded from computer-readable storage media to various computing / processing devices, or downloaded via a network, such as the Internet, local area network, wide area network, and / or wireless network, to an external computer or external storage device. The network may include copper transmission cables, fiber optic transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards them to the computer-readable storage media in the respective computing / processing device.
[0163] The computer program instructions used to perform the operations of this application may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages such as Smalltalk, C++, etc., and conventional procedural programming languages such as the "C" language or similar programming languages. The computer-readable program instructions may be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or may be connected to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, electronic circuits, such as programmable logic circuits, field-programmable gate arrays (FPGAs), or programmable logic arrays (PLAs), are personalized by utilizing state information from the computer-readable program instructions. These electronic circuits can execute the computer-readable program instructions to implement various aspects of this application.
[0164] The computer program product described herein can be implemented specifically through hardware, software, or a combination thereof. In one alternative embodiment, the computer program product is specifically embodied in a computer storage medium; in another alternative embodiment, the computer program product is specifically embodied in a software product, such as a software development kit (SDK), etc.
[0165] Various aspects of this application are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0166] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processor of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner; thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.
[0167] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.
[0168] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction containing one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions marked in the blocks may occur in a different order than those marked in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.
[0169] Example embodiments have been disclosed herein, and while specific terminology has been used, it is for general illustrative purposes only and should not be construed as limiting. In some instances, it will be apparent to those skilled in the art that features, characteristics, and / or elements described in conjunction with particular embodiments may be used alone, or in combination with features, characteristics, and / or elements described in conjunction with other embodiments, unless otherwise expressly indicated. Therefore, those skilled in the art will understand that various changes in form and detail may be made without departing from the scope of this application as set forth by the appended claims.
Claims
1. A method of detecting diversity of text data, characterized by, include: Obtain the original text data of m categories corresponding to m category labels, wherein the original text data of the m categories belong to the same business scenario; Extract j seed point data from the original text data of each category to obtain j sets of seed points; each set of seed points contains m seed point data, and the m seed point data correspond to m categories respectively; where m and j are natural numbers; Based on the m seed point data contained in each seed point set, perform j clustering operations on the original text data of the m categories to obtain j clustering results; wherein, each clustering result contains m clusters, and the m clusters correspond to the m categories respectively; For each category label, the j clusters corresponding to the category in the j-th clustering results are determined as the category clusters of the category; For each category, the data diversity detection result of the category is obtained based on the clustered text data contained in the category cluster and the original text data corresponding to the category label of the category.
2. The method of claim 1, wherein, The step of obtaining the data diversity detection result for each category, based on the clustered text data contained in the category clusters of the category and the original text data corresponding to the category label of the category, includes: For each category cluster of the category, obtain the number of consistent data between the clustered text data in the category cluster and the original text data corresponding to the category label of the category; The clustering bias value of the category is calculated based on the number of consistent clusters in each category and the number of samples of the original text data corresponding to the category label of the category. Based on the comparison between the clustering deviation value and the preset clustering deviation threshold, it is determined whether the data diversity of the category meets the preset diversity conditions. The consistency count is used to characterize the number of identical texts in the clustered text data of the category cluster and the original text data corresponding to the category label of the category.
3. The method according to claim 1, characterized in that, The step of obtaining the data diversity detection result for each category, based on the clustered text data contained in the category clusters of the category and the original text data corresponding to the category label of the category, includes: For each category cluster of the stated category, obtain the number of cluster texts in the clustered text data within that category cluster; The quantity deviation value of the category is calculated based on the number of cluster texts in each category cluster and the number of samples of the original text data corresponding to the category label. Based on the comparison between the quantity deviation value and the preset quantity deviation threshold, it is determined whether the data diversity of the category meets the preset diversity conditions.
4. The method according to claim 1, characterized in that, After obtaining the detection results of the data diversity of the aforementioned category, the method further includes: Categories whose data diversity does not meet the preset diversity conditions are identified as categories that need to be enhanced; Perform data augmentation category processing on the text data contained in the category to be augmented.
5. The method according to claim 4, characterized in that, The data augmentation category processing performed on the text data contained in the category to be augmented includes: Obtain the intersection of j category clusters of the category to be enhanced, and calculate the semantic similarity between every two text data contained in the intersection; Delete text data in the intersection that have a semantic similarity greater than a preset similarity threshold to reduce the amount of text data in the category to be enhanced.
6. The method according to claim 5, characterized in that, The data augmentation category processing performed on the text data contained in the category to be augmented includes: Obtain the union of the j-category clusters of the category to be enhanced; Based on the union, determine the data to be enhanced, and perform data enhancement processing on the data to be enhanced; The data to be enhanced includes: text data in the original text data corresponding to the category label of the category to be enhanced that does not belong to the union, and / or text data in the union that does not belong to the intersection.
7. The method according to claim 4, characterized in that, The data augmentation process performed on the text data contained in the category to be augmented includes: Extract multiple text data from m categories other than the category to be enhanced to form a comparison text set; Calculate the average text distance between each text data in the category to be enhanced and the text data in the reference text set; Based on the average text distance of each text data item, multiple text data items in the category to be enhanced are divided into at least two text data sets; wherein each text data set corresponds to a different text distance interval; For each text data set, the number of texts to be added to the text data set is determined based on the text distance interval corresponding to the text data set. Data augmentation processing is then performed on the text data set to add texts to the text data set corresponding to the number of texts to be added. The larger the average text distance, the smaller the number of texts to be added to the corresponding text dataset.
8. A text data diversity detection device, characterized in that, include: The acquisition module is suitable for acquiring the original text data of m categories corresponding to m category labels; The extraction module is suitable for extracting j seed point data from the original text data of each category, resulting in j sets of seed points; Each set of seed points contains m seed point data, and the m seed point data correspond to m categories respectively; where m and j are natural numbers; The clustering module is adapted to perform j clustering processes on the original text data of the m categories based on the m seed point data contained in each seed point set, to obtain j clustering results; wherein each clustering result contains m clusters, and the m clusters correspond to the m categories respectively; The determination module is adapted to determine the j clusters corresponding to the category in the j-th clustering results as the category clusters of the category for each category label; The detection module is adapted to detect the data diversity of each category based on the clustered text data contained in the category clusters of the category and the original text data corresponding to the category label of the category.
9. An electronic device, characterized in that, include: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores one or more computer programs that can be executed by the at least one processor, the one or more computer programs being executed by the at least one processor to enable the at least one processor to perform the method as described in any one of claims 1-7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, The computer program, when executed by a processor, implements the method as described in any one of claims 1-7.