Methods, apparatuses, and electronic devices for text processing

By generating and utilizing category graphs and acetic vectors, the problem of low text classification accuracy in the prior art is solved, and a higher text classification accuracy is achieved.

CN116127061BActive Publication Date: 2025-06-27MASHANG CONSUMER FINANCE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211334354.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-10-28
Publication Date
2025-06-27
Estimated Expiration
2042-10-28

AI Technical Summary

Technical Problem

The existing unsupervised or semi-supervised clustering methods cannot effectively guide the semantic meaning of sentence vectors in text classification, resulting in poor classification accuracy.

Method used

By obtaining text data carrying category tags, dividing them into multiple data sets, a first category map and a second category map are generated, and the first sense vector is determined so as to classify text data without category tags.

Benefits of technology

The classification accuracy of text classification is improved, and the classification basis is used to characterize the semantic information of text and the category diagram representing semantic categories.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116127061B_ABST
    Figure CN116127061B_ABST
Patent Text Reader

Abstract

An embodiment of the present application discloses a method, device, and electronic device for text processing. The method includes: obtaining first text data carrying N category labels and second text data not carrying category labels; dividing the first text data into multiple data sets, the multiple data sets including at least one first data set and at least one second data set, and the first data set and the second data set are non-overlapping; generating M first category graphs corresponding to M category labels by using the data in the first data set; determining K first sememe vectors corresponding to the M category labels by using M quantities, and the first sememe vectors are used to indicate target semantic information for classifying text; generating N second category graphs corresponding to the N category labels by using the data in the multiple data sets, and the semantic category range of the second category graph is larger than the semantic category range of the first category graph; classifying the second text data according to the K first sememe vectors and the N second category graphs.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of natural language processing technology, and in particular, to a method, apparatus, and electronic device for text processing. Background Art

[0002] Text clustering classifies texts according to the fact that similar documents in the same category have a greater similarity, while documents in different categories have a smaller similarity. Unsupervised or semi-supervised methods are often used to complete text clustering because these methods do not require a training process and do not require manual annotation of document categories in advance. Therefore, they have a certain degree of flexibility and a high level of automated processing ability, and have become an important means for effectively organizing texts.

[0003] In related technologies, the unsupervised or semi-supervised clustering method essentially determines whether sentence vectors belong to the same category by calculating the distance between sentence vectors, and it cannot guide what specific meanings the sentence vectors express. As a result, the classification accuracy after text classification is poor. Summary of the Invention

[0004] This application provides a method, apparatus, and electronic device for text processing to improve the accuracy of text classification.

[0005] In a first aspect, this application provides a method for text processing, including: obtaining first text data carrying N category labels and second text data not carrying category labels, where N is an integer greater than 1; dividing the first text data into multiple data sets, the multiple data sets including at least one first data set and at least one second data set, and the first data set and the second data set do not intersect; generating M first category graphs corresponding to M category labels by using the data in the first data set, the first data set including text data of the M category labels, and the category graph is used to represent the semantic category represented by the text data of its corresponding category label, where M is an integer less than or equal to N; determining K first sememe vectors corresponding to the M category labels by using M quantities, the M quantities being the quantities of the data in the second data set in the M first category graphs respectively, and the first sememe vector is used to indicate the target semantic information for text classification, where K is a positive integer; generating N second category graphs corresponding to the N category labels by using the data in the multiple data sets, and the semantic category range of the second category graph is larger than the semantic category range of the first category graph; classifying the second text data according to the K first sememe vectors and the N second category graphs.

[0006] Second aspect, the present application provides an apparatus for text processing, including: an acquisition module, configured to acquire first text data carrying N category labels and second text data not carrying category labels, where N is an integer greater than 1; a division module, configured to divide the first text data into multiple data sets, the multiple data sets including at least one first data set and at least one second data set, and the first data set and the second data set are non-overlapping; a generation module, configured to generate M first category graphs corresponding to M category labels by using the data in the first data set, the first data set including text data of the M category labels, and the category graph is used to represent the semantic category represented by the text data of its corresponding category label, where M is an integer less than or equal to N; a determination module, configured to determine K first sememe vectors corresponding to the M category labels by using M quantities, the M quantities being the quantities of the data in the second data set within the M first category graphs respectively, and the first sememe vector is used to indicate the target semantic information for text classification, where K is a positive integer; the generation module is further configured to generate N second category graphs corresponding to the N category labels by using the data in the multiple data sets, and the range of the semantic category of the second category graph is greater than the range of the semantic category of the first category graph; a processing module, configured to perform classification processing on the second text data according to the K first sememe vectors and the N second category graphs.

[0007] Third aspect, the present application provides an electronic device, including: a processor; a memory for storing executable instructions of the processor; wherein, the processor is configured to execute the instructions to implement the method as described in the first aspect.

[0008] Fourth aspect, the present application provides a computer-readable storage medium, when the instructions in the storage medium are executed by a processor of an electronic device, enabling the electronic device to execute the method as described in the first aspect.

[0009] It can be seen that by obtaining the first text data carrying N category labels and the second text data not carrying category labels, where N is an integer greater than 1; dividing the first text data into multiple data sets, the multiple data sets include at least one first data set and at least one second data set, and the first data set and the second data set do not intersect; generating M first category graphs corresponding to M category labels by using the data in the first data set, the first data set includes the text data of M category labels, and the category graph is used to represent the semantic category represented by the text data of its corresponding category label, where M is an integer less than or equal to N; determining K first sememe vectors corresponding to M category labels by using M quantities, the M quantities are the quantities of the data in the second data set within the M first category graphs respectively, and the first sememe vector is used to indicate the target semantic information for text classification, where K is a positive integer; generating N second category graphs corresponding to N category labels by using the data in the multiple data sets, and the semantic category range of the second category graph is larger than the semantic category range of the first category graph; classifying the second text data according to the K first sememe vectors and the N second category graphs. In this way, in the embodiments of the present application, the sememe vectors representing the semantic information of the text are determined through the first category graph, and when classifying the second text data to be classified, the sememe vectors representing the semantic information of the text and the category graph representing the semantic category are used as the basis for text classification, thereby improving the classification accuracy of text classification. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] The drawings described herein are used to provide a further understanding of the present specification and form a part of the present specification. The illustrative embodiments of the present specification and their descriptions are used to explain the present specification and do not constitute an improper limitation to the present specification. In the drawings:

[0011] Figure 1 is a schematic flowchart of a text processing method provided by an embodiment of the present application;

[0012] Figure 2 is a schematic flowchart of another text processing method provided by an embodiment of the present application;

[0013] Figure 3 is a schematic structural diagram of a text processing device provided by an embodiment of the present application;

[0014] Figure 4 is a schematic structural diagram of an electronic device provided by an embodiment of the present specification. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0015] To make the objectives, technical solutions, and advantages of this specification clearer, the technical solutions of this specification will be clearly and completely described below in conjunction with specific embodiments of this specification and the corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this specification, rather than all of them. Based on the embodiments in this specification, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of this application.

[0016] The terms "first", "second", etc. in this specification and the claims are used to distinguish similar objects, rather than to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of this application can be implemented in an order other than those illustrated or described herein. In addition, the "and / or" in this specification and the claims indicates at least one of the connected objects, and the character " / " generally indicates an "or" relationship between the associated objects before and after.

[0017] As mentioned above, the unsupervised or semi-supervised clustering method essentially determines whether sentence vectors belong to the same class by calculating the distances between sentence vectors. In practical applications, whether texts can be classified into the same class is related to a certain common semantic feature expressed in the texts. However, the unsupervised or semi-supervised methods in the related art cannot guide what specific meanings the sentence vectors express. Thus, the unsupervised or semi-supervised methods have no interpretability in terms of semantic distance, and the classification accuracy after classifying texts is relatively poor.

[0018] To improve the classification accuracy of text classification. Embodiments of the present application aim to provide a text processing solution, which includes: obtaining first text data carrying N category labels and second text data without category labels, where N is an integer greater than 1; dividing the first text data into multiple data sets, the multiple data sets including at least one first data set and at least one second data set, and the first data set and the second data set do not intersect; generating M first category graphs corresponding to M category labels using the data in the first data set, the first data set including text data of M category labels, and the category graph is used to represent the semantic category represented by the text data corresponding to its category label, where M is an integer less than or equal to N; determining K first sememe vectors corresponding to M category labels using M quantities, the M quantities being the quantities of the data in the second data set in the M first category graphs respectively, and the first sememe vector is used to indicate the target semantic information for classifying the text, where K is a positive integer; generating N second category graphs corresponding to N category labels using the data in the multiple data sets, and the semantic category range of the second category graph is greater than the semantic category range of the first category graph; classifying the second text data according to the K first sememe vectors and the N second category graphs. In this way, in the embodiments of the present application, the sememe vectors representing the semantic information of the text are determined through the first category graph. When classifying the second text data to be classified, the sememe vectors representing the semantic information of the text and the category graph representing the semantic category are used as the basis for text classification, thereby improving the classification accuracy of text classification.

[0019] It should be understood that the text processing methods provided in the embodiments of the present application can all be executed by an electronic device or software installed in the electronic device, and can specifically be executed by a terminal device or a server device. Among them, the text processing method can be executed by the same electronic device, or can also be executed by different electronic devices.

[0020] The following will describe in detail the technical solutions provided in each embodiment of this specification with reference to the accompanying drawings.

[0021] Please refer to Figure 1 , which is a schematic flowchart of a text processing method provided in an embodiment of this specification, applied to an electronic device. The method may include:

[0022] Step S101, obtaining first text data carrying N category labels and second text data without category labels. Among them, N is an integer greater than 1.

[0023] Specifically, the category labels can be classification labels manually marked for the first text data. The first text data is labeled data, which carries a label representing its category, and its format can be (the first text data, label). The types of category labels can vary according to different classification tasks. Here, N represents the types of category labels. For example, for an emotion classification task, the types of its category labels can be positive labels, neutral labels, and negative labels, and the corresponding value of N is 3. The second text data is unlabeled data, which only has a large amount of text data. For the first text data and the second text data, according to different task classification requirements, text data that meets the classification requirements of this type of task can be obtained from different data sources.

[0024] Step S103, divide the first text data into multiple data sets.

[0025] Among them, the multiple data sets include at least one first data set and at least one second data set, and the first data set and the second data set do not intersect.

[0026] Specifically, for the first text data, different text data can carry different category labels. For a large amount of the first text data, it can be divided into multiple data sets. For each data set, the amount of the first text data with the same category label included in each data set can be the same, or the amount of the first text data with various category labels included in each data set can be the same proportion. More specifically, first count the different category labels of the first text data, then count the amount of the first text data with different category labels, and then predetermine the number of data sets to be divided. The amount of the first text data with each category label divided in each data set is the ratio of the total amount of the first text data with that category label to the number of data sets. For example, taking the first text data with 3 category labels label1, label2, and label3 as an example, the amounts of the 3 category labels are label1 = m1, label2 = m2, and label3 = m3 respectively. Assuming that it is to be divided into n data sets, randomly process the data with different category labels respectively, and then divide the first text data with each category label according to each data set. The amount of the first text data with different category labels included in each data set is label1 = m1 / n, label2 = m2 / n, and label3 = m3 / n respectively. Among them, the multiple divided data sets are divided into a first data set and a second data set. The data in the first data set is used to determine the first category graph, and the data in the second data set is used to determine the important semantic vector that has an important impact on text classification.

[0027] In a possible implementation, dividing the first text data into multiple data sets may be dividing the first text data into two data sets. Among them, the data in the first data set of the two data sets is used to generate the first category graph, and the data in the second data set of the two data sets is used to determine the first sememe vector. Among them, the amount of data of the first text data corresponding to the same category label included in the two data sets is the same. In this way, dividing the first text data into two data sets, one data set is used to generate the category graph, and the other part is used to determine the sememe vector, which not only simplifies the process of dividing the data set, but also divides it into two data sets, and the use of the data in each data set is more targeted, the data set division efficiency is high, the data use efficiency is high, and the amount of data in the two data sets remains the same, avoiding the problem of insufficient data volume in a certain data set.

[0028] Step S105: Use the data in the first data set to generate M first category graphs corresponding to M category labels.

[0029] Among them, the first data set includes the text data of M category labels, and the first category graph is used to represent the semantic category represented by the text data of its corresponding category label, and M is an integer less than or equal to N.

[0030] Specifically, the first category graph refers to the category representing the semantics of a certain category label, which is a sememe vector used to distinguish the degree of important influence on text classification. For example, for category labels label1, label2, label3, they each have their own category graph. When generating the category graphs of each category label, the category graph can be generated according to the sememe vectors of each text sentence in the first text data corresponding to each category label in the first data set. Among them, the type of category labels in the first data set may be the same as the type of category labels of the first text data, and the type of category labels in the first data set may also be less than the type of category labels of the first text data. In order to ensure the accuracy and integrity of the subsequent unlabeled data, in the embodiments of the present application, the type M of category labels in the first data set is the same as the type N of category labels of the first text data.

[0031] In a possible implementation, as Figure 2 shown, the specific implementation of using the data in data set i to generate the category graph corresponding to category label j in step S105 is: where data set i is the first data set, and category label j is one of the M category labels in the first data set.

[0032] Step S201: Segment the text data corresponding to category label j to obtain X first words.

[0033] Among them, X first words carry the sentence identifier of the text sentence they belong to and the position information in the text sentence they belong to, where X is a positive integer.

[0034] Specifically, for tokenizing the first text data in the first dataset, a pre-trained language representation model (Bidirectional Encoder Representation from Transformers, BERT) can be used to tokenize the first text data, splitting each text sentence in the first text data into multiple first words (word vectors). After using the BERT model to split each text sentence into multiple first words, for each first word, it carries the sentence identifier segmentation embedding of the text sentence it is in and the position information position embedding of its position in that text sentence.

[0035] Step S203, obtain the initial word vectors of the X first words through a word vector table, and obtain the second sememe vectors of the X first words through a sememe vector table.

[0036] Specifically, the word vector table is obtained by training the BERT model on a large amount of text data, enabling the BERT model to learn each word, obtaining word vectors and generating a word vector table. The word vector table stores multiple words and the corresponding word vectors for each word. In actual use, the initial word vectors word_embedding corresponding to the X first words can be found from the word vector table.

[0037] The sememe vector table is generated by a sememe vector production framework. The sememe vector production framework generates various sememe vectors and stores them in the sememe vector table. The second sememe vector semely_word_embeddin represents the semantic features of the first word.

[0038] Step S205, determine the final word vector of each first word according to the initial word vector, the second sememe vector, the sentence identifier, and the position information of each first word.

[0039] Specifically, perform vector addition on the sentence identification segmentation embedding, position information position embedding, initialized word vector word_embedding, and the second sememe vector semely_word_embedding obtained above to obtain the final word vector embedding_A for each first word, that is, embedding_A = word_embedding + semely_word_embedding + segmentation embedding + position embedding.

[0040] Step S207: Concatenate the final word vectors associated with each text sentence in the text data of dataset i to obtain the sentence vector for each text sentence.

[0041] Specifically, after obtaining the final word vectors of each first word in each text sentence, concatenate the final word vectors of each first word in each text sentence according to the semantic relationship before and after to obtain the sentence vector for each text sentence. Among them, in order to take into account the semantic interaction and restriction relationship between the final word vectors of the text sentence, so that the vector representation of the final word vector is more comprehensive and conforms to the context, multi-head additive attention can be used to perform self-attention on the embedding of the final word vector to obtain the lexical representation that takes into account the semantic relationship in this text sentence, that is, obtain the standard that takes into account the semantic relationship in this text sentence, and the obtained final word vector takes into account the semantic relationship in this text sentence.

[0042] Step S209: Generate a category graph corresponding to category label j according to the sentence vector corresponding to category label j.

[0043] Specifically, when drawing the category graph, the sentence vector of any text sentence can be used as the center point, and the first category graph can be generated according to the distance between the sentence vectors of the remaining text sentences and this center point, or the center point can be determined according to a predetermined rule, and the first category graph can be generated according to the distance between the sentence vectors of the remaining text sentences and this center point.

[0044] In a possible implementation, generating a category graph corresponding to the category label j according to the sentence vector corresponding to the category label j includes: determining the sum of distances associated with each text sentence corresponding to the category label j, where the sum of distances is used to represent the sum of distances from the sentence vector of the associated text sentence to the sentence vectors of other text sentences corresponding to the category label j; using the sentence vector of the text sentence corresponding to the minimum sum of distances among the sums of distances as the center point corresponding to the category label j; determining the spatial positions of multiple points in the coordinate system corresponding to the center point, where the multiple points are points formed by the sentence vectors of other text sentences in the first dataset except the text sentence corresponding to the minimum sum of distances; and connecting the multiple spatial positions in sequence to generate the category graph corresponding to the category label j.

[0045] For example, assume that there are text data of three category labels in the first dataset, namely label1, label2, and label3. Taking the category label label1 as an example, assume that there are 10 text data. Calculate the sum of distances between the sentence vector of each text data and the sentence vectors of the other 9 text data. Then the sums of distances of the 10 text data are denoted as sum1, sum2, sum3, sum4, …… sum10 respectively. Use the data point data3 corresponding to the minimum value among sum1 to sum10 (such as sum3) as the center point of this category label. Then determine the spatial positions of the remaining 9 data relative to the data point data3, and connect the spatial positions where the remaining 9 data are located in sequence to form the first category graph, and the first category graph can be a polygon. In this way, each word vector in the sentence vector of each text sentence corresponds to a sememe vector, and each sentence vector of each text sentence also corresponds to a sememe vector. The first category graph generated according to the sentence vector of each text sentence incorporates more semantic features, so that the important sememe vectors that have an important influence on text classification can be determined according to the first category graph, making the result of text classification more accurate.

[0046] Step S107: Use M quantities to determine K first sememe vectors corresponding to M category labels.

[0047] Wherein, the M quantities are the quantities of the data in the second dataset within the M first category graphs respectively. The first sememe vector is used to indicate the target semantic information for text classification, and K is a positive integer.

[0048] Wherein, the first sememe vector represents the semantic information that has an important influence on text classification.

[0049] For text classification, there are semantic vectors that have an important impact on text classification and semantic vectors that have a relatively small impact on text classification. After determining the first category graph for each category label, the important semantic vectors that have an important impact on text classification are then determined, facilitating accurate text classification of unlabeled data by combining the important semantic vectors. When determining the important semantic vectors, the number of the first text data in another part of the dataset that falls within the first category graph is used to determine the semantic vectors that have an important impact on text classification. If the number of data falling within the first category graph is large, it indicates that this semantic vector has an important impact on the text classification of the category label corresponding to the first category graph.

[0050] In a possible implementation, the determination method for the M quantities is as follows: perform word segmentation on the text data corresponding to the M category labels in the second dataset to obtain multiple second words, where the multiple second words carry the sentence identifiers of the text sentences they belong to and the position information in the text sentences they belong to; respectively obtain the initial word vectors of the multiple second words through a word vector table, and obtain the third semantic vectors of the multiple second words through a semantic vector table; determine the final word vector of each second word according to the initial word vector, third semantic vector, sentence identifier, and position information of each second word; splice the final word vectors associated with each text sentence of the text data in the second dataset to obtain the sentence vector of each text sentence; determine the distances from multiple points to the center points within the M first category graphs respectively, and determine the M quantities of the sentence vectors of each text sentence in the second dataset within the M first category graphs according to the distances, where the multiple points are the points formed by the sentence vectors corresponding to the M category labels in the second dataset.

[0051] Among them, the method for determining whether the data vectors in the second dataset fall within the first category graph can be: when the distance is not greater than the first threshold, it indicates that the sentence vector corresponding to the distance is within the first category graph; the specific implementation method for determining the K first semantic vectors within the first category graph according to the M quantities is: when the M quantity of the target sentence vector within the first category graph is greater than the second threshold, determine the third semantic vector corresponding to the target sentence vector as the first semantic vector, and the target sentence vector is the sentence vector containing the same third semantic vector.

[0052] Specifically, the process of generating the sentence vectors of each text sentence in the first text data in the second dataset in the embodiments of the present application has the same or similar implementation as the process of generating the sentence vectors of each text sentence in the first text dataset in the first dataset described above, and they can be referred to each other. The embodiments of the present application will not elaborate here. When determining the first sememe vector, it can be determined according to the distance between each sentence vector in the second dataset and the center point within each first category graph of each class. According to the distance, it is judged whether the first text data in the second dataset falls inside or outside the first category graph (that is, when the distance is not greater than the first threshold, it falls inside the first category graph). According to the number of sentence vectors corresponding to the first text data in the second sub-dataset that fall inside the first category graph, it is judged whether the same sememe vector corresponding to these sentence vectors is an important sememe or a non-important sememe for text classification. If the number of sentence vectors corresponding to the same sememe vector that fall inside the first category graph is relatively large (that is, greater than the second threshold), then it can be determined that this sememe vector is an important sememe corresponding to the category label, otherwise, it is a non-important sememe. Among them, an important sememe refers to a sememe that has an important indicative effect on the text category of text classification, that is, a sememe that needs to be considered key during classification. Among them, when calculating the distance between each sentence vector in the second dataset and the center point within each first category graph of each class, the similarity calculation of sentence vectors can be used, that is, the dot product between sentence vectors, and the similarity of sentence vectors is the distance between two sentence vectors. It should be noted that the first threshold and the second threshold can be determined according to the actual situation, and the embodiments of the present application do not limit this here.

[0053] In this way, by using the distance to determine whether the sentence vector falls inside the first category graph, the accuracy of the determined important sememe vector is improved. If the number of sentence vectors corresponding to the same sememe vector that fall inside the first category graph is relatively large, it indicates that this sememe vector has an important influence on text classification. By using the first category graph representing the semantic range to determine the important sememe vector with an important influence degree on text classification, the accuracy of the determined important sememe vector is further improved. When classifying the second text data to be classified, the important sememe vector representing the semantic information of the text and the category graph representing the semantic category are used as the basis for text classification, thereby further improving the classification accuracy of text classification.

[0054] Step S109: Generate N second category graphs corresponding to N category labels by using the data in multiple datasets.

[0055] Among them, the second category graph is used to represent the semantic category represented by the text data corresponding to its corresponding category label, and the semantic category range of the second category graph is larger than the semantic category range of the first category graph.

[0056] Specifically, the second category graph refers to the category representing the semantics of a certain type of label, which is used to classify the second text data in the above steps. Since the second category graph is a semantic category graph drawn based on the first text data of all current N category labels, the size of its semantic category graph should be larger than the semantic category of the first category graph.

[0057] In a possible implementation, the category graph corresponding to the category label j is generated using the data in the data set i. If the data set i is multiple data sets, then the category label j is one of the N category labels. The specific implementation of step S109 is as follows: The text data corresponding to the category label j is tokenized to obtain X first words. The X first words carry the sentence identifier of the text sentence they belong to and the position information in the text sentence they belong to, where X is a positive integer. The initial word vectors of the X first words are obtained through the word vector table respectively, and the second semantic vectors of the X first words are obtained through the sememe vector table respectively. According to the initial word vector, the second semantic vector, the sentence identifier, and the position information of each first word, the final word vector of each first word is determined. The final word vectors associated with each text sentence of the text data in the data set i are concatenated to obtain the sentence vector of each text sentence. According to the sentence vector corresponding to the category label j, the category graph corresponding to the category label j is generated. By generating the second category graph using the data in multiple data sets, the semantic range of the second category graph is wider than that of the first category graph, which can further make the classification result of the unlabeled data more accurate.

[0058] Specifically, the process of generating the second category graph has the same or similar implementation as the process of generating the above first category graph, and they can be referred to each other. This application embodiment will not elaborate here.

[0059] Step S111, classify the second text data according to the K first semantic vectors and the N second category graphs.

[0060] Specifically, the second text data as unlabeled data is used to generate the sentence vector of each text sentence of the second text data in the same way as the sentence vector of the first text data above. Then, the sentence vector of each text sentence in the second text data is matched with its corresponding semantic vector, and then grouped according to different semantic vectors. It should be noted that the same text sentence may have multiple semantic vectors. Therefore, the same text sentence may be divided into different groups.

[0061] In a possible implementation, step S109 includes: obtaining the sentence vectors of each text sentence in the second text data; calculating the average distance of the sentence vectors of each text sentence in the second text data on K first sememe vectors; when the average distance is not greater than a third threshold, determining that the text sentences with a distance not greater than the third threshold belong to the text category corresponding to the second category graph.

[0062] Specifically, for the second category graph corresponding to a certain category label, it has K important sememe vectors (first sememe vectors) that have an important influence on classification. Calculate the distance between the sentence vector of each unlabeled data and each important sememe vector, and average the distances of each unlabeled data on each important sememe vector. If the obtained average distance is less than or equal to a pre-set threshold, it is determined that the unlabeled data falls within the second category graph of this category; otherwise, it is determined that it falls outside the second category graph. Those that fall within the second category graph belong to the category corresponding to this category graph, and those that fall outside the second category graph do not belong to the category corresponding to this category graph. If it falls within two second category graphs at the same time, the text data can be defined as a confused category, which can be returned to manual annotation or, by setting a distance threshold, forced to determine that the confused data belongs to a certain category or directly determined as a dual-category.

[0063] Further, the implementation manner of the sentence vector of each text sentence in the second text data is the same as or similar to the implementation manner of the sentence vector of each text sentence in the above-mentioned first text data, and they can be referred to each other. This application embodiment will not elaborate here. To calculate the distance of the sentence vector of each text sentence in the second text data on K first sememe vectors respectively, the similarity calculation of vectors can be used, that is, vector dot product. The obtained vector similarity is the distance of the sentence vector of each text sentence on the first sememe vector. When there are multiple first sememe vectors, calculate the average value of the distances of the sentence vector of each text sentence on multiple first sememe vectors. When the average value of the distances is not greater than a third threshold, determine that the text sentences with an average distance not greater than the third threshold belong to the text category corresponding to the second category graph. It should be noted that the third threshold can be determined according to the actual situation, and this application embodiment does not limit it here.

[0064] Through the technical solution provided by this application embodiment, in this application embodiment, the sememe vectors representing the semantic information of the text are determined through the first category graph. When classifying the second text data to be classified, the sememe vectors representing the semantic information of the text and the second category graph representing the semantic category are used as the basis for text classification, thereby improving the classification accuracy of text classification.

[0065] In addition, corresponding to the above Figure 1 shown text processing method, this application embodiment also provides a text processing device. Figure 3It is a schematic structural diagram of a text processing device 300 provided by an embodiment of the present application, including: an acquisition module 301, configured to acquire first text data carrying N category labels and second text data not carrying category labels, where N is an integer greater than 1; a division module 302, configured to divide the first text data into multiple data sets, the multiple data sets including at least one first data set and at least one second data set, and the first data set and the second data set do not intersect; a generation module 303, configured to generate M first category diagrams corresponding to M category labels by using the data in the first data set, the first data set including text data of M category labels, and the category diagram is used to represent the semantic category represented by the text data corresponding to its corresponding category label, where M is an integer less than or equal to N; a determination module 304, configured to determine K first sememe vectors corresponding to M category labels by using M quantities, the M quantities being the quantities of the data in the second data set in the M first category diagrams respectively, and the first sememe vector is used to indicate the target semantic information for classifying the text, where K is a positive integer; the generation module 303 is further configured to generate N second category diagrams corresponding to N category labels by using the data in the multiple data sets, and the semantic category range of the second category diagram is greater than the semantic category range of the first category diagram; a processing module 305, configured to classify the second text data according to the K first sememe vectors and the N second category diagrams.

[0066] The text processing device provided by the embodiment of the present application determines the sememe vector representing the semantic information of the text through the first category diagram. When classifying the second text data to be classified, the sememe vector representing the semantic information of the text and the category diagram representing the semantic category are used as the basis for text classification, thereby improving the classification accuracy of text classification.

[0067] In a possible implementation manner, the generation module 303 is further configured to perform word segmentation on the text data corresponding to the category label j to obtain X first words, the X first words carrying the sentence identifier of the text sentence where they are located and the position information in the text sentence where they are located, where X is a positive integer; respectively obtain the initial word vectors of the X first words through a word vector table, and respectively obtain the second sememe vectors of the X first words through a sememe vector table; determine the final word vector of each first word according to the initial word vector, the second sememe vector, the sentence identifier and the position information of each first word; splice the final word vectors associated with each text sentence of the text data in the data set i to obtain the sentence vector of each text sentence; generate a category diagram corresponding to the category label j according to the sentence vector corresponding to the category label j; where, if the data set i is the first data set, the category label j is one of the M category labels; if the data set i is multiple data sets, the category label j is one of the N category labels.

[0068] In a possible implementation, the generating module 303 is further configured to determine the sum of distances associated with each text sentence corresponding to the category label j, where the sum of distances is used to represent the sum of the distances from the sentence vector of the associated text sentence to the sentence vectors of other text sentences corresponding to the category label j; use the sentence vector of the text sentence corresponding to the minimum sum of distances among the sums of distances as the center point corresponding to the category label j; determine the spatial positions of multiple points in the coordinate system corresponding to the center point, where the multiple points are points formed by the sentence vectors of other text sentences in the first dataset except for the text sentence corresponding to the minimum sum of distances; connect the multiple spatial positions in sequence to generate a category graph corresponding to the category label j.

[0069] In a possible implementation, the determining module 304 is further configured to perform word segmentation on the text data corresponding to M category labels in the second dataset respectively to obtain multiple second words, where the multiple second words carry the sentence identifier of the text sentence they belong to and the position information in the text sentence they belong to; respectively obtain the initial word vectors of the multiple second words through the word vector table, and obtain the third semantic element vectors of the multiple second words through the semantic element vector table; determine the final word vector of each second word according to the initial word vector, the third semantic element vector, the sentence identifier and the position information of each second word; splice the final word vectors associated with each text sentence of the text data in the second dataset to obtain the sentence vector of each text sentence; determine the distances from the multiple points to the center points within the M first category graphs respectively, and determine the M quantities of the sentence vectors of each text sentence in the second dataset within the M first category graphs according to the distances, where the multiple points are points formed by the sentence vectors corresponding to the M category labels in the second dataset.

[0070] In a possible implementation, the determining module 304 is further configured to indicate that the sentence vector corresponding to the distance is within the first category graph when the distance is not greater than the first threshold; the specific implementation manner of determining the K first semantic element vectors within the first category graph includes: when the M quantity of the target sentence vector within the first category graph is greater than the second threshold, determine the third semantic element vector corresponding to the target sentence vector as the first semantic element vector, where the target sentence vector is a sentence vector containing the same third semantic element vector.

[0071] In a possible implementation, the processing module 305 is further configured to obtain the sentence vector of each text sentence in the second text data; calculate the average distance between the sentence vector of each text sentence in the second text data and the K first semantic element vectors; when the average distance is not greater than the third threshold, determine that the text sentence with the average distance not greater than the third threshold belongs to the text category corresponding to the second category graph.

[0072] In a possible implementation, the partitioning module 302 is further configured to partition the first text data into two data sets. Among them, the data in the first data set of the two data sets is used to generate the first category graph, and the data in the second data set of the two data sets is used to determine the first sememe vector. The data volumes of the data corresponding to the same category label in the first data set and the second data set are the same.

[0073] Obviously, the text processing device disclosed in the embodiments of the present application can be used as the execution subject of the text processing method shown in the above embodiments, and thus can implement the functions realized by the text processing method in the above embodiments. Since the principles are the same, they will not be elaborated here.

[0074] Figure 4 is a schematic structural diagram of an electronic device according to an embodiment of this specification. Please refer to Figure 4 , at the hardware level, the electronic device includes a processor, and optionally also includes an internal bus, a network interface, and a memory. Among them, the memory may include a memory, such as a high-speed random access memory (Random-Access Memory, RAM), and may also include a non-volatile memory, such as at least one disk memory, etc. Of course, the electronic device may also include other hardware required for other services.

[0075] The processor, network interface, and memory can be interconnected through an internal bus, and the internal bus can be an ISA (Industry Standard Architecture) bus, a PCI (Peripheral Component Interconnect) bus, or an EISA (Extended Industry Standard Architecture) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of representation, Figure 4 only a bidirectional arrow is used in [the figure] to represent it, but it does not mean that there is only one bus or one type of bus.

[0076] The memory is used to store programs. Specifically, the program may include program code, and the program code includes computer operation instructions. The memory may include a memory and a non-volatile memory, and provide instructions and data to the processor.

[0077] The processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it, forming a text processing device at the logical level. The processor executes the program stored in the memory and is specifically configured to execute the text processing method mentioned in any of the above method embodiments.

[0078] The method executed by the text processing device disclosed in the above embodiments as described in this specification Figures 1 to 2 can be applied to or implemented by a processor. The processor may be an integrated circuit chip with signal processing capabilities. During implementation, the steps of the above method can be completed by the integrated logic circuit in the hardware of the processor or instructions in software form. The above-mentioned processor may be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it may also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as being executed and completed by the hardware decoding processor, or executed and completed by a combination of the hardware and software modules in the decoding processor. The software module may be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps of the above method.

[0079] It should be understood that the electronic device in the embodiments of the present application can implement the functions of the text processing device in Figures 1 to 2 the embodiments shown. Since the principles are the same, they will not be elaborated herein in the embodiments of the present application.

[0080] Of course, in addition to the software implementation method, the electronic device in this specification does not exclude other implementation methods, such as a logic device or a combination of software and hardware, etc. That is to say, the execution subject of the following processing flow is not limited to each logic unit, and can also be hardware or a logic device.

[0081] The embodiments of the present application also propose a computer-readable storage medium. The computer-readable storage medium stores one or more programs. The one or more programs include instructions that, when executed by a portable electronic device including a plurality of application programs, can enable the portable electronic device to execute the text processing method in any of the above embodiments.

[0082] The above description is of specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the acts or steps recited in the claims may be performed in a different order than in the embodiments and still achieve the desired result. Additionally, the processes depicted in the figures do not necessarily require the particular order or sequential order shown to achieve the desired result. In certain implementations, multitasking and parallel processing are also possible or may be advantageous.

[0083] In summary, the above are only the preferred embodiments of this specification and are not intended to limit the scope of protection of this specification. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of this specification shall be included within the scope of protection of this specification.

[0084] The systems, devices, modules, or units illustrated in the above embodiments may be specifically implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, the computer may be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or any combination of these devices.

[0085] Computer-readable media includes both permanent and non-permanent, removable and non-removable media and can be implemented by any method or technology for storing information. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices, or any other non-transitory media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media, such as modulated data signals and carrier waves.

[0086] It should also be noted that the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, commodity or device comprising a series of elements not only includes those elements but also other elements not expressly listed, or elements inherent to such process, method, commodity or device. Without further limitation, an element defined by the phrase "comprising an..." does not exclude the presence of additional identical elements in the process, method, commodity or device comprising the element.

[0087] Each embodiment in this specification is described in a progressive manner, and for the parts that are the same or similar among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and for the relevant parts, reference can be made to the description of the method embodiment.

Claims

1. A method for text processing, characterized in that, Including: Obtaining first text data carrying N category labels and second text data not carrying category labels, where N is an integer greater than 1; Dividing the first text data into multiple data sets, the multiple data sets including at least one first data set and at least one second data set, and the first data set and the second data set are non-overlapping; Generating M first category graphs corresponding to M category labels by using the data in the first data set, the first data set including the text data of the M category labels, and the category graph being used to represent the semantic category represented by the text data of its corresponding category label, where M is an integer less than or equal to N; Determining K first sememe vectors corresponding to the M category labels by using M quantities, the M quantities being the quantities of the data in the second data set within the M first category graphs respectively, and the first sememe vector being used to indicate the target semantic information for text classification, where K is a positive integer; Generating N second category graphs corresponding to the N category labels by using the data in the multiple data sets, and the range of the semantic category of the second category graph is greater than the range of the semantic category of the first category graph; Classifying the second text data according to the K first sememe vectors and the N second category graphs.

2. The method for text processing according to claim 1, characterized in that, The specific implementation manner of generating a category graph corresponding to category label j by using the data in data set i is as follows: Performing word segmentation on the text data corresponding to category label j to obtain X first words, the X first words carrying the sentence identifier of the text sentence where they are located and the position information in the text sentence where they are located, where X is a positive integer; Respectively obtaining the initial word vectors of the X first words through a word vector table, and respectively obtaining the second sememe vectors of the X first words through a sememe vector table; Determining the final word vector of each of the first words according to the initial word vector, the second sememe vector, the sentence identifier and the position information of each of the first words; Concatenating the final word vectors associated with each text sentence of the text data in data set i to obtain the sentence vector of each text sentence; Generating a category graph corresponding to category label j according to the sentence vector corresponding to category label j; Wherein, if data set i is the first data set, then category label j is one of the M category labels; if data set i is the multiple data sets, then category label j is one of the N category labels.

3. The method for text processing according to claim 2, wherein The generating a category graph corresponding to category label j according to the sentence vector corresponding to category label j includes: Determining the sum of distances associated with each text sentence corresponding to category label j, the sum of distances being used to represent the sum of the distances from the sentence vector of the text sentence it is associated with to the sentence vectors of other text sentences corresponding to category label j; Taking the sentence vector of the text sentence corresponding to the minimum sum of distances among the sums of distances as the center point corresponding to category label j; Determine the spatial positions of multiple points in the coordinate system corresponding to the central point, where the multiple points are points formed by the sentence vectors of other text sentences in the first dataset except for the minimum distance and the corresponding text sentence; Connect the multiple spatial positions in sequence to generate a category graph corresponding to the category label j.

4. The method for text processing according to claim 1, wherein, The determination method of the M quantities is as follows: Perform word segmentation on the text data corresponding to the M category labels in the second dataset respectively to obtain multiple second words, and the multiple second words carry the sentence identifier of the text sentence where they are located and the position information in the text sentence where they are located; Obtain the initial word vectors of the multiple second words through the word vector table respectively, and obtain the third semantic element vectors of the multiple second words through the semantic element vector table; Determine the final word vector of each second word according to the initial word vector, the third semantic element vector, the sentence identifier and the position information of each second word; Concatenate the final word vectors associated with each text sentence of the text data in the second dataset to obtain the sentence vector of each text sentence; Determine the distances from multiple points to the central points within the M first category graphs respectively, and determine the M quantities of the sentence vectors of each text sentence in the second dataset within the M first category graphs according to the distances, where the multiple points are points formed by the sentence vectors corresponding to the M category labels in the second dataset.

5. The method for text processing according to claim 4, characterized in that, When the distance is not greater than the first threshold, it indicates that the sentence vector corresponding to the distance is within the first category graph; The specific implementation method of determining the K first semantic element vectors within the first category graph according to the M quantities is as follows: when the M quantity of the target sentence vector within the first category graph is greater than the second threshold, determine the third semantic element vector corresponding to the target sentence vector as the first semantic element vector, and the target sentence vector is multiple sentence vectors containing the same third semantic element vector.

6. The method for text processing according to claim 1, wherein The classification processing of the second text data according to the K first semantic element vectors and the N second category graphs includes: Obtain the sentence vector of each text sentence in the second text data; Calculate the average distance between the sentence vector of each text sentence in the second text data and the K first semantic element vectors; When the average distance is not greater than the third threshold, determine that the text sentence with the average distance not greater than the third threshold belongs to the text category corresponding to the second category graph.

7. The method for text processing according to claim 1, wherein The division of the first text data into multiple datasets includes: Divide the first text data into two datasets. Among them, the data in the first dataset of the two datasets is used to generate the first category graph, and the data in the second dataset of the two datasets is used to determine the first semantic element vectors. The data volumes corresponding to the same category label in the first dataset and the second dataset are the same.

8. An apparatus for text processing, characterized in that, Including: An acquisition module, configured to acquire first text data carrying N category labels and second text data not carrying category labels, where N is an integer greater than 1; A partitioning module, configured to partition the first text data into a plurality of data sets, the plurality of data sets including at least one first data set and at least one second data set, and the first data set and the second data set being disjoint; A generating module, configured to generate M first category graphs corresponding to M category labels by using the data in the first data set, the first data set including text data of the M category labels, and the category graph being used to represent the semantic category represented by the text data of its corresponding category label, where M is an integer less than or equal to N; A determining module, configured to determine K first sememe vectors corresponding to the M category labels by using M quantities, the M quantities being the quantities of the data in the second data set in the M first category graphs respectively, and the first sememe vector being used to indicate target semantic information for text classification, where K is a positive integer; The generating module is further configured to generate N second category graphs corresponding to the N category labels by using the data in the plurality of data sets, and the range of the semantic category of the second category graph is larger than the range of the semantic category of the first category graph; A processing module, configured to perform classification processing on the second text data according to the K first sememe vectors and the N second category graphs.

9. An electronic device, characterized in that, Comprising: A processor; A memory for storing instructions executable by the processor; Wherein, the processor is configured to execute the instructions to implement the text processing method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, When the instructions in the storage medium are executed by the processor of the electronic device, the electronic device is enabled to execute the text processing method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Text category determination method, related device and equipment

    CN113821590A

  • Label generation method, system and device based on semantic similarity model and medium

    CN114443850A