A text processing method and device, computer equipment and a storage medium
Patent Information
- Application Number
- CN202211330309.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-10-27
- Publication Date
- 2026-09-15
- Estimated Expiration
- 2042-10-27
AI Technical Summary
[0004]本发明实施例提供一种文本处理方法、装置、计算机设备及存储介质,以解决人工标注文本效率与质量低下的问题
[0056]The aforementioned text processing method, apparatus, computer equipment, and storage medium calculate the similarity of multiple unlabeled texts to identify semantically similar texts and text clusters. Texts with similar semantics are then assigned the same label, avoiding the error of human judgment during manual annotation that could lead to semantically similar texts being labeled differently. This effectively reduces data noise and improves annotation quality. Furthermore, the use of a program to pre-annotate the text significantly improves annotation efficiency. The pre-annotated texts are then sorted, grouping semantically clear, easily annotated, and semantically similar texts together. This facilitates subsequent verification by separating semantically clear and similar texts from semantically unclear and difficult-to-annotate texts, further improving annotation quality.
Smart Images

Figure CN115577715B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence, and more particularly to a text processing method, apparatus, computer device, and storage medium. Background Technology
[0002] With the advancement of technology, artificial intelligence is increasingly involved in people's lives, making human-computer interaction ever more important. In handling human-computer interaction, the ability to recognize the meaning of natural language is particularly crucial. Typically, artificial intelligence itself cannot recognize natural language; natural language text needs to be tagged so that machines can understand its specific meaning.
[0003] Therefore, text annotation is a necessary process and an indispensable part of the realization of artificial intelligence. Artificial intelligence requires a large amount of carefully annotated data. Currently, much text annotation is done manually, which often results in low annotation quality and low efficiency. Especially with large amounts of data, manual annotation is not only inefficient, but annotators often fail to identify similar data, assigning different labels to semantically similar text, leading to text noise, poor annotation quality, and interference with subsequent training. Summary of the Invention
[0004] This invention provides a text processing method, apparatus, computer device, and storage medium to solve the problem of low efficiency and quality of manual text annotation.
[0005] Firstly, a text processing method is provided, including:
[0006] Obtain the text to be processed, which includes multiple unlabeled texts;
[0007] The similarity of the multiple unlabeled texts is calculated to obtain the text similarity and similar text clusters.
[0008] The unlabeled text is pre-labeled based on the text similarity to obtain pre-labeled text;
[0009] The pre-labeled texts are sorted by combining the density of the similar text clusters and the text similarity.
[0010] The pre-annotated text is validated according to its sorting order to obtain the target annotated text.
[0011] In conjunction with the first aspect, in one possible implementation, the similarity calculation of the plurality of unlabeled texts includes:
[0012] Vectorize the multiple unlabeled texts to obtain the text vector corresponding to each unlabeled text.
[0013] Set an expected data volume, where the expected data volume refers to the expected number of unlabeled texts;
[0014] Determine the number of the multiple unlabeled texts;
[0015] When the number of unlabeled texts exceeds the expected data volume, a clustering analysis algorithm with iterative solution is used to calculate the similarity between every two text vectors.
[0016] When the number of unlabeled texts is less than or equal to the expected data volume, a density clustering algorithm is used to calculate the similarity between every two text vectors.
[0017] In conjunction with the first aspect, in one possible implementation, the vectorization of the plurality of unlabeled texts to obtain a text vector corresponding to each unlabeled text includes:
[0018] Determine the context to which the text to be processed belongs, where the context refers to the source classification of the text to be processed;
[0019] When the text to be processed exists in the given context, unsupervised contrastive learning is used to transform the multiple unlabeled texts into vectors to obtain the text vector corresponding to each unlabeled text.
[0020] When the text to be processed does not belong to the specified scenario, a Siamese network based on a bidirectional language representation model is used to perform vector transformation on the multiple unlabeled texts, and the text vector corresponding to each unlabeled text is obtained.
[0021] In conjunction with the first aspect, in one possible implementation, before pre-labeling the unlabeled text based on the text similarity to obtain the pre-labeled text, the method further includes:
[0022] Determine whether the unlabeled text has a general text annotation model;
[0023] When the unlabeled text exists in the general text annotation model, the unlabeled text can be pre-annotated based on the general text annotation model, or the unlabeled text can be pre-annotated based on the text similarity.
[0024] When the unlabeled text does not exist in the general text annotation model, the text similarity is directly used to pre-annotate the unlabeled text.
[0025] In conjunction with the first aspect, in one possible implementation, the step of ranking the pre-labeled text by combining the density of the similar text clusters and the text similarity includes:
[0026] Based on the text similarity, find the center point of the text similarity distribution;
[0027] Based on the density of the similar text clusters, the similar text clusters are sorted from compact to loose.
[0028] Measure the distance between the pre-annotated text and the center point;
[0029] Based on the sorting results of the similar text clusters and combined with the distance from smallest to largest, the pre-labeled texts are sorted.
[0030] In conjunction with the first aspect, in one possible implementation, finding the center point of the text similarity distribution based on the text similarity includes:
[0031] Determine the distribution of the text similarity;
[0032] If the distribution is spherical, then calculate the average vector of all the pre-annotated texts to obtain the center point of the text similarity distribution;
[0033] If the distribution is non-spherical, then find the region with the highest density in the distribution.
[0034] Calculate the average vector value of the pre-annotated text within the region to obtain the center point within the region;
[0035] The center point within the region is taken as the center point of the text similarity distribution.
[0036] In conjunction with the first aspect, in one possible implementation, sorting the similar text clusters from compact to loose according to their density includes:
[0037] Obtain the cluster center point of the similar text clusters;
[0038] Set a desired value, where the desired value refers to the distance from the center point within the cluster;
[0039] Extract the pre-annotated text within the desired value;
[0040] Calculate the vector variance between the pre-annotated text within the expected value and the centroid of the cluster to obtain the density value of the similar text clusters;
[0041] The similar text clusters are sorted from smallest to largest based on the density value.
[0042] In conjunction with the first aspect, in one possible implementation, the step of validating the pre-annotated text according to its sorting order to obtain the target annotated text includes:
[0043] The pre-annotated text is verified sequentially from front to back according to the sorting order.
[0044] If the verification is successful, the target labeled text will be obtained;
[0045] If the verification fails, the target labeled text will be manually annotated.
[0046] Secondly, a text processing apparatus is provided, comprising:
[0047] The acquisition module is used to acquire the text to be processed, which includes multiple unlabeled texts;
[0048] The calculation module is used to calculate the similarity of the multiple unlabeled texts to obtain text similarity and similar text clusters;
[0049] The pre-labeling module is used to pre-label the unlabeled text based on the text similarity to obtain pre-labeled text;
[0050] The sorting module is used to sort the pre-labeled text by combining the density of the similar text clusters and the text similarity.
[0051] The verification module is used to verify the pre-annotated text according to the sorting order of the pre-annotated text to obtain the target annotated text.
[0052] Thirdly, a computer device is provided, including a memory, a transceiver, a processor, and a bus system, characterized in that it includes:
[0053] The memory is used to store programs;
[0054] The processor is used to execute the program stored in the memory. When the processor executes the program stored in the memory, the processor is used to perform the steps of implementing the above-described text processing method.
[0055] Fourthly, a computer-readable storage medium is provided, including instructions that, when executed on a computer, cause the computer to perform steps implementing the above-described text processing method.
[0056] The aforementioned text processing method, apparatus, computer equipment, and storage medium calculate the similarity of multiple unlabeled texts to identify semantically similar texts and text clusters. Texts with similar semantics are then assigned the same label, avoiding the error of human judgment during manual annotation that could lead to semantically similar texts being labeled differently. This effectively reduces data noise and improves annotation quality. Furthermore, the use of a program to pre-annotate the text significantly improves annotation efficiency. The pre-annotated texts are then sorted, grouping semantically clear, easily annotated, and semantically similar texts together. This facilitates subsequent verification by separating semantically clear and similar texts from semantically unclear and difficult-to-annotate texts, further improving annotation quality. Attached Figure Description
[0057] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments of the present invention will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0058] Figure 1 This is a schematic diagram of an application environment for a text processing method according to an embodiment of the present invention;
[0059] Figure 2 This is a flowchart illustrating a text processing method according to an embodiment of the present invention;
[0060] Figure 3 This is a flowchart illustrating a text processing method according to an embodiment of the present invention;
[0061] Figure 4 This is a flowchart illustrating a text processing method according to an embodiment of the present invention;
[0062] Figure 5 This is a flowchart illustrating a text processing method according to an embodiment of the present invention;
[0063] Figure 6 This is a flowchart illustrating a text processing method according to an embodiment of the present invention;
[0064] Figure 7 This is a schematic block diagram of a text processing device according to an embodiment of the present invention;
[0065] Figure 8 This is a schematic diagram of a computer device according to an embodiment of the present invention. Detailed Implementation
[0066] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0067] The text processing method provided in this embodiment of the invention can be applied to, for example, Figure 1 In this application environment, terminal devices communicate with the server via a network to obtain the text to be processed. These terminal devices can be, but are not limited to, various personal computers, laptops, smartphones, tablets, and portable wearable devices. The server can be a standalone server or a server cluster consisting of multiple servers.
[0068] In one embodiment, such as Figure 2 As shown, a text processing method is provided, which is applied to... Figure 1 Taking the server in the example, the following steps are included:
[0069] S10: Obtain the text to be processed, which includes multiple unlabeled texts.
[0070] The system receives a certain number and scale of text resources as text to be processed. The unprocessed text will contain one or more unlabeled texts, i.e., untagged text. Methods for receiving the text to be processed include, but are not limited to, user input and web scraping. Text resources include, but are not limited to, user input and web text content. Furthermore, the language of the text content is not limited.
[0071] S20: Calculate the similarity of the multiple unlabeled texts to obtain the text similarity and similar text clusters.
[0072] After obtaining the text to be processed, the similarity of the unlabeled text within the text is calculated. Based on the calculation results, text similarity and similar text clusters are obtained. The similarity calculation methods include, but are not limited to, cosine similarity and Jaccard similarity coefficient. Text similarity includes, but is not limited to, the distance between two texts and the vector similarity between two texts. Similar text clusters include, but are not limited to, the set of texts with the closest similarity.
[0073] S30: Pre-label the unlabeled text based on the text similarity to obtain pre-labeled text.
[0074] After calculating the text similarity, the unlabeled texts are pre-labeled based on the text similarity, so that each unlabeled text is given a pre-labeled label. The text pre-labeling models include, but are not limited to, SNN (Shared Nearest Neighbor) and K-means (K-means clustering algorithm).
[0075] S40: Sort the pre-labeled texts by combining the density of the similar text clusters and the text similarity.
[0076] After obtaining the pre-annotated text, the density of the text clusters is determined based on their similarity. More compact clusters are considered higher-quality clusters, allowing high-quality text to be grouped together. Additionally, based on text similarity, the texts with the most similar semantics are identified, ensuring that texts with similar semantics are grouped together.
[0077] S50: According to the sorting order of the pre-annotated text, the pre-annotated text is verified to obtain the target annotated text.
[0078] After obtaining the pre-annotated text in order, high-quality, semantically clear, and semantically similar pre-annotated texts can be prioritized for verification, while some semantically unclear or difficult-to-annotate pre-annotated texts can be processed later, making the entire verification and annotation process gradual. The above sorting order includes, but is not limited to, arranging semantically clear and semantically similar texts from semantically unclear texts to those from semantically unclear texts, and arranging semantically unclear texts from those from semantically clear and semantically similar texts.
[0079] This embodiment calculates the similarity of multiple unlabeled texts to identify semantically similar texts and text clusters. Texts with similar semantics are assigned the same label, avoiding the problem of semantically similar texts being labeled differently due to human error during manual annotation. This effectively reduces data noise and improves annotation quality. Furthermore, the use of a program to pre-annotate the text significantly improves annotation efficiency. The pre-annotated texts are then sorted, grouping semantically clear and easily annotated text together. This facilitates subsequent verification by separating semantically clear and similar texts from semantically unclear and difficult-to-annotate texts, further improving annotation quality.
[0080] like Figure 3 As shown, step S20, which involves calculating the similarity of the multiple unlabeled texts, specifically includes the following steps:
[0081] S21: Vectorize the multiple unlabeled texts to obtain the text vector corresponding to each unlabeled text.
[0082] S22: Set an expected data volume, where the expected data volume refers to the expected number of unlabeled texts.
[0083] S23: Determine the number of the multiple unlabeled texts.
[0084] S24: When the number of unlabeled texts is greater than the expected data volume, a clustering analysis algorithm with iterative solution is used to calculate the similarity between every two text vectors.
[0085] S25: When the number of the plurality of unlabeled texts is less than or equal to the expected amount of data, a density clustering algorithm is used to calculate the similarity between every two text vectors.
[0086] In step S21, the unlabeled text data needs to be converted into text vectors for subsequent similarity calculation.
[0087] In steps S22-S25, an expected data volume is set, namely the expected number of unlabeled texts.
[0088] If the number of unlabeled texts exceeds the expected data volume, it is determined that the number of unlabeled texts is large, and an iterative clustering analysis algorithm is used. Iterative clustering analysis algorithms include, but are not limited to, K-means and Mini Batch K-Means (Mini Batch K-means clustering algorithm). This embodiment uses the Mini Batch K-Means algorithm to calculate the similarity of unlabeled texts with a large data volume.
[0089] When the number of unlabeled texts is less than the expected data volume, it is determined that the number of unlabeled texts is small. In this case, the quality requirements for the calculation results are high, so a density-based clustering algorithm is used. Density-based clustering algorithms include, but are not limited to, DBSCAN (Density Based Spatial Clustering of Application with Noise) and OPTICS (Ordering Points To Identify the Clustering Structure). In this embodiment, the SNN-DBSCAN (Shared Nearest Neighbor-Density Based Spatial Clustering of Application with Noise) algorithm is selected to calculate the similarity of unlabeled texts with a small data volume and high quality requirements.
[0090] In step S25, this embodiment uses the SNN-DBSCAN algorithm, which is a combination of SNN and DBSCAN. The two parameters from DBSCAN, MinPts (Min Points) and EPts (Edge Points), are used to calculate similarity based on SNN. The calculation process is as follows:
[0091] Take any two text vectors from a set of multiple text vectors, and find the Top K (K-means optimal solution) of the adjacent text vectors between the two text vectors. Determine the number of neighbors shared by the two text vectors, which is the number of SNNs. The formula for the proportion is as follows:
[0092]
[0093] Hyperparameters are unknown variables, but unlike parameters during training, they influence the trained parameters and require manual input and adjustment by the trainer to optimize the model's performance. For two text vectors to be considered reachable, the number of reachable points calculated based on EPts must be at least greater than EPts. When multiple text vectors exist, if the number of reachable points calculated based on EPts for a particular text vector exceeds MinPts, then that text vector can be considered a core point.
[0094] For example, given two text vectors A and B, the Top K values between A and B are labeled as the number of shared neighbors in the SNN, or a percentage. Assuming there are m neighbors shared by A and B, and the hyperparameter is set to k, then the number of SNNs is m, and their percentage is... For points A and B to be considered reachable, the number of SNNs calculated based on EPts must be at least greater than EPts. If the number of reachable points calculated based on EPts for A exceeds MinPts, then A can be considered a core point.
[0095] In this embodiment, for cases with a small amount of data and high quality requirements, the SNN-DBSCAN algorithm is adopted, which can achieve relatively accurate calculation results without requiring too many parameter configurations. Other clustering algorithms often require setting many parameters, reducing efficiency. For example, K-means requires specifying the number of clusters, and DBSCAN requires specifying two density parameters, EPts and MinPts. However, using SNN-DBSCAN can transform the poorly interpretable EPts and MinPts parameters into an intuitive TopK using SNN, increasing the algorithm's ease of use while maintaining performance.
[0096] like Figure 4As shown, step S21, which involves vectorizing the multiple unlabeled texts to obtain the text vector corresponding to each unlabeled text, specifically includes the following steps:
[0097] S211: Determine the context to which the text to be processed belongs, wherein the context to which the text to be processed belongs refers to the source classification of the text to be processed.
[0098] S212: When the text to be processed exists in the given context, unsupervised contrastive learning is used to transform the multiple unlabeled texts into vectors to obtain the text vector corresponding to each unlabeled text.
[0099] S213: When the text to be processed does not belong to the specified scenario, a Siamese network based on a bidirectional language representation model is used to perform vector transformation on the multiple unlabeled texts, and the text vector corresponding to each unlabeled text is obtained.
[0100] In steps S211-S213, unlabeled text is converted into corresponding text vectors based on a text representation model for subsequent similarity calculation. Text representation models include, but are not limited to, bag-of-words models and topic models. This embodiment selects a large-scale pre-trained model, which, due to its network structure and data volume, can obtain high-quality text vectors.
[0101] In step S211, it is necessary to determine the context to which the text to be processed belongs. Text to be processed with a context refers to corpus with a clear context, including but not limited to article comments, news headlines, etc.
[0102] In step S212, where the text to be processed has a clearly defined context, there will be a large amount of unlabeled text within that context. Unsupervised contrastive learning is then used to transform the unlabeled text into vectors. For example, in a Weibo comment scenario, a large amount of unlabeled text is obtained. These unlabeled texts inherently share similarities, so the goal is to obtain high-quality Weibo text vectors. Therefore, unsupervised contrastive learning is chosen based on this large amount of unlabeled Weibo text to obtain text vectors more suitable for the Weibo scenario. Unsupervised contrastive learning methods include, but are not limited to, BYOL (Bootstrap Your Own Latent) and SimSiam (Simple Siamese). This embodiment preferably uses SimCSE (Simple Contrastive Learning of Sentence Embeddings) as the vector transformation method for the current path.
[0103] In step S213, where the text to be processed does not have a clearly defined scene, there will be a large amount of unlabeled text that does not belong to any scene. At this point, a Siamese network based on a bidirectional language representation model is used to transform the unlabeled text into vectors. Siamese networks based on bidirectional language representation models include, but are not limited to, SBERT (Sentence Embeddings using Siamese BERT-Networks, also known as Sentence-BERT). This embodiment uses SBERT for text vector transformation. SBERT is trained using a large-scale open-source dataset and often has good transformation results for text in general scenarios.
[0104] In this embodiment, the source context of the unlabeled text is first determined. An appropriate algorithm is selected based on the existence of a context. This allows for higher-quality text vectors when a context exists, based on the inherent similarity of the unlabeled text. For cases where a context does not exist, a large-scale open corpus is used for training to achieve better vector conversion results for texts in general contexts. Specifically, when a context exists, this embodiment uses the SimCSE unsupervised learning method to perform context-based data augmentation on the original unlabeled text. This data augmentation can be considered a minimal form of data expansion, constructing positive examples for subsequent training. SimCSE, in unsupervised vector representation, constructs positive examples for contrastive learning in a simple way, achieving results comparable to supervised learning. In cases where a context does not exist, this embodiment uses SBERT to vectorize the unlabeled text. SBERT utilizes the Siamese Network to improve the slow inference speed and high resource requirements of BERT (Bidirectional Encoder Representations from Transformers). SBERT's sub-networks all use the BERT model, and the two BERT models share parameters, which allows BERT to better capture the relationships between texts and generate higher-quality text vectors.
[0105] Before step S30, that is, before pre-annotating the unlabeled text based on the text similarity to obtain the pre-annotated text, the text processing method further includes the following steps:
[0106] S61: Determine whether the unlabeled text has a general text annotation model.
[0107] S62: When the unlabeled text exists in the general text annotation model, the unlabeled text can be pre-annotated based on the general text annotation model, or the unlabeled text can be pre-annotated based on the text similarity.
[0108] S63: When the unlabeled text does not exist in the general text annotation model, the text similarity is directly used to pre-annotate the unlabeled text.
[0109] In steps S61-S63, similar to the classic intelligent annotation method, when there is a corresponding general text annotation model for unlabeled text or when there is open data, this embodiment will use the general text annotation model to annotate the unlabeled text. The acquisition channels of the general text annotation model include, but are not limited to, training based on open data of similar scenarios, training based on annotated small amounts of data, etc.
[0110] In this embodiment, different general text annotation models can be selected to pre-annotate unlabeled text for different situations, thereby improving the quality of annotation. This embodiment also provides a method for pre-annotating unlabeled text based on text similarity when a general text annotation model is unavailable or cannot be used, ensuring the completeness of the process in this embodiment.
[0111] like Figure 5 As shown, step S40, which involves ranking the pre-annotated texts by combining the density of the similar text clusters and the text similarity, specifically includes the following steps:
[0112] S41: Based on the text similarity, find the center point of the text similarity distribution.
[0113] S42: Sort the similar text clusters from compact to loose according to their density.
[0114] S43: Measure the distance between the pre-annotated text and the center point.
[0115] S44: Sort the pre-labeled texts according to the sorting results of the similar text clusters and in combination with the distance from small to large.
[0116] In steps S41-S44, the distance between each pre-labeled text and the center point of the text similarity distribution is measured one by one. Then, the pre-labeled texts are sorted according to the distance from smallest to largest. This way, the closer a text is to the center point, the higher it will be in the ranking, making it easier to grasp the core semantic information of the current text cluster during subsequent verification. The methods used to measure the distance include, but are not limited to, Manhattan distance and cosine distance; this embodiment uses cosine distance for distance calculation.
[0117] This embodiment uses cosine distance for distance calculation. The distance score is between 0 and 1, which aligns with the intuitive meaning of similarity and has high interpretability. Furthermore, due to data noise and algorithm accuracy, some text clusters will inevitably lack clear semantic information after a similarity calculation. Therefore, text clusters with clear semantics are prioritized for annotation, while those with unclear semantics or difficult annotation are placed later. This creates a gradual annotation process, which improves annotation efficiency and quality.
[0118] like Figure 6 As shown, step S41, which involves finding the center point of the text similarity distribution based on the text similarity, specifically includes the following steps:
[0119] S411: Determine the distribution state of the text similarity.
[0120] S412: If the distribution is spherical, calculate the average vector of all the pre-annotated texts to obtain the center point of the text similarity distribution.
[0121] S413: If the distribution is non-spherical, then find the region with the highest density in the distribution.
[0122] S414: Calculate the average vector value of the pre-annotated text within the region to obtain the center point within the region.
[0123] S415: Take the center point of the region as the center point of the text similarity distribution.
[0124] In steps S411-S415, the center point of the text similarity distribution is calculated to facilitate subsequent sorting of the pre-labeled texts. Calculating the center point requires first determining the distribution of text similarity. When the distribution is spherical, the average vector of all pre-labeled texts or the center point provided by K-means can be used directly as the center point. When the distribution is non-spherical, neither the average vector nor K-means methods are very effective. Therefore, in this case, the region with the highest density in the distribution is first identified. Methods for identifying the region with the highest density include, but are not limited to, Mean Shift and DBSCAN. This embodiment uses Mean Shift to find the region with the highest density. After identifying the region with the highest density, the center point of the region is calculated using the aforementioned average vector or K-means method. Finally, the center point of the region is used as the center point of the text similarity distribution.
[0125] In this embodiment, different methods are selected based on the distribution state when calculating the center point of the text similarity distribution. In the case of non-spherical distribution, the region with the highest density is calculated first, and then the center point of the region is calculated as the center point of the entire text similarity. This effectively ensures the effect and quality of the center point calculation results.
[0126] like Figure 5 As shown, step S42, which involves sorting the similar text clusters from compact to loose according to their density, specifically includes the following steps:
[0127] S421: Obtain the cluster center point of the similar text clusters.
[0128] S422: Set a desired value, which refers to the distance from the center point within the cluster.
[0129] S423: Extract the pre-annotated text within the desired value.
[0130] S424: Calculate the vector variance between the pre-annotated text within the expected value and the cluster center point to obtain the density value of the similar text clusters.
[0131] S425: Sort the similar text clusters according to the density value from small to large.
[0132] In step S421, the intra-cluster center point of the similar text cluster is the center point calculated in step S41 above. That is to say, the center point of each text cluster is the center point of the corresponding similar text.
[0133] In addition, in step S423, the pre-labeled text within the expected value refers to the Top N (fast filtering) pre-labeled texts within the expected distance from the cluster center point, where Top N refers to quickly selecting the largest or smallest pre-labeled texts.
[0134] In this embodiment, axes S421-S425 sort the text clusters themselves, and the sorting order is based on the compactness of the text clusters. This allows more compact text clusters to be ranked higher, meaning that higher quality text clusters can be processed first, which effectively improves the efficiency and quality of annotation.
[0135] In step S50, the pre-annotated text is validated according to its sorting order to obtain the target annotated text. This specifically includes the following steps:
[0136] S51: The pre-annotated text is verified sequentially from front to back according to the sorting order.
[0137] S52: If the verification is successful, the target annotation text is obtained.
[0138] S53: If the verification fails, the target text will be manually annotated.
[0139] In step S51, the pre-annotated texts are sorted according to their arrangement order. This sorting process processes the semantically clear and semantically unclear pre-annotated texts one by one in sequence, allowing the verification work to proceed step by step. The sorting order includes both sequential and reverse sorting, and the sorting methods include, but are not limited to, arranging from semantically clear to semantically unclear, and from difficult to annotate to easy to annotate. In this embodiment, the sorting method used is to arrange the semantically clear and easy-to-annotate pre-annotated texts first for priority processing, while arranging the semantically unclear and difficult-to-annotate pre-annotated texts later.
[0140] In this embodiment, high-quality pre-labeled text with clear semantics is given priority. This allows for a gradual verification process and effectively avoids the situation where semantically similar texts are labeled differently during manual annotation, which could affect subsequent training. This improves the quality and efficiency of annotation.
[0141] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0142] In one embodiment, a text processing apparatus is provided, which corresponds one-to-one with the text processing methods described in the above embodiments. For example... Figure 7 As shown, the text processing device includes an acquisition module, a calculation module, a pre-annotation module, a sorting module, and a verification module. Detailed descriptions of each functional module are as follows:
[0143] The acquisition module is used to acquire the text to be processed, which includes multiple unlabeled texts;
[0144] The calculation module is used to calculate the similarity of the multiple unlabeled texts to obtain text similarity and similar text clusters;
[0145] The pre-labeling module is used to pre-label the unlabeled text based on the text similarity to obtain pre-labeled text;
[0146] The sorting module is used to sort the pre-labeled text by combining the density of the similar text clusters and the text similarity.
[0147] The verification module is used to verify the pre-annotated text according to the sorting order of the pre-annotated text to obtain the target annotated text.
[0148] For specific limitations regarding the text processing device, please refer to the limitations of the text processing method above, which will not be repeated here. Each module in the aforementioned text processing device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in the computer device in hardware form, or stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to each module.
[0149] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 8 As shown, the computer device includes a processor, memory, network interface, and database connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and database. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The database stores data required for executing text processing methods. The network interface communicates with external terminals via a network connection. The computer program is executed by the processor to implement the text processing methods.
[0150] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to perform the following steps:
[0151] Obtain the text to be processed, which includes multiple unlabeled texts;
[0152] The similarity of the multiple unlabeled texts is calculated to obtain the text similarity and similar text clusters.
[0153] The unlabeled text is pre-labeled based on the text similarity to obtain pre-labeled text;
[0154] The pre-labeled texts are sorted by combining the density of the similar text clusters and the text similarity.
[0155] The pre-annotated text is validated according to its sorting order to obtain the target annotated text.
[0156] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, the computer program performing the following steps when executed by a processor:
[0157] Obtain the text to be processed, which includes multiple unlabeled texts;
[0158] The similarity of the multiple unlabeled texts is calculated to obtain the text similarity and similar text clusters.
[0159] The unlabeled text is pre-labeled based on the text similarity to obtain pre-labeled text;
[0160] The pre-labeled texts are sorted by combining the density of the similar text clusters and the text similarity.
[0161] The pre-annotated text is validated according to its sorting order to obtain the target annotated text.
[0162] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0163] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0164] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.
Claims
1. A text processing method characterized by, include: Obtain the text to be processed, which includes multiple unlabeled texts; The similarity of the multiple unlabeled texts is calculated to obtain the text similarity and similar text clusters. The unlabeled text is pre-labeled based on the text similarity to obtain pre-labeled text; The pre-labeled texts are sorted by combining the density of the similar text clusters and the text similarity. The pre-annotated text is validated according to its sorting order to obtain the target annotated text; The process of ranking the pre-labeled text by combining the density of the similar text clusters and the text similarity includes: Based on the text similarity, find the center point of the text similarity distribution; Based on the density of the similar text clusters, the similar text clusters are sorted from compact to loose. Measure the distance between the pre-annotated text and the center point; Based on the sorting results of the similar text clusters and combined with the distance from smallest to largest, the pre-labeled texts are sorted.
2. The text processing method as described in claim 1, characterized in that, The similarity calculation of the plurality of unlabeled texts includes: Vectorize the multiple unlabeled texts to obtain the text vector corresponding to each unlabeled text. Set an expected data volume, where the expected data volume refers to the expected number of unlabeled texts; Determine the number of the multiple unlabeled texts; When the number of unlabeled texts exceeds the expected data volume, a clustering analysis algorithm with iterative solution is used to calculate the similarity between every two text vectors. When the number of unlabeled texts is less than or equal to the expected data volume, a density clustering algorithm is used to calculate the similarity between every two text vectors.
3. The text processing method as described in claim 2, characterized in that, The vectorization of the multiple unlabeled texts to obtain a text vector corresponding to each unlabeled text includes: Determine the context to which the text to be processed belongs, where the context refers to the source classification of the text to be processed; When the text to be processed exists in the given context, unsupervised contrastive learning is used to transform the multiple unlabeled texts into vectors to obtain the text vector corresponding to each unlabeled text. When the text to be processed does not belong to the specified scenario, a Siamese network based on a bidirectional language representation model is used to perform vector transformation on the multiple unlabeled texts, and the text vector corresponding to each unlabeled text is obtained.
4. The text processing method as described in claim 1, characterized in that, Before pre-labeling the unlabeled text based on the text similarity to obtain the pre-labeled text, the method further includes: Determine whether the unlabeled text has a general text annotation model; When the unlabeled text exists in the general text annotation model, the unlabeled text can be pre-annotated based on the general text annotation model, or the unlabeled text can be pre-annotated based on the text similarity. When the unlabeled text does not exist in the general text annotation model, the text similarity is directly used to pre-annotate the unlabeled text.
5. The text processing method as described in claim 1, characterized in that, Finding the center point of the text similarity distribution based on the text similarity includes: Determine the distribution of the text similarity; If the distribution is spherical, then calculate the average vector of all the pre-annotated texts to obtain the center point of the text similarity distribution; If the distribution is non-spherical, then find the region with the highest density in the distribution. Calculate the average vector value of the pre-annotated text within the region to obtain the center point within the region; The center point within the region is taken as the center point of the text similarity distribution.
6. The text processing method as described in claim 1, characterized in that, The step of sorting the similar text clusters from compact to loose according to their density includes: Obtain the cluster center point of the similar text clusters; Set a desired value, where the desired value refers to the distance from the center point within the cluster; Extract the pre-annotated text within the desired value; Calculate the vector variance between the pre-annotated text within the expected value and the centroid of the cluster to obtain the density value of the similar text clusters; The similar text clusters are sorted from smallest to largest based on the density value.
7. The text processing method as described in claim 1, characterized in that, The step of validating the pre-annotated text according to its sorting order to obtain the target annotated text includes: The pre-annotated text is verified sequentially from front to back according to the sorting order. If the verification is successful, the target labeled text will be obtained; If the verification fails, the target labeled text will be manually annotated.
8. A text processing apparatus, characterized in that, include: The acquisition module is used to acquire the text to be processed, which includes multiple unlabeled texts; The calculation module is used to calculate the similarity of the multiple unlabeled texts to obtain text similarity and similar text clusters; The pre-labeling module is used to pre-label the unlabeled text based on the text similarity to obtain pre-labeled text; The sorting module is used to sort the pre-labeled text by combining the density of the similar text clusters and the text similarity. The verification module is used to verify the pre-annotated text according to the sorting order of the pre-annotated text to obtain the target annotated text; The process of ranking the pre-labeled text by combining the density of the similar text clusters and the text similarity includes: Based on the text similarity, find the center point of the text similarity distribution; Based on the density of the similar text clusters, the similar text clusters are sorted from compact to loose. Measure the distance between the pre-annotated text and the center point; Based on the sorting results of the similar text clusters and combined with the distance from smallest to largest, the pre-labeled texts are sorted.
9. A computer device, comprising a memory, a transceiver, a processor, and a bus system, characterized in that, include: The memory is used to store programs; The processor is configured to execute a program stored in the memory, and when the processor executes the program stored in the memory, the processor is configured to perform the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Text classification integrated hierarchical clustering analysis-based automatic tag generation method
CN107180075A
Intelligent question and answer method, device and equipment
CN113377936A