Text data enhancement method, device and system and computer readable storage medium

Through the amplification and clustering of text data, the problem of data augmentation in the existing technology relying on manual writing of instruction data is solved, and high-quality and diverse new data generation is achieved, which improves the generalization ability of the model.

CN119961666APending Publication Date: 2025-05-09TCL TECHNOLOGY GROUP CORPORATION
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311445283.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-01
Publication Date
2025-05-09

AI Technical Summary

Technical Problem

Existing data augmentation methods rely on manual writing of instruction data, resulting in limited amount, diversity and creativity in generating new data, reducing the generalization ability of the model.

Method used

By acquiring text data, performing amplification operations and clustering, target text data is generated, and avoiding relying on manual writing of instruction data.

Benefits of technology

Reduce data enhancement costs, improve the quality and diversity of new data generation, and enhance the generalization ability of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119961666A_ABST
    Figure CN119961666A_ABST
Patent Text Reader

Abstract

The invention discloses a text data enhancement method, device and system and a computer readable storage medium. The method comprises the steps of obtaining text data; performing amplification operation on the text data to obtain enhanced text data; and clustering the enhanced text data according to the feature data corresponding to the enhanced text data to obtain target text data. The target text data is obtained by performing amplification operation and clustering on the text data, and new data is generated without depending on manually written instruction data, so that the data enhancement cost is reduced, the quantity, diversity and creativity of the target text data can be improved, and the quality of the generated target text data can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence technology, and in particular to a text data enhancement method, device, system and computer-readable storage medium. Background Art

[0002] Data augmentation (DA) is a method to increase the amount of data by adding slightly modified copies of existing data or creating new synthetic data from existing data. Existing data augmentation methods can be roughly divided into paraphrase-based methods, noise-based methods, and sampling-based methods.

[0003] However, existing data augmentation methods rely on manually written instruction data to generate new data, which is not only costly, but also often limited in quantity, diversity, and creativity. This results in low quality of the generated new data, and training models for downstream tasks based on these new data will reduce the generalization ability of the models for downstream tasks. Therefore, how to improve the quality of generated new data while reducing the cost of data augmentation is an urgent problem to be solved. Summary of the invention

[0004] The embodiments of the present disclosure provide a text data enhancement method, device, system and computer-readable storage medium, which can reduce the cost of data enhancement while improving the quality of generated new data.

[0005] In a first aspect, an embodiment of the present disclosure provides a text data enhancement method, the method comprising:

[0006] Get text data;

[0007] Performing an amplification operation on the text data to obtain enhanced text data;

[0008] The enhanced text data is clustered according to the feature data corresponding to the enhanced text data to obtain target text data.

[0009] In a second aspect, an embodiment of the present disclosure provides a text data enhancement device comprising:

[0010] An acquisition unit, used for acquiring text data;

[0011] an amplification unit, used for performing an amplification operation on the text data to obtain enhanced text data;

[0012] The clustering unit is used to cluster the enhanced text data according to the feature data corresponding to the enhanced text data to obtain target text data.

[0013] In a third aspect, an embodiment of the present disclosure also provides a text data enhancement system, including a memory storing multiple instructions; a processor loads instructions from the memory to execute the steps of any text data enhancement method provided by the embodiment of the present disclosure.

[0014] In a fourth aspect, an embodiment of the present disclosure further provides a computer-readable storage medium, which stores a plurality of instructions suitable for loading by a processor to execute the steps of any text data enhancement method provided in an embodiment of the present disclosure.

[0015] In a fifth aspect, the embodiments of the present disclosure also provide a computer program product, including a computer program or instructions, which, when executed by a processor, implement the steps of any text data enhancement method provided in the embodiments of the present disclosure.

[0016] The scheme of the embodiment of the application is adopted to obtain text data; perform an amplification operation on the text data to obtain enhanced text data; cluster the enhanced text data according to the feature data corresponding to the enhanced text data to obtain target text data. By performing an amplification operation and clustering on the text data to obtain the target text data, it does not rely on manually written instruction data to generate new data, reduces the data enhancement cost, and can improve the quantity, diversity and creativity of the target text data, and improve the quality of the generated target text data. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present disclosure. For those skilled in the art, other drawings can be obtained based on these drawings without creative work.

[0018] Figure 1 It is a flowchart of a first embodiment of the text data enhancement method provided in the embodiments of the present disclosure;

[0019] Figure 2 is a schematic diagram of a process for generating enhanced text data in the initial stage provided in an embodiment of the present disclosure;

[0020] Figure 3 is a schematic diagram of a process for generating enhanced text data in a subsequent stage provided in an embodiment of the present disclosure;

[0021] Figure 4 is a flowchart of a second embodiment of the text data enhancement method provided in the embodiments of the present disclosure;

[0022] Figure 5 is a schematic diagram of constructing a contrast loss function and a reconstruction loss function provided in an embodiment of the present disclosure;

[0023] Figure 6 is a structural schematic diagram of a text data enhancement device provided in an embodiment of the present disclosure;

[0024] Figure 7 It is a structural diagram of the text data enhancement system provided in the embodiment of the present disclosure. DETAILED DESCRIPTION

[0025] The technical solutions in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the drawings in the embodiments of the present disclosure. Obviously, the described embodiments are only part of the embodiments of the present disclosure, rather than all of the embodiments. Based on the embodiments in the present disclosure, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present disclosure. At the same time, in the description of the embodiments of the present disclosure, the terms "first", "second", etc. are only used to distinguish the descriptions and cannot be understood as indicating or implying relative importance. Therefore, the features defined as "first" and "second" may explicitly or implicitly include one or more features. In the description of the embodiments of the present disclosure, the meaning of "multiple" is two or more, unless otherwise clearly and specifically defined.

[0026] Embodiments of the present disclosure provide a text data enhancement method, device, system, and computer-readable storage medium.

[0027] Specifically, the embodiments of the present disclosure will be described from the perspective of a text data enhancement system, and the text data enhancement system can be specifically integrated in a text data enhancement device, that is, the text data enhancement method of the embodiments of the present disclosure can be executed by the text data enhancement system.

[0028] The text data enhancement method provided in the embodiments of the present disclosure can be applied to a text data enhancement device, which may include devices such as a mobile terminal, a PC terminal, etc.

[0029] The following is a detailed description in conjunction with the accompanying drawings. In the embodiments of the present disclosure, the execution subject is an example of a text data enhancement system. It should be noted that the description order of the following embodiments is not intended to limit the preferred order of the embodiments. Although the logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in an order different from that shown in the accompanying drawings.

[0030] Please refer to Figure 1 The specific process of the first embodiment of the text data enhancement method includes the following steps:

[0031] Step 101, obtaining text data;

[0032] Step 102, performing an amplification operation on the text data to obtain enhanced text data;

[0033] Step 103: clustering the enhanced text data according to the feature data corresponding to the enhanced text data to obtain target text data.

[0034] In this embodiment, when the text data needs to be enhanced to obtain more training text data for training the corresponding model, the text data enhancement system obtains the text data, and performs an amplification operation on the text data, and then filters the obtained amplified text data to delete low-quality or similar amplified text data, and finally obtains enhanced text data after looping the above steps multiple times; the text data enhancement system extracts feature data corresponding to the enhanced text data, and clusters the enhanced text data according to the feature data, determines the cluster center, and cyclically updates the cluster center to finally obtain the target text data.

[0035] The text data enhancement system of this embodiment acquires text data, performs an amplification operation on the text data, and obtains enhanced text data; extracts feature data corresponding to the enhanced text data, and clusters the enhanced text data according to the feature data to obtain target text data. By performing an amplification operation and clustering on the text data to obtain the target text data, it does not rely on manually written instruction data to generate new data, reduces the data enhancement cost, and can improve the quantity, diversity and creativity of the target text data, and improve the quality of the generated target text data.

[0036] Specifically, each step is described in detail below:

[0037] Step 101, obtaining text data;

[0038] In this step, the text data enhancement system obtains text data. Specifically, there is a data pool in the text data enhancement system, which includes a data pool for storing text data and a data pool for storing newly generated amplified text data. In the initial stage, since there is no newly generated amplified text data, the data pool only contains text data, that is, the data pool for storing newly generated amplified text data is empty. At this time, the text data enhancement system obtains text data from the data pool for storing text data; after amplifying and filtering the text data once, some newly generated amplified text data are obtained, and these newly generated amplified text data will be stored in the data pool for storing the newly generated amplified text data, and in the subsequent cycle stage, the text data enhancement system obtains text data from the data pool for storing text data and the data pool for storing the newly generated amplified text data, respectively, and participates in the subsequent amplification and filtering steps.

[0039] Step 102, performing an amplification operation on the text data to obtain enhanced text data;

[0040] In this step, after acquiring the text data, the text data enhancement system performs an amplification operation on the text data, and then filters the obtained amplified text data to delete low-quality or similar amplified text data, and finally obtains enhanced text data after cycling the above steps for multiple times. Specifically, in the initial stage, since there is no newly generated enhanced text data, the data pool only contains text data, that is, the data pool storing the newly generated amplified text data is empty. After the text data enhancement system performs amplification and filtering on the text data once, it obtains some newly generated amplified text data, which will be stored in the data pool storing the newly generated amplified text data, and in the subsequent cycle stage, the text data enhancement system uses the newly generated amplified text data as text data to participate in the subsequent amplification and filtering steps until the number of newly generated amplified text data in the data pool storing the newly generated amplified text data reaches the target number, and then the enhanced text data is obtained.

[0041] Specifically, step 102 includes:

[0042] Step 1021, performing an amplification operation on a preset amount of text data to obtain amplified text data;

[0043] In this step, the text data enhancement system selects a preset number of text data, and performs an amplification operation on the preset number of text data to obtain amplified text data; specifically, in the initial stage, the data pool only contains text data, that is, the data pool storing the newly generated amplified text data is empty. At this time, the text data enhancement system selects a preset number of text data, and performs an amplification operation on the preset number of text data to obtain amplified text data; in the subsequent stage, since the text data enhancement system obtains some newly generated amplified text data, that is, the data pool storing the newly generated amplified text data is not empty, at this time, the text data enhancement system selects a number of text data from the data pool storing the text data according to preset rules, and selects a number of amplified text data from the data pool storing the newly generated amplified text data, so that the sum of the number of selected text data and the number of selected amplified text data is the preset number, and then performs an amplification operation on the selected text data and the amplified text data to obtain new amplified text data.

[0044] Exemplarily, in the initial stage, 8 text data are stored in the data pool storing text data. It is assumed that the text data enhancement system amplifies and filters the 8 text data to obtain 4 amplified text data, and stores the 4 amplified text data in the data pool storing the newly generated amplified text data. In the subsequent stage, the text data enhancement system selects 6 text data from the data pool storing text data according to preset rules, and selects 2 amplified text data from the data pool storing the newly generated amplified text data, and then performs amplification operations on the selected text data and the amplified text data to obtain new amplified text data.

[0045] Further, step 1021 includes:

[0046] Step 10211, inputting the instruction data corresponding to the preset number of text data into the pre-created augmentation model in sequence to generate augmentation instruction data;

[0047] In this step, the text data enhancement system sequentially inputs the data corresponding to a preset number of text data into a pre-created augmentation model to generate augmented instruction data based on the initial instruction data; specifically, the pre-created augmentation model is a large language model (LLM model) integrated with Self-Instruct, and the text data enhancement system sequentially inputs the initial instruction data corresponding to a preset number of text data into the large language model integrated with Self-Instruct to generate augmented instruction data.

[0048] It should be noted that each text data includes corresponding instruction data, input data and output data. The instruction data is used to input the model in the training process so that the model knows the task to be performed; the input data is used to input the model in the training process so that the model knows what input is needed to perform the corresponding task; the output data is used to verify the output value obtained by the model performing the task corresponding to the instruction data according to the input data in the training process. However, in most cases, there is no strict boundary between the instruction data and the input data. For example, "write an article about preventing influenza" can be directly input into the model as instruction data, where the instruction data contains the input data, and then the output data is obtained; or "write an article with the following text as the theme" can be used as instruction data, and "preventing influenza" can be sent to the model as input data to obtain the corresponding output data. In order to ensure the diversity of data formats, the instruction data includes both instruction data that requires additional input data (that is, there is a boundary between the instruction data and the input data) and instruction data that does not require additional input data (that is, there is no boundary between the instruction data and the input data, and the two are integrated).

[0049] Step 10212, determining the task identifier corresponding to the amplification instruction data, and generating amplification input data and amplification output data corresponding to the amplification instruction data according to the task identifier;

[0050] In this step, the initial instruction data in the text data will include the corresponding task identifier. When the augmented instruction data is generated by the pre-created augmentation model, the task identifier corresponding to the augmented instruction data will also be generated. The text data enhancement system generates augmented input data and augmented output data corresponding to the augmented instruction data in different ways according to the task identifier.

[0051] Further, step 10212 includes:

[0052] Step 102121, comparing the task identifier with a preset target task identifier;

[0053] In this step, a target task identifier is set in the text data enhancement system before performing text data enhancement. The task identifier is used to indicate which task the text data is suitable for training the model to perform, and the target task identifier indicates the task that the model needs to perform. Before generating the augmented input data and augmented output data corresponding to the augmented instruction data, the text data enhancement system compares the task identifier with the preset target task identifier.

[0054] Step 102122: if the task identifier is the same as the preset target task identifier, generating amplification output data according to the amplification instruction data, and generating amplification input data according to the amplification output data;

[0055] In this step, if the text data enhancement system determines that the task identifier is the same as the preset target task identifier, that is, the augmented text data corresponding to the augmented instruction data may be used to train the corresponding model, then the text data enhancement system uses the output priority method for the augmented instruction data, that is, first generates augmented output data according to the augmented instruction data, and then generates augmented input data according to the augmented output data. It should be noted that usually the input data is found according to the instruction data, and then the corresponding output data is generated. This generation order is similar to how to use a model to respond to instructions and inputs. It is more likely to generate inputs with a certain tendency and is not suitable for training models. Therefore, for augmented instruction data whose task identifier is the same as the preset target task identifier, the augmented output data is first generated according to the augmented instruction data, and then the augmented input data is generated according to the augmented output data.

[0056] Step 102123, if the task identifier is different from the preset target task identifier, then generate amplification input data according to the amplification instruction data, and generate amplification output data according to the amplification input data.

[0057] In this step, if the text data enhancement system determines that the task identifier is different from the preset target task identifier, that is, the augmented text data corresponding to the augmented instruction data will not be used to train the corresponding model, then the text data enhancement system uses the input priority method for the augmented instruction data, that is, first generates augmented input data based on the augmented instruction data, and then generates augmented output data based on the augmented input data.

[0058] Step 10213, obtaining amplified text data according to the amplification instruction data, the amplification input data and the amplification output data.

[0059] In this step, the text data enhancement system obtains the augmented text data according to the augmented instruction data, the augmented input data and the augmented output data. It should be noted that the augmented instruction data is used to input the model during the training process so that the model knows the tasks to be performed; the augmented input data is used to input the model during the training process so that the model knows what input is needed to perform the corresponding tasks; the augmented output data is used to verify the output value obtained by the model performing the tasks corresponding to the instruction data according to the input data during the training process. In order to ensure the diversity of data formats, the augmented instruction data includes both augmented instruction data that requires additional augmented input data (i.e., there is a boundary between the augmented instruction data and the augmented input data) and augmented instruction data that does not require additional augmented input data (i.e., there is no boundary between the augmented instruction data and the augmented input data, and the two are integrated).

[0060] Step 1022, adding the amplified text data that meets the preset conditions to the text data;

[0061] In this step, after obtaining the amplified text data, the text data enhancement system compares the amplified text data with all the text data stored in the data pool, and adds the amplified text data whose comparison results meet the preset conditions to the data pool storing the newly generated amplified text data, so that the text data enhancement system uses the amplified text data in the data pool storing the newly generated amplified text data in subsequent cycles.

[0062] Exemplarily, the text data includes initial instruction data, initial input data and initial output data, and the augmented text data includes augmented instruction data, augmented input data and augmented output data; after obtaining the new augmented text data, the text data enhancement system compares the new augmented instruction data in each new augmented text data with the initial instruction data of each text data stored in the data pool and the augmented instruction data of each augmented text data stored in the data pool, and compares the new augmented input data and the new augmented output data in each new augmented text data with the initial input data and initial output data of each text data stored in the data pool and the augmented input data and the augmented output data of each augmented text data stored in the data pool, respectively; if it is determined that the similarity between the new augmented instruction data in the new augmented text data and the initial instruction data of each text data stored in the data pool and the augmented instruction data of each augmented text data is less than 0.7, and the new augmented input data and the new augmented output data in the augmented text data are different from the initial input data and initial output data of each text data stored in the data pool and the augmented input data and the augmented output data of each augmented text data, respectively, then the augmented text data is added to the data pool storing the newly generated augmented text data.

[0063] Step 1023, until the amount of the amplified text data included in the text data reaches a preset threshold, enhanced text data is obtained.

[0064] In this step, in the subsequent stage, the text data enhancement system continuously cycles: select a number of text data from the data pool storing text data, and select a number of amplified text data from the data pool storing newly generated amplified text data, so that the sum of the number of selected text data and the number of selected amplified text data is a preset number, and then the selected text data and the amplified text data are amplified to obtain new amplified text data, and the new amplified text data is compared with all text data stored in the data pool, and the new amplified text data whose comparison results meet the preset conditions is added to the data pool storing the newly generated amplified text data. Until the number of amplified text data stored in the data pool storing the newly generated amplified text data reaches a preset threshold, the enhanced text data is obtained. It should be noted that since the text data enhancement system treats the newly generated amplified text data as text data for processing in the subsequent stage, the text data enhancement system can obtain the enhanced text data until the number of amplified text data contained in the text data reaches the preset threshold.

[0065] In a feasible example, the specific implementation process of step 1021 to step 1023 is as follows:

[0066] Reference Figure 2In the initial stage, the data pool only contains text data, that is, the data pool storing the newly generated augmented text data is empty. The text data enhancement system selects a preset number of text data from the data pool storing the text data, and sequentially inputs the initial instruction data corresponding to the preset number of text data into the large language model (LLM) integrated with Self-Instruct to generate augmented instruction data and task identifiers corresponding to the augmented instruction data. The text data enhancement system determines whether the task identifier corresponding to the augmented instruction data is the same as the target task identifier. If they are the same, the corresponding augmented input data and augmented output data are generated by the LLM using the output priority method. If they are different, the corresponding augmented input data and augmented output data are generated by the LLM using the input priority method. After obtaining the augmented text data, the text data enhancement system compares the augmented text data with the text data stored in the data pool, and adds the augmented text data whose comparison results meet the preset conditions to the data pool storing the newly generated augmented text data.

[0067] refer to Figure 3 In the subsequent stage, the data pool contains text data and augmented text data, that is, the data pool storing the newly generated augmented text data is not empty. The text data enhancement system selects several text data and several augmented text data, and the sum of the number of text data and the selected augmented text data is a preset number. The initial instruction data corresponding to the several text data and the several augmented text data are sequentially input into the large language model (LLM) integrated with Self-Instruct to generate augmented instruction data and task identifiers corresponding to the augmented instruction data. The text data enhancement system determines whether the task identifier corresponding to the augmented instruction data is the same as the target task identifier. If they are the same, the corresponding augmented input data and augmented output data are generated by the LLM using the output priority method. If they are different, the corresponding augmented input data and augmented output data are generated by the LLM using the input priority method. After obtaining the augmented text data, the text data enhancement system compares the augmented text data with the text data and augmented text data stored in the data pool, and adds the augmented text data whose comparison results meet the preset conditions to the data pool storing the newly generated augmented text data. Then, the text data enhancement system executes the input in a loop. Figure 3 The steps shown are performed until it is determined that the amount of amplified text data in the data pool storing the newly generated amplified text data reaches a preset threshold, thereby obtaining enhanced text data.

[0068] Step 103: clustering the enhanced text data according to the feature data corresponding to the enhanced text data to obtain target text data.

[0069] In this step, after obtaining the enhanced text data, the text data enhancement system extracts the feature data corresponding to each enhanced text data, and clusters the enhanced text data according to the feature data to obtain the target text data. Specifically, the text data enhancement system first connects the enhanced instruction data, enhanced input data, and enhanced output data corresponding to each data to form a text, and then uses the ELMo pre-trained model to extract features from each text to obtain the embedding vector corresponding to each text, and then uses the K-means clustering algorithm to cluster these embedding vectors. After several iterations, the enhanced text data corresponding to the final cluster center is used as the target text data.

[0070] Specifically, step 103 includes:

[0071] Step 1031, selecting a preset number of target feature data from the feature data corresponding to the enhanced text data as cluster centers;

[0072] In this step, the text data enhancement system extracts the feature data corresponding to each enhanced text data, randomly selects a preset number of feature data from all the feature data as target feature data, and uses these target feature data as cluster centers; wherein the preset number can be set according to actual conditions.

[0073] Step 1032, calculating the distance between the remaining feature data except the target feature data and the cluster center;

[0074] In this step, after determining the cluster center, the text data enhancement system will calculate the distance between each feature data remaining except the target feature data and each cluster center. The formula for calculating the distance is as follows:

[0075]

[0076] in, represents the cluster center corresponding to the kth cluster at the t-1th iteration, represents the i-th embedding vector h at the t-th iteration i The cluster it belongs to, ||·|| represents the L2 norm.

[0077] Step 1033, clustering the enhanced text data according to the distance, and updating the cluster center corresponding to each cluster;

[0078] In this step, the text data enhancement system divides the enhanced text data corresponding to each feature data into the cluster corresponding to the cluster center closest to it according to the distance between each remaining feature data and each cluster center, and updates the cluster center corresponding to each cluster.

[0079]

[0080] in, represents the cluster center corresponding to the kth cluster at the tth iteration, represents the subscript of all embedding vectors in the k-th cluster at the t-th iteration, and ||·|| represents the L2 norm.

[0081] Step 1034, until the cluster center corresponding to each cluster no longer changes, the enhanced text data corresponding to the final cluster center is used as the target text data.

[0082] In this step, after updating the cluster center corresponding to each cluster, the text data enhancement system executes the steps of calculating the distance between each feature data outside the cluster center and each updated cluster center, and dividing the enhanced text data corresponding to each feature data into the cluster corresponding to the updated cluster center closest to it. Until the cluster center corresponding to each cluster no longer changes, the enhanced text data corresponding to the final cluster center is used as the target text data. Furthermore, when the number of cycles reaches the preset number, the enhanced text data corresponding to the cluster center updated in the last iteration is used as the target text data.

[0083] The text data enhancement system of this embodiment acquires text data, performs an amplification operation on the text data, and obtains enhanced text data; extracts feature data corresponding to the enhanced text data, and clusters the enhanced text data according to the feature data to obtain target text data. By performing an amplification operation and clustering on the text data to obtain the target text data, it does not rely on manually written instruction data to generate new data, reduces the data enhancement cost, and can improve the quantity, diversity and creativity of the target text data, and improve the quality of the generated target text data.

[0084] Please refer to Figure 4 A second embodiment of a text data enhancement method is proposed. The second embodiment is specifically that after clustering the enhanced text data according to the feature data corresponding to the enhanced text data to obtain the target text data, the method includes:

[0085] Step 104, using the text data and the target text data as training text data;

[0086] In this step, after obtaining the target text data, the text data enhancement system uses the text data and the target text data as training text data. It can be understood that the text data is the initial original text data, and the target text data is the newly generated text data after data enhancement.

[0087] Step 105, extracting first feature data and second feature data corresponding to the training instruction data and the training input data in the training text data;

[0088] In this step, the text data enhancement system divides the training text data into two parts, one part consists of training instruction data and training input data, and the other part consists of training output data. The text data enhancement system encodes the training instruction data and the training input data to obtain the first feature data and the second feature data corresponding to the training instruction data and the training input data.

[0089] Exemplarily, the text data augmentation system divides the training text data into two parts, one of which consists of training instruction data and training input data, denoted as {x1,x2,L,x N}, and the other part consists of training output data, denoted as {y1,y2,L,y N}, where N represents the amount of data contained in the training text data. The text data enhancement system uses the random dropout mechanism in the encoder layer of the T5 model to encode the training instruction data and the training input data twice to obtain two sets of embedding vectors {z1,z2,L,z N}and The first feature data is the embedding vector {z1,z2,L,z N}, the second feature data is the embedding vector

[0090] Step 106: construct a target loss function based on the first feature data, the second feature data and the training output data, and train a preset model based on the training text data and the target loss function.

[0091] In this step, the text data enhancement system constructs a first loss function for the first feature data and the second feature data, decodes the first feature data and the second feature data, and constructs a second loss function and a third loss function based on the decoded first feature data, the second feature data and the training output data, and then constructs a target loss function based on the first loss function, the second loss function and the third loss function, and trains a preset model based on the training text data and the target loss function to minimize the target loss function, until the training is stopped when the maximum number of iterations is reached or the total loss drops to a preset value.

[0092] Specifically, constructing a target loss function based on the first feature data, the second feature data and the training output data includes:

[0093] Step 1061, constructing a positive sample pair and a negative sample pair based on the first feature data and the second feature data, and constructing a first loss function based on the positive sample pair and the negative sample pair;

[0094] In this step, the text data enhancement system constructs positive sample pairs and negative sample pairs based on the first feature data and the second feature data, and constructs a first loss function based on the positive sample pairs and the negative sample pairs; specifically, Figure 5 As shown in the figure, the first loss function is the contrast loss function; the text data enhancement system uses the random dropout mechanism in the encoder layer of the T5 model to encode the training instruction data and the training input data twice to obtain two sets of embedding vectors {z1,z2,L,z N}and The first feature data is the embedding vector {z1,z2,L,z N}, the second feature data is the embedding vector in Constitute a positive sample pair, After constructing the positive and negative sample pairs, we can get the following contrast loss function L cl :

[0095]

[0096]

[0097] in, is the similarity of the positive sample pair, is the sum of the similarities between the positive sample pair and all negative sample pairs, τ is the temperature parameter, represents the distance between two vectors (the larger the value, the closer the semantic information of the two samples is), ||·|| represents the L2 norm, and N represents the amount of data contained in the training text data. Minimizing the contrast loss function is equivalent to minimizing the similarity of the negative sample pair while maximizing the similarity of the positive sample pair, which is conducive to shortening the distance between similar samples, separating different samples, and then learning the common features between similar samples.

[0098] Step 1062, respectively decode the first feature data and the second feature data to obtain first decoded feature data and second decoded feature data;

[0099] In this step, the text data enhancement system decodes the first feature data and the second feature data respectively to obtain first decoded feature data and second decoded feature data; specifically, Figure 5 As shown in the figure, the text data augmentation system embeds two sets of vectors {z1,z2,L,z N}and Send it to the decoder layer of the T5 model and get two sets of outputs and Among them, the first decoding feature data is The second decoding feature data is

[0100] Step 1063: constructing a second loss function based on the first decoded feature data and the training output data, and constructing a third loss function based on the second decoded feature data and the training output data;

[0101] In this step, the text data enhancement system constructs a second loss function based on the first decoded feature data and the training output data, and constructs a third loss function based on the second decoded feature data and the training output data; specifically, Figure 5 As shown, the first decoded feature data is and the second decoding feature data is Respectively with the training output data {y1,y2,L,y N} to reconstruct and obtain two reconstruction losses L rec1 and L rec2 , the specific formula is as follows:

[0102]

[0103]

[0104] in, It represents the output corresponding to the i-th training sample after the first dropout mechanism. represents the output corresponding to the i-th training sample after the second dropout mechanism, y i represents the original output of the i-th training sample, ||·|| represents the L2 norm, and N represents the amount of data contained in the training text data.

[0105] Step 1064: construct a target loss function based on the first loss function, the second loss function and the third loss function.

[0106] In this step, the text data enhancement system constructs a target loss function based on the first loss function, the second loss function and the third loss function; specifically, the first loss function is the contrast loss function L cl , the second loss function and the third loss function are two reconstruction loss functions L rec1 , L rec2 , the text data enhancement system calculates the contrast loss function L cl And two reconstruction loss functions L rec1 , L rec2 The target loss function is obtained by summing up. The specific formula is as follows:

[0107]

[0108] Among them, L clis the contrast loss function corresponding to the training data using two different dropout mechanisms, L rec1 is the reconstruction loss function corresponding to the training data after the first dropout mechanism, L rec2 is the reconstruction loss function corresponding to the training data after the second dropout mechanism.

[0109] The text data enhancement system of this embodiment uses the initial text data and the target text data as training text data; extracts the first feature data and the second feature data corresponding to the training instruction data and the training input data in the training text data; constructs a target loss function based on the first feature data, the second feature data and the training output data, and trains a preset model based on the training text data and the target loss function. By constructing a contrast loss function and a reconstruction loss function, contrast learning is introduced into the process of training the model, so that the model can learn the information of the sample itself through the reconstruction loss, and can also learn the information of the negative samples in the same batch of training text data through the contrast loss, thereby improving the generalization ability of the trained model.

[0110] This embodiment also provides a text data enhancement device, which can specifically include devices such as mobile terminals, PC terminals, etc. Figure 6 As shown, the text data enhancement device may include:

[0111] An acquisition unit 1001 is used to acquire text data;

[0112] An amplification unit 1002, configured to perform an amplification operation on the text data to obtain enhanced text data;

[0113] The clustering unit 1003 is used to cluster the enhanced text data according to the feature data corresponding to the enhanced text data to obtain target text data.

[0114] In an optional example, the amplification unit is further used for:

[0115] Performing an amplification operation on a preset amount of text data to obtain amplified text data;

[0116] Adding the amplified text data that meets the preset conditions to the text data;

[0117] Until the amount of the amplified text data contained in the text data reaches a preset threshold, enhanced text data is obtained.

[0118] In an optional example, the amplification unit is further used for:

[0119] Inputting the instruction data corresponding to the preset number of text data into the pre-created augmentation model in sequence to generate augmentation instruction data;

[0120] Determine a task identifier corresponding to the amplification instruction data, and generate amplification input data and amplification output data corresponding to the amplification instruction data according to the task identifier;

[0121] Amplified text data is obtained according to the amplification instruction data, the amplification input data and the amplification output data.

[0122] In an optional example, the amplification unit is further used for:

[0123] Comparing the task identifier with a preset target task identifier;

[0124] If the task identifier is the same as the preset target task identifier, generating amplified output data according to the amplification instruction data, and generating amplified input data according to the amplification output data;

[0125] If the task identifier is different from the preset target task identifier, the amplification input data is generated according to the amplification instruction data, and the amplification output data is generated according to the amplification input data.

[0126] In an optional example, the clustering unit is further used to:

[0127] Selecting a preset number of target feature data from the feature data corresponding to the enhanced text data as cluster centers;

[0128] Calculating the distance between the remaining feature data except the target feature data and the cluster center;

[0129] Clustering the enhanced text data according to the distance, and updating the cluster center corresponding to each cluster;

[0130] Until the cluster center corresponding to each cluster no longer changes, the enhanced text data corresponding to the final cluster center is used as the target text data.

[0131] In an optional example, the text data enhancement device further includes a training unit, and the training unit is used to:

[0132] Using the text data and the target text data as training text data;

[0133] Extracting first feature data and second feature data corresponding to the training instruction data and the training input data in the training text data;

[0134] A target loss function is constructed based on the first feature data, the second feature data and the training output data, and a preset model is trained based on the training text data and the target loss function.

[0135] In an optional example, the training unit is further configured to:

[0136] Constructing a positive sample pair and a negative sample pair based on the first feature data and the second feature data, and constructing a first loss function based on the positive sample pair and the negative sample pair;

[0137] Decoding the first feature data and the second feature data respectively to obtain first decoded feature data and second decoded feature data;

[0138] Constructing a second loss function based on the first decoded feature data and the training output data, and constructing a third loss function based on the second decoded feature data and the training output data;

[0139] A target loss function is constructed based on the first loss function, the second loss function and the third loss function.

[0140] By adopting the scheme of this embodiment, target text data is obtained by performing augmentation operations and clustering on text data, which does not rely on manually written instruction data to generate new data, reduces the data enhancement cost, and can improve the quantity, diversity and creativity of target text data, and improve the quality of generated target text data; and by constructing a contrast loss function and a reconstruction loss function, contrast learning is introduced into the process of training the model, so that the model can learn the information of the sample itself through the reconstruction loss, and can also learn the information of negative samples in the same batch of training text data through the contrast loss, thereby improving the generalization ability of the trained model.

[0141] Accordingly, the present disclosure also provides a text data enhancement system, such as Figure 7 As shown, Figure 7 A schematic diagram of the structure of a text data enhancement system provided in an embodiment of the present disclosure. The text data enhancement system 1100 includes a processor 1101 having one or more processing cores, a memory 1102 having one or more computer-readable storage media, and a computer program stored in the memory 1102 and executable on the processor. The processor 1101 is electrically connected to the memory 1102. Those skilled in the art will appreciate that the text data enhancement system structure shown in the figure does not constitute a limitation on the text data enhancement system, and may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently.

[0142] The processor 1101 is the control center of the text data enhancement system 1100, and uses various interfaces and lines to connect various parts of the entire electronic device 1100. By running or loading software programs and / or units stored in the memory 1102, and calling data stored in the memory 1102, the processor 1101 executes various functions of the electronic device 1100 and processes data, thereby monitoring the text data enhancement system 1100 as a whole. The processor 1101 can be a processor CPU, a graphics processor GPU, a network processor (Network Processor, NP), etc., and can implement or execute the various methods, steps and logic block diagrams disclosed in the embodiments of the present disclosure.

[0143] In the embodiment of the present disclosure, the processor 1101 in the text data enhancement system 1100 will load instructions corresponding to the processes of one or more application programs into the memory 1102 according to the following steps, and the processor 1101 will run the application programs stored in the memory 1102 to implement various functions, such as:

[0144] Get text data;

[0145] Performing an amplification operation on the text data to obtain enhanced text data;

[0146] The enhanced text data is clustered according to the feature data corresponding to the enhanced text data to obtain target text data.

[0147] In an optional example, it also includes:

[0148] Performing an amplification operation on a preset amount of text data to obtain amplified text data;

[0149] Adding the amplified text data that meets the preset conditions to the text data;

[0150] Until the amount of the amplified text data contained in the text data reaches a preset threshold, enhanced text data is obtained.

[0151] In an optional example, it also includes:

[0152] Inputting the instruction data corresponding to the preset number of text data into the pre-created augmentation model in sequence to generate augmentation instruction data;

[0153] Determine a task identifier corresponding to the amplification instruction data, and generate amplification input data and amplification output data corresponding to the amplification instruction data according to the task identifier;

[0154] Amplified text data is obtained according to the amplification instruction data, the amplification input data and the amplification output data.

[0155] In an optional example, it also includes:

[0156] Comparing the task identifier with a preset target task identifier;

[0157] If the task identifier is the same as the preset target task identifier, generating amplified output data according to the amplification instruction data, and generating amplified input data according to the amplification output data;

[0158] If the task identifier is different from the preset target task identifier, the amplification input data is generated according to the amplification instruction data, and the amplification output data is generated according to the amplification input data.

[0159] In an optional example, it also includes:

[0160] Selecting a preset number of target feature data from the feature data corresponding to the enhanced text data as cluster centers;

[0161] Calculating the distance between the remaining feature data except the target feature data and the cluster center;

[0162] Clustering the enhanced text data according to the distance, and updating the cluster center corresponding to each cluster;

[0163] Until the cluster center corresponding to each cluster no longer changes, the enhanced text data corresponding to the final cluster center is used as the target text data.

[0164] In an optional example, it also includes:

[0165] Using the text data and the target text data as training text data;

[0166] Extracting first feature data and second feature data corresponding to the training instruction data and the training input data in the training text data;

[0167] A target loss function is constructed based on the first feature data, the second feature data and the training output data, and a preset model is trained based on the training text data and the target loss function.

[0168] In an optional example, it also includes:

[0169] Constructing a positive sample pair and a negative sample pair based on the first feature data and the second feature data, and constructing a first loss function based on the positive sample pair and the negative sample pair;

[0170] Decoding the first feature data and the second feature data respectively to obtain first decoded feature data and second decoded feature data;

[0171] Constructing a second loss function based on the first decoded feature data and the training output data, and constructing a third loss function based on the second decoded feature data and the training output data;

[0172] A target loss function is constructed based on the first loss function, the second loss function and the third loss function.

[0173] As a result, new data can be generated without relying on manually written instruction data, which reduces the cost of data enhancement, and can increase the quantity, diversity and creativity of target text data, improve the quality of generated target text data, and improve the generalization ability of the trained model.

[0174] The specific implementation of the above operations can be found in the previous embodiments, which will not be described in detail here.

[0175] Optional, such as Figure 7 As shown, the text data enhancement system 1100 also includes: a touch screen 1103, a radio frequency circuit 1104, an audio circuit 1105, an input unit 1106, and a power supply 1107. The processor 1101 is electrically connected to the touch screen 1103, the radio frequency circuit 1104, the audio circuit 1105, the input unit 1106, and the power supply 1107. Those skilled in the art can understand that Figure 4 The text data enhancement system structure shown in the figure does not constitute a limitation on the electronic device, and may include more or less components than shown in the figure, or combine certain components, or arrange the components differently.

[0176] The touch display screen 1103 can be used to display a graphical user interface and receive operation instructions generated by the user acting on the graphical user interface. The touch display screen 1103 may include a display panel and a touch panel. Among them, the display panel can be used to display information input by the user or information provided to the user and various graphical user interfaces of the electronic device, and these graphical user interfaces can be composed of graphics, text, icons, videos and any combination thereof. Optionally, the display panel can be configured in the form of a liquid crystal display (LCD, Liquid Crystal Display), an organic light-emitting diode (OLED, Organic Light-EmittingDiode) and the like. The touch panel can be used to collect the user's touch operation on or near it (such as the user uses any suitable object or attachment such as a finger, a stylus, etc. on the touch panel or near the touch panel), and generate corresponding operation instructions, and the operation instructions execute corresponding programs. Optionally, the touch panel may include two parts: a touch detection device and a touch controller. Among them, the touch detection device detects the user's touch orientation, detects the signal brought by the touch operation, and transmits the signal to the touch controller; the touch controller receives the touch information from the touch detection device, converts it into the touch point coordinates, and then sends it to the processor 1101, and can receive the command sent by the processor 1101 and execute it. The touch panel can cover the display panel. When the touch panel detects a touch operation on or near it, it is transmitted to the processor 1101 to determine the type of touch event, and then the processor 1101 provides a corresponding visual output on the display panel according to the type of touch event. In the embodiment of the present disclosure, the touch panel and the display panel can be integrated into the touch display screen 1103 to realize the input and output functions. However, in some embodiments, the touch panel and the touch panel can be used as two independent components to realize the input and output functions. That is, the touch display screen 1103 can also be used as a part of the input unit 1106 to realize the input function.

[0177] The radio frequency circuit 1104 may be used to send and receive radio frequency signals, so as to establish wireless communication with a network device or other electronic devices through wireless communication, and to send and receive signals with the network device or other electronic devices.

[0178] The audio circuit 1105 can be used to provide an audio interface between the user and the electronic device through a speaker and a microphone. The audio circuit 1105 can transmit the electrical signal converted from the received audio data to the speaker, which is converted into a sound signal for output; on the other hand, the microphone converts the collected sound signal into an electrical signal, which is received by the audio circuit 1105 and converted into audio data, and then the audio data is output to the processor 1101 for processing, and then sent to another electronic device through the radio frequency circuit 1104, or the audio data is output to the memory 1102 for further processing. The audio circuit 1105 may also include an earplug jack to provide communication between an external headset and an electronic device.

[0179] The input unit 1106 may be used to receive input numbers, character information or user feature information (such as fingerprint, iris, facial information, etc.), and generate keyboard, mouse, joystick, optical or trackball signal input related to user settings and function control.

[0180] The power supply 1107 is used to supply power to various components of the electronic device 1100. Optionally, the power supply 1107 can be logically connected to the processor 1101 through a power management system, so that the power management system can manage charging, discharging, and power consumption. The power supply 1107 can also include one or more DC or AC power supplies, recharging systems, power failure detection circuits, power converters or inverters, power status indicators, and other arbitrary components.

[0181] although Figure 7 Not shown, the electronic device 1100 may also include a camera, a sensor, a wireless fidelity module, a Bluetooth module, etc., which will not be described in detail here.

[0182] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0183] A person of ordinary skill in the art will appreciate that all or part of the steps in the various methods of the above embodiments may be completed by instructions, or by controlling related hardware through instructions. The instructions may be stored in a computer-readable storage medium and loaded and executed by a processor.

[0184] To this end, the present disclosure provides a computer-readable storage medium, which stores multiple computer programs, which can be loaded by a processor to execute any text data enhancement method provided by the present disclosure. The computer program can execute the following steps of the text data enhancement method:

[0185] Get text data;

[0186] Performing an amplification operation on the text data to obtain enhanced text data;

[0187] The enhanced text data is clustered according to the feature data corresponding to the enhanced text data to obtain target text data.

[0188] In an optional example, it also includes:

[0189] Performing an amplification operation on a preset amount of text data to obtain amplified text data;

[0190] Adding the amplified text data that meets the preset conditions to the text data;

[0191] Until the amount of the amplified text data contained in the text data reaches a preset threshold, enhanced text data is obtained.

[0192] In an optional example, it also includes:

[0193] Inputting the instruction data corresponding to the preset number of text data into the pre-created augmentation model in sequence to generate augmentation instruction data;

[0194] Determine a task identifier corresponding to the amplification instruction data, and generate amplification input data and amplification output data corresponding to the amplification instruction data according to the task identifier;

[0195] Amplified text data is obtained according to the amplification instruction data, the amplification input data and the amplification output data.

[0196] In an optional example, it also includes:

[0197] Comparing the task identifier with a preset target task identifier;

[0198] If the task identifier is the same as the preset target task identifier, generating amplified output data according to the amplification instruction data, and generating amplified input data according to the amplification output data;

[0199] If the task identifier is different from the preset target task identifier, the amplification input data is generated according to the amplification instruction data, and the amplification output data is generated according to the amplification input data.

[0200] In an optional example, it also includes:

[0201] Selecting a preset number of target feature data from the feature data corresponding to the enhanced text data as cluster centers;

[0202] Calculating the distance between the remaining feature data except the target feature data and the cluster center;

[0203] Clustering the enhanced text data according to the distance, and updating the cluster center corresponding to each cluster;

[0204] Until the cluster center corresponding to each cluster no longer changes, the enhanced text data corresponding to the final cluster center is used as the target text data.

[0205] In an optional example, it also includes:

[0206] Using the text data and the target text data as training text data;

[0207] Extracting first feature data and second feature data corresponding to the training instruction data and the training input data in the training text data;

[0208] A target loss function is constructed based on the first feature data, the second feature data and the training output data, and a preset model is trained based on the training text data and the target loss function.

[0209] In an optional example, it also includes:

[0210] Constructing a positive sample pair and a negative sample pair based on the first feature data and the second feature data, and constructing a first loss function based on the positive sample pair and the negative sample pair;

[0211] Decoding the first feature data and the second feature data respectively to obtain first decoded feature data and second decoded feature data;

[0212] Constructing a second loss function based on the first decoded feature data and the training output data, and constructing a third loss function based on the second decoded feature data and the training output data;

[0213] A target loss function is constructed based on the first loss function, the second loss function and the third loss function.

[0214] As a result, new data can be generated without relying on manually written instruction data, which reduces the cost of data enhancement, and can increase the quantity, diversity and creativity of target text data, improve the quality of generated target text data, and improve the generalization ability of the trained model.

[0215] The specific implementation of the above operations can be found in the previous embodiments, which will not be described in detail here.

[0216] The computer-readable storage medium may include: a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.

[0217] Since the computer program stored in the computer-readable storage medium can execute any of the text data enhancement methods provided in the embodiments of the present disclosure, the beneficial effects that can be achieved by any of the text data enhancement methods provided in the embodiments of the present disclosure can be achieved. Please refer to the previous embodiments for details and will not be repeated here.

[0218] According to one aspect of the present disclosure, a computer program product or a computer program is also provided, the computer program product or the computer program including computer instructions, the computer instructions being stored in a computer-readable storage medium. A processor of an electronic device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the electronic device executes the methods provided in various optional implementations in the above-mentioned embodiments.

[0219] In the above-mentioned text data enhancement device, computer-readable storage medium, text data enhancement system, and computer program product embodiments, the description of each embodiment has its own emphasis. For parts not described in detail in a certain embodiment, reference can be made to the relevant description of other embodiments. Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working process and beneficial effects of the above-described text data enhancement device, computer-readable storage medium, computer program product, text data enhancement system, and corresponding units can refer to the description of the text data enhancement method in the above embodiment, and will not be repeated here.

[0220] The above is a detailed introduction to a text data enhancement method, device, system, computer-readable storage medium, and computer program product provided by the embodiments of the present disclosure. Specific examples are used in this article to illustrate the principles and implementation methods of the present disclosure. The description of the above embodiments is only used to help understand the method of the present disclosure and its core ideas. At the same time, for those skilled in the art, according to the ideas of the present disclosure, there will be changes in the specific implementation methods and application scopes. In summary, the content of this specification should not be understood as a limitation on the present disclosure.

Claims

1. A text data enhancement method, characterized in that: The text data enhancement method comprises: Get text data; Performing an amplification operation on the text data to obtain enhanced text data; The enhanced text data is clustered according to the feature data corresponding to the enhanced text data to obtain target text data.

2. The text data enhancement method according to claim 1, characterized in that: The step of performing an amplification operation on the text data to obtain enhanced text data includes: Performing an amplification operation on a preset amount of text data to obtain amplified text data; Adding the amplified text data that meets the preset conditions to the text data; Until the amount of the amplified text data contained in the text data reaches a preset threshold, enhanced text data is obtained.

3. The text data enhancement method according to claim 2, characterized in that: The step of performing an amplification operation on a preset number of text data to obtain amplified text data includes: Inputting the instruction data corresponding to the preset number of text data into the augmentation model in sequence to generate augmentation instruction data; Determine a task identifier corresponding to the amplification instruction data, and generate amplification input data and amplification output data corresponding to the amplification instruction data according to the task identifier; Amplified text data is obtained according to the amplification instruction data, the amplification input data and the amplification output data.

4. The text data enhancement method according to claim 3, characterized in that: The step of generating the amplification input data and the amplification output data corresponding to the amplification instruction data according to the task identifier includes: Comparing the task identifier with a preset target task identifier; If the task identifier is the same as the preset target task identifier, generating amplified output data according to the amplification instruction data, and generating amplified input data according to the amplification output data; If the task identifier is different from the preset target task identifier, the amplification input data is generated according to the amplification instruction data, and the amplification output data is generated according to the amplification input data.

5. The text data enhancement method according to claim 1, characterized in that: The step of clustering the enhanced text data according to the feature data corresponding to the enhanced text data to obtain target text data includes: Selecting a preset number of target feature data from the feature data corresponding to the enhanced text data as cluster centers; Calculating the distance between the remaining feature data except the target feature data and the cluster center; Clustering the enhanced text data according to the distance, and updating the cluster center corresponding to each cluster; Until the cluster center corresponding to each cluster no longer changes, the enhanced text data corresponding to the final cluster center is used as the target text data.

6. The text data enhancement method according to any one of claims 1 to 5, characterized in that: After clustering the enhanced text data according to the feature data corresponding to the enhanced text data to obtain the target text data, the method includes: Using the text data and the target text data as training text data; Extracting first feature data and second feature data corresponding to the training instruction data and the training input data in the training text data; A target loss function is constructed based on the first feature data, the second feature data and the training output data, and a preset model is trained based on the training text data and the target loss function.

7. The text data enhancement method according to claim 6, characterized in that: Constructing a target loss function based on the first feature data, the second feature data, and the training output data, including: Constructing a positive sample pair and a negative sample pair based on the first feature data and the second feature data, and constructing a first loss function based on the positive sample pair and the negative sample pair; Decoding the first feature data and the second feature data respectively to obtain first decoded feature data and second decoded feature data; Constructing a second loss function based on the first decoded feature data and the training output data, and constructing a third loss function based on the second decoded feature data and the training output data; A target loss function is constructed based on the first loss function, the second loss function and the third loss function.

8. A text data enhancement device, characterized in that: The device comprises: An acquisition unit, used for acquiring text data; an amplification unit, used for performing an amplification operation on the text data to obtain enhanced text data; The clustering unit is used to cluster the enhanced text data according to the feature data corresponding to the enhanced text data to obtain target text data.

9. A text data enhancement system, characterized in that: It comprises a processor and a memory, wherein the memory stores a plurality of instructions; the processor loads instructions from the memory to execute the steps of the text data enhancement method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a plurality of instructions, and the instructions are suitable for being loaded by a processor to execute the steps of the text data enhancement method according to any one of claims 1 to 7.