A data processing method, apparatus, storage medium, and computer device

The method improves natural language processing model training by combining word embeddings and mixed labels to enhance data diversity, addressing the limitations of insufficient training data and semantic gaps, thereby increasing model efficiency and output variety.

CN113590803BActive Publication Date: 2025-07-15TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110209713.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-02-24
Publication Date
2025-07-15
Estimated Expiration
2041-02-24

AI Technical Summary

Technical Problem

In the prior art, due to insufficient training data, the training data is low in diversity, and it is difficult to learn replies between similar semantics. The model output diversity is poor, and simply increasing the number of samples will lead to increased training time and poor model diversity.

Method used

By obtaining the vocabulary collection of source samples and label samples, calculating the similar scores of the target vocabulary and synonyms, converting the vocabulary and synonyms into word vector sets, and mixing vectors, replacing the word vectors of the target vocabulary for training, generating mixed labels to iteratively train model parameters, and improving model diversity.

Benefits of technology

Without increasing the number of samples, the data processing efficiency and diversity of model output are improved, the data enhancement effect is achieved, and the training efficiency and diversity of the model are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113590803B_ABST
    Figure CN113590803B_ABST
Patent Text Reader

Abstract

An embodiment of the present application discloses a data processing method, which includes obtaining a vocabulary set corresponding to a source sample and a label sample; obtaining a target word and its synonyms, and calculating a similarity score between the target word and its synonyms; converting the words and synonyms in the vocabulary set into a word vector set, and performing vector mixing on the target word and the corresponding synonyms to obtain a mixed word vector; replacing the word vector of the corresponding target word with the mixed word vector and inputting it into a preset model for training; generating a mixed label, obtaining the difference between the word probability distribution of the mixed label and the word prediction probability distribution of the mixed label output by the preset model, and iteratively training the model parameters of the preset model according to the difference to obtain a trained preset model. In this way, the efficiency of data processing is improved, and the diversity of the output of the trained model is increased.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence technology, and in particular, to a data processing method, apparatus, storage medium, and computer device. Background Art

[0002] Natural Language Processing (NLP) is an important direction in the fields of computer science and artificial intelligence. It studies various theories and methods that can enable effective communication between humans and computers using natural language. With the development of computer technology and artificial intelligence technology, people's requirements for natural language processing technology have also been continuously increasing. However, in language training models, due to insufficient training data, the diversity of training data is low, and the semantics between different training data often vary greatly, making it difficult to learn responses between similar semantics, which reduces the diversity of the model output. Therefore, it is very necessary to use data augmentation to improve training performance.

[0003] In the current prior art, data augmentation is often performed by simply increasing the number of samples. On the one hand, the increased samples will lead to an increase in training time. On the other hand, even if multiple relatively similar samples are added, due to random sampling, it is difficult to train them within the same batch, resulting in poor data augmentation effects and poor diversity of the model output obtained from training. Summary of the Invention

[0004] Embodiments of the present application provide a data processing method, apparatus, storage medium, and computer device. It can improve the efficiency of data processing and the diversity of the model output after training.

[0005] Facilitate data conversion in different configuration environments and improve the efficiency of data configuration.

[0006] A data processing method includes:

[0007] Obtain a vocabulary set corresponding to source samples and label samples;

[0008] Obtain a target vocabulary and synonyms of the target vocabulary, and calculate a similarity score between the target vocabulary and its synonyms, where the target vocabulary is at least one vocabulary selected from the vocabulary set;

[0009] Convert each vocabulary and its synonyms in the vocabulary set into a set of word vectors, and perform vector mixing on the target vocabulary and its corresponding synonyms according to the similarity score to obtain a mixed word vector of the target vocabulary;

[0010] Replace the word vector of the corresponding target vocabulary with the mixed word vector, and input the set of word vectors after replacement into a preset model for training;

[0011] Generate a mixed label, obtain the difference between the word probability distribution of the mixed label and the word prediction probability distribution of the mixed label output by the preset model, and iteratively train the model parameters of the preset model according to the difference to obtain a trained preset model.

[0012] Correspondingly, an embodiment of the present application provides a data processing device, including:

[0013] A word segmentation unit for obtaining a vocabulary set corresponding to a source sample and a label sample;

[0014] An acquisition unit for obtaining a target word and a synonym of the target word, and calculating a similarity score between the target word and the synonym of the target word, where the target word is at least one word selected from the vocabulary set;

[0015] A mixing unit for converting each word and its synonym in the vocabulary set into a word vector set, and vectorially mixing the target word and its corresponding synonym according to the similarity score to obtain a mixed word vector of the target word;

[0016] A replacement unit for replacing the word vector of the corresponding target word with the mixed word vector, and inputting the word vector set after replacement into a preset model for training;

[0017] A training unit for generating a mixed label, obtaining the difference between the word probability distribution of the mixed label and the word prediction probability distribution of the mixed label output by the preset model, and iteratively training the model parameters of the preset model according to the difference to obtain a trained preset model.

[0018] In one embodiment, the mixing unit includes:

[0019] A conversion subunit for converting each word and its corresponding synonym in the vocabulary set into a word vector set through a word embedding layer of a preset model;

[0020] A calculation subunit for obtaining the weights of the word vectors of the target word and its corresponding synonym according to the similarity score;

[0021] A mixing subunit for weighted mixing of the word vectors of the target word and its corresponding synonym according to the weights to obtain a mixed word vector of the target word.

[0022] In one embodiment, the calculation subunit is used for:

[0023] Accumulating the similarity scores of the target word and its corresponding synonym to obtain the total score of the target word;

[0024] Calculate the ratio of the similarity score between the target vocabulary and its corresponding near-synonym to the total score to obtain the weight of the target vocabulary and its corresponding near-synonym.

[0025] In one embodiment, the training unit includes:

[0026] A determination subunit, configured to determine the target label of the preset model according to the vocabulary set of the label sample;

[0027] A construction subunit, configured to perform soft label construction on the target vocabulary corresponding to the vocabulary set included in the target label based on the target near-synonym of the target vocabulary in the vocabulary set of the label sample and the similarity score corresponding to the target near-synonym;

[0028] A combination subunit, configured to combine the target label and the soft label to obtain the mixed label of the preset model.

[0029] In one embodiment, the construction subunit is configured to:

[0030] Obtain the expected probabilities of the target vocabulary and the target near-synonym in the label sample according to the target near-synonym of the target vocabulary in the vocabulary set of the label sample and the similarity score corresponding to the target near-synonym;

[0031] Based on the expected probabilities of the target vocabulary and the target near-synonym, obtain the word probability distribution of the target vocabulary;

[0032] Obtain the word probability distribution of the target vocabulary, and perform soft label construction on the target vocabulary corresponding to the vocabulary set included in the target label.

[0033] In one embodiment, the training unit includes:

[0034] An acquisition subunit, configured to acquire the word probability distribution of each vocabulary in the mixed label and the word prediction probability distribution of the corresponding vocabulary in the mixed label output by the preset model;

[0035] An input subunit, configured to input the word probability distribution and the word prediction probability distribution into the loss function of the preset model to obtain the target loss;

[0036] A training subunit, configured to iteratively train the model parameters of the preset model according to the target loss, and obtain the trained preset model when the target loss meets the convergence condition.

[0037] In one embodiment, the input subunit is configured to:

[0038] Input the word probability distribution and the word prediction probability distribution into the loss function of the preset model to obtain the loss of each vocabulary in the mixed label output by the preset model;

[0039] Accumulate the losses of each word in the mixed labels output by the preset model to obtain a total loss value;

[0040] Perform an averaging process on the total loss value to obtain a target loss.

[0041] In one embodiment, the data processing device further includes:

[0042] A filtering unit, configured to perform part-of-speech analysis on the words in the vocabulary set, and filter the words with the target part of speech in the vocabulary set according to the part-of-speech analysis result;

[0043] A determining unit, configured to determine a preset replacement ratio, and determine the number of target words in the filtered vocabulary set according to the preset replacement ratio;

[0044] A selecting unit, configured to randomly select words from the filtered vocabulary set according to the number of the target words, and obtain target words according to the random selection result.

[0045] In one embodiment, the data processing device further includes:

[0046] A deleting subunit, configured to delete the synonyms of the target word and the target word whose similarity scores are not greater than a preset threshold.

[0047] In addition, an embodiment of the present application further provides a text generation method, and the method includes:

[0048] Receive user request information, where the user request information includes text data input by the user;

[0049] Input the text data into a trained preset model, where the model parameters of the preset model are trained by using any data processing method provided by the embodiment of the present application;

[0050] Determine the output result of the trained preset model as target text data.

[0051] In addition, an embodiment of the present application further provides a storage medium, which stores multiple instructions, and the instructions are suitable for being loaded by a processor to execute the steps in any data processing method provided by the embodiment of the present application.

[0052] In addition, an embodiment of the present application further provides a computer device, including a processor and a memory, where the memory stores an application program, and the processor is configured to run the application program in the memory to implement the data processing method provided by the embodiment of the present application.

[0053] The embodiments of the present application also provide a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a storage medium. A processor of a computer device reads the computer instructions from the storage medium, and the processor executes the computer instructions, so that the computer device executes the steps in the data processing method provided by the embodiments of the present application.

[0054] In the embodiments of the present application, a vocabulary set corresponding to a source sample and a label sample is obtained; target vocabulary and synonyms of the target vocabulary are obtained, and a similarity score between the target vocabulary and the synonyms of the target vocabulary is calculated; each vocabulary in the vocabulary set and the corresponding synonyms are converted into a word vector set, and the target vocabulary and the synonyms are vector-mixed according to the similarity score to obtain a mixed word vector of the target vocabulary; the mixed word vector is used to replace the word vector of the corresponding target vocabulary, and the replaced word vector set is input into a preset model for training; a mixed label is generated, and the difference between the word probability distribution of the mixed label and the word prediction probability distribution of the mixed label output by the preset model is obtained, and the model parameters of the preset model are iteratively trained according to the difference to obtain a trained preset model. In this way, without increasing the number of samples, the obtained source samples and the corresponding label samples are segmented to obtain a vocabulary set, and synonyms and similarity scores of the target vocabulary in the vocabulary set are obtained. By converting each vocabulary in the vocabulary set and the corresponding synonyms into word vectors, mixing the target vocabulary and the corresponding synonyms and replacing the word vector of the corresponding target vocabulary and inputting them into a preset model for training, and then iteratively training the preset model according to the difference between the determined mixed label and the output of the preset model, a trained preset model is obtained, effectively realizing data augmentation, improving the efficiency of data processing, and increasing the diversity of the output of the trained model. BRIEF DESCRIPTION OF THE DRAWINGS

[0055] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those skilled in the art, without creative efforts, other drawings can be obtained according to these drawings.

[0056] Figure 1 FIG. is a schematic diagram of an implementation scenario of a data processing method provided by an embodiment of the present application;

[0057] Figure 2 FIG. is a flowchart of a data processing method provided by an embodiment of the present application;

[0058] Figure 3 FIG. is another flowchart of a data processing method provided by an embodiment of the present application;

[0059] Figure 4 It is a schematic diagram of an application scenario of a data processing method provided by an embodiment of the present application;

[0060] Figure 5 It is another schematic diagram of an application scenario of a data processing method provided by an embodiment of the present application;

[0061] Figure 6 It is a schematic diagram of the structure of a data processing device provided by an embodiment of the present application;

[0062] Figure 7 It is a schematic diagram of the structure of a computer device provided by an embodiment of the present application. Detailed implementation manners

[0063] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present application.

[0064] The embodiments of the present application provide a data processing method, device, storage medium, and computer device. Among them, the data processing device can be integrated in the computer device, and the computer device can be a server or a terminal device, etc.

[0065] For better illustration of the embodiments of the present application, please refer to the following terms for reference:

[0066] Artificial Intelligence (AI) is to use a digital computer or a machine controlled by a digital computer to simulate, extend, and expand human intelligence, a theory, method, technology, and application system that can perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science. It attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Artificial intelligence is also to study the design principles and implementation methods of various intelligent machines to enable the machine to have the functions of perception, reasoning, and decision-making.

[0067] Artificial intelligence technology is an interdisciplinary subject, involving a wide range of fields, including both hardware-level technologies and software-level technologies. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0068] Among them, natural language processing is an important direction in the fields of computer science and artificial intelligence. It studies various theories and methods that can achieve effective communication between humans and computers in natural language. Natural language processing is a science that integrates linguistics, computer science, and mathematics. Therefore, the research in this field will involve natural language, that is, the language people use in daily life, so it has a close connection with the research of linguistics. Natural language processing technologies usually include text processing, semantic understanding, machine translation, robot question answering, knowledge graphs, and other technologies.

[0069] Among them, Word Embedding is a general term for language model and representation learning technologies in natural language processing (NLP). It refers to embedding a high-dimensional space with the dimension of the number of all words into a much lower-dimensional continuous vector space, and each word or phrase is mapped to a vector in the real number domain.

[0070] Data augmentation is a relatively broad concept. It can refer to enhancing the quality of training data, increasing the diversity of data, or simply increasing the quantity of data. However, its fundamental purpose is still to enable the artificial intelligence model to be trained to have better performance in the domain where the dataset is located.

[0071] With the research and progress of artificial intelligence technology, artificial intelligence technology has been studied and applied in multiple fields, such as common smart homes, smart wearable devices, virtual assistants, smart speakers, smart marketing, driverless, autonomous driving, drones, robots, smart healthcare, smart customer service, etc. It is believed that with the development of technology, artificial intelligence technology will be applied in more fields and play an increasingly important role.

[0072] Please refer to Figure 1 , taking the integration of the data processing device in a computer device as an example. Figure 1This is a schematic diagram of the implementation environment scenario of the data processing method provided by the embodiments of this application, including server A and terminal B. Among them, server A can be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, network acceleration services (Content Delivery Network, CDN), and big data and artificial intelligence platforms. Server A can obtain the vocabulary sets corresponding to the source samples and label samples; obtain the target vocabulary and the synonyms of the target vocabulary, and calculate the similarity scores between the target vocabulary and its synonyms. The target vocabulary is at least one vocabulary selected from the vocabulary set; convert each vocabulary and its synonyms in the vocabulary set into a word vector set, and mix the vectors of the target vocabulary and its corresponding synonyms according to the similarity scores to obtain the mixed word vector of the target vocabulary; replace the word vector of the corresponding target vocabulary with the mixed word vector, and input the replaced word vector set into a preset model for training; generate a mixed label, obtain the difference between the word probability distribution of the mixed label and the word prediction probability distribution of the mixed label output by the preset model, and iteratively train the model parameters of the preset model according to the difference to obtain the trained preset model.

[0073] Terminal B can be various computer devices that can perform data input, such as smart phones, tablet computers, laptop computers, desktop computers, etc., but is not limited thereto. Terminal B and server A can be directly or indirectly connected through wired or wireless communication methods. Server A can receive the data uploaded by terminal B to perform corresponding data processing operations, which are not limited in this application.

[0074] It should be noted that Figure 1 The schematic diagram of the implementation environment scenario of the data processing method shown is only an example. The implementation environment scenario of the data processing method described in the embodiments of this application is to more clearly illustrate the technical solutions of the embodiments of this application, and does not constitute a limitation on the technical solutions provided by the embodiments of this application. Those of ordinary skill in the art know that with the evolution of data processing and the emergence of new business scenarios, the technical solutions provided by this application are equally applicable to similar technical problems.

[0075] The solution provided by the embodiments of this application involves technologies such as natural language processing in artificial intelligence, and is specifically described through the following embodiments. It should be noted that the description order of the following embodiments does not limit the preferred order of the embodiments.

[0076] This embodiment will be described from the perspective of a data processing device, which can be specifically integrated in a computer device. The computer device can be a server or a terminal, and this application does not limit it here.

[0077] Please refer to Figure 2 , Figure 2 which is a schematic flowchart of the data processing method provided by the embodiment of this application. The data processing method includes:

[0078] In step 101, obtain the vocabulary sets corresponding to the source sample and the label sample.

[0079] Among them, the label sample has a corresponding relationship with the source sample. The label sample can be the answer corresponding to the source sample or the required processing result. For example, the source sample can be "What did you have for lunch", and correspondingly, the label sample can be "I had a sandwich" or "I had rice", etc.; the source sample can also be a piece of text data, and correspondingly, the label sample can be the text data obtained after translating this piece of text data into the target language; the source sample can also be an article, and correspondingly, the label sample can be the abstract corresponding to this article, etc.

[0080] Perform word segmentation processing on the obtained source sample and label sample through text preprocessing. Among them, different source samples and label samples in different languages or different fields can have different word segmentation processing methods. For example, for English, word segmentation can be performed through the spaces between words, for Chinese, a word segmentation tool can be used for processing, and for source samples and label samples in professional fields, word segmentation algorithms can be designed according to the characteristics of samples in different professional fields to achieve word segmentation processing, etc., to obtain the vocabulary sets after word segmentation processing of the source sample and the label sample. For example, for the source sample "What did you have for lunch" and the corresponding label sample "I had a sandwich", through word segmentation processing, the vocabulary set is obtained. The vocabulary set of this source sample can be "you, lunch, eat, of, what", and the vocabulary set of this label sample can be "I, eat, had, sandwich", etc.

[0081] Training sample data such as the source sample and the corresponding label sample can be obtained from the memory connected to the data processing device, or from other data storage terminals. It can also be obtained from the memory of an entity terminal, or from a virtual storage space such as a dataset or a corpus. In some embodiments, the training sample data can be obtained from one storage location or from multiple storage locations. For example, the training sample data can be saved on a blockchain, and the data processing device obtains the above training sample data from the blockchain. The data processing device can centrally obtain the training sample data in a time period in response to a certain training sample data acquisition instruction, or continuously obtain the training sample data according to a certain data acquisition logic.

[0082] In step 102, the target word and its synonyms are obtained, and the similarity score between the target word and its synonyms is calculated.

[0083] To achieve the effect of data augmentation, the target words in the vocabulary set are expanded with synonyms. Specifically, the target words in the vocabulary set and their synonyms are obtained, and the similarity score between the target words and their synonyms is calculated. The target word is at least one word selected from the vocabulary set that needs to be expanded with synonyms. Here, the synonyms can be words with the same or similar meanings as the target word, such as "happy" and "joyful", etc., or words that can be substituted for each other in some contexts. For example, when describing what was eaten at noon, it can be said that rice was eaten at noon, or it can be said that a hamburger was eaten at noon. In this context, "rice" and "hamburger" are synonyms. Among them, the larger the similarity score, the higher the similarity to the target word. For example, for the vocabulary set "you, lunch, eat, of, what", the "lunch" in this vocabulary set can be selected as the target word. By obtaining the synonyms of "lunch", the synonym expansion of the target word "lunch" is achieved, thus obtaining synonyms such as "dinner", "breakfast", and "rice", and calculating the similarity scores between the target word "lunch" and its synonyms such as "dinner", "breakfast", and "rice". Among them, the similarity score between the target word "lunch" and "lunch" will also be calculated. For example, the similarity scores between the target word "lunch" and its synonyms "lunch", "dinner", "breakfast", and "rice" can be 1, 0.5, 0.3, and 0.1, etc.

[0084] In one embodiment, the synonyms of the target word can be obtained through a synonym prediction model, and the similarity score between the target word and its synonyms is calculated. For example, a synonym prediction model based on the fastText model or the WordNet database can be used to obtain the synonyms of the target word and calculate the similarity score between the target word and its synonyms. Or a pre-trained language model such as the BERT (Bidirectional Encoder Representation from Transformers) model can be used to obtain the synonyms of the target word and calculate the similarity score between the target word and its synonyms, etc.

[0085] In one embodiment, by performing part-of-speech analysis on the words in the vocabulary set, the words in the vocabulary set that do not need to be expanded with synonyms can be filtered out. The parts of speech of the words that do not need to be expanded with synonyms can be prepositions, articles, etc. For example, for the vocabulary set "you, lunch, eat, of, what", "you", "of" and "what" are words with parts of speech such as prepositions and articles. Since replacing these words cannot enhance the diversity of the data and may even destroy the original semantics of the sample, the words with these parts of speech are filtered out, and then selected from the filtered vocabulary set to obtain the target vocabulary that needs to be expanded with synonyms.

[0086] In one embodiment, the number of target words to be selected can be determined based on the number of words included in the sample data or based on factors such as the number of sample data, wherein a replacement ratio can be set and the number of target words in the filtered vocabulary set can be determined based on the replacement ratio, thereby selecting target words in the filtered vocabulary set based on the number of target words determined by the replacement ratio.

[0087] In one embodiment, in order to obtain synonyms with a relatively high similarity to the target vocabulary and improve the efficiency of training, a threshold can be set to delete synonyms with a similarity score not greater than the threshold, thereby obtaining synonyms with similarity that meets the requirements. For example, the threshold can be set to 0.2, and synonyms with a similarity score not greater than 0.2 are deleted. Assuming that the similarity scores between the target vocabulary "lunch" and synonyms such as "lunch", "dinner", "breakfast" and "rice" are 1, 0.5, 0.3 and 0.1, the synonym "rice" with a similarity score less than 0.1 is deleted.

[0088] In step 103, each word and synonym in the vocabulary set is converted into a word vector set, and the target word and the corresponding synonym are vector-mixed according to the similarity score to obtain a mixed word vector of the target word.

[0089] Among them, in order to input the words in the vocabulary set and the synonyms corresponding to the target word into the model for training, each word in the vocabulary set and the synonyms of the target word can be converted into word vectors in the word embedding layer of the preset model, and the word vector of the target word and the word vectors of the corresponding synonyms are vector-mixed according to the similarity scores to obtain the mixed word vector of the target word. Among them, the preset model can be a generative model, where the generative model refers to a model that can randomly generate observed data, and is a model that learns the joint probability distribution P(X,Y) through sample data, that is, the probability that the source sample X and the label sample Y appear together, and then obtains the conditional probability distribution P(Y / X) as the prediction model. For example, the generative model can be a dialogue generation model. The generative model learns the joint probability distribution through the source sample and the corresponding label sample, that is, the probability that the source sample and the label sample appear together, and then obtains the conditional probability distribution, that is, the probability that the label sample appears when the source sample appears, so as to be used as the prediction model. Among them, word embedding is a method of converting words in text into digital vectors. In order to analyze and calculate them using standard machine learning algorithms, these vectors converted into numbers need to be used as input in digital form. The word embedding process is to embed a high-dimensional space with the dimension of the number of all words into a much lower-dimensional continuous vector space. Each word is represented as a real number vector in the predefined vector space, and each word is mapped to a vector. For example, in a text containing words such as "lunch", "dinner", and "rice", these words are mapped into the vector space. The vector corresponding to "lunch" can be (0.1, 0.2, 0.3), the vector corresponding to "dinner" can be (0.2, 0.2, 0.4), and the vector corresponding to "rice" can be (0.3, -0.4, -0.2). Therefore, by converting words into word vectors through word embedding, the machine can calculate words. For example, the similarity between words can be obtained by calculating the cosine value of the angle between different word vectors.

[0090] Among them, it can be assumed that each target word and the corresponding synonyms form a mixed set as follows:

[0091] C = ((c0, s0), (c1, s1), (c2, s2), …, (c k , s k ))

[0092] Among them, c0 is the target word itself, s0 = 1, representing that the similarity score of the target word itself is 1, and c1 to c k are the synonyms of the target word obtained, and s1 to s k represent the similarity scores between the target word and the synonyms. Among them, it is assumed that the word vector corresponding to each word in the mixed set C is e i ∈R d, where i represents the i-th word in the mixture set, d represents the dimension, and R d represents a d-dimensional vector space.

[0093] In one embodiment, the weight for vector mixing of the word vector of the target word and the corresponding word vectors of the near-synonyms of the target word can be obtained according to the similarity scores between the target word and its near-synonyms. For example, the weight for vector mixing can be obtained according to the expected probability of the target word and each corresponding near-synonym. The expected probability is the probability that the preset model outputs the target word and its near-synonyms. The formula for the expected probability is:

[0094]

[0095] where p(c i ) represents the expected probability of each word in the mixture set C, s i represents the similarity score corresponding to the i-th word in the mixture set C, and s j represents the similarity score of the j-th word from s0 to s k . Thus, the weight for vector mixing is obtained. According to this weight, the target word and its corresponding near-synonyms can be weighted and added to obtain the mixed word vector of the target word:

[0096]

[0097] In one embodiment, it is also possible to input the target word, the near-synonyms of the target word, and the corresponding similarity scores into a neural network model for training, and calculate the loss of the model through a loss function. When the loss of the model meets the convergence condition, a trained neural network model is obtained. According to the trained neural network model, the weight for vector mixing is obtained, and thus, according to this weight, the target word and its corresponding near-synonyms are weighted and added to obtain the mixed word vector of the target word.

[0098] In step 104, the mixed word vector is used to replace the word vector of the corresponding target word, and the set of word vectors after replacement is input into the preset model for training.

[0099] Among them, in order to enable training of the sample data obtained after expanding near-synonyms without increasing the number of samples, the mixed word vector obtained by vector mixing can be used to replace the word vector of the corresponding target word. For example, for a source sample with a vocabulary set of "you, lunch, eat, of, what", "lunch" is selected as the target word. The mixed word vector of this target word is obtained by vector mixing the word vectors corresponding to each word in the mixture set C of this target word. Thus, the mixed vector is used to replace the word vector of the target word "lunch", and the set of word vectors after replacement is input into the preset model for training.

[0100] In the current prior art, to obtain a training model with better performance, a large amount of sample data needs to be input into the model for training. However, in practical applications, there are often not enough sample data available for model training. Therefore, to solve the problem of lack of sample data, those skilled in the art often need to perform data augmentation on the limited sample data. However, in existing data augmentation methods, most are based on the token level for data augmentation. For example, according to the distribution probability of tokens or synonyms, token replacement is directly performed at the text level, or according to the language model, some randomly selected tokens are predicted to expand the data volume, or the existing sample data is rewritten to increase the sample data, and so on. This method of simply increasing the number of sample data to achieve data augmentation, on the one hand, the increased sample data will lead to an increase in the training time of the model. On the other hand, there is still only a fixed label, so that the model still learns in a one-to-one mode during the training process. Even if multiple relatively similar samples are added, due to random sampling, it is difficult to learn within the same batch, making it difficult for the model to train on samples with similar semantics. Therefore, the diversity of the trained model is poor.

[0101] To solve the above problems, the embodiments of the present application provide a data processing method. By performing word segmentation on the obtained source samples and corresponding label samples to obtain a vocabulary set, and obtaining the synonyms and similarity scores of the target words in the vocabulary set, converting each word in the vocabulary set and its corresponding synonyms into word vectors, and mixing the target word and its corresponding synonyms and replacing the word vector of the corresponding target word and inputting it into a preset model for training, data augmentation is effectively achieved without increasing the number of samples. At the same time, by introducing a mixed label, the rationality that multiple source samples can correspond to multiple label samples is retained, thereby realizing the many-to-many training of the model and improving the diversity of the output of the trained model. The specific implementation process is as follows.

[0102] In step 105, a mixed label is generated, the difference between the word probability distribution of the mixed label and the word prediction probability distribution of the mixed label output by the preset model is obtained, and the model parameters of the preset model are iteratively trained according to the difference to obtain the trained preset model.

[0103] To correspond to the hybrid word vectors of the target vocabulary, the embodiments of the present application introduce hybrid labels to achieve many-to-many training and increase the diversity of model outputs. In the embodiments of the present application, corresponding hybrid labels are generated according to the vocabulary set of the label samples, the target synonyms of the target vocabulary in the label samples, and the similarity scores corresponding to the target synonyms. For example, for a label sample with a vocabulary set of "I, ate, a, sandwich", "sandwich" can be selected as the target vocabulary, and the synonyms of the target vocabulary and the corresponding similarity scores can be obtained through synonym expansion, so as to obtain a hybrid set C = ((sandwich, 1), (hamburger, 0.6), (salad, 0.2)). According to the vocabulary set of the label sample, the target label "I, ate, a, sandwich" is obtained.

[0104] Furthermore, according to the target synonyms of the target vocabulary in the label samples and the similarity scores corresponding to the target synonyms, a soft label is constructed based on the target vocabulary included in the target label. For example, the target vocabulary, target synonyms, and corresponding similarity scores in the hybrid set C of the target vocabulary in the label sample are calculated through the expected probability formula in step 103 above to obtain the expected probability of each vocabulary in the hybrid set of the target vocabulary. Then, according to the expected probability of the target vocabulary and the expected probability of the synonyms of the target vocabulary, a soft label is constructed based on the target vocabulary corresponding to the vocabulary set included in the target label. Herein, a soft label refers to a label composed of multiple words with carrying probabilities, and the probabilities carried by these words add up to 1. For example, the soft label of the target vocabulary "sandwich" can be composed of "sandwich" and the corresponding probability "0.5", "hamburger" and the corresponding probability "0.3", and "salad" and the corresponding probability "0.2", so as to obtain the word probability distribution of the target vocabulary. Combining the target label with the soft label of the target vocabulary, the hybrid label of the label sample is obtained. Herein, the hybrid label carries the word probability distribution, and the word probability distribution carried by the hybrid label can be expressed as "(I, 1), (ate, 1), (a, 1), [(sandwich, 0.5), (hamburger, 0.3), (salad, 0.2)]". Among them, for "I", "ate", and "a" in the target label, since no corresponding synonym expansion is performed, the expected probabilities of these words can be 1, and the word probability distribution of each word is obtained accordingly.

[0105] Among them, in order to optimize the model parameters of the preset model, the word prediction probability distribution of each word in the mixed label output by the preset model can be obtained and calculated with the word probability distribution of the corresponding word in the mixed label, so as to obtain the difference between the word probability distribution of the mixed label and the word prediction probability distribution of the mixed label output by the preset model. This difference can be calculated by a loss function. In one embodiment, for the words with soft labels constructed in the mixed label, the cross-entropy loss function can be used to calculate the difference. The calculation formula of this cross-entropy loss function is as follows:

[0106]

[0107] where c i is the i-th word in the mixed set C of target words, and p(c i ) represents the expected probability of c i in the mixed set C, and g(c i ) is the probability that the preset model outputs c i . Thus, the difference of each soft label is obtained by calculating this cross-entropy loss function. In one embodiment, for the words without soft labels constructed in the mixed label, the following logarithmic loss function formula can be used to calculate the difference:

[0108] L = -logg(c i )

[0109] The difference of each word in the mixed label is calculated by the loss function, the differences of each word are accumulated, and the accumulated result is divided by the number of words in the mixed label to obtain the difference of this mixed label. For example, for the mixed label expressed as "(I, 1), (eat, 1), (already, 1), [(sandwich, 0.5), (hamburger, 0.3), (salad, 0.2)]", the first difference can be calculated for the soft label composed of the sandwich, hamburger, salad and their corresponding expected probabilities through the calculation formula of the cross-entropy loss function, the second difference can be calculated for "I", "eat" and "already" through the logarithmic loss function formula, and the difference of the current preset model is obtained by accumulating and averaging the first difference and the second difference. The model parameters of the preset model are iteratively trained according to this difference to optimize the model parameters of the preset model. When the preset model meets the convergence condition, the trained preset model is obtained. Thus, by introducing the mixed label for data augmentation, one source sample can correspond to multiple label samples, so as to perform many-to-many training of the preset model, ensuring that the trained preset model can output diverse results.

[0110] As can be seen from the above, in the embodiment of the present application, a vocabulary set corresponding to a source sample and a label sample is obtained; a target vocabulary and a synonym of the target vocabulary are obtained, and a similarity score between the target vocabulary and the synonym of the target vocabulary is calculated; each vocabulary and synonym in the vocabulary set are converted into a word vector set, and the target vocabulary and the corresponding synonym are vector-mixed according to the similarity score to obtain a mixed word vector of the target vocabulary; the mixed word vector is used to replace the word vector of the corresponding target vocabulary, and the replaced word vector set is input into a preset model for training; a mixed label is generated, the difference between the word probability distribution of the mixed label and the word prediction probability distribution of the mixed label output by the preset model is obtained, and the model parameters of the preset model are iteratively trained according to the difference to obtain a trained preset model. In this way, without increasing the number of samples, the vocabulary set of the obtained source sample and the corresponding label sample is obtained, the synonyms and corresponding similarity scores of the target vocabulary in the vocabulary set are obtained, each vocabulary and the corresponding synonym in the vocabulary set are converted into word vectors, the target vocabulary and the corresponding synonym are mixed and the word vector of the corresponding target vocabulary is replaced and input into the preset model for training, and then the preset model is iteratively trained by the difference between the mixed label and the output of the preset model to obtain a trained preset model, effectively realizing data augmentation, improving the efficiency of data processing, and increasing the diversity of the output of the trained model.

[0111] According to the method described in the above embodiment, the following will give an example for further detailed description.

[0112] In this embodiment, it will be described by taking the data processing device being specifically integrated in a computer device as an example. Among them, the data processing method takes a server as the execution subject, and at the same time uses a synonym prediction model to obtain the synonyms and corresponding similarity scores of the target vocabulary.

[0113] As Figure 3 shown, Figure 3 is another flowchart of the data processing method provided by the embodiment of the present application. The specific process is as follows:

[0114] In step 201, the server obtains a vocabulary set corresponding to a source sample and a label sample, performs part-of-speech analysis on the vocabulary in the vocabulary set, and filters the vocabulary with the target part of speech in the vocabulary set according to the part-of-speech analysis result.

[0115] Among them, the server obtains source samples and corresponding labeled samples. For example, the corresponding sample data can be obtained through a corpus or a dataset, etc. Correspondingly, the obtained source samples and corresponding labeled samples can be tokenized through text preprocessing to obtain a vocabulary set. For example, for the sample data where the source sample is "What did you have for lunch?" and the corresponding labeled sample is "I had a sandwich", the vocabulary set of the source sample obtained through tokenization is "you, lunch, have, for, what", and the vocabulary set of this labeled sample can be "I, had, a, sandwich". Then, part-of-speech analysis is performed on the words in the vocabulary set to obtain the part-of-speech analysis results of each word in the vocabulary set. According to the part-of-speech analysis results, the words with the target part of speech in the vocabulary set are filtered. Among them, the target part of speech can be parts of speech such as prepositions and articles. For example, words such as "you", "for", and "what" in the vocabulary set "you, lunch, have, for, what", replacing these words cannot enhance the diversity of the data and may even damage the original semantics of the sample. Therefore, the words of these parts of speech are filtered, and then selection is made from the words remaining after removing these words from the vocabulary set to obtain the target words that need to be expanded with synonyms.

[0116] In step 202, the server determines a preset replacement ratio, determines the number of target words in the filtered vocabulary set according to the preset replacement ratio, randomly selects words from the filtered vocabulary set according to the number of target words, and obtains the target words according to the random selection result.

[0117] Among them, the number of target words can be limited by setting a preset replacement ratio. The server can determine the preset replacement ratio, and this replacement ratio can be 50%, or 100%, or 20%, etc. The number of target words in the filtered vocabulary set is determined according to the preset replacement ratio. For example, for the source sample with the vocabulary set "you, lunch, have, for, what", after filtering out the words with the target part of speech, the words "lunch, have" are obtained. A replacement ratio of 50% can be set, then for the words "lunch, have", the number of target words can be obtained as 1 according to this replacement ratio. One target word is randomly selected from the words "lunch, have" in the filtered vocabulary set according to the number of target words, and the target word is obtained according to the random selection result. For example, this target word can be "lunch" or "have". In this way, by setting a preset replacement ratio to limit the number of target words to meet the needs of model training and improve the efficiency of preset model training.

[0118] In step 203, the server obtains the target words and their synonyms, calculates the similarity scores between the target words and their synonyms, and deletes the synonyms with similarity scores not greater than the preset threshold among the target words and their synonyms.

[0119] Among them, in order to achieve data augmentation, synonyms of the target word can be obtained through a synonym prediction model or a pre-trained language model, and the similarity score between the target word and the synonyms can be calculated. To improve the training efficiency, synonyms with low similarity can be deleted. For example, by setting a preset threshold, synonyms with a similarity score not greater than the preset threshold among the target word and its synonyms can be deleted to remove synonyms with low similarity. In one embodiment, please refer to Figure 4 , Figure 4 FIG. is a schematic diagram of an application scenario of a data processing method provided by an embodiment of the present application, which is one of the application scenario diagrams provided to more clearly illustrate the technical solution of the embodiment of the present application and does not constitute a limitation on the technical solution provided by the embodiment of the present application. Among them, in Figure 4 In the provided application scenario diagram, the synonym prediction model 110 is used to expand the synonyms of the target words in the vocabulary set H including the source samples and the vocabulary set R of the label samples, and delete the synonyms with a similarity score not greater than the preset threshold. Thus, the synonyms "dinner" and "breakfast" of the target word "lunch" in the vocabulary set corresponding to the source samples can be obtained, and the corresponding similarity scores 0.5 and 0.3 can be obtained, as well as the similarity score of the target word is 1. Among them, the larger the similarity score, the higher the similarity to the target word. At the same time, the synonyms of the target word "sandwich" in the vocabulary set R of the label samples are obtained through the synonym prediction model 110, such as "hamburger" and "salad". Correspondingly, the similarity scores of the target words "sandwich", "hamburger" and "salad" can be obtained as 1, 0.6 and 0.4 respectively.

[0120] In step 204, the server converts each word and its synonyms in the vocabulary set into a word vector set through the word embedding layer of the preset model, and accumulates the similarity scores of the target word and its corresponding synonyms to obtain the total score of the target word.

[0121] Among them, in order to input the words in the vocabulary set into the model for analysis and calculation to achieve the purpose of training, the server can convert each word and its corresponding synonyms in the vocabulary set into a word vector set through the word embedding layer of the preset model. In order to input the word vector corresponding to the target word and the word vector corresponding to the synonym of the target word into the model for training without increasing the number of samples, the corresponding target word and its synonyms can be vector-mixed. Among them, please continue to refer to Figure 4, the server converts the words in the vocabulary sets H and R and the synonyms of the target word therein into word vectors through the word embedding layer 131 in the dialogue generation model 130. Among them, in order to obtain the weights for vector mixing of the target word and its synonyms, the similarity scores of the target word and its corresponding synonyms can be accumulated to obtain the total score of the target word. For example, please continue to refer to Figure 4 , for the target word "sandwich" in the vocabulary set R and its corresponding synonyms "hamburger" and "salad", accumulate the similarity scores of these words, that is, accumulate the similarity scores 1, 0.6, and 0.4 to get the total score 2, and then obtain the weights of the target word and its corresponding synonyms according to the ratio of the similarity score of the target word and its corresponding synonyms to this total score. For the specific implementation, please continue to refer to the following steps.

[0122] In step 205, the server calculates the ratio of the similarity score and the total score of the target word and its corresponding synonyms to obtain the weights of the target word and its corresponding synonyms, and performs weighted mixing on the word vectors of the target word and its corresponding synonyms according to the weights to obtain the mixed word vector of the target word.

[0123] Among them, the server calculates the ratio of the similarity score and the total score of the target word and its corresponding synonyms. For example, in one embodiment, please continue to refer to Figure 4 , calculate the ratio of the similarity score and the total score 2 of the target word "sandwich" in the vocabulary set R and its corresponding synonyms "hamburger" and "salad". Thus, the weights of the target word and its corresponding synonyms are obtained respectively. The weights of "sandwich", "hamburger", and "salad" are 0.5, 0.3, and 0.2. Perform weighted mixing on the word vectors of the target word and its corresponding synonyms according to the above weights, that is, according to the following formula

[0124]

[0125] perform weighted addition, where c i is the i-th word in the mixed set C = ((sandwich, 1), (hamburger, 0.6), (salad, 0.2)) of the target word "sandwich", p(c i ) is the expected probability of the i-th word in the mixed set C of the target word "sandwich", that is, the weight of the word vectors of the target word and its corresponding synonyms, e i is the word vector corresponding to the i-th word in the mixed set C of the target word "sandwich", so as to obtain the mixed word vector of the target word.

[0126] In step 206, the server replaces the word vector of the corresponding target word with the mixed word vector, and inputs the replaced word vector set into the preset model for training.

[0127] Among them, the server replaces the word vector of the target vocabulary with the mixed word vector of the target vocabulary, and inputs the set of word vectors after replacement into a preset model for training. For example, please continue to refer to Figure 4 Replace the word vector of the target vocabulary in the set of word vectors corresponding to the vocabulary set H with the mixed word vector of the target vocabulary of the source sample, so as to input the set of word vectors after replacement into the dialogue generation model 130. The encoder encodes the vocabulary set H of the source sample to obtain the corresponding hidden variable and sends it to the decoder. The decoder receives the hidden variable sent by the encoder for training.

[0128] In step 207, the server determines the target label of the preset model according to the vocabulary set of the label sample, and obtains the expected probabilities of the target vocabulary and the target near-synonyms in the label sample according to the target near-synonyms of the target vocabulary in the vocabulary set of the label sample and the corresponding similarity scores.

[0129] Among them, the server determines the target label of the preset model according to the vocabulary set of the label sample. For example, please continue to refer to Figure 4 The target label of the preset model can be determined as "I, ate, a, sandwich" according to the vocabulary set R of the label sample "I ate a sandwich". In order to correspond the target vocabulary included in the target label with the mixed word vector of the target vocabulary, corresponding soft labels are constructed for the target vocabulary in this application embodiment. The expected probabilities of the target vocabulary and the target near-synonyms in the label sample can be obtained according to the target near-synonyms of the target vocabulary in the vocabulary set of the label sample and the corresponding similarity scores. For example, the expected probabilities of the target vocabulary and the target near-synonyms in the label sample can be calculated according to the following formula:

[0130]

[0131] Among them, s i represents the similarity score of the i-th vocabulary in the mixed set C of the target vocabulary. For example, please refer to Figure 4 s i can represent the similarity score of the i-th vocabulary in the mixed set 140 of the target vocabulary "sandwich". The expected probability of each word in the mixed set 140 of "sandwich" is calculated according to the calculation formula of the expected probability, and thus a soft label is constructed for the target vocabulary "sandwich" included in the target label. The specific implementation can continue to refer to the following steps.

[0132] In step 208, the server obtains the word probability distribution of the target vocabulary based on the expected probabilities of the target vocabulary and the target near-synonyms, obtains the word probability distribution of the target vocabulary, constructs a soft label based on the target vocabulary corresponding to the vocabulary set included in the target label, and combines the target label and the soft label to obtain the mixed label of the preset model.

[0133] To correspond to the mixed word vectors input to the preset model, the embodiments of the present application introduce mixed labels based on target labels and soft labels to implement many-to-many training of the preset model. Among them, the server obtains the word probability distribution of the target word based on the expected probabilities of the target word and the target near-synonyms. The target near-synonym is a near-synonym of the target word in the vocabulary set of the label samples. In one embodiment, please continue to refer to Figure 4 , based on the expected probabilities of the target word "sandwich" and the target near-synonyms "hamburger" and "salad", the word probability distribution of the target word "sandwich" is obtained, which can be expressed as "sandwich 0.5, hamburger 0.3, salad 0.2". Thus, a soft label is constructed based on the target word corresponding to the vocabulary set included in the target label. Combining the target label and the soft label 121 of the target word "sandwich", the mixed label 120 of the preset model is obtained.

[0134] In step 209, the server obtains the word probability distribution of each word in the mixed label and the word prediction probability distribution of the corresponding word in the mixed label output by the preset model.

[0135] Among them, in the training stage, the decoder sequentially receives the input word vectors corresponding to the mixed label and simultaneously outputs the expected probability of the vocabulary corresponding to the next word vector, so as to obtain the word prediction probability distribution of the corresponding word in the mixed label 120 output by the preset model.

[0136] In one embodiment, the decoder sequentially receives the set of word vectors of the label samples after the mixed word vectors are replaced from the word embedding layer 131. For example, the decoder receives the starting vocabulary <bos>, the output obtains the expected probability that the next word is "I". Then, the decoder receives the starting word <bos>The word vector corresponding to "I", and the expected probability of the next word being "eat" is output. The decoder receives the starting word <bos>, the word vectors corresponding to "I" and "eat", output the expected probability that the next word is "le", and the decoder receives the starting word <bos>, the word vectors corresponding to "I", "eat", and "le", output the expected probabilities of the next words being "sandwich", "hamburger", and "salad", so as to obtain the word prediction probability distribution of the corresponding words in the mixed label 120 output by the dialogue generation model 130.

[0137] In step 210, the server inputs the word probability distribution and the word prediction probability distribution into the loss function of the preset model, obtains the loss of each word in the mixed label output by the preset model, and accumulates the losses of each word in the mixed label output by the preset model to obtain the total loss value.

[0138] Among them, in order to calculate the loss of the model to optimize the preset model according to the loss, the word probability distribution of each word in the mixed label and the word prediction probability distribution of the corresponding word in the mixed label output by the preset model can be obtained, and then the word probability distribution and the word prediction probability distribution are input into the loss function of the preset model. Among them, please continue to refer to Figure 4 , for the words "sandwich", "hamburger", and "salad" in the mixed label 120 for which the soft label 121 is constructed, the following cross-entropy loss function can be used to calculate the loss:

[0139]

[0140] Obtain the loss of the target word and its corresponding near-synonyms in the soft label. Among them, c i is the i-th word in the mixed set C of the target word "sandwich", and p(c i ) represents the expected probability of c i in the mixed set C of "sandwich", and g(c i ) is the probability that the preset model 130 outputs as c i . For other words in the mixed label 120, the following logarithmic loss function can be used to calculate the loss:

[0141] L = -logg(c i )

[0142] Thus, the loss of each word in the mixed label 120 output by the dialogue generation model 130 is obtained, and the losses of each word in the mixed label 120 output by the dialogue generation model 130 are accumulated to obtain the total loss value.

[0143] In step 211, the server averages the total loss value to obtain the target loss, and iteratively trains the model parameters of the preset model according to the target loss. When the target loss meets the convergence condition, the trained preset model is obtained.

[0144] The server averages the total loss value. For example, please continue to refer to Figure 4 , the losses of each word in the mixed label 120 are accumulated, and then the total loss value is divided by 4 to achieve average processing, obtaining the target loss of the current dialogue generation model. The model parameters of the dialogue generation model 130 are iteratively trained according to the target loss to optimize the preset model. When the target loss meets the convergence condition, that is, when the minimum value is obtained, the trained dialogue generation model is obtained. Among them, by introducing the mixed word vector and the mixed label, the dialogue generation model 130 realizes many-to-many training, improving the diversity of the output of the trained dialogue generation model.

[0145] In some embodiments, the above-mentioned trained preset model can be applied to the text generation scenario, specifically as follows:

[0146] Receive user request information, where the user request information includes the text data input by the user;

[0147] Input the text data into the trained preset model, and the model parameters of the preset model are obtained by training using various optional data processing methods provided in the above embodiments;

[0148] Determine the output result of the trained preset model as the target text data.

[0149] Among them, receive the request information input by the user, and the request information includes the text data input by the user. For example, it can be a piece of chat history, an article, or the text to be translated, etc. Input the text data into the trained preset model, and determine the output result of the trained preset model as the target text data. Among them, the target text data can be the reply corresponding to the chat history input by the user, the article abstract output according to the article input by the user, or the text in another language output according to the text input by the user, etc. As can be seen from the above, the text generation method provided in the embodiments of the present application can be used for dialogue generation. For example, generate corresponding replies according to the text content input by the user. The text generation method provided in the embodiments of the present application can be used for abstract generation. For example, generate corresponding abstracts according to an article input by the user. In addition, the text generation method provided in the embodiments of the present application can be used for machine translation, by outputting the text in another language corresponding to the text input by the user, etc.

[0150] Among them, the preset model can be a probabilistic generative model, which also needs to be trained with a certain amount of historical data before use to determine the model parameters of the preset model. For different model functions, different training sample data are required. For example, when the function of the text generation model is to reply to the conversation input by the user, the training sample data for training can be conversation sentences. For example, the persona-chat dataset can be used. When the function of the text generation model is to generate a summary of the article input by the user, the training sample data for training can be the article and the corresponding summary. When the function of the text generation model is to translate the text input by the user, the training sample data for training can be texts in different languages. The process of model training, that is, the process of training model parameters, can be trained by any of the data processing methods provided in the foregoing embodiments.

[0151] Specifically, the data processing method provided in the embodiments of the present application can be applied to machine question answering, such as a chatbot. Please refer to Figure 5 , Figure 5 FIG. is a schematic diagram of a specific application scenario of the data processing method provided in the embodiments of the present application. Among them, the robot can receive the text information input by the user and generate a corresponding reply. In the prior art, most dialogue generation models are trained based on only one fixed reply. This one-to-one training results in a relatively low generation diversity of this model, and the chat effect of the chatbot obtained by applying this model is poor. However, the chatbot obtained by applying the data processing method provided in the embodiments of the present application can obtain multiple replies according to the text information input by the user, and the generated diversity is relatively high. For example, please continue to refer to Figure 5 , for the user input of "What did you have for lunch", the robot can get replies such as "I had a sandwich", "I had a hamburger", or "I had a salad" and the corresponding generation probabilities 0.5, 0.3, and 0.2, etc. Thus, the final reply can be randomly selected from multiple replies and fed back to the user. The greater the generation probability of each reply, the greater the probability of being selected. For example, the reply "I had a sandwich" can be sent to the user, thereby realizing the process of machine question answering.

[0152] As can be seen from the above, in the embodiment of the present application, the server obtains the vocabulary sets corresponding to the source sample and the label sample, performs part-of-speech analysis on the vocabulary in the vocabulary sets, filters the vocabulary with the target part-of-speech in the vocabulary sets according to the part-of-speech analysis result; the server determines a preset replacement ratio, determines the number of target vocabulary in the vocabulary set after filtering according to the preset replacement ratio, randomly selects the vocabulary in the vocabulary set after filtering according to the number of target vocabulary, and obtains the target vocabulary according to the result of the random selection; the server obtains the target vocabulary and its synonyms, calculates the similarity score between the target vocabulary and its synonyms, and deletes the synonyms with similarity scores not greater than the preset threshold among the target vocabulary and its synonyms; the server converts each vocabulary and its synonyms in the vocabulary set into a set of word vectors through the word embedding layer of the preset model, accumulates the similarity scores of the target vocabulary and its corresponding synonyms, and obtains the total score of the target vocabulary; the server calculates the ratio of the similarity score and the total score of the target vocabulary and its corresponding synonyms to obtain the weight of the target vocabulary and its corresponding synonyms, and performs weighted mixing on the word vectors of the target vocabulary and its corresponding synonyms according to the weight to obtain the mixed word vector of the target vocabulary; the server replaces the word vector of the corresponding target vocabulary with the mixed word vector, and inputs the set of word vectors after replacement into the preset model for training; the server determines the target label of the preset model according to the vocabulary set of the label sample, and obtains the expected probability of the target vocabulary and its target synonyms in the label sample according to the target synonyms of the target vocabulary in the vocabulary set of the label sample and the corresponding similarity scores; the server obtains the word probability distribution of the target vocabulary based on the expected probability of the target vocabulary and its target synonyms, obtains the word probability distribution of the target vocabulary, constructs a soft label based on the target vocabulary corresponding to the vocabulary set included in the target label, and combines the target label and the soft label to obtain the mixed label of the preset model; the server obtains the word probability distribution of each vocabulary in the mixed label and the word prediction probability distribution of the corresponding vocabulary in the mixed label output by the preset model; the server inputs the word probability distribution and the word prediction probability distribution into the loss function of the preset model to obtain the loss of each vocabulary in the mixed label output by the preset model, and accumulates the losses of each vocabulary in the mixed label output by the preset model to obtain the total loss value; the server performs an average process on the total loss value to obtain the target loss of the current preset model, and iteratively trains the model parameters of the preset model according to the target loss. When the target loss meets the convergence condition, the trained preset model is obtained.In this way, without increasing the number of samples, near synonyms of the target word in the vocabulary set and similarity scores are obtained. By converting each word and its near synonyms in the vocabulary set into word vectors, and mixing the target word and its corresponding near synonyms and replacing the word vector of the corresponding target word, the mixed word vector is input into a preset model for training. Then, by using the difference between the mixed label obtained by combining the target label and the soft label and the output of the preset model to iteratively train the preset model, when the difference meets the convergence condition, the trained preset model is obtained, effectively achieving the data augmentation effect, thereby improving the data processing efficiency and increasing the diversity of the output of the trained preset model.

[0153] To better implement the above method, an embodiment of the present application further provides a data processing device, which can be integrated in a network device, such as a server or a terminal device, etc. The terminal may include a tablet computer, a notebook computer, and / or a personal computer, etc.

[0154] For example, as Figure 6 shown, the data processing device may include a word segmentation unit 301, an acquisition unit 302, a mixing unit 303, a replacement unit 304, and a training unit 305, as follows:

[0155] The word segmentation unit 301 is used to obtain the vocabulary set corresponding to the source sample and the label sample;

[0156] The acquisition unit 302 is used to obtain the target word and its near synonyms, and calculate the similarity score between the target word and its near synonyms. The target word is at least one word selected from the vocabulary set;

[0157] The mixing unit 303 is used to convert each word and its corresponding near synonyms in the vocabulary set into a word vector set, and mix the target word and its corresponding near synonyms according to the similarity score to obtain the mixed word vector of the target word;

[0158] The replacement unit 304 is used to replace the word vector of the corresponding target word with the mixed word vector, and input the replaced word vector set into the preset model for training;

[0159] The training unit 305 is used to generate a mixed label, obtain the difference between the word probability distribution of the mixed label and the word prediction probability distribution of the mixed label output by the preset model, and iteratively train the model parameters of the preset model according to the difference to obtain the trained preset model.

[0160] In one embodiment, the mixing unit 303 includes:

[0161] The conversion sub-unit is used to convert each word and its corresponding near synonyms in the vocabulary set into a word vector set through the word embedding layer of the preset model;

[0162] A calculation subunit, configured to obtain weights of word vectors of a target word and corresponding near synonyms according to similarity scores;

[0163] A mixing subunit, configured to perform weighted mixing on the word vectors of the target word and corresponding near synonyms according to the weights to obtain a mixed word vector of the target word.

[0164] In one embodiment, the calculation subunit is configured to:

[0165] Accumulate the similarity scores of the target word and corresponding near synonyms to obtain a total score of the target word;

[0166] Calculate a ratio of the similarity score of the target word and corresponding near synonyms to the total score to obtain weights of the target word and corresponding near synonyms.

[0167] In one embodiment, the training unit 305 includes:

[0168] A determination subunit, configured to determine a target label of a preset model according to a vocabulary set of a label sample;

[0169] A construction subunit, configured to perform soft label construction on the target word included in the target label based on the target near synonyms of the target word in the vocabulary set of the label sample and the corresponding similarity scores of the target near synonyms;

[0170] A combination subunit, configured to combine the target label and the soft label to obtain a mixed label of the preset model.

[0171] In one embodiment, the construction subunit is configured to:

[0172] According to the target near synonyms of the target word in the vocabulary set of the label sample and the corresponding similarity scores of the target near synonyms, obtain the expected probabilities of the target word and the target near synonyms in the label sample;

[0173] Based on the expected probabilities of the target word and the target near synonyms, obtain a word probability distribution of the target word;

[0174] Obtain the word probability distribution of the target word, and perform soft label construction on the target word included in the target label based on the word probability distribution.

[0175] In one embodiment, the training unit 305 includes:

[0176] An acquisition subunit, configured to acquire the word probability distribution of each word in the mixed label and the word prediction probability distribution of the corresponding word in the mixed label output by the preset model;

[0177] An input subunit, configured to input the word probability distribution and the word prediction probability distribution into a loss function of the preset model to obtain a target loss;

[0178] A training subunit, configured to iteratively train the model parameters of a preset model according to a target loss, and obtain the trained preset model when the target loss meets the convergence condition.

[0179] In one embodiment, an input subunit, configured to:

[0180] Input a word probability distribution and a word prediction probability distribution into a loss function of a preset model, and obtain the loss of each word in the mixed label output by the preset model;

[0181] Accumulate the losses of each word in the mixed label output by the preset model to obtain a total loss value;

[0182] Perform an averaging process on the total loss value to obtain a target loss.

[0183] In one embodiment, the data processing device further includes:

[0184] A filtering unit, configured to perform part-of-speech analysis on the words in a vocabulary set, and filter the words with a target part of speech in the vocabulary set according to the part-of-speech analysis result;

[0185] A second determination unit, configured to determine a preset replacement ratio, and determine the number of target words in the filtered vocabulary set according to the preset replacement ratio;

[0186] A selection unit, configured to randomly select words from the filtered vocabulary set according to the number of target words, and obtain target words according to the random selection result.

[0187] In one embodiment, the data processing device further includes:

[0188] A deletion subunit, configured to delete synonyms of a target word and the target word whose similarity scores are not greater than a preset threshold.

[0189] In specific implementation, the above units can be implemented as independent entities, or can be combined arbitrarily to be implemented as the same or several entities. For the specific implementation of the above units, reference can be made to the foregoing method embodiments, which will not be elaborated herein.

[0190] As can be seen from the above, in the embodiment of the present application, the tokenization unit 301 obtains the vocabulary sets corresponding to the source samples and label samples; the acquisition unit 302 obtains the target vocabulary and the synonyms of the target vocabulary, and calculates the similarity scores between the target vocabulary and its synonyms; the mixing unit 303 converts each vocabulary and its synonyms in the vocabulary set into a word vector set, and mixes the target vocabulary and its corresponding synonyms according to the similarity scores to obtain the mixed word vector of the target vocabulary; the replacement unit 304 replaces the word vector of the corresponding target vocabulary with the mixed word vector, and inputs the replaced word vector set into a preset model for training; the training unit 305 generates a mixed label, obtains the difference between the word probability distribution of the mixed label and the word prediction probability distribution of the mixed label output by the preset model, and iteratively trains the model parameters of the preset model according to the difference to obtain the trained preset model. In this way, the vocabulary sets corresponding to the obtained source samples and label samples are obtained, the synonyms and similarity scores of the target vocabulary in the vocabulary set are obtained, each vocabulary and its corresponding synonyms in the vocabulary set are converted into word vectors, the target vocabulary and its corresponding synonyms are mixed and the word vector of the corresponding target vocabulary is replaced and input into the preset model for training, and then the preset model is iteratively trained according to the difference between the determined mixed label and the output of the preset model to obtain the trained preset model, effectively realizing data augmentation, thereby improving the efficiency of data processing and increasing the diversity of the output of the trained model.

[0191] The embodiment of the present application also provides a computer device, as Figure 7 shown, which shows the structural schematic diagram of the computer device involved in the embodiment of the present application. Specifically:

[0192] The computer device may include a processor 401 with one or more processing cores, a memory 402 with one or more computer-readable storage media, a power supply 403, an input unit 404 and other components. Those skilled in the art can understand that Figure 7 the structure of the computer device shown in

[0193] The processor 401 is the control center of the computer device, connecting various parts of the entire computer device through various interfaces and circuits. By running or executing software programs and / or modules stored in the memory 402, and by invoking the data stored in the memory 402, it executes various functions of the computer device and processes data, thereby performing overall detection and control of the computer device. Optionally, the processor 401 may include one or more processing cores; preferably, the processor 401 may integrate an application processor and a modem processor. Among them, the application processor mainly processes the operating system, user interface, application programs, etc., and the modem processor mainly processes wireless communication. It can be understood that the above-mentioned modem processor may not be integrated into the processor 401 either.

[0194] The memory 402 can be used to store software programs and modules. The processor 401 executes various functional applications and data processing by running the software programs and modules stored in the memory 402. The memory 402 mainly includes a program storage area and a data storage area. Among them, the program storage area can store the operating system, application programs required for at least one function (such as the sound playback function, image playback function, etc.); the data storage area can store data created according to the use of the computer device. In addition, the memory 402 may include high-speed random access memory, and may also include non-volatile memory, such as at least one magnetic disk storage device, flash memory device, or other non-volatile solid-state storage devices. Correspondingly, the memory 402 may also include a memory controller to provide the processor 401 with access to the memory 402.

[0195] The computer device also includes a power supply 403 that powers each component. Preferably, the power supply 403 can be logically connected to the processor 401 through a power management system, so as to realize functions such as management of charging, discharging, and power consumption management through the power management system. The power supply 403 may also include any components such as one or more DC or AC power supplies, a recharge system, a power failure detection circuit, a power converter or inverter, and a power status indicator.

[0196] The computer device may also include an input unit 404, which can be used to receive input digital or character information, and generate keyboard, mouse, joystick, optical or trackball signal inputs related to user settings and function control.

[0197] Although not shown, the computer device may also include a display unit, etc., which will not be elaborated here. Specifically, in this embodiment, the processor 401 in the computer device will load the executable files corresponding to the processes of one or more application programs into the memory 402 according to the following instructions, and the processor 401 will run the application programs stored in the memory 402 to realize various functions as follows:

[0198] Obtain the vocabulary sets corresponding to the source samples and label samples; obtain the similarity scores between the target vocabulary and the synonyms of the target vocabulary, where the target vocabulary is at least one vocabulary selected from the vocabulary sets; convert each vocabulary and its synonyms in the vocabulary sets into a set of word vectors, and perform vector mixing on the target vocabulary and its corresponding synonyms according to the similarity scores to obtain the mixed word vectors of the target vocabulary; replace the word vectors of the corresponding target vocabulary with the mixed word vectors, and input the set of word vectors after replacement into a preset model for training; generate mixed labels, obtain the difference between the word probability distribution of the mixed labels and the word prediction probability distribution of the mixed labels output by the preset model, and perform iterative training on the model parameters of the preset model according to the difference to obtain the trained preset model.

[0199] For the specific implementation of each of the above operations, reference may be made to the previous embodiments, which will not be elaborated here. It should be noted that the computer device provided in the embodiments of the present application belongs to the same concept as the data processing method in the above embodiments, and the specific implementation process can be seen in the above method embodiments, which will not be elaborated here.

[0200] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructions, or by instructions controlling related hardware. The instructions can be stored in a computer-readable storage medium and loaded and executed by a processor.

[0201] Therefore, the embodiments of the present application provide a storage medium, in which multiple instructions are stored, and the instructions can be loaded by a processor to execute the steps in any one of the data processing methods provided by the embodiments of the present application. For example, the instructions can execute the following steps:

[0202] Obtain the vocabulary sets corresponding to the source samples and label samples; obtain the target vocabulary and the synonyms of the target vocabulary, and calculate the similarity scores between the target vocabulary and the synonyms of the target vocabulary; convert each vocabulary and its synonyms in the vocabulary sets into a set of word vectors, and perform vector mixing on the target vocabulary and its corresponding synonyms according to the similarity scores to obtain the mixed word vectors of the target vocabulary; replace the word vectors of the corresponding target vocabulary with the mixed word vectors, and input the set of word vectors after replacement into a preset model for training; generate mixed labels, obtain the difference between the word probability distribution of the mixed labels and the word prediction probability distribution of the mixed labels output by the preset model, and perform iterative training on the model parameters of the preset model according to the difference to obtain the trained preset model.

[0203] For the specific implementation of each of the above operations, reference may be made to the previous embodiments, which will not be elaborated here.

[0204] Among them, the storage medium may include: read-only memory (ROM, Read Only Memory), random access memory (RAM, Random Access Memory), magnetic disk or optical disk, etc.

[0205] Since the instructions stored in the storage medium can execute the steps in any of the data processing methods provided in the embodiments of the present application, the beneficial effects that can be achieved by any of the data processing methods provided in the embodiments of the present application can be realized. For details, see the previous embodiments and will not be repeated here.

[0206] Among them, according to an aspect of the present application, a computer program product or a computer program is provided. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the methods provided in the various alternative implementations provided in the above embodiments.

[0207] The above has introduced in detail a data processing method, device, storage medium, and computer device provided in the embodiments of the present application. Specific examples are used in this article to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application; at the same time, for those skilled in the art, according to the idea of the present application, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present application.< / bos> < / bos> < / bos> < / bos>

Claims

1. A data processing method, characterized in that Including: Obtain the vocabulary sets corresponding to the source samples and label samples; Obtain the target vocabulary and its synonyms, and calculate the similarity scores between the target vocabulary and its synonyms, where the target vocabulary is at least one vocabulary selected from the vocabulary sets; Convert each vocabulary and its synonyms in the vocabulary sets into a set of word vectors, and mix the vectors of the target vocabulary and its corresponding synonyms according to the similarity scores to obtain the mixed word vector of the target vocabulary; Replace the word vector of the corresponding target vocabulary with the mixed word vector, and input the set of word vectors after replacement into a preset model for training; Generate a mixed label, obtain the difference between the word probability distribution of the mixed label and the word prediction probability distribution of the mixed label output by the preset model, and iteratively train the model parameters of the preset model according to the difference to obtain the trained preset model; Before obtaining the target vocabulary and its synonyms and calculating the similarity scores between the target vocabulary and its synonyms, it further includes: Perform part-of-speech analysis on the vocabulary in the vocabulary sets, and filter the vocabulary with the target part of speech in the vocabulary sets according to the part-of-speech analysis results; Determine a preset replacement ratio, and determine the number of target vocabulary in the filtered vocabulary sets according to the preset replacement ratio; Randomly select the vocabulary in the filtered vocabulary sets according to the number of the target vocabulary, and obtain the target vocabulary according to the random selection results.

2. The data processing method according to claim 1, wherein The step of converting each vocabulary and its synonyms in the vocabulary sets into a set of word vectors, and mixing the vectors of the target vocabulary and its corresponding synonyms according to the similarity scores to obtain the mixed word vector of the target vocabulary includes: Convert each vocabulary and its synonyms in the vocabulary sets into a set of word vectors through the word embedding layer of the preset model; Obtain the weights of the word vectors of the target vocabulary and its corresponding synonyms according to the similarity scores; Perform weighted mixing on the word vectors of the target vocabulary and its corresponding synonyms according to the weights to obtain the mixed word vector of the target vocabulary.

3. The data processing method according to claim 2, wherein The step of obtaining the weights of the word vectors of the target vocabulary and its corresponding synonyms according to the similarity scores includes: Accumulate the similarity scores of the target vocabulary and its corresponding synonyms to obtain the total score of the target vocabulary; Calculate the ratio of the similarity score of the target vocabulary and its corresponding synonyms to the total score to obtain the weights of the target vocabulary and its corresponding synonyms.

4. The data processing method according to claim 1, characterized in that The step of generating the mixed label includes: Determine the target label of the preset model according to the vocabulary sets of the label samples; Based on the target synonyms of the target vocabulary in the vocabulary sets of the label samples and the corresponding similarity scores, perform soft label construction on the target vocabulary corresponding to the vocabulary sets included in the target label; Combine the target label and the soft label to obtain the mixed label of the preset model.

5. The data processing method according to claim 4, characterized in that, The step of performing soft label construction on the target vocabulary corresponding to the vocabulary sets included in the target label based on the target synonyms of the target vocabulary in the vocabulary sets of the label samples and the corresponding similarity scores includes: Obtain the expected probabilities of the target word and its target near-synonyms in the label sample according to the target near-synonyms of the target word in the vocabulary set of the label sample and the similarity scores corresponding to the target near-synonyms; Based on the expected probabilities of the target word and its target near-synonyms, obtain the word probability distribution of the target word; Obtain the word probability distribution of the target word, and construct a soft label based on the target word corresponding to the vocabulary set included in the target label.

6. The data processing method according to claim 1, wherein The obtaining of the difference between the word probability distribution of the mixed label and the word prediction probability distribution of the mixed label output by the preset model, and the iterative training of the model parameters of the preset model according to the difference to obtain the trained preset model includes: Obtain the word probability distribution of each word in the mixed label and the word prediction probability distribution of the corresponding word in the mixed label output by the preset model; Input the word probability distribution and the word prediction probability distribution into the loss function of the preset model to obtain the target loss; Iteratively train the model parameters of the preset model according to the target loss, and when the target loss meets the convergence condition, obtain the trained preset model.

7. The data processing method according to claim 6, wherein The inputting the word probability distribution and the word prediction probability distribution into the loss function of the preset model to obtain the target loss includes: Input the word probability distribution and the word prediction probability distribution into the loss function of the preset model to obtain the loss of each word in the mixed label output by the preset model; Accumulate the losses of each word in the mixed label output by the preset model to obtain the total loss value; Perform an averaging process on the total loss value to obtain the target loss.

8. The data processing method according to claim 1, wherein, After obtaining the target word and its near-synonyms and calculating the similarity scores between the target word and its near-synonyms, it further includes: Delete the near-synonyms whose similarity scores are not greater than the preset threshold among the target word and its near-synonyms.

9. A text generation method, characterized in that, The method includes: Receive user request information, where the user request information includes text data input by the user; Input the text data into the trained preset model, and the model parameters of the preset model are trained by using the data processing method described in any one of claims 1 to 8; Determine the output result of the trained preset model as the target text data.

10. A data processing device, characterized in that, It includes: A word segmentation unit for obtaining the vocabulary sets corresponding to the source sample and the label sample; A filtering unit for performing part-of-speech analysis on the words in the vocabulary set, and filtering the words with the target part of speech in the vocabulary set according to the part-of-speech analysis result; A determination unit for determining a preset replacement ratio, and determining the number of target words in the vocabulary set after filtering according to the preset replacement ratio; A selection unit for randomly selecting words from the vocabulary set after filtering according to the number of target words, and obtaining the target word according to the random selection result; An obtaining unit for obtaining the target word and its near-synonyms, and calculating the similarity scores between the target word and its near-synonyms, where the target word is at least one word selected from the vocabulary set; A mixing unit, configured to convert each word and its near-synonym in a vocabulary set into a word vector set, and mix the target word and its corresponding near-synonym according to the similarity score to obtain a mixed word vector of the target word; A replacement unit, configured to replace the word vector of the corresponding target word with the mixed word vector, and input the word vector set after replacement into a preset model for training; A training unit, configured to generate a mixed label, obtain the difference between the word probability distribution of the mixed label and the word prediction probability distribution of the mixed label output by the preset model, and iteratively train the model parameters of the preset model according to the difference to obtain a trained preset model.

11. The device according to claim 10, characterized in that, The mixing unit includes: A conversion sub-unit, configured to convert each word and its corresponding near-synonym in a vocabulary set into a word vector set through the word embedding layer of a preset model; A calculation sub-unit, configured to obtain the weights of the word vectors of the target word and its corresponding near-synonym according to the similarity score; A mixing sub-unit, configured to perform weighted mixing on the word vectors of the target word and its corresponding near-synonym according to the weights to obtain a mixed word vector of the target word.

12. The device according to claim 11, characterized in that, The calculation sub-unit is configured to: Accumulate the similarity scores of the target word and its corresponding near-synonym to obtain the total score of the target word; Calculate the ratio of the similarity score of the target word and its corresponding near-synonym to the total score to obtain the weights of the target word and its corresponding near-synonym.

13. A storage medium, characterized in that, The storage medium stores multiple instructions, and the instructions are suitable for being loaded by a processor to execute the steps in the data processing method according to any one of claims 1 to 8.

14. A computer device, characterized in that, It includes a memory and a processor; the memory stores an application program, and the processor is configured to run the application program in the memory to execute the steps in the data processing method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • A multi-semantic supervised word vector training method and device

    CN109165288A

  • Push information generation method and device

    CN110427617A