Data set generation method and device of discrete sequence data, electronic equipment and medium
By generating high-quality discrete sequence datasets through adversarial training of the SeqGAN model, the problem of insufficient quality of discrete sequence data generation in existing technologies is solved, the training effect of large AI models is improved, and it is suitable for various sequence data application scenarios.
Patent Information
- Application Number
- CN202510889880.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-30
- Publication Date
- 2025-10-17
AI Technical Summary
Existing data generation technologies cannot meet the demand for high-quality discrete sequence data for training large AI models, resulting in poor model training performance.
The SeqGAN model is used to train the first and second sequence data to generate a high-quality discrete sequence dataset. The SeqGAN model consists of a generator and a discriminator in a generative adversarial network. It learns data dependencies and generates rules through self-supervision to achieve data cleaning and repair.
It improves the quality of generated discrete sequence datasets, enhances the success rate and efficiency of training large AI models, and is applicable to various sequence data types, including natural language processing, financial trend prediction, medical time series analysis, and bioinformatics.
Smart Images

Figure CN120804073A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of data generation, and in particular to a discrete sequence data set generation method and device, electronic equipment and a medium. BACKGROUND
[0002] In the field of artificial intelligence, a high-quality data set is the core basis for AI (Artificial Intelligence) large model training, and its quality directly determines the effect and performance of model training. In the training process of an AI model, only a high-quality data set can provide effective data support for large model training, promote the positive improvement of model performance, and thus realize precise prediction and efficient decision-making of the model in various tasks.
[0003] However, in the data preparation link of the early stage of AI large model training, there are many problems to be solved. On the one hand, the actual acquired data set often has low data quality, containing a large amount of noisy data, incorrectly labeled data and other dirty data. These low-quality data will interfere with the training process of the model, hinder the achievement of the training goal of the large model, and reduce the training efficiency and final performance of the model. On the other hand, although the generation technology of high-quality data sets can achieve good results when dealing with time-continuous data, for discrete sequence data sets, the existing generation technology cannot meet the needs of AI large model training, resulting in unsatisfactory training effect of the model when dealing with discrete sequence data related tasks, and it is difficult to achieve the ideal performance indicators.
[0004] Therefore, how to solve the quality problems existing in the current data preparation, and improve the generation quality of discrete sequence data sets, has become a key technical difficulty in improving the success rate of AI large model training, optimizing the training efficiency and resource utilization. SUMMARY
[0005] The present application provides a discrete sequence data set generation method, device, electronic equipment and medium, which cleans the existing discrete sequence data set and improves the data generation quality.
[0006] According to an aspect of the present application, a discrete sequence data set generation method is provided, which comprises:
[0007] obtaining first sequence data and second sequence data; wherein the first sequence data is a discrete sequence data set after cleaning processing; and the second sequence data is an original sequence data set;
[0008] training a to-be-trained rule generation model based on the first sequence data and the second sequence data to obtain a target rule generation model; wherein the target rule generation model is a SeqGAN model, and the SeqGAN model is composed of a generator and a discriminator in a generative adversarial network;
[0009] generating a data set of discrete sequence data according to the target rule generation model.
[0010] According to another aspect of the present application, a data set generation apparatus of discrete sequence data is provided, which comprises:
[0011] a data acquisition module configured to acquire first sequence data and second sequence data; wherein the first sequence data is a discrete sequence data set after cleaning processing, and the second sequence data is an original sequence data set;
[0012] a target rule generation model obtaining module configured to train a to-be-trained rule generation model based on the first sequence data and the second sequence data to obtain a target rule generation model; wherein the target rule generation model is a SeqGAN model, and the SeqGAN model is composed of a generator and a discriminator in a generative adversarial network;
[0013] a data set generation module configured to generate a data set of discrete sequence data according to the target rule generation model.
[0014] According to another aspect of the present application, an electronic device is provided, which comprises:
[0015] at least one processor; and a memory connected to the at least one processor in communication; wherein the memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor to enable the at least one processor to execute the data set generation method of discrete sequence data according to any one of the embodiments of the present application.
[0016] According to another aspect of the present application, a computer readable storage medium is provided, which stores computer instructions for enabling a processor to execute the data set generation method of discrete sequence data according to any one of the embodiments of the present application when executed.
[0017] The technical solution of the embodiments of the present invention obtains first and second sequence data and trains a target rule generation model based on these two types of data, thereby obtaining a target rule generation model. This model is then used to generate a dataset of discrete sequence data. This technical solution performs data cleaning on existing discrete sequence datasets, effectively improving the quality of data generation. This not only addresses quality issues in current data preparation, but also optimizes and upgrades the quality of discrete sequence dataset generation.
[0018] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present invention, nor is it intended to limit the scope of the present invention. Other features of the present invention will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0020] Figure 1 This is a flowchart of a method for generating a dataset of discrete sequence data according to the first embodiment of the present invention;
[0021] Figure 2 A schematic diagram of the model architecture provided in Example 1 of this application;
[0022] Figure 3 This is a schematic diagram of rule generation based on a generator provided in Example 1 of the present application;
[0023] Figure 4 A schematic diagram of a process for generating a dataset of discrete sequence data provided in the second embodiment of the present invention;
[0024] Figure 5 A schematic diagram of the structure of a device for generating a dataset of discrete sequence data provided in the third embodiment of the present invention;
[0025] Figure 6 It is a structural diagram of an electronic device for implementing the method for generating a data set of discrete sequence data according to an embodiment of the present invention. DETAILED DESCRIPTION
[0026] In the following, the technical solutions in the embodiments of the present application will be described clearly and completely with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all the other embodiments obtained by a person of ordinary skill in the art without creative work should belong to the protection scope of the present application.
[0027] It should be noted that the terms "first", "second" and the like in the description and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily indicate a specific order or sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the application described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "comprise" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or device comprising a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but can include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0028] Embodiment one
[0029] Figure 1 A flowchart of a method for generating a data set of discrete sequence data according to the first embodiment of the present application is shown. The embodiment can be applicable to the case of cleaning discrete sequence data. The method can be executed by a data set generation device for discrete sequence data, which can be realized in the form of hardware and / or software, and can be configured in a device. For example, the device can be a background server or other device with communication and computing capabilities. As shown in the figure, the method comprises: Figure 1
[0030] S110, obtaining first sequence data and second sequence data; wherein the first sequence data is a cleaned discrete sequence data set; and the second sequence data is an original sequence data set.
[0031] In the present scheme, the first sequence data is a cleaned discrete sequence data set, which is an ordered sequence composed of discrete data. Each data point has a clear classification or limited value, rather than a continuous numerical interval.
[0032] In this embodiment, by cleaning the original sequence data set, a cleaned discrete sequence data set is obtained, that is, the second sequence data is cleaned to obtain the first sequence data. The preprocessing of the original discrete sequence data eliminates noise, corrects errors, standardizes formats, and improves data quality. Common operations include missing value processing, outlier filtering, format standardization, deduplication and noise reduction. Missing value processing is used to process incomplete or unrecorded values in the data set to ensure data integrity. Outlier filtering is used to identify and eliminate extreme values that do not conform to business logic or statistical rules to avoid bias in analysis results. Format standardization is used to unify data formats, types and structures to ensure data consistency and computability. De-duplication and noise reduction are used to eliminate duplicate data and high-frequency invalid information to improve data quality.
[0033] In this scheme, the first sequence data and the second sequence data can be obtained from different databases. The second sequence data can also be extracted from the database, and after cleaning, the first sequence data is determined.
[0034] S120, training the to-be-trained rule generation model based on the first sequence data and the second sequence data to obtain a target rule generation model; wherein the target rule generation model is a SeqGAN model, and the SeqGAN model is composed of a generator and a discriminator in a generative adversarial network.
[0035] In this scheme, Figure 2 The model architecture diagram provided by Embodiment One of the present application is shown in Figure 2 A typical application case of SeqGAN is the Garf framework, which is a self-supervised and data-driven data cleaning framework. The Garf framework learns the relationships in the data using SeqGAN and generates rules for data cleaning. These rules are presented in an interpretable form, improving the transparency and interpretability of the model. When generating repair rules for raw sequence data, the core mechanism of the Garf framework is to use a rule generation model to learn the dependencies between data. This framework can learn and induce a rule system suitable for data repair by mining the temporal correlations, logical constraints and pattern characteristics hidden in sequence data, thereby achieving accurate repair of missing, erroneous or abnormal parts in raw sequence data. This rule generation method based on dependency learning effectively improves the accuracy and adaptability of data repair, providing reliable technical support for preprocessing and quality optimization of sequence data. Specifically, it includes:
[0036] The target rule generation model is based on SeqGAN and includes G s generator and D s discriminator. Among them, G sThe generator is responsible for learning and generating sequences that conform to the distribution of the original data. s The discriminator is used to distinguish the authenticity of generated data from real data, assisting the generator in optimization.
[0037] Discrete sequence data is embedded into a vector form that the model can process. Training logic: drive model learning in a self-supervised manner, the generator attempts to generate sequences that approximate the real distribution, and the discriminator constantly identifies the authenticity of generated data. Both are dynamically optimized through adversarial training. Training goal: make the distribution of generated data consistent with that of real data, and ensure that the generator G s Capture the dependencies between data.
[0038] After training, the generator G s The modeling ability of data dependencies is converted into candidate rules.
[0039] This rule can achieve two major functions: dirty data detection: by comparing the distribution difference between generated data and original data, locate data that does not conform to the expected pattern. Data repair: use the learned dependencies to fill, correct or replace abnormal data to conform to the overall rules of the sequence.
[0040] The rule generation process abstracts data dependencies into executable repair logic through the adversarial learning framework of SeqGAN, forming a complete link from data modeling to rule implementation, and ultimately achieving automated cleaning and repair of original sequence data.
[0041] Specifically, the first sequence data and the second sequence data can be used as input to train the to-be-trained rule generation model to obtain a target rule generation model.
[0042] Optionally, the to-be-trained rule generation model is trained based on the first sequence data and the second sequence data to obtain a target rule generation model, including:
[0043] A candidate rule is generated during the training of the to-be-trained rule generation model, wherein the candidate rule is a rule for repairing the data set.
[0044] In this scheme, discrete sequence data is embedded into a vector form that the model can process. Drive model learning in a self-supervised manner, the generator attempts to generate sequences that approximate the real distribution, and the discriminator constantly identifies the authenticity of generated data. Both are dynamically optimized through adversarial training. Make the distribution of generated data consistent with that of real data, and ensure that the generator G s Capture the dependencies between data. After training, the generator G s The modeling ability of data dependencies is converted into candidate rules.
[0045] The candidate rule generation process abstracts the data dependency relationship into executable repair logic through the adversarial learning framework of SeqGAN, forms a complete link from data modeling to rule landing, and finally realizes the automatic cleaning and repair of the original sequence data.
[0046] Optionally, the candidate rule is generated in the training process of the rule generation model to be trained, including:
[0047] The first sequence data and the second sequence data are input into the rule generation model to be trained for training, and the parameters of the rule generation model to be trained are adjusted in the training process to generate the candidate rule.
[0048] In this embodiment, the first sequence data and the second sequence data are input into the rule generation model to be trained for training, and the parameters of the model are optimized and adjusted in the training process to generate the candidate rule.
[0049] Specifically, Figure 3 is a schematic diagram of the rule generation based on the generator provided by the first embodiment of the present application, as Figure 3 shown, for a set of discrete sequence data R=(a1,K,a n ), the candidate rule form on R is: [A L ,v(A L )]→[A R ,v(A R )], wherein A L is a set of attributes, v(A L ) is the value of attribute A L ; similarly, A R is also a set of attributes, and v(A R ) is the value of another set of attributes. A L and A R are usually referred to as left attributes and right attributes, and AV L =[A L ,v(A L )], AV R =[A R ,v(A R )]. When SeqGAN converges, the generated discrete sequence data follows the true relationship controlled by G s . When the input value of G s is v(A L ), the predicted value of the generated discrete sequence data is v(A R ). If the value of v(A R ) is equal to the true value, it can be considered that v(A L ) determines v(A R ).
[0050] To generate candidate rules, you need to define AV L and AV R Attributes and values. Given a discrete sequence data tuple with n attributes, Figure 3 The i-th attribute value AV i =(a i ,v i ) as AV R The result of , where i ranges from 2 to n. Given an AV L The attribute value pairs are taken as input, based on G s Generator predicts AV R . Initially, AV L ={AV i-1}. If the predicted value is not equal to v i ,AV L →AV R Obviously it is wrong. When the error occurs, you can use AV L Previous attribute value AV L ={AV i-2 ,AV i-1 Supplementary AV L Repeat the above process until the predicted value is equal to the true value, or the first attribute value pair AV1 is included in AV L middle.
[0051] To solve the problem that discrete sequence data may miss rules due to order irrelevant and tuples parsed from left to right, the model can be trained twice with different attribute orders, that is, the model is trained in forward and backward order respectively.
[0052] By generating candidate rules and performing data cleaning on existing discrete sequence data sets, the quality of data generation can be improved.
[0053] Optionally, after generating candidate rules during the training of the to-be-trained rule generation model, the method further includes:
[0054] For candidate rules, the number of tuples is defined as n;
[0055] During the rule generation process, the number of tuples matched by each candidate rule is determined;
[0056] Remove candidate rules with matching tuple number n < 2; wherein, matching tuple number n = 0 indicates that the rule cannot match a real tuple, and n = 1 indicates that the rule can only match one error tuple.
[0057] In the present scheme, although a large number of candidate rules have been generated through the previous operation, these rules cannot be directly applied to the data cleaning work. The reason is that the generated rules may have the problem of insufficient accuracy, or contain redundant attributes. The core goal of the optimization link is to solve such problems encountered in the candidate rule generation process.
[0058] Further, since the predicted value is derived from the probability distribution in the policy network, it is easy to derive inaccurate rules, and such rules may be disturbed by dirty data. A filter mechanism can be introduced for optimization. Specifically, for each candidate rule, the number of tuples can be defined as n. In the rule generation process, few errors are exactly the same, which leads to a commonality of many inaccurate rules, that is, they cannot match the real tuple (n = 0) or can only match one error tuple (n = 1). The rule of n = 1 has no contribution to the repaired data set, even if it is correct. Because n = 1 cannot be used for other tuples, and the only matched tuple is correct, rules with n < 2 can be removed.
[0059] For the rule r: {AV k ,K,AV i-1}→AV i , screen AV i-1 , and predict v i again. If the predicted value is still equal to v i , it can be considered that AV i-1 is redundant and should be deleted; otherwise, AV i-1 is useful and remains unchanged. Next, screen AV i-2 , and repeat this process until AV R does not depend on any subset of AV L . At the same time, according to the algorithm, the generator can be trained in the gth step, and the discriminator can be trained in the dth step. In the learning process based on SeqGAN, the generation of candidate rules is completely self-supervised, and it can be considered that the target data set is training itself. The generated candidate rules can break the dimensional and language restrictions.
[0060] By optimizing the generated candidate rules, the data cleaning efficiency of the existing discrete sequence data set is significantly enhanced, and the data generation quality is improved.
[0061] S130, generating a data set of discrete sequence data according to the target rule generation model.
[0062] In the present scheme, sequence data is obtained, the sequence data is input into the target rule generation model, and a data set of discrete sequence data is output.
[0063] The technical scheme of the embodiment of the present application obtains first sequence data and second sequence data, trains a rule generation model to be trained based on the two types of data, thereby obtaining a target rule generation model, and generates a data set of discrete sequence data relying on the model. By executing the technical scheme, SeqGAN combines reinforcement learning and a generative adversarial network across domains, retains the fitting ability of the generative adversarial network to data distribution, and solves the gradient transmission problem of discrete sequence generation by means of the decision optimization mechanism of reinforcement learning. The Garf framework further applies this technology to a data governance scene, realizes a closed loop of data mode learning and cleaning logic generation by means of self-supervised learning and interpretable rule generation, and provides a solution with efficiency and interpretability for automatic data cleaning.
[0064] Embodiment two
[0065] Figure 4 The schematic diagram of the data set generation process of discrete sequence data provided by the second embodiment of the present application is a detailed description of the data set generation process of discrete sequence data between the present embodiment and the above-mentioned embodiment. As shown in the figure, Figure 4 The method comprises the following steps:
[0066] S410, first sequence data and second sequence data are obtained; wherein the first sequence data is a discrete sequence data set after cleaning processing; and the second sequence data is an original sequence data set.
[0067] S420, the first sequence data is taken as input to train a generator in a rule generation model to be trained, thereby obtaining target first sequence data.
[0068] Specifically, the first sequence data is taken as input to train the generator in the rule generation model, thereby obtaining the target first sequence data.
[0069] Optionally, the first sequence data is taken as input to train the generator in the rule generation model to be trained, thereby obtaining the target first sequence data, which comprises the following steps:
[0070] Determining initial parameters of the generator in the rule generation model to be trained;
[0071] Taking the first sequence data as input, the generator in the rule generation model to be trained is trained, thereby obtaining the target first sequence data.
[0072] Based on the target first sequence data and the pre-determined initial first sequence data, a loss function is calculated.
[0073] According to the loss function, the initial parameters of the generator in the rule generation model to be trained are updated.
[0074] In the present scheme, the initial parameters of the generator in the rule generation model to be trained can be initialized based on a random theta parameter policy network. The initial parameters include input weight matrices (such as weight matrices input to the forget gate, input to the input gate, input to the cell state, and input to the output gate), hidden layer weight matrices (hidden layer to the forget gate, hidden layer to the input gate, hidden layer to the cell state, and hidden layer to the output gate), and bias vectors, etc.
[0075] Further, the preprocessed first sequence data is taken as input, and the generator parameters are optimized through iterative training to generate the target first sequence data.
[0076] The loss function is a mathematical function used to measure the difference between the model prediction results and the true target in machine learning and deep learning, and its core function is to provide a clear target for model optimization and guide the adjustment direction of model parameters. For example, the loss function can be mean square error, mean absolute error, etc.
[0077] Specifically, according to the pre-determined loss function calculation formula, the target first sequence data and the pre-determined initial first sequence data are combined and operated to calculate the loss function.
[0078] In the present embodiment, the loss function value is used to calculate the gradient of each initial parameter in the model through the back propagation algorithm. The gradient represents the change rate of the loss function under the current parameter, which can indicate the direction and amplitude of parameter adjustment. Then, according to the calculated gradient, a suitable optimization algorithm is selected, such as stochastic gradient descent (SGD), adaptive moment estimation (Adam), etc., to update the initial parameters.
[0079] Further, in the process of updating the parameters, multiple iterations of training are required. Each iteration repeats the above process. Through continuous iterative training, the parameters of the model are gradually adjusted, so that the prediction data of the model becomes more and more close to the real data, and the loss function value gradually decreases. When the loss function value reaches a pre-set smaller value, or after multiple iterations the change of the loss function value is no longer obvious, it is considered that the model has converged.
[0080] By adjusting the parameters, the learning ability of the generator can be continuously improved, and finally accurate discrete sequence data can be generated.
[0081] Optionally, the loss function is an expected final reward value; and the expected final reward value determination process comprises:
[0082] Obtaining the state space, action space and reward function of the reinforcement learning task;
[0083] Initializing the parameters of the loss function;
[0084] In the reinforcement learning training process, training data is generated based on the state transition of the agent in the environment, the action execution, and the reward obtained according to a preset policy;
[0085] The training data is input into the loss function, and a current loss value is calculated;
[0086] According to the current loss value, the parameters of the loss function are adjusted by an optimization algorithm to minimize the loss value;
[0087] The above training data generation, loss value calculation and parameter adjustment steps are repeated until the loss function converges or a preset termination condition is met. The problem of updating discrete values can be solved by reinforcement learning.
[0088] In this scheme, in the reinforcement learning framework, the state space, action space and reward function of the task need to be defined first. The state space describes the environment features that the agent can perceive, the action space defines the operations that the agent can perform, and the reward function quantifies the long-term value of each state-action pair. After initializing the parameters of the loss function, the training process enters the loop iteration phase. In each iteration, the agent interacts with the environment according to the current policy, generates a state transition sequence, executes the action, and obtains the corresponding reward. These interaction data are collected as training samples for optimizing the policy. The training data is input into the loss function to calculate the loss value under the current parameters. The loss function usually measures the difference between the policy prediction and the actual cumulative reward. Then, the parameters of the loss function are adjusted by an optimization algorithm, aiming to minimize the loss value. The data generation, loss calculation and parameter adjustment steps are repeated until the loss function converges or a preset termination condition is met. This iterative optimization process enables the agent to gradually learn the optimal policy, effectively solving the discrete value optimization problem.
[0089] Specifically, at time step t, the state s is the current sequence (v1, K, v t-1 ), a is v t the next value to be selected. Therefore, the generator G θ (v t |V 1:t-1 ) is random, and the state transition is determined after the action is selected, so Gs can be updated based on the generator gradient and the Monte Carlo search based on the expected final reward. This search is based on the expected final reward, which is estimated by G s the possibility of deceiving D s . The purpose of G θ is to generate a discrete sequence that maximizes the expected final reward value The action value function Q can be expressed as:
[0090]
[0091] wherein, is the sample result of N times of Monte Carlo search, the parameter θ of the generator is updated as follows: θ←θ+α h ▽J(θ), wherein α h is the learning rate of the hth step.
[0092] When SeqGAN converges, the ability to generate discrete sequence data is obtained, wherein the attributes and values follow the same relationship as the real data. Whenever the input has the same attribute-value pair, the output always remains the same result, and it can be considered that the input functionally determines the output, which can be regarded as a dependency relationship.
[0093] By adjusting the parameters, the learning ability of the generator can be continuously improved, and finally accurate discrete sequence data can be generated.
[0094] S430, taking the target first sequence data and the second sequence data as input, training the discriminator in the to-be-trained rule generation model to obtain a target rule generation model.
[0095] Further, the target first sequence data and the second sequence data are spliced as input, and the discriminator in the to-be-trained rule generation model is trained to obtain a target rule generation model.
[0096] S440, generating a data set of discrete sequence data according to the target rule generation model.
[0097] In this embodiment, SeqGAN innovatively uses a strategy gradient-based optimization method to act on the generator, so that the model can directly evaluate the generator output, and dynamically adjust the parameters according to the feedback result. With the help of Monte Carlo search technology, the reward signal generated by the GAN discriminator is effectively transmitted to the intermediate state action step, so as to realize the optimization of the generator. Through the continuous adversarial training mechanism, SeqGAN can continuously enhance the learning efficiency of the generator, and finally achieve the goal of accurately generating discrete sequence data.
[0098] Further, the improved SeqGAN is used to generate high-quality discrete data sequences, which has the following advantages:
[0099] SeqGAN model exhibits excellent flexibility, and its core advantage lies in its universality for multiple sequence data types. It is not only suitable for natural language processing scenarios, but also can efficiently process non-text sequence data such as time series and biological sequences. This cross-domain adaptation capability makes it have application potential in financial trend prediction, medical time series analysis, and bioinformatics sequence modeling. By dynamically adapting the sequence characteristics of different data, SeqGAN breaks the application boundaries of traditional models and provides a unified sequence generation solution for cross-disciplinary problems, fully demonstrating the generalization value of generative adversarial networks in sequence modeling.
[0100] SeqGAN uses a policy gradient method to optimize the generator in the discrete decision optimization field, effectively solving the problem of discrete sequence generation. This method breaks through the adaptation bottleneck of traditional continuous optimization for discrete data, can directly evaluate the generator output, and dynamically adjusts the parameters based on the evaluation feedback to form a closed-loop mechanism of generation, evaluation, and optimization, providing a more targeted solution for discrete sequence generation tasks.
[0101] SeqGAN proposes an end-to-end differentiable framework, which provides a key technical breakthrough for discrete sequence data generation tasks. This feature builds a differentiable computation graph, allowing the model to directly use the backpropagation algorithm to optimize parameters, effectively solving the problem of gradient transmission difficulty in discrete output space. Compared to traditional generation models that rely on reinforcement learning or multi-stage training when dealing with discrete data such as text and sequence decisions, this framework converts the generation process into a differentiable computation flow, significantly simplifying the training complexity and providing a more efficient optimization path for sequence generation tasks, especially for scenarios that require high sequence continuity such as natural language generation and dialogue systems.
[0102] SeqGAN exhibits excellent performance advantages in multiple benchmark tests, with its generated sequences achieving a dual breakthrough in quality and diversity. Specifically, the model can produce more realistic sequence data, while enhancing sample diversity to expand the application scope of the model, fully verifying its technical leadership in sequence generation tasks.
[0103] SeqGAN, as the first innovative method that applies GAN to discrete sequence generation tasks, opens up a new path for the generation of discrete sequence data. Its core value lies in breaking through the gradient propagation barrier faced by traditional GAN when dealing with discrete data (such as text, sequence decision, etc.), and by combining reinforcement learning mechanism, the generator is optimized by policy gradient algorithm, so that the model can effectively learn the probability distribution of sequence generation in discrete space. This pioneering attempt not only expands the application boundary of GAN, but also provides an extremely inspiring technical paradigm for natural language generation, dialogue system, recommendation system and other tasks that rely on sequence modeling, and promotes the key breakthrough of discrete sequence generation research from theoretical exploration to practical application.
[0104] SeqGAN uses policy gradient to update the generator, which enables it to learn based on the final reward signal, which is crucial for generating high-quality discrete sequence data.
[0105] SeqGAN draws on the strategy of reinforcement learning (RL) to solve the problem of applying GAN to discrete data, especially the problem of data discontinuity in the process of word index and word vector conversion.
[0106] SeqGAN models the data generator as a random policy in reinforcement learning, bypasses the generator differentiation problem by directly performing gradient policy update, so that the model can update the generation strategy based on the evaluation of the complete sequence.
[0107] The technical scheme of the embodiment of the application acquires first sequence data and second sequence data, trains the rule generation model to be trained based on the two types of data, obtains the target rule generation model, and generates the data set of discrete sequence data relying on the model. By executing the technical scheme, the existing discrete sequence data set can be cleaned and the data generation quality can be improved.
[0108] Embodiment three
[0109] Figure 5 The structure diagram of the discrete sequence data set generation device provided by the third embodiment of the application is shown in FIG. 3. Figure 5 As shown in the figure, the device comprises:
[0110] The data acquisition module 510 is configured to acquire first sequence data and second sequence data; wherein the first sequence data is a discrete sequence data set after cleaning processing; and the second sequence data is an original sequence data set.
[0111] The target rule generation model obtaining module 520 is configured to train a rule generation model to be trained based on the first sequence data and the second sequence data, and obtain a target rule generation model; the target rule generation model is a SeqGAN model, and the SeqGAN model is composed of a generator and a discriminator in a generative adversarial network.
[0112] The discrete sequence data set generation module 530 is configured to generate a discrete sequence data set according to the target rule generation model.
[0113] Optionally, the target rule generation model obtaining module 520 comprises:
[0114] The target first sequence data obtaining unit is configured to take the first sequence data as input, train the generator in the rule generation model to be trained, and obtain target first sequence data.
[0115] The target rule generation model obtaining unit is configured to take the target first sequence data and the second sequence data as input, train the discriminator in the rule generation model to be trained, and obtain the target rule generation model.
[0116] Optionally, the target first sequence data obtaining unit is specifically configured to:
[0117] Determine initial parameters of the generator in the rule generation model to be trained.
[0118] Take the first sequence data as input, train the generator in the rule generation model to be trained, and obtain the target first sequence data.
[0119] Calculate a loss function based on the target first sequence data and a predetermined initial first sequence data.
[0120] Update the initial parameters of the generator in the rule generation model to be trained according to the loss function.
[0121] Optionally, the loss function is an expected final reward value; and the expected final reward value determination process comprises:
[0122] Obtain a state space, an action space and a reward function of a reinforcement learning task.
[0123] Initialize parameters of the loss function.
[0124] In a reinforcement learning training process, generate training data based on a preset strategy according to state transition, action execution and rewards obtained by an agent in an environment.
[0125] Input the training data into the loss function, and calculate a current loss value.
[0126] According to the current loss value, parameters of the loss function are adjusted by an optimization algorithm to minimize the loss value.
[0127] The above training data generation, loss value calculation and parameter adjustment steps are repeated until the loss function converges or a preset termination condition is met.
[0128] Optionally, the target rule generation model obtaining module 520 comprises:
[0129] The candidate rule generation unit is configured to generate a candidate rule in a training process of the rule generation model to be trained, wherein the candidate rule is a rule for repairing the data set.
[0130] Optionally, the candidate rule generation unit is specifically configured to:
[0131] The first sequence data and the second sequence data are input into the rule generation model to be trained for training, and parameters of the rule generation model to be trained are adjusted in the training process to generate the candidate rule.
[0132] Optionally, the target rule generation model obtaining module 520 is further configured to:
[0133] For the candidate rule, the number of tuples is defined as n.
[0134] In the rule generation process, the number of tuples matched by each candidate rule is determined.
[0135] The candidate rule with the number of matched tuples n < 2 is removed, wherein the number of matched tuples n = 0 indicates that the rule cannot match a real tuple, and n = 1 indicates that the rule can only match one error tuple.
[0136] The discrete sequence data data set generation apparatus provided in the embodiments of the present application can execute the discrete sequence data data set generation method provided in any of the embodiments of the present application, and has the corresponding function modules and beneficial effects of the execution method.
[0137] Embodiment four
[0138] Figure 6A structural diagram of an electronic device 10 that can be used to implement embodiments of the present application is shown. The electronic device is intended to represent various forms of digital computers, such as laptops, desktops, tablets, personal digital assistants, servers, blade servers, mainframes, and other appropriate computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular telephones, smartphones, wearable devices (e.g., headsets, glasses, watches, etc.), and other similar computing devices. The components shown here, their connections and relationships, and their functions, are meant to be examples only, and are not intended to limit the implementations of the present application described and / or claimed in this document.
[0139] As shown in Figure 6 The electronic device 10 includes at least one processor 11, and a memory, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc., communicatively connected to the at least one processor 11, where the memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes according to the computer programs stored in the read-only memory (ROM) 12 or loaded into the random access memory (RAM) 13 from the storage unit 18. In the RAM 13, various programs and data required for the operation of the electronic device 10 can also be stored. The processor 11, the ROM 12, and the RAM 13 are connected to each other through a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0140] Various components in the electronic device 10 are connected to the I / O interface 15, including an input unit 16, such as a keyboard, a mouse, etc., an output unit 17, such as various types of displays, speakers, etc., a storage unit 18, such as a magnetic disk, an optical disk, etc., and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices through a computer network, such as the Internet, and / or various telecommunication networks.
[0141] The processor 11 can be various general and / or special purpose processing components with processing and computing capabilities. Some examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The processor 11 performs various methods and processes described above, such as the data set generation method for discrete sequence data.
[0142] In some embodiments, the dataset generation method of discrete sequence data can be implemented as a computer program tangibly embodied in a computer readable storage medium, e.g., storage unit 18. In some embodiments, parts or all of the computer program can be loaded and / or installed onto electronic device 10 via, e.g., ROM 12 and / or communication unit 19. When the computer program is loaded onto RAM 13 and executed by processor 11, one or more steps of the above-described dataset generation method of discrete sequence data can be performed. Alternatively, in other embodiments, processor 11 can be configured to perform the dataset generation method of discrete sequence data by other means, e.g., with the aid of firmware.
[0143] Various implementations of the systems and techniques described above can be realized in digital electronic circuitry, integrated circuitry, a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), a system on a chip (SOC), a complex programmable logic device (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.
[0144] Computer programs used to implement the methods of the present application can be written in any combination of one or more programming languages. These computer programs can be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the computer program, when executed by the processor, implements the functions / acts specified in the flowcharts and / or block diagrams. The computer program can be executed entirely on a machine, partially on a machine, partially on a machine as part of a standalone software package, partially on a machine and partially on a remote machine or entirely on a remote machine or server.
[0145] In the context of the present application, a computer-readable storage medium can be a tangible medium that can contain or store a computer program for use by or in connection with an instruction execution system, apparatus, or device. A computer-readable storage medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. Alternatively, a computer-readable storage medium can be a machine-readable signal medium. More specific examples of a machine-readable storage medium will include one or more lines of a program of instructions in a transitory signal, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0146] To provide for interaction with a user, the systems and techniques described here can be implemented on an electronic device having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the electronic device. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, including acoustic, speech, or tactile input.
[0147] The systems and techniques described here can be implemented in a computing system that includes a back end component (e.g., as a data server), or that includes a middleware component (e.g., an application server), or that includes a front end component (e.g., a user computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the systems and techniques described here), or any combination of such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.
[0148] The computing system can include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a host product in the cloud computing service system, to solve the defects of large management difficulty and weak business scalability in traditional physical host and VPS service.
[0149] It should be understood that the various forms of flow shown above can be reordered, additional steps added, or steps deleted. For example, the steps described in the present application can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions of the present application can be achieved, which are not limited herein.
[0150] The above detailed description does not constitute a limitation on the protection scope of the present application. Those skilled in the art should understand that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modifications, equivalent replacements and improvements made within the spirit and principles of the present application shall be included in the protection scope of the present application.
Claims
1. A method for generating a dataset of discrete sequence data, characterized in that: include: Acquire first sequence data and second sequence data; wherein the first sequence data is a discrete sequence data set after cleaning; and the second sequence data is an original sequence data set; Training a to-be-trained rule generation model based on the first sequence data and the second sequence data to obtain a target rule generation model; wherein the target rule generation model is a SeqGAN model, which is composed of a generator and a discriminator in a generative adversarial network; A dataset of discrete sequence data is generated according to the target rule generation model.
2. The method according to claim 1, characterized in that Training a to-be-trained rule generation model based on the first sequence data and the second sequence data to obtain a target rule generation model includes: Using the first sequence data as input, training the generator in the training rule generation model to obtain target first sequence data; The target first sequence data and the second sequence data are used as inputs to train the discriminator in the rule generation model to be trained to obtain a target rule generation model.
3. The method according to claim 2, characterized in that The first sequence data is used as input to train a generator in a training rule generation model to obtain target first sequence data, including: Determine the initial parameters of the generator in the rule generation model to be trained; Using the first sequence data as input, training the generator in the training rule generation model to obtain target first sequence data; Calculating a loss function based on the target first sequence data and predetermined initial first sequence data; Based on the loss function, the initial parameters of the generator in the rule generation model to be trained are updated.
4. The method according to claim 3, characterized in that The loss function is the expected final reward value; the process of determining the expected final reward value includes: Obtain the state space, action space, and reward function of the reinforcement learning task; Initialize the parameters of the loss function; During reinforcement learning training, training data is generated based on the preset strategy according to the state transitions, actions performed, and rewards obtained by the agent in the environment; Input the training data into the loss function to calculate the current loss value; According to the current loss value, adjusting the parameters of the loss function through an optimization algorithm to minimize the loss value; Repeat the above steps of training data generation, loss value calculation and parameter adjustment until the loss function converges or meets the preset termination condition.
5. The method according to claim 1, wherein Training a to-be-trained rule generation model based on the first sequence data and the second sequence data to obtain a target rule generation model includes: Candidate rules are generated during the training process of the training rule generation model; wherein the candidate rules are rules for repairing the data set.
6. The method according to claim 5, characterized in that Candidate rules are generated during the training process of the training rule generation model, including: The first sequence data and the second sequence data are input into a rule generation model to be trained for training. During the training process, parameters of the rule generation model to be trained are adjusted to generate candidate rules.
7. The method according to claim 5, characterized in that After generating candidate rules during the training of the rule generation model to be trained, the method further includes: For candidate rules, the number of tuples is defined as n; During the rule generation process, the number of tuples matched by each candidate rule is determined; Remove candidate rules with matching tuple number n < 2; wherein, matching tuple number n = 0 indicates that the rule cannot match a real tuple, and n = 1 indicates that the rule can only match one wrong tuple.
8. A device for generating a data set of discrete sequence data, characterized in that: include: A data acquisition module, configured to acquire first sequence data and second sequence data; wherein the first sequence data is a discrete sequence data set after cleaning; and the second sequence data is an original sequence data set; a target rule generation model obtaining module, configured to train the to-be-trained rule generation model based on the first sequence data and the second sequence data to obtain a target rule generation model; wherein the target rule generation model is a SeqGAN model, which is composed of a generator and a discriminator in a generative adversarial network; The discrete sequence data set generation module is used to generate a discrete sequence data set according to the target rule generation model.
9. An electronic device, characterized in that: The electronic device comprises: At least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor to enable the at least one processor to perform the method for generating a dataset of discrete sequence data according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the method for generating a dataset of discrete sequence data according to any one of claims 1 to 7 when executed.