Data generation method and device, equipment and storage medium

By sampling from the target feature space and processing feature representations using encoding and diffusion units, the challenge of generating accurate and high-quality data in generative models is addressed, enabling the generation of high-quality data samples that are similar to real data.

CN121765366APending Publication Date: 2026-03-31BEIJING ZITIAO NETWORK TECH CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-09-29
Publication Date
2026-03-31

AI Technical Summary

Technical Problem

Existing generative models suffer from inherent uncertainties and limitations in their understanding of the distribution and structure of the target data when generating accurate and high-quality data, resulting in poor generation performance.

Method used

The target data sample is generated by sampling a first feature representation from the target feature space, processing the feature representation using encoding and diffusion units, and providing a second feature representation to a pre-trained language model.

Benefits of technology

High-quality data samples that are highly similar to real data were generated, preserving the core characteristics of the data and ensuring diversity and realism.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121765366A_ABST
    Figure CN121765366A_ABST
Patent Text Reader

Abstract

The embodiment of the invention relates to a data generation method and device, equipment and a computer readable storage medium. The method proposed herein comprises: sampling a first feature representation from a target feature space, the target feature space being determined by processing a set of training samples with a coding unit; processing the first feature representation with a diffusion unit to determine a second feature representation; and providing the second feature representation for a pre-trained language model to generate a target data sample. In this way, according to the embodiment of the invention, the sample generation efficiency and the quality of the generated sample can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The exemplary embodiments disclosed herein relate generally to the field of computers, and more particularly to methods, apparatus, devices and computer-readable storage media for data generation. Background Technology

[0002] With the development of computer technology, generative models have been widely applied to the generation of various modalities. For example, language models can synthesize desired data based on prompts from user input. However, due to the inherent uncertainties in prompt engineering and the limitations of models in understanding the distribution and structure of target data, generating accurate and high-quality data through prompts remains extremely challenging. Summary of the Invention

[0003] In a first aspect of this disclosure, a method for generating data is provided. The method includes: sampling a first feature representation from a target feature space, the target feature space being determined by processing a set of training samples using an encoding unit; processing the first feature representation using a diffusion unit to determine a second feature representation; and providing the second feature representation to a pre-trained language model to generate target data samples.

[0004] In a second aspect of this disclosure, an apparatus for data generation is provided. The apparatus includes: a sampling module configured to sample a first feature representation from a target feature space, the target feature space being determined by processing a set of training samples using an encoding unit; a processing module configured to process the first feature representation using a diffusion unit to determine a second feature representation; and a generation module configured to provide the second feature representation to a pre-trained language model to generate target data samples.

[0005] In a third aspect of this disclosure, an electronic device is provided. The device includes at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit. When executed by the at least one processing unit, the instructions cause the device to perform the method of the first aspect.

[0006] In a fourth aspect of this disclosure, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program that can be executed by a processor to implement the method of the first aspect.

[0007] It should be understood that the content described in this content section is not intended to limit the key or essential features of the embodiments of this disclosure, nor is it intended to restrict the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0008] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. In the drawings, the same or similar reference numerals denote the same or similar elements, wherein:

[0009] Figure 1 A schematic diagram is shown of an example environment in which embodiments of the present disclosure may be implemented;

[0010] Figure 2 A flowchart illustrating an example process for data generation according to some embodiments of this disclosure is shown;

[0011] Figure 3 A schematic block diagram of an example data synthesis system according to some embodiments of the present disclosure is shown;

[0012] Figure 4 A schematic structural block diagram of an example apparatus for data generation according to some embodiments of the present disclosure is shown; and

[0013] Figure 5 A block diagram of an electronic device capable of implementing several embodiments of the present disclosure is shown. Detailed Implementation

[0014] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.

[0015] It should be noted that the headings of any section / subsection provided herein are not limiting. Various embodiments are described throughout this document, and embodiments of any type may be included under any section / subsection. Furthermore, embodiments described in any section / subsection may be combined in any way with any other embodiments described in the same section / subsection and / or different sections / subsections.

[0016] In the description of embodiments of this disclosure, the term "comprising" and similar terms should be understood as open-ended inclusion, i.e., "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The term "some embodiments" should be understood as "at least some embodiments". Other explicit and implicit definitions may also be included below. The terms "first", "second", etc., may refer to different or the same objects. Other explicit and implicit definitions may also be included below.

[0017] The embodiments of this disclosure may involve user data, data acquisition, and / or use. All of these aspects comply with applicable laws, regulations, and relevant provisions. In the embodiments of this disclosure, all data collection, acquisition, processing, manipulation, forwarding, and use are conducted with the user's knowledge and confirmation. Accordingly, in implementing the embodiments of this disclosure, the type, scope of use, and usage scenarios of any data or information that may be involved should be communicated to the user and their authorization obtained in accordance with relevant laws and regulations through appropriate means. The specific methods of notification and / or authorization may vary depending on the actual situation and application scenario, and the scope of this disclosure is not limited in this respect.

[0018] In this specification and the embodiments, any processing of personal information will be carried out only under the premise of legality (such as obtaining the consent of the personal information subject, or being necessary for the performance of a contract), and will only be carried out within the scope stipulated or agreed upon. A user's refusal to process personal information other than that necessary for basic functions will not affect the user's use of basic functions.

[0019] As mentioned above, generating accurate and high-quality data through prompts remains extremely challenging due to the inherent uncertainties of prompting engineering and the limitations of models in understanding the distribution and structure of target data.

[0020] Embodiments of this disclosure propose a data generation scheme. According to this scheme, a first feature representation can be sampled from a target feature space, which is determined by processing a set of training samples using an encoding unit. Further, the first feature representation can be processed using a diffusion unit to determine a second feature representation. Accordingly, the second feature representation can be provided to a pre-trained language model to generate target data samples.

[0021] Through feature space modeling and denoising diffusion processes, embodiments of this disclosure can preserve the core features of the data and ensure the diversity and realism of the synthesized data samples. Therefore, embodiments of this disclosure can generate high-quality data samples that are highly similar to real data.

[0022] The following section provides a detailed description of various example implementations of this scheme, with reference to the accompanying drawings.

[0023] Example Environment

[0024] Figure 1 A schematic diagram of an example environment 100 in which embodiments of the present disclosure can be implemented is shown. For example... Figure 1 As shown, example environment 100 may include electronic device 110.

[0025] In this example environment 100, electronic device 110 can deploy data synthesis system 120 to generate data samples by sampling feature representations from a feature space determined through training. The specific structure and processing of data synthesis system 120 will be referenced below. Figure 2 and Figure 3 Detailed description.

[0026] Electronic device 110 can be any type of mobile terminal, fixed terminal, or portable terminal, including mobile phones, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, media computers, multimedia tablets, handheld computers, portable gaming terminals, VR / AR devices, personal communication system (PCS) devices, personal navigation devices, personal digital assistants (PDAs), audio / video players, digital cameras / camcorders, positioning devices, television receivers, radio receivers, e-book devices, gaming devices, or any combination thereof, including accessories and peripherals of these devices or any combination thereof. In some embodiments, electronic device 110 can also support any type of user-facing interface (such as "wearable" circuitry).

[0027] Electronic device 110 can also be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks, and big data and artificial intelligence platforms. Electronic device 110 may include, for example, computing systems / servers, such as mainframes, edge computing nodes, computing devices in cloud environments, and so on.

[0028] It should be understood that the structure and function of the various elements in environment 100 are described for illustrative purposes only and do not imply any limitation on the scope of this disclosure.

[0029] The following description will continue with reference to the accompanying drawings, which will provide some exemplary embodiments of this disclosure.

[0030] Example process

[0031] Figure 2 A flowchart of an example process 200 for data generation according to some embodiments of the present disclosure is shown. Process 200 can be implemented at electronic device 110. Reference is made below. Figure 1 To describe process 200.

[0032] As shown in the figure, in box 210, electronic device 110 samples a first feature representation from a target feature space, which is determined by processing a set of training samples using an encoding unit.

[0033] The following will be referenced Figure 3 To describe the specific process of data generation. Figure 3 A schematic block diagram of an example data synthesis system 120 according to some embodiments of the present disclosure is shown. Figure 3 As shown, the data synthesis system 120 may include an encoding unit 310, a pre-trained language model 350, and a diffusion unit 325.

[0034] In some embodiments, the encoding unit 310 and the pre-trained language model 350 can constitute a variational autoencoder (VAE). Further, the data synthesis system 120 can train the VAE using training samples. During VAE training, the parameters of the language model 350 can remain fixed, and the parameters of the encoding unit 310 can be adjusted based on the VAE loss.

[0035] In some embodiments, the encoding unit 310 may be implemented using a pre-trained language model, and the size of the language model serving as the encoding unit 310 may be smaller than, for example, the size of the language model 350 serving as the decoding unit in the VAE. As an example, the encoding unit 310 may be implemented using a language model such as injected BERT (Bidirectional Encoder Representations from Transformers).

[0036] As shown in the figure, during the training process of the VAE, the data synthesis system 120 can use the encoding unit 310 to be trained to determine the first training feature representation of the training samples. Further, the data synthesis system 120 can use the pre-trained language model 350 to process the first training feature representation to determine the first training loss of the variational autoencoder (VAE) composed of the encoding unit and the pre-trained language model.

[0037] Furthermore, the data synthesis system 120 can fix the parameters of the language model 350 and adjust the parameters of the encoding unit 310 based on the first training loss. As an example, the first training loss of the VAE can include reconstruction loss and KL loss:

[0038] ELBO β =L rec -βL kl (1)

[0039]

[0040] L kl =DKL (q φ (z|x)||p(z)) (3)

[0041] Formula (1) represents the training objective of VAE, namely the Evidence Lower Bound (ELBO), which can be based on the reconstruction loss L. rec And KL loss L kl The value is determined, where β is a weighting coefficient used to balance the reconstruction quality and the smoothness of the feature space.

[0042] Formula (2) represents the calculation process of reconstruction loss, where q φ (z|x) represents the posterior distribution of the feature vector z generated by the encoding unit given input x, p θ (x|z) represents the probability that the decoding unit (i.e., the language model) reconstructs the input x given the feature vector z. The expectation operator is used to represent the expected value of the reconstruction probability.

[0043] Formula (3) represents the calculation process of KL loss, where D KL This indicates the calculation of the Kullback-Leibler divergence, q θ (z|x) See the explanation of formula (2), p(z) represents the prior distribution of the eigenvector z, for example, the standard normal distribution.

[0044] Therefore, the data synthesis system 120 can use training samples to train the VAE composed of the encoding unit 310 and the language model 350, thereby obtaining the trained encoding unit 310. Furthermore, the data synthesis system 120 can use the trained encoding unit 310 to process a set of training samples 305, thereby determining the target feature space 315 corresponding to the set of training samples 305.

[0045] In some embodiments, as mentioned above, the encoding unit 310 may utilize a language model. Accordingly, the encoding unit 310 may, for example, process the training text content corresponding to the set of training samples 305 to determine the target feature space. As an example, the training samples 305 may include table samples, which may be converted into corresponding text content for input into the encoding unit 310.

[0046] Furthermore, such as Figure 3 As shown, the data synthesis system 120 can sample the first feature representation 320 from the target feature space 315.

[0047] In block 220, electronic device 110 uses a diffusion unit to process a first feature representation to determine a second feature representation.

[0048] like Figure 3 As shown, the diffusion unit 325 may include a noise-adding module for performing noise-adding processing 330 and a noise-reducing module for performing noise-reducing processing 335. Specifically, the diffusion unit 325 may use the noise-adding module to perform noise-adding processing 330 on the sampled first feature representation 320 to generate a noisy feature representation 340. Further, the diffusion unit 325 may use the noise-reducing module to perform noise-reducing processing 335 on the noisy feature representation 340 to determine a second feature representation 345.

[0049] As an example, noise addition processing 330 can be expressed as formula (4), and noise reduction processing 335 can be expressed as formula (5):

[0050]

[0051] In formula (4), z0 represents the initial first feature representation, t represents the time step of the forward diffusion process, and σ(t)| represents the time-dependent noise scaling function, which determines the amount of noise added at time step t. This represents the noise term sampled from the standard normal distribution. In formula (5), This represents the derivative of the noise scaling function with respect to time. This represents the logarithmic probability density function relative to the hidden variable z. t The gradient.

[0052] In some embodiments, training of the diffusion unit 325 may be performed after training of the VAE is completed. Specifically, the data synthesis system 120 may sample from the training feature space, which is determined using the trained encoding units, to determine a second training feature representation. Further, the data synthesis system 120 may process the second training feature representation using the diffusion unit 325 to determine a second training loss associated with the diffusion unit 325.

[0053] As an example, the second training loss can be expressed as formula (6):

[0054]

[0055] in, The expectation operator is used to represent the expected value under a given distribution, where t~p(t) represents the time point sampled from the time distribution p(t), and z0~p(z0) represents the initial hidden feature sampled from the initial feature distribution. Denotes the noise term sampled from the standard normal distribution, ∈ θ (z t ,t) represents a given hidden feature z t A network that predicts noise using a time step t, with parameter θ.

[0056] Furthermore, the data synthesis system 120 can adjust the parameters of the diffusion unit 325 based on the second training loss.

[0057] In box 230, electronic device 110 provides a second feature representation to a pre-trained language model to generate target data samples.

[0058] Continue to refer to Figure 3 The electronic device 110 can provide a second feature representation 345 to the language model 350 to generate target data samples 355.

[0059] In some embodiments, the second feature representation 345 can be injected into the language model 350 through an appropriate pattern to control the processing of the language model 350. Specifically, the data synthesis system 120 can map the second feature representation to a target token embedding, for example, H. latent Furthermore, the data synthesis system 120 can inject the target token embedding H into the pre-trained language model. latent To generate target data samples.

[0060] In some embodiments, the data synthesis system 120 can embed a target token as a soft hint token for the language model using a prefix injection pattern. Specifically, the soft hint token can be added before a preset token (e.g., a BOS token) of the language model.

[0061] In this injection mode, the second feature representation 345 is mapped by a higher-level multilayer perceptron (MLP) into a set of soft cue tag embeddings H. latent H latent It can be used as a guide vector and concatenated before the start token (BOS token) of the language model to help the language model better understand the generation target during the decoding process.

[0062] In some embodiments, the data synthesis system 120 can embed the target token into the key-value cache of the language model through a cache injection (also known as memory injection) mode.

[0063] As an example, a data synthesis system can use H latent This is injected as a key-value (KV) memory of the past into each layer of the language model. This pattern leverages the key-value caching technique used during language model decoding, by mapping the H... latent It is concatenated with KV cache to inject memory information at multiple levels.

[0064] In some embodiments, the data synthesis system 120 may use an embedding injection mode to inject a target token embedding into the token embedding space of the language model to combine it with the original token embedding of the language model.

[0065] As an example, the second feature representation 345 can be directly mapped to the token embedding space to correspond with the original token embedding H. emb Combined, forming a new embedded H emb +H latent This approach allows for the injection of information from the second feature representation by directly modifying the embedding layer of the language model.

[0066] In some embodiments, the language model 350 can be configured to output target text content based on the second feature representation 345 to generate target data samples 355 corresponding to the target text content. Taking a chart sample as an example, the target text content generated by the language model 350 can be further transformed to generate corresponding chart samples.

[0067] In some embodiments, training samples 305 and / or generated target data samples 355 may include any suitable type of data samples, examples of which may include, but are not limited to, text samples, code samples, chart samples, tool samples, etc.

[0068] Therefore, through feature space modeling and denoising diffusion processes, the embodiments of this disclosure can preserve the core features of the data and ensure the diversity and realism of the synthesized data samples. Thus, the embodiments of this disclosure can generate high-quality data samples that are highly similar to real data.

[0069] Example devices and equipment

[0070] Embodiments of this disclosure also provide corresponding apparatus for implementing the above methods or processes. Figure 4 A schematic structural block diagram of an example apparatus 400 for data generation according to certain embodiments of the present disclosure is shown. Apparatus 400 may be implemented as or included in electronic device 110. Various modules / components in apparatus 400 may be implemented by hardware, software, firmware, or any combination thereof.

[0071] like Figure 4 As shown, the apparatus 400 includes a sampling module 410 configured to sample a first feature representation from a target feature space, the target feature space being determined by processing a set of training samples using an encoding unit; a processing module 420 configured to process the first feature representation using a diffusion unit to determine a second feature representation; and a generation module 430 configured to provide the second feature representation to a pre-trained language model to generate target data samples.

[0072] In some embodiments, the encoding unit is trained based on the following process: determining a first training feature representation of the training samples using the encoding unit to be trained; processing the first training feature representation using a pre-trained language model to determine a first training loss of the variational autoencoder (VAE) consisting of the encoding unit and the pre-trained language model; and adjusting the parameters of the encoding unit based on the first training loss.

[0073] In some embodiments, the processing module 420 is configured to: process a first feature representation using a noise-adding module of a diffusion unit to generate a noisy feature representation; and process the noisy feature representation using a denoising module of a diffusion unit to generate a second feature representation.

[0074] In some embodiments, the diffusion unit is trained based on the following process: sampling from a training feature space to determine a second training feature representation, the training feature space being determined using a trained encoding unit; processing the second training feature representation using the diffusion unit to determine a second training loss associated with the diffusion unit; and adjusting the parameters of the diffusion unit based on the second training loss.

[0075] In some embodiments, the generation module 430 is further configured to: map the second feature representation to a target token embedding; and inject the target token embedding into a pre-trained language model to generate a target data sample.

[0076] In some embodiments, the generation module 430 is further configured to: inject the target token embedding as a soft hint token of the language model to be added before the preset token of the language model; inject the target token embedding into the key-value cache of the language model; and inject the target token embedding into the token embedding space of the language model to be combined with the original token embedding of the language model.

[0077] In some embodiments, the target feature space is determined by processing the training text content corresponding to a set of training samples using coding units.

[0078] In some embodiments, the language model is configured to output target text content based on a second feature representation to generate target data samples corresponding to the target text content.

[0079] In some embodiments, the target data sample includes at least one of the following: text sample, code sample, chart sample, and tool sample.

[0080] Figure 5 A block diagram of an electronic device 500 in which one or more embodiments of the present disclosure may be implemented is shown. It should be understood that... Figure 5The electronic device 500 shown is merely exemplary and should not be construed as limiting the functionality and scope of the embodiments described herein. Figure 5 The electronic device 500 shown can be used to achieve Figure 1 Electronic devices 110.

[0081] like Figure 5 As shown, electronic device 500 is in the form of a general-purpose electronic device. Components of electronic device 500 may include, but are not limited to, one or more processors or processing units 510, memory 520, storage device 530, one or more communication units 540, one or more input devices 550, and one or more output devices 560. Processing unit 510 may be a physical or virtual processor and is capable of performing various processes according to programs stored in memory 520. In a multiprocessor system, multiple processing units execute computer-executable instructions in parallel to improve the parallel processing capability of electronic device 500.

[0082] Electronic device 500 typically includes multiple computer storage media. Such media can be any accessible media that is accessible to electronic device 500, including but not limited to volatile and non-volatile media, removable and non-removable media. Memory 520 can be volatile memory (e.g., registers, cache, random access memory (RAM)), non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. Storage device 530 can be removable or non-removable media and can include machine-readable media, such as flash drives, disks, or any other media that can be used to store information and / or data and can be accessed within electronic device 500.

[0083] Electronic device 500 may further include additional removable / non-removable, volatile / non-volatile storage media. Although not explicitly stated... Figure 5 As shown, disk drives for reading from or writing to removable, non-volatile disks (e.g., "floppy disks") and optical disk drives for reading from or writing to removable, non-volatile optical disks can be provided. In these cases, each drive can be connected to a bus (not shown) via one or more data media interfaces. Memory 520 may include computer program product 525 having one or more program modules configured to perform various methods or actions of various embodiments of this disclosure.

[0084] Communication unit 540 enables communication with other electronic devices via a communication medium. Additionally, the functionality of components of electronic device 500 can be implemented using a single computing cluster or multiple computing machines capable of communicating via communication connections. Therefore, electronic device 500 can operate in a networked environment using logical connections to one or more other servers, network personal computers (PCs), or another network node.

[0085] Input device 550 can be one or more input devices, such as a mouse, keyboard, trackball, etc. Output device 560 can be one or more output devices, such as a monitor, speaker, printer, etc. Electronic device 500 can also communicate with one or more external devices (not shown) via communication unit 540 as needed. These external devices include storage devices, display devices, etc., and can communicate with one or more devices that enable user interaction with electronic device 500, or with any device that enables electronic device 500 to communicate with one or more other electronic devices (e.g., network card, modem, etc.). Such communication can be performed via input / output (I / O) interface (not shown).

[0086] According to an exemplary implementation of this disclosure, a computer-readable storage medium is provided that stores computer-executable instructions thereon, wherein the computer-executable instructions are executed by a processor to implement the methods described above. According to an exemplary implementation of this disclosure, a computer program product is also provided, which is tangibly stored on a non-transitory computer-readable medium and includes computer-executable instructions, which are executed by a processor to implement the methods described above.

[0087] Various aspects of this disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatuses, devices, and computer program products implemented according to this disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.

[0088] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processing unit of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.

[0089] Computer-readable program instructions can be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions that execute on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.

[0090] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

[0091] Various implementations of this disclosure have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed implementations. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described implementations. The terminology used herein is chosen to best explain the principles, practical applications, or improvements to technology in the market, or to enable others skilled in the art to understand the various implementations disclosed herein.

Claims

1. A method of data generation, comprising: sampling a first feature representation from a target feature space, the target feature space being determined by processing a set of training samples with an encoding unit; processing the first feature representation with a diffusion unit to determine a second feature representation; and providing the second feature representation to a pre-trained language model to generate a target data sample.

2. The method of claim 1, wherein the encoding unit is trained based on the following process: determining a first training feature representation of a training sample with an encoding unit to be trained; processing the first training feature representation with the pre-trained language model to determine a first training loss of a variational autoencoder (VAE) composed of the encoding unit and the pre-trained language model; and adjusting parameters of the encoding unit based on the first training loss.

3. The method of claim 1, wherein processing the first feature representation with a diffusion unit to determine a second feature representation comprises: processing the first feature representation with a noise adding module of the diffusion unit to generate a noise-added feature representation; and processing the noise-added feature representation with a noise removing module of the diffusion unit to generate the second feature representation.

4. The method of claim 1, wherein the diffusion unit is trained based on the following process: determining a second training feature representation from a training feature space determined with the trained encoding unit; processing the second training feature representation with the diffusion unit to determine a second training loss associated with the diffusion unit; and adjusting parameters of the diffusion unit based on the second training loss.

5. The method of claim 1, wherein providing the second feature representation to a pre-trained language model to generate a target data sample comprises: mapping the second feature representation as a target token embedding; and injecting the target token embedding to the pre-trained language model to generate the target data sample.

6. The method of claim 5, wherein injecting the set of token embeddings to the pre-trained language model comprises one of: injecting the target token embedding as a soft prompt token of the language model to be added before preset marker tokens of the language model; injecting the target token embedding into a key-value cache of the language model; injecting the target token embedding into a token embedding space of the language model to be combined with original token embeddings of the language model.

7. The method of claim 1, wherein the target feature space is determined by processing training textual content corresponding to the set of training samples with the encoding unit.

8. The method of claim 7, wherein the language model is configured to output target textual content based on the second feature representation for generating the target data sample corresponding to the target textual content.

9. The method of claim 1, wherein the target data sample comprises at least one of: a text sample, a code sample, a graph sample, a tool sample. ​ ​ ​ 10. An apparatus for data generation, comprising: a sampling module configured to sample a first feature representation from a target feature space, the target feature space being determined by processing a set of training samples with an encoding unit; a processing module configured to process the first feature representation with a diffusion unit to determine a second feature representation; and a generation module configured to provide the second feature representation to a pre-trained language model to generate a target data sample.

11. An electronic device, comprising: at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions, when executed by the at least one processing unit, cause the electronic device to perform the method according to any one of claims 1-9.

12. A computer-readable storage medium having stored thereon a computer program, the computer program being executable by a processor to implement the method according to any one of claims 1-9. ​ ​