Data generation method and device, equipment and medium

By combining the target language model and the target codebook, the problem of high cost of multimodal data generation in existing technologies is solved, and unified processing of image and text data is achieved, improving the convenience and flexibility of data generation.

CN121935334APending Publication Date: 2026-04-28BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING ZITIAO NETWORK TECH CO LTD
Filing Date
2024-10-24
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

In existing technologies, data generation requires setting up separate graph-to-text, text-to-graph, or text-to-text models, which is costly and complex, and makes it difficult to achieve flexible processing of multimodal data through a single model.

Method used

Using a target language model and a target codebook, the sequence number of the data to be processed is obtained, and the corresponding discrete feature sequence number is generated based on the target language model. Finally, the target data is generated. The target codebook contains multiple text and image features, realizing unified processing of image modality and text modality.

Benefits of technology

It enables the convenient generation of image and text data using the same model, reducing costs, improving the flexibility and convenience of data generation, and avoiding the competitive impact between different modalities of data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121935334A_ABST
    Figure CN121935334A_ABST
Patent Text Reader

Abstract

The embodiment of the invention relates to a data generation method and device, equipment and a medium. The method comprises the steps of obtaining to-be-processed first data; determining a first serial number corresponding to the first data based on a target codebook corresponding to the target language model; the first serial number comprises a serial number of a discrete feature corresponding to the first data in the target codebook; the target codebook comprises a plurality of discrete features, and the plurality of discrete features comprise a plurality of text features and a plurality of image features; based on the first serial number, utilizing a target language model to generate a second serial number; the second serial number comprises a serial number of a discrete feature corresponding to second data to be generated in the target codebook; second data is generated based on the second serial number. According to the embodiment of the invention, flexible processing of various modal data can be realized by means of the same target language model based on the same codebook, the required cost is relatively low, and the convenience of data generation is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of computer technology, and in particular to a data generation method, apparatus, device, and medium. Background Technology

[0002] Currently, data generation technology has been widely applied in various fields such as entertainment, education, and home furnishing. This technology can generate corresponding descriptive text based on images, generate corresponding images based on image descriptive text, and generate corresponding response text based on text. The inventors discovered that related technologies struggle to achieve these effects using a single model; each modality of data requires a corresponding network module for processing, resulting in high costs and complexity. Summary of the Invention

[0003] In order to solve the above-mentioned technical problems, or at least partially solve the above-mentioned technical problems, this disclosure provides a data generation method, apparatus, device and medium.

[0004] This disclosure provides a data generation method, the method comprising: acquiring first data to be processed; determining a first sequence number corresponding to the first data based on a target codebook corresponding to a target language model; wherein the first sequence number includes the sequence number of a discrete feature corresponding to the first data in the target codebook; the target codebook includes multiple discrete features, the multiple discrete features including multiple text features and multiple image features; generating a second sequence number based on the first sequence number using the target language model; wherein the second sequence number includes the sequence number of a discrete feature corresponding to second data to be generated in the target codebook; generating second data based on the second sequence number; wherein the first modality corresponding to the first data is any one of an image modality and a text modality, and the second modality corresponding to the second data is any one of the image modality and the text modality.

[0005] Optionally, the serial numbers corresponding to the multiple text features in the target codebook are arranged sequentially, and the serial numbers corresponding to the multiple image features in the target codebook are arranged sequentially; and the smallest serial number among the serial numbers corresponding to the multiple image features is placed after the largest serial number among the serial numbers corresponding to the multiple text features.

[0006] Optionally, when the first modality is an image modality, determining the first sequence number corresponding to the first data based on the target codebook corresponding to the target language model includes: determining the target sequence number corresponding to the first data based on the image codebook corresponding to the target generation model; wherein the target sequence number includes the sequence number of the image feature corresponding to the first data in the image codebook; determining the first sequence number corresponding to the first data based on the target sequence number and the target codebook corresponding to the target language model; wherein the number of image features contained in the target codebook is equal to the number of image features contained in the image codebook, and the difference between the first sequence number and the target sequence number is the total number of text features contained in the target codebook.

[0007] Optionally, when the second modality is an image modality, generating the second data based on the second sequence number includes: determining a third sequence number based on the image codebook corresponding to the target generation model and the second sequence number; wherein the third sequence number contains the sequence number of the image feature corresponding to the second data to be generated in the image codebook, the number of image features contained in the target codebook is equal to the number of image features contained in the image codebook, and the difference between the second sequence number and the third sequence number is the total number of text features contained in the target codebook; determining the image feature corresponding to the third sequence number from the image codebook; and generating the second data based on the image feature corresponding to the third sequence number.

[0008] Optionally, the target language model and the target codebook are obtained through the following methods: A training sample set is acquired; the training sample set includes a first sample pair, a second sample pair, and a third sample pair; wherein, the first sample pair is a sample pair consisting of an image sample and a corresponding descriptive text sample; the second sample pair is a sample pair consisting of a text sample describing the image to be generated and a corresponding generated image sample; the third sample pair is a sample pair consisting of a target text sample and a corresponding response text sample; the sequence number of the discrete features of each sample in the training sample set in the initial augmented codebook corresponding to the preset language model is acquired; wherein, the initial augmented codebook is obtained by adding multiple random image features to the text codebook corresponding to the preset language model before adjustment, and the total number of the multiple random image features is equal to the total number of image features contained in the image codebook corresponding to the target generation model; based on the sequence number corresponding to each sample in the training sample set, the parameters of the preset language model and the random image features in the initial augmented codebook are adjusted; the target language model is obtained based on the preset language model after parameter adjustment, and the target codebook is obtained based on the initial augmented codebook after feature adjustment.

[0009] Optionally, adjusting the parameters of the preset language model and the random image features in the initial augmented codebook includes: adjusting the parameters of the preset language model and the random image features in the initial augmented codebook using an autoregressive training method.

[0010] Optionally, at least two input controls and at least two output controls are provided on the target interactive interface; wherein, different input controls correspond to different data modalities, and different output controls correspond to different data modalities; the step of obtaining the first data to be processed includes: obtaining the first data to be processed based on the data received by the target input control among the at least two input controls; the method further includes: determining the target output control corresponding to the second modal from the at least two output controls, and displaying the second data through the target output control.

[0011] This disclosure also provides a data generation apparatus, comprising: a first data acquisition module for acquiring first data to be processed; a first sequence number determination module for determining a first sequence number corresponding to the first data based on a target codebook corresponding to a target language model; wherein the first sequence number includes the sequence number of a discrete feature corresponding to the first data in the target codebook; the target codebook includes multiple discrete features, the multiple discrete features including multiple text features and multiple image features; a second sequence number generation module for generating a second sequence number based on the first sequence number using the target language model; wherein the second sequence number includes the sequence number of a discrete feature corresponding to the second data to be generated in the target codebook; and a second data generation module for generating second data based on the second sequence number; wherein the first modality corresponding to the first data is any one of an image modality and a text modality, and the second modality corresponding to the second data is any one of the image modality and the text modality.

[0012] This disclosure also provides an electronic device, the electronic device comprising: a processor; a memory for storing executable instructions of the processor; the processor being configured to read the executable instructions from the memory and execute the instructions to implement the data generation method provided in this disclosure.

[0013] This disclosure also provides a computer-readable storage medium storing a computer program for executing the data generation method provided in this disclosure.

[0014] This disclosure also provides a computer program product, including a computer program that, when executed by a processor, implements the data generation method provided in this disclosure.

[0015] The technical solution provided in this disclosure can determine the first sequence number (containing the sequence number of the discrete feature corresponding to the first data in the target codebook) based on the target codebook corresponding to the target language model (which contains both multiple text features and multiple image features). Based on the first sequence number, a second sequence number (containing the sequence number of the discrete feature corresponding to the second data to be generated in the target codebook) is generated using the target language model. The second data can then be generated. The first modality corresponding to the first data and the second modality corresponding to the second data can both be either an image modality or a text modality. This method can achieve flexible processing of multiple modalities of data based on the same codebook and using the same target language model. It can achieve rich data generation effects such as image-to-text, text-to-image, and text-to-text generation using the same target language model, with low cost and greatly improved convenience of data generation.

[0016] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0017] The accompanying drawings, which are incorporated in and form a part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure.

[0018] To more clearly illustrate the technical solutions in the embodiments of this disclosure or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0019] Figure 1 A flowchart illustrating a data generation method provided in an embodiment of this disclosure;

[0020] Figure 2 This is a schematic flowchart of an image generation method provided in an embodiment of the present disclosure;

[0021] Figure 3 A flowchart illustrating a text generation method provided in an embodiment of this disclosure;

[0022] Figure 4 A flowchart illustrating a text generation method provided in an embodiment of this disclosure;

[0023] Figure 5 This is a schematic diagram of the structure of a data generation apparatus provided in an embodiment of the present disclosure;

[0024] Figure 6This is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure. Detailed Implementation

[0025] To better understand the above-mentioned objectives, features, and advantages of this disclosure, the solutions disclosed herein will be further described below. It should be noted that, unless otherwise specified, the embodiments and features described herein can be combined with each other.

[0026] Numerous specific details are set forth in the following description in order to provide a full understanding of this disclosure, but this disclosure may also be implemented in other ways different from those described herein; obviously, the embodiments in the specification are only some, and not all, of the embodiments of this disclosure.

[0027] The inventors discovered through research that existing data generation methods are inadequate. For example, they require separate models such as image-to-text, text-to-image, or text-to-text models to achieve the desired image-to-text, text-to-image, and text-to-text effects respectively, resulting in high costs. Even for existing models capable of processing multimodal data simultaneously, it is still necessary to configure corresponding decoders and other network modules for each modality within the model. This is not only costly but also leads to competition between modalities, making it difficult to achieve satisfactory generation results. To improve at least one of the above problems, this disclosure provides a data generation method, apparatus, device, and medium, which are described in detail below:

[0028] Figure 1 This is a flowchart illustrating a data generation method provided in an embodiment of the present disclosure. The method can be executed by a data generation device, which can be implemented using software and / or hardware, and is generally integrated into an electronic device. Figure 1 As shown, the method mainly includes the following steps S102 to S108:

[0029] Step S102: Obtain the first data to be processed. The first data can be an image or text, depending on the user input. This embodiment does not limit the specific content of the first data. For example, if the first data is an image, it can be any image or an image containing a target object. If the first data is text, it can be text used to describe the image information to be generated, or it can be dialogue text to be responded to.

[0030] Step S104: Based on the target codebook corresponding to the target language model, determine the first sequence number corresponding to the first data; wherein, the first sequence number contains the index of the discrete feature corresponding to the first data in the target codebook; the target codebook contains multiple discrete features, including multiple text features and multiple image features. The image features and text features have the same feature dimension. Furthermore, the first data typically corresponds to multiple discrete features in the target codebook, and the first sequence number is the permutation set of the indices of each of the multiple discrete features corresponding to the first data; it can also be described as the first sequence number containing multiple discretized codebook numbers.

[0031] The codebook can be used to discretize and represent the input data of the model. It should be noted that existing language models use text codebooks, that is, codebooks that only contain text features. Based on this, the language model mainly processes text data. However, the embodiments of this disclosure improve upon this by expanding multiple image features on top of the multiple text features contained in the original text codebook, thereby enabling the language model to process multiple modalities of data, such as image data and text data.

[0032] This disclosure does not limit the number of text features and image features. For example, the text features can be those found in the text codebook of an existing language model, totaling 32,000. The number of image features can be equal to the number of image features in the image codebook of an existing generative network, totaling 8,192. The existing generative network can be, for example, a generative adversarial network implemented using vector quantization. Based on the above example, the target codebook can contain a total of 40,192 features.

[0033] The target codebook may contain multiple discrete features. In some embodiments, the target codebook may explicitly specify the sequence number (also called the number) corresponding to each discrete feature, or record the association between discrete features and their corresponding sequence numbers. In other embodiments, the target codebook may contain only discrete features, and the sequence number of each discrete feature can be determined based on its sorting position during subsequent processing. In some embodiments, the sequence numbers corresponding to multiple text features in the target codebook are arranged sequentially, and the sequence numbers corresponding to multiple image features in the target codebook are arranged sequentially; moreover, the smallest sequence number among the multiple image features is placed after the largest sequence number among the multiple text features. To minimize codebook modifications and improve the efficiency of target codebook acquisition, the amplified image features can be directly placed after the multiple text features in the original text codebook to efficiently obtain the target codebook. It should be noted that both the target codebook and the target language model using the target codebook need to be trained and adjusted based on the initial language model and the initial augmented codebook. For example, the initial augmented codebook includes multiple text features and multiple random image features from the original text codebook. The initial language model and the initial augmented codebook are trained and adjusted based on the training samples, and finally, usable target codebook and target language model can be obtained.

[0034] In this way, regardless of the modality of the first data, it can be directly converted into a discrete codebook number for representation, which is simpler and easier for subsequent model processing, ensuring the data generation effect. Specifically, data of different modalities can be converted into discrete numbers through the same codebook for unified processing, and there is no need to set up decoders and other network modules for different modalities in the same model. This can effectively avoid the situation where different modalities of data compete and affect the generation effect.

[0035] Step S106: Based on the first sequence number, generate a second sequence number using the target language model; wherein the second sequence number contains the index of the discrete feature corresponding to the second data to be generated in the target codebook. The target language model can generate data based on the first sequence number and output the second sequence number corresponding to the second data to be generated, representing the second data in a discretized form.

[0036] Step S108: Generate second data based on the second serial number; wherein, the first modality corresponding to the first data is either an image modality or a text modality, and the second modality corresponding to the second data is either an image modality or a text modality, and the first modality corresponding to the first data and the second modality corresponding to the second data may be the same as or different from each other. In practical applications, the second serial number can be decoded to obtain the second data.

[0037] The above methods can be based on the same codebook and use the same target language model to achieve flexible processing of multiple modal data. With the same target language model, rich data generation effects such as image-to-text, text-to-image, and text-to-text can be achieved. The required cost is low and the convenience of data generation is greatly improved.

[0038] When the first modality is text modality, by referring to the existing language model to convert text data into sequence numbers in the corresponding text codebook, the first sequence number corresponding to the first data can be directly obtained based on the target codebook corresponding to the target language model. When the first modality is image modality, this embodiment provides an implementation method for determining the first sequence number corresponding to the first data based on the target codebook corresponding to the target language model, which can be specifically executed as follows: Steps (1) to (2) are as follows:

[0039] Step (1): Based on the image codebook corresponding to the target generation model, determine the target sequence number corresponding to the first data; wherein, the target sequence number includes the sequence number of the image feature corresponding to the first data in the image codebook, such as the target sequence number can be presented in the form of [23,123,5112,566,123,623,145,232,425……].

[0040] For example, the target generation model may include a generative adversarial network based on vector quantization technology. That is, the target generation network can use vector quantization technology to compress and process image data, converting the input image into a discrete representation, which can then be used to generate new images. The target generation model can directly adopt an existing model. In the embodiments of this disclosure, when processing the first data of the image modality, in order to efficiently and accurately obtain the first sequence number in the target codebook corresponding to the user input image (first data), the target sequence number in the image codebook of the target generation model corresponding to the first data can be determined first. That is, based on the target generation model, the first data is first converted into a discrete image feature codebook number (target sequence number), so that the first sequence number corresponding to the first data in the target codebook can be determined subsequently based on the correspondence between the target codebook and the image codebook.

[0041] Step (2): Based on the target sequence number and the target codebook corresponding to the target language model, determine the first sequence number corresponding to the first data; wherein, the number of image features contained in the target codebook is equal to the number of image features contained in the image codebook, and the difference between the first sequence number and the target sequence number is the total number of text features contained in the target codebook.

[0042] In practical applications, the number of image features contained in the target codebook is equal to the number of image features contained in the image codebook, such as 8192 in both. However, it should be noted that the image features contained in the target codebook are not identical to those contained in the image codebook; they are only related. The image features contained in the target codebook are features obtained through training and adjustment. In this embodiment, the number of image features contained in the target codebook is set to be equal to the number of image features contained in the image codebook, and is placed after the text features contained in the target codebook. Therefore, the difference between the sequence number of a certain image feature in the target codebook and its corresponding sequence number in the image codebook is equal to the total number of text features contained in the target codebook. In this embodiment, the corresponding target sequence number can be obtained first through an existing target generation model, and then the total number of text features can be added to quickly determine the first sequence number corresponding to the first data. For example, assuming the total number of text features contained in the target codebook is 32000, and the sequence number of a certain discrete feature in the image codebook corresponding to the first data is 32, then the corresponding sequence number in the target codebook is 32032. The first data usually corresponds to the sequence number of multiple discrete features in the image codebook. Adding 32000 to each sequence number will give you the first sequence number corresponding to the first data.

[0043] When the second modality is text-based, following the existing method of generating text data based on the sequence number corresponding to the text codebook in language models, the second data can be directly generated based on the second sequence number. When the second modality is image-based, the steps for generating the second data based on the second sequence number can be performed as follows: steps a to c.

[0044] Step a: Based on the image codebook and second sequence number corresponding to the target generation model, determine the third sequence number. The third sequence number contains the index of the image feature corresponding to the second data to be generated in the image codebook. The number of image features in the target codebook is equal to the number of image features in the image codebook, and the difference between the second and third sequence numbers is the total number of text features in the target codebook. That is, after obtaining the second sequence number, it can be converted into the third sequence number corresponding to the second data to be generated in the image codebook. For example, assuming the total number of text features is 32000, subtracting 32000 from each index in the second sequence number yields the third sequence number. Using this method, image generation can be directly performed based on the existing target generation model, fully ensuring the image generation effect. There is no need to add additional network modules such as image decoding to the target language model, greatly reducing structural modifications to the language model. Whether the data is text-based or image-based, the target language model can discretize and process it based on the same codebook, making the interaction between image and text features simpler and more convenient.

[0045] Step b: Determine the image features corresponding to the third sequence number from the image codebook. Based on the correspondence between each image feature and its corresponding sequence number in the image codebook, the corresponding image features can be directly obtained by filtering based on the known third sequence number.

[0046] Step c: Generate the second data based on the image features corresponding to the third sequence number. This can be achieved by decoding the image features corresponding to the third sequence number. Specifically, refer to existing target generation models that generate corresponding images based on known codebook sequence numbers; details will not be elaborated here.

[0047] For example, the target language model and the target codebook are obtained through the following steps A to D:

[0048] Step A: Obtain the training sample set. The training sample set includes a first sample pair, a second sample pair, and a third sample pair. The first sample pair consists of an image sample and its corresponding descriptive text sample. The second sample pair consists of a text sample describing the image to be generated and its corresponding generated image sample. The third sample pair consists of a target text sample and its corresponding response text sample. This training sample set can also be called the full dataset. It should also be noted that the training sample set may contain other sample pairs, such as a fourth sample pair, which consists of a first image sample and its corresponding generated image sample. This also facilitates the model's ability to perform image-to-image tasks. The specific combination of sample pairs in the training sample set can be flexibly set according to requirements and is not limited here.

[0049] Step B involves obtaining the sequence number of the discrete features of each sample in the training sample set within the initial augmented codebook corresponding to the preset language model. The initial augmented codebook is obtained by adding multiple random image features to the text codebook corresponding to the preset language model before adjustment. The total number of random image features is equal to the total number of image features contained in the image codebook corresponding to the target generation model. For example, if the total number of image features in the existing image codebook corresponding to the target generation model is 8192, then the total number of random image features is also 8192. The text codebook corresponding to the preset language model before adjustment can directly use the existing language model's text codebook, which, for example, can contain 32,000 text features, and no modification is required here.

[0050] Step C involves adjusting the parameters of the preset language model and the random image features in the initial expanded codebook based on the sequence numbers corresponding to each sample in the training sample set. This adjustment process is also the training process of the model and codebook. During training, the text features in the original text codebook can be used to minimize changes to the language model, effectively reduce training difficulty, and efficiently obtain a target language model that can process both images and text simultaneously.

[0051] It should be noted that the embodiments of this disclosure do not require direct training using a training sample set. For example, instead of explicitly inputting images into the model, only the corresponding sequence number of the image needs to be input. This approach allows the language model to process both image and text data uniformly based on discretized sequence numbers and output sequence numbers. Subsequently, the required data can be obtained by parsing based on the sequence numbers output by the language model and the codebook. In practical applications, for text samples in the training sample set, the corresponding sequence number can be directly determined based on the initial augmented codebook. For image samples in the training sample set, the sequence number corresponding to them in the initial augmented codebook is equal to the sum of their sequence number in the aforementioned image codebook and the total number of text features in the initial augmented codebook.

[0052] Furthermore, in practical applications, autoregressive training can be used to adjust the parameters of the preset language model and the random image features in the initial expanded codebook. Autoregressive training, which uses previous elements to predict the next element, can be understood as follows: since both image and text features are represented by discretized indices in the codebook, not only can they be trained jointly, but autoregressive training can also be used to adjust the model and codebook.

[0053] During training, the first sample pair mentioned above corresponds to the image-to-text task. For example, based on the sequence number corresponding to the image sample, a preset language model outputs the corresponding sequence number prediction result. Then, based on the difference between the sequence number prediction result and the sequence number corresponding to the descriptive text sample of the image sample, the model and codebook are adjusted so that the sequence number prediction result output by the adjusted model can be close to the sequence number corresponding to the descriptive text sample. Similarly, the training methods for the second sample pair, corresponding to the text-to-image task, and the third sample pair, corresponding to the text-to-text task, are similar. In addition, the training method for the case where the training sample set includes the fourth sample pair, corresponding to the image-to-image task, is also similar and will not be elaborated here. The above multiple sample pairs can be used to jointly train the model on full data, ultimately enabling the preset language model with adjusted parameters to have good processing capabilities for the above multiple tasks. Based on the input of any of the above multiple tasks, it can output the expected result.

[0054] Step D involves obtaining the target language model based on the parameter-adjusted preset language model, and obtaining the target codebook based on the feature-adjusted initial augmented codebook. Specifically, the parameter-adjusted preset language model can be directly used as the target language model, and the feature-adjusted initial augmented codebook can be used as the target codebook.

[0055] Using the above methods, various tasks such as graph-to-text, text-to-graph, and text-to-text can be performed relatively conveniently with the help of the obtained target language model and the corresponding target codebook.

[0056] Based on the foregoing, for ease of understanding, this disclosure provides three specific implementation examples, which can be referred to as follows:

[0057] See Figure 2 The flowchart of an image generation method shown mainly includes the following steps S202 to S210:

[0058] Step S202: Obtain target text; wherein, the target text is used to describe information about the image to be generated.

[0059] Step S204: Determine the first sequence number corresponding to the target text based on the target codebook corresponding to the target language model.

[0060] Step S206: Based on the first sequence number, generate a second sequence number using the target language model; wherein the second sequence number contains the sequence number of the discrete feature corresponding to the image to be generated in the target codebook.

[0061] Step S208: Determine the third sequence number based on the image codebook and the second sequence number corresponding to the target generation model.

[0062] Step S210: Determine the image features corresponding to the third sequence number from the image codebook, and generate the target image based on the image features corresponding to the third sequence number. The target image is also the image that matches the target text. For example, if the target text describes the facial information that the image to be generated needs to present, then the target image is a facial image that matches the facial information described in the target text.

[0063] Using the above methods, text-to-graph tasks can be performed based on the target language model, thus meeting users' text-to-graph needs.

[0064] See Figure 3 The flowchart of a text generation method shown mainly includes the following steps S302 to S310:

[0065] Step S302: Obtain the target image. This embodiment of the disclosure does not limit the content of the target image.

[0066] Step S304: Based on the image codebook corresponding to the target generation model, determine the target sequence number corresponding to the target image; wherein, the target sequence number includes the sequence number of the image feature corresponding to the target image in the image codebook.

[0067] Step S306: Determine the first sequence number corresponding to the target image based on the target sequence number and the target codebook corresponding to the target language model.

[0068] Step S308: Based on the first sequence number, generate a second sequence number using the target language model; wherein the second sequence number contains the sequence number of the discrete feature corresponding to the text to be generated in the target codebook.

[0069] Step S310: Determine the text features corresponding to the second sequence number from the target codebook, and generate the target text based on the text features corresponding to the second sequence number. The target text is the content description text of the target image.

[0070] Using the above methods, graph-to-text tasks can be performed based on the target language model, thus meeting users' graph-to-text needs.

[0071] See Figure 4 The flowchart of a text generation method shown mainly includes the following steps S402 to S408:

[0072] Step S402: Obtain the first text; wherein, the first text is the text to be responded to.

[0073] Step S404: Determine the first sequence number corresponding to the first text based on the target codebook corresponding to the target language model.

[0074] Step S406: Based on the first sequence number, generate a second sequence number using the target language model; wherein the second sequence number contains the sequence number of the discrete feature corresponding to the text to be generated in the target codebook.

[0075] Step S408: Determine the text features corresponding to the second sequence number from the target codebook, and generate the second text based on the text features corresponding to the second sequence number. The second text is the text that responds to the first text.

[0076] Using the above methods, text-to-text tasks can be performed based on the target language model, thus meeting users' text-to-text needs.

[0077] It should be noted that the above Figures 2-4 These are merely examples of text-to-image, image-to-text, and text-to-text relationships. In practical applications, other tasks such as image-to-image generation can also be achieved using the target language model, which will not be listed here. Additionally, Figures 2-4 The specific implementation methods of the steps in the above can be referred to the relevant content, and will not be repeated here.

[0078] In some implementation examples, at least two input controls and at least two output controls are set on the target interactive interface; wherein, different input controls correspond to different data modalities, and different output controls correspond to different data modalities; obtaining the first data to be processed includes: obtaining the first data to be processed based on the data received by the target input control among the at least two input controls; based on the foregoing, the above method further includes: determining the target output control corresponding to the second modal from the at least two output controls, and displaying the second data through the target output control. Through the above settings, users can flexibly upload the first data as needed on the target interactive interface and view the second data generated based on the model, which is very convenient and fast, greatly improving the user's interactive experience. Moreover, the backend only needs to use the same model to realize any combination of interaction between the above images and text, which requires low cost and is easier to implement.

[0079] In summary, the data generation method provided in this disclosure can flexibly process multiple modalities of data based on the same codebook and using the same target language model. It can achieve rich data generation effects such as image-to-text, text-to-image, and text-to-text generation through the same target language model, without requiring separate models for each data generation effect. Furthermore, in this disclosure, data from different modalities can be converted into discretized sequences using the same codebook for unified processing. This facilitates unified processing of multiple modalities of data and eliminates the need to set up separate decoders or other network modules for different modalities within the same model. This effectively avoids competition between different modalities of data that could affect the generation effect. In addition, the above method has low cost and greatly improves the convenience of data generation.

[0080] Corresponding to the aforementioned data generation method, this disclosure further provides a data generation apparatus. Figure 5 This is a schematic diagram of a data generation device provided in an embodiment of the present disclosure. The device can be implemented by software and / or hardware, and is generally integrated into an electronic device, such as... Figure 5 As shown, the data generation device includes:

[0081] The first data acquisition module 502 is used to acquire the first data to be processed.

[0082] The first sequence number determination module 504 is used to determine the first sequence number corresponding to the first data based on the target codebook corresponding to the target language model; wherein, the first sequence number includes the sequence number of the discrete feature corresponding to the first data in the target codebook; the target codebook includes multiple discrete features, including multiple text features and multiple image features.

[0083] The second sequence number generation module 506 is used to generate a second sequence number based on the first sequence number using the target language model; wherein the second sequence number contains the sequence number of the discrete feature corresponding to the second data to be generated in the target codebook.

[0084] The second data generation module 508 is used to generate second data based on the second serial number; wherein the first modality corresponding to the first data is any modality between the image modality and the text modality, the second modality corresponding to the second data is any modality between the image modality and the text modality, and the first modality corresponding to the first data and the second modality corresponding to the second data are the same or different.

[0085] The aforementioned device can flexibly process multiple modal data based on the same codebook and with the help of the same target language model. It can achieve rich data generation effects such as image-to-text, text-to-image, and text-to-text through the same target language model, with low cost and greatly improved convenience of data generation.

[0086] In some implementations, the serial numbers corresponding to multiple text features in the target codebook are arranged sequentially, and the serial numbers corresponding to multiple image features in the target codebook are arranged sequentially; and the smallest serial number among the serial numbers corresponding to the multiple image features is placed after the largest serial number among the serial numbers corresponding to the multiple text features.

[0087] In some implementations, when the first modality is an image modality, the first sequence number determination module is specifically used to: determine the target sequence number corresponding to the first data based on the image codebook corresponding to the target generation model; wherein the target sequence number includes the sequence number of the image feature corresponding to the first data in the image codebook; and determine the first sequence number corresponding to the first data based on the target sequence number and the target codebook corresponding to the target language model; wherein the number of image features contained in the target codebook is equal to the number of image features contained in the image codebook, and the difference between the first sequence number and the target sequence number is the total number of text features contained in the target codebook.

[0088] In some implementations, when the second modality is an image modality, the second sequence number generation module is specifically used to: determine a third sequence number based on the image codebook corresponding to the target generation model and the second sequence number; wherein the third sequence number includes the sequence number of the image feature corresponding to the second data to be generated in the image codebook, the number of image features contained in the target codebook is equal to the number of image features contained in the image codebook, and the difference between the second sequence number and the third sequence number is the total number of text features contained in the target codebook; determine the image feature corresponding to the third sequence number from the image codebook; and generate the second data based on the image feature corresponding to the third sequence number.

[0089] In some embodiments, the apparatus further includes a model and codebook acquisition module, configured to obtain the target language model and target codebook by: acquiring a training sample set; the training sample set including a first sample pair, a second sample pair, and a third sample pair; wherein, the first sample pair is a sample pair consisting of an image sample and a corresponding descriptive text sample; the second sample pair is a sample pair consisting of a text sample describing the image to be generated and a corresponding generated image sample; the third sample pair is a sample pair consisting of a target text sample and a corresponding response text sample; acquiring various samples from the training sample set... The sequence number of discrete features in the initial augmented codebook corresponding to the preset language model; wherein, the initial augmented codebook is obtained by adding multiple random image features to the text codebook corresponding to the preset language model before adjustment, and the total number of the multiple random image features is equal to the total number of image features contained in the image codebook corresponding to the target generation model; based on the sequence number corresponding to each sample in the training sample set, the parameters of the preset language model and the random image features in the initial augmented codebook are adjusted; the target language model is obtained based on the preset language model after parameter adjustment, and the target codebook is obtained based on the initial augmented codebook after feature adjustment.

[0090] In some implementations, the model and codebook acquisition module is specifically used to: adjust the parameters of the preset language model and the random image features in the initial expanded codebook using an autoregressive training method.

[0091] In some implementations, at least two input controls and at least two output controls are displayed on the target interactive interface; wherein, different input controls correspond to different data modalities, and different output controls correspond to different data modalities; the first data acquisition module is specifically used to: obtain first data to be processed based on the data received by the target input control among the at least two input controls; the device further includes a second data display module, used to determine the target output control corresponding to the second modal from the at least two output controls, and set the second data through the target output control.

[0092] The data generation apparatus provided in this disclosure can execute the data generation method provided in any embodiment of this disclosure, and has the corresponding functional modules and beneficial effects of executing the method.

[0093] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the above-described device embodiments can be referred to the corresponding process in the method embodiments, and will not be repeated here.

[0094] This disclosure provides an electronic device, which includes: a storage device storing a computer program thereon; and a processing device for executing the computer program in the storage device to implement the steps of any method of this disclosure.

[0095] The following is for reference. Figure 6 This document illustrates a structural schematic diagram of an electronic device 600 suitable for implementing embodiments of the present disclosure. The terminal devices in these embodiments may include, but are not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 6 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.

[0096] like Figure 6 As shown, electronic device 600 may include a processing device (e.g., a central processing unit, a graphics processor, etc.) 601, which can perform various appropriate actions and processes according to a program stored in read-only memory (ROM) 602 or a program loaded from storage device 608 into random access memory (RAM) 603. RAM 603 also stores various programs and data required for the operation of electronic device 600. Processing device 601, ROM 602, and RAM 603 are interconnected via bus 604. Input / output (I / O) interface 605 is also connected to bus 604.

[0097] Typically, the following devices can be connected to I / O interface 605: input devices 606 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 607 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 608 including, for example, magnetic tapes, hard disks, etc.; and communication devices 609. Communication device 609 allows electronic device 600 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 6 An electronic device 600 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.

[0098] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 609, or installed from a storage device 608, or installed from a ROM 602. When the computer program is executed by the processing device 601, it performs the functions defined in the methods of embodiments of this disclosure.

[0099] In addition to the methods and devices described above, embodiments of this disclosure can also be computer program products, comprising computer program instructions that, when executed by a processor, cause the processor to perform the methods provided in the embodiments of this disclosure. The computer program product can be written in any combination of one or more programming languages ​​to perform the operations of the embodiments of this disclosure. These programming languages ​​include object-oriented programming languages ​​such as Java and C++, as well as conventional procedural programming languages ​​such as C or similar languages. The program code can be executed entirely on a user computing device, partially on a user device, as a standalone software package, partially on a user computing device and partially on a remote computing device, or entirely on a remote computing device or server.

[0100] Furthermore, embodiments of this disclosure may also be computer-readable storage media storing computer program instructions that, when executed by a processor, cause the processor to perform the data generation method provided in embodiments of this disclosure.

[0101] The computer-readable storage medium may be any combination of one or more readable media. A readable medium may be a readable signal medium or a readable storage medium. A readable storage medium may, for example, include, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatuses, or devices, or any combination thereof. More specific examples of readable storage media (a non-exhaustive list) include: electrical connections having one or more wires, portable disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0102] This disclosure also provides a computer program product, including a computer program / instructions that, when executed by a processor, implement the data generation method of this disclosure.

[0103] It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this disclosure in an appropriate manner in accordance with relevant laws and regulations, and their authorization should be obtained.

[0104] For example, upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the software or hardware, such as the electronic device, application, server, or storage medium performing the operations of this disclosed technical solution, based on the prompt message.

[0105] As an optional but non-limiting implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose whether to "agree" or "disagree" to provide personal information to the electronic device.

[0106] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this disclosure.

[0107] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0108] The above description is merely a specific embodiment of this disclosure, enabling those skilled in the art to understand or implement it. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this disclosure. Therefore, this disclosure is not to be limited to the embodiments described herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A data generation method, characterized in that, include: Get the first data to be processed; Based on the target codebook corresponding to the target language model, a first sequence number corresponding to the first data is determined; wherein, the first sequence number includes the sequence number of the discrete feature corresponding to the first data in the target codebook; the target codebook includes multiple discrete features, and the multiple discrete features include multiple text features and multiple image features; Based on the first sequence number, a second sequence number is generated using the target language model; wherein, the second sequence number contains the sequence number of the discrete feature corresponding to the second data to be generated in the target codebook; Second data is generated based on the second sequence number; wherein, the first modality corresponding to the first data is either an image modality or a text modality, and the second modality corresponding to the second data is either an image modality or a text modality.

2. The method according to claim 1, characterized in that, The serial numbers corresponding to the multiple text features in the target codebook are arranged sequentially, and the serial numbers corresponding to the multiple image features in the target codebook are arranged sequentially; and the smallest serial number among the serial numbers corresponding to the multiple image features is placed after the largest serial number among the serial numbers corresponding to the multiple text features.

3. The method according to claim 2, characterized in that, When the first modality is an image modality, determining the first sequence number corresponding to the first data based on the target codebook corresponding to the target language model includes: Based on the image codebook corresponding to the target generation model, the target sequence number corresponding to the first data is determined; wherein, the target sequence number includes the sequence number of the image feature corresponding to the first data in the image codebook; Based on the target sequence number and the target codebook corresponding to the target language model, a first sequence number corresponding to the first data is determined; wherein, the number of image features contained in the target codebook is equal to the number of image features contained in the image codebook, and the difference between the first sequence number and the target sequence number is the total number of text features contained in the target codebook.

4. The method according to claim 2, characterized in that, When the second modality is an image modality, generating the second data based on the second sequence number includes: Based on the image codebook corresponding to the target generation model and the second sequence number, a third sequence number is determined; wherein, the third sequence number contains the sequence number of the image feature corresponding to the second data to be generated in the image codebook, the number of image features contained in the target codebook is equal to the number of image features contained in the image codebook, and the difference between the second sequence number and the third sequence number is the total number of text features contained in the target codebook; Determine the image features corresponding to the third sequence number from the image codebook; The second data is generated based on the image features corresponding to the third serial number.

5. The method according to any one of claims 1 to 4, characterized in that, The target language model and the target codebook are obtained in the following way: Obtain a training sample set; the training sample set includes a first sample pair, a second sample pair, and a third sample pair; wherein, the first sample pair is a sample pair consisting of an image sample and a corresponding descriptive text sample; the second sample pair is a sample pair consisting of a text sample describing the image to be generated and a corresponding generated image sample; the third sample pair is a sample pair consisting of a target text sample and a corresponding response text sample. Obtain the sequence number of the discrete features of each sample in the training sample set in the initial augmented codebook corresponding to the preset language model; wherein, the initial augmented codebook is obtained by adding multiple random image features to the text codebook corresponding to the preset language model before adjustment, and the total number of the multiple random image features is equal to the total number of image features contained in the image codebook corresponding to the target generation model; Based on the sequence number corresponding to each sample in the training sample set, adjust the parameters of the preset language model and the random image features in the initial amplified codebook; The target language model is obtained based on the preset language model after parameter adjustment, and the target codebook is obtained based on the initial augmented codebook after feature adjustment.

6. The method according to claim 5, characterized in that, The adjustment of the parameters of the preset language model and the random image features in the initial expanded codebook includes: The parameters of the preset language model and the random image features in the initial augmented codebook are adjusted using an autoregressive training method.

7. The method according to claim 1, characterized in that, The target interactive interface has at least two input controls and at least two output controls; different input controls correspond to different data modalities, and different output controls correspond to different data modalities. The step of obtaining the first data to be processed includes: obtaining the first data to be processed based on the data received by the target input control among the at least two input controls; The method further includes: determining a target output control corresponding to the second modality from the at least two output controls, and displaying the second data through the target output control.

8. A data generation apparatus, characterized in that, include: The first data acquisition module is used to acquire the first data to be processed. The first sequence number determination module is used to determine the first sequence number corresponding to the first data based on the target codebook corresponding to the target language model; wherein, the first sequence number includes the sequence number of the discrete feature corresponding to the first data in the target codebook; the target codebook includes multiple discrete features, and the multiple discrete features include multiple text features and multiple image features; The second sequence number generation module is used to generate a second sequence number based on the first sequence number using the target language model; wherein, the second sequence number contains the sequence number of the discrete feature corresponding to the second data to be generated in the target codebook; The second data generation module is used to generate second data based on the second serial number; wherein the first modality corresponding to the first data is any one of the image modality and the text modality, and the second modality corresponding to the second data is any one of the image modality and the text modality.

9. An electronic device, characterized in that, The electronic device includes: A storage device on which computer programs are stored; A processing device for executing the computer program in the storage device to implement the steps of the data generation method according to any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The storage medium stores a computer program for executing the data generation method according to any one of claims 1-7.

11. A computer program product, characterized in that, It includes a computer program that, when executed by a processor, implements the data generation method according to any one of claims 1-7.