Method, apparatus, electronic device, and storage medium for generating interleaved graphics and text
By training the interlaced graphic and text generation model to improve image consistency, the problem of poor image consistency in the prior art is solved, and the interlaced graphic and text generation with higher consistency and coherence is achieved.
Patent Information
- Application Number
- CN202510170201.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-17
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-02-17
AI Technical Summary
The consistency between images in interlaced graphics generated in the prior art is poor.
By training the initial interleaved graphic generation model based on consistency loss, the consistency loss is determined by using the similarity between the first feature of the sample intermediate feature map corresponding to the current sample image token group and the second feature map of the previous sample intermediate feature map, and the generation model is optimized to improve the consistency between images.
Improve the consistency and coherence between images in the generated interlaced pictures and texts, and enhance the effect of combining pictures and texts.
Smart Images

Figure CN119625139B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular, to a method, device, electronic device and storage medium for generating interleaved text and images. Background Art
[0002] In recent years, with the rapid development of large models in natural language processing and computer vision, multi-modal generation technology has received extensive attention.
[0003] In related technologies, usually the user demand information is input into an autoregressive model, and interleaved text and images are generated through the autoregressive model, but the consistency between the images in the generated interleaved text and images is poor. Summary of the Invention
[0004] The present invention provides a method, device, electronic device and storage medium for generating interleaved text and images, so as to solve the defect that the consistency between the images in the generated interleaved text and images in the prior art is poor.
[0005] The present invention provides a method for generating interleaved text and images, including the following steps.
[0006] Obtain an interleaved text and image generation instruction, where the interleaved text and image generation instruction includes user demand information;
[0007] Input the user demand information into an interleaved text and image generation model to obtain the interleaved text and images output by the interleaved text and image generation model;
[0008] Wherein, the interleaved text and image generation model is obtained by training an initial interleaved text and image generation model based on a consistency loss, and the consistency loss is determined based on the sample similarity between the first feature of the sample intermediate feature map corresponding to the current first sample image token group and the second features of the previous sample intermediate feature maps, and the current first sample image token group is obtained by inputting the interleaved sample text and sample images into the initial multi-modal large model in the initial interleaved text and image generation model.
[0009] According to the method for generating interleaved text and images provided by the present invention, the step of inputting the user demand information into an interleaved text and image generation model to obtain the interleaved text and images output by the interleaved text and image generation model includes:
[0010] Input the user demand information into the multi-modal large model in the interleaved text and image generation model to obtain the current text token group and at least one current image token group output by the multi-modal large model;
[0011] For each of the current image token groups, determine the similarity between the first feature map corresponding to the current image token group and the second feature maps corresponding to the previous image token groups;
[0012] Determine a target current image token group from all the current image token groups based on each of the similarities, generate a current target image based on the target current image token group, and generate a current target text based on the current text token group;
[0013] Generate the interleaved text and images based on each of the target images and each of the target texts.
[0014] According to a method for generating interleaved text and images provided by the present invention, the inputting the user requirement information into a multi-modal large model in an interleaved text and image generation model to obtain a current text token group and at least one current image token group output by the multi-modal large model includes:
[0015] Input the user requirement information into the multi-modal large model, and obtain the current text token group and the at least one current image token group output by the multi-modal large model in an autoregressive manner.
[0016] According to a method for generating interleaved text and images provided by the present invention, the interleaved text and image generation model is trained based on the following method:
[0017] Input the interleaved sample text and the sample images into the initial multi-modal large model to obtain at least one interleaved first sample text token group and a first sample image token group output by the initial multi-modal large model;
[0018] Input the current first sample image token group into an image decoder in the initial interleaved text and image generation model to obtain a current sample intermediate feature map;
[0019] Extract a first feature of the current sample intermediate feature map through an apparent feature extraction model in the initial interleaved text and image generation model, determine a sample similarity between the first feature and second features of each previous sample intermediate feature map, and determine a consistency loss based on each of the sample similarities;
[0020] Determine a classification loss based on the first sample text token group and the first sample image token group;
[0021] Iteratively optimize the parameters of the initial interleaved text and image generation model based on the consistency loss and the classification loss to obtain the interleaved text and image generation model.
[0022] A method for generating interleaved text and images provided by the present invention, wherein inputting the interleaved sample text and the sample image into an initial multi-modal large model in an initial interleaved text and image generation model, and obtaining at least one interleaved first sample text token group and a first sample image token group output by the initial multi-modal large model, includes:
[0023] Based on the interleaved sample text and the sample image, determining a sample token sequence, where the sample token sequence includes an interleaved second sample text token group and a second sample image token group;
[0024] Inputting the interleaved second sample text token group and the second sample image token group into a self-attention module in a first group of modules of the initial multi-modal large model, and obtaining self-attention features output by the self-attention module in an autoregressive manner;
[0025] Inputting the self-attention features into a forward feedback module in the first group of modules, and obtaining hidden layer features output by the forward feedback module in an autoregressive manner;
[0026] Inputting the hidden layer features into a second group of modules, and obtaining hidden layer features output by the second group of modules until obtaining the hidden layer features output by the last group of modules;
[0027] Based on the hidden layer features output by the last group of modules, obtaining the interleaved first sample text token group and the first sample image token group.
[0028] A method for generating interleaved text and images provided by the present invention, wherein inputting the interleaved second sample text token group and the second sample image token group into a self-attention module in a first group of modules of the initial multi-modal large model, and obtaining self-attention features output by the self-attention module in an autoregressive manner, includes:
[0029] Inputting the interleaved second sample text token group and the second sample image token group into the self-attention module in the first group of modules, and obtaining text intermediate hidden layer features corresponding to each second sample text token group and image intermediate hidden layer features corresponding to each second sample image token group;
[0030] Determining a memory matrix in an autoregressive manner based on each of the text intermediate hidden layer features, each of the image intermediate hidden layer features, a text update intensity, and an image update intensity;
[0031] Based on the memory matrix, determining the self-attention features.
[0032] A method for generating interleaved text and images provided by the present invention, the method further includes:
[0033] Receive a modification instruction for the interleaved graphic text, where the modification instruction includes user requirement modification information;
[0034] Input the user requirement modification information, the user requirement information, and the interleaved graphic text into the interleaved graphic text generation model to obtain the modified interleaved graphic text output by the interleaved graphic text generation model.
[0035] The present invention also provides a device for generating interleaved graphic text, including:
[0036] An acquisition unit for acquiring an interleaved graphic text generation instruction, where the interleaved graphic text generation instruction includes user requirement information;
[0037] A first generation unit for inputting the user requirement information into an interleaved graphic text generation model to obtain the interleaved graphic text output by the interleaved graphic text generation model;
[0038] Wherein, the interleaved graphic text generation model is obtained by training an initial interleaved graphic text generation model based on a consistency loss, and the consistency loss is determined based on the sample similarity between the first feature of the sample intermediate feature map corresponding to the current first sample image token group and the second features of the previous sample intermediate feature maps, and the current first sample image token group is obtained by inputting the interleaved sample text and the sample image into the initial multi-modal large model in the initial interleaved graphic text generation model.
[0039] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, where when the processor executes the computer program, the method for generating interleaved graphic text as described in any one of the above is implemented.
[0040] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the method for generating interleaved graphic text as described in any one of the above is implemented.
[0041] The present invention also provides a computer program product, including a computer program, and when the computer program is executed by a processor, the method for generating interleaved graphic text as described in any one of the above is implemented.
[0042] The present invention provides a method, apparatus, electronic device, and storage medium for generating interleaved graphics and texts. The user requirement information included in the interleaved graphics and texts generation instruction is input into a trained interleaved graphics and texts generation model to obtain the interleaved graphics and texts output by the interleaved graphics and texts generation model. Among them, the interleaved graphics and texts generation model is obtained by training an initial interleaved graphics and texts generation model based on a consistency loss, and the consistency loss is determined based on the sample similarity between the first feature of the sample intermediate feature map corresponding to the current first sample image token group and the second features of the previous sample intermediate feature maps. It can be seen that the present invention trains the interleaved graphics and texts generation model based on the consistency loss determined by the sample similarity between the first feature of the sample intermediate feature map corresponding to the current first sample image token group and the second features of the previous sample intermediate feature maps, taking into account the similarity between the sample intermediate feature maps, so as to improve the consistency between the images in the interleaved graphics and texts generated by the interleaved graphics and texts generation model. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0044] Figure 1 It is one of the schematic flowcharts of the method for generating interleaved graphics and texts provided by an embodiment of the present invention.
[0045] Figure 2 It is the second schematic flowchart of the method for generating interleaved graphics and texts provided by an embodiment of the present invention.
[0046] Figure 3 It is the schematic flowchart of the method for training the interleaved graphics and texts generation model provided by an embodiment of the present invention.
[0047] Figure 4 It is the schematic framework diagram of the initial interleaved graphics and texts generation model provided by an embodiment of the present invention.
[0048] Figure 5 It is the third schematic flowchart of the method for generating interleaved graphics and texts provided by an embodiment of the present invention.
[0049] Figure 6 It is one of the example diagrams of the interleaved graphics and texts and the modification of the interleaved graphics and texts provided by an embodiment of the present invention.
[0050] Figure 7 It is the second example diagram of the interleaved graphics and texts and the modification of the interleaved graphics and texts provided by an embodiment of the present invention.
[0051] Figure 8This is the third example diagram of the interleaved graphics and text and the modification of the interleaved graphics and text provided by the embodiments of the present invention.
[0052] Figure 9 This is a schematic structural diagram of a generating device for interleaved graphics and text provided by the embodiments of the present invention.
[0053] Figure 10 This is a schematic physical structure diagram of an electronic device provided by the embodiments of the present invention. Detailed implementation manners
[0054] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below with reference to the accompanying drawings in the present invention. Obviously, the described embodiments are some but not all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present invention without making creative efforts shall fall within the protection scope of the present invention.
[0055] The following Figures 1 - 8 describes the method for generating interleaved graphics and text of the present invention. The execution subject of the method for generating interleaved graphics and text may be an electronic device such as a terminal, a tablet computer, a computer, etc., or a generating device for interleaved graphics and text provided in the electronic device. The generating device for interleaved graphics and text may be implemented by software, hardware or a combination of both.
[0056] Figure 1 This is one of the flow schematic diagrams of the method for generating interleaved graphics and text provided by the embodiments of the present invention. As Figure 1 shown, the method for generating interleaved graphics and text includes the following steps:
[0057] Step 101, obtain an interleaved graphics and text generation instruction, where the interleaved graphics and text generation instruction includes user requirement information.
[0058] Exemplarily, when a user wants to generate interleaved graphics and text, the user can input user requirement information to an application program loaded with the method for generating interleaved graphics and text, and click a generation control, so that an electronic device installed with the application program receives an interleaved graphics and text generation instruction including the user requirement information.
[0059] Step 102: Input the user requirement information into the interleaved graphic-text generation model to obtain the interleaved graphic-text output by the interleaved graphic-text generation model. Among them, the interleaved graphic-text generation model is obtained by training the initial interleaved graphic-text generation model based on the consistency loss, and the consistency loss is determined based on the sample similarity between the first feature of the sample intermediate feature map corresponding to the current first sample image token group and the second features of the previous sample intermediate feature maps. The current first sample image token group is obtained by inputting the interleaved sample text and the sample image into the initial multi-modal large model in the initial interleaved graphic-text generation model.
[0060] Exemplarily, when an interleaved graphic-text generation instruction is obtained, the interleaved graphic-text generation instruction is parsed to obtain user requirement information. The user requirement information is input into a pre-trained interleaved graphic-text generation model to obtain the interleaved graphic-text corresponding to the user requirement information output by the interleaved graphic-text generation model, and the generated interleaved graphic-text is presented to the user.
[0061] The method for generating interleaved graphic-text provided by the present invention inputs the user requirement information included in the interleaved graphic-text generation instruction into the trained interleaved graphic-text generation model to obtain the interleaved graphic-text output by the interleaved graphic-text generation model. Among them, the interleaved graphic-text generation model is obtained by training the initial interleaved graphic-text generation model based on the consistency loss, and the consistency loss is determined based on the sample similarity between the first feature of the sample intermediate feature map corresponding to the current first sample image token group and the second features of the previous sample intermediate feature maps. It can be seen that the present invention trains the interleaved graphic-text generation model based on the consistency loss determined by the sample similarity between the first feature of the sample intermediate feature map corresponding to the current first sample image token group and the second features of the previous sample intermediate feature maps, taking into account the similarity between the sample intermediate feature maps, so as to improve the consistency between the images in the interleaved graphic-text generated by the interleaved graphic-text generation model. This method for generating interleaved graphic-text can be applied to the generation or editing tasks of picture books, and can also be applied to the generation of instructions, recipes, teaching plans, etc. It is illustrated with pictures and texts and is closer to the daily generation of users, with a wide range of application scenarios.
[0062] In one embodiment, Figure 2 is the second flowchart of the method for generating interleaved graphic-text provided by the embodiment of the present invention. As Figure 2 shown, the above step 102 inputs the user requirement information into the interleaved graphic-text generation model to obtain the interleaved graphic-text output by the interleaved graphic-text generation model, which can be specifically implemented through the following steps:
[0063] Step 1021: Input the user requirement information into the multi-modal large model in the interleaved text and image generation model to obtain the current text token group and at least one current image token group output by the multi-modal large model.
[0064] Exemplarily, input the user requirement information into the text tokenizer of the interleaved text and image generation model to obtain each tokenized token output by the text tokenizer. Form a token sequence composed of the tokenized tokens and input it into the multi-modal large model. The multi-modal large model outputs a token through the self-attention module and the feed-forward module included in each of several groups of modules. Then, combine this token with the previous token sequence and input it into the multi-modal large model. The multi-modal large model outputs the next token, and so on, autoregressively generating until the end identifier is output, ending the output of tokens. Finally, obtain a new token sequence, which includes an interleaved text token group and image token groups. Each of the image token groups includes at least one group of image tokens. Traverse the new token sequence to obtain the current text token group and at least one current image token group.
[0065] Step 1022: For each of the current image token groups, determine the similarity between the first feature map corresponding to the current image token group and the second feature maps corresponding to the previous image token groups.
[0066] Exemplarily, for each current image token group, input the current image token group into the image decoder in the interleaved text and image generation model. The intermediate layer of the image decoder outputs the first feature map corresponding to the current image token group. Extract the features of the first feature map through the apparent feature extraction model in the interleaved text and image generation model, and calculate the similarity between the features of the first feature map and the features of the second feature maps corresponding to the previous image token groups. Here, the similarity can be cosine similarity or Pearson similarity, etc.
[0067] Step 1023: Based on each of the similarities, determine the target current image token group from all the current image token groups, generate the current target image based on the target current image token group, and generate the current target text based on the current text token group.
[0068] Exemplarily, when obtaining each similarity, all similarities are sorted in descending order, the top k similarities are taken, and the top k similarities are sent to the image decoder, so that the image decoder selects a target similarity from the top k similarities based on Top-k sampling, determines the current image token group corresponding to the target similarity as the target current image token group, and decodes the target current image token group to obtain the current target image; the current text token group is also input into the text decoder in the interleaved text-image generation model, and the text decoder decodes the current text token group to obtain the current target text.
[0069] Step 1024: Generate the interleaved text-image based on each of the target images and each of the target texts.
[0070] Exemplarily, when obtaining the target text corresponding to each text token group and the target image corresponding to each image token group, each target image and each target text are combined to generate an interleaved text-image.
[0071] In this embodiment, the target current image token group is determined from all current image token groups based on each similarity, the current target image is generated based on the target current image token group, and the current target text is generated based on the current text token group. Finally, the interleaved text-image is generated based on each target image and each target text, further considering the similarity between the first feature map corresponding to the current image token group and the second feature maps corresponding to the previous image token groups respectively, thereby further improving the consistency and coherence between the images in the generated interleaved text-image.
[0072] In one embodiment, in the above step 1021, inputting the user requirement information into the multi-modal large model in the interleaved text-image generation model to obtain the current text token group and at least one current image token group output by the multi-modal large model can be specifically implemented in the following manner:
[0073] Input the user requirement information into the multi-modal large model, and obtain the current text token group and the at least one current image token group output by the multi-modal large model through an autoregressive manner.
[0074] Exemplarily, the user requirement information is input into the text tokenizer of the interleaved text-image generation model to obtain each token output by the text tokenizer. The first token is input into the multi-modal large model to obtain the text token corresponding to the first token output by the multi-modal large model. Then, the text token corresponding to the first token and the second token are input into the multi-modal large model to obtain the text token corresponding to the second token. By repeating this process, the current text token group is obtained in an autoregressive manner, and at least one current image token group is generated based on the current text token group. After traversing all the tokens, multiple interleaved text token groups and image token groups can be obtained.
[0075] In this embodiment, the current text token group and at least one current image token group output by the multi-modal large model are obtained in an autoregressive manner, which improves the coherence between the text corresponding to the current text token group and the text corresponding to the previous text token group, and improves the coherence between the image corresponding to the current image token group and the image corresponding to the previous image token group.
[0076] In one embodiment, Figure 3 is a schematic flowchart of the training method of the interleaved text-image generation model provided by the embodiment of the present invention. As Figure 3 shown, the interleaved text-image generation model is trained based on the following method:
[0077] Step 301: Input the interleaved sample text and the sample image into the initial multi-modal large model to obtain at least one interleaved first sample text token group and first sample image token group output by the initial multi-modal large model.
[0078] Exemplarily, Figure 4 is a schematic framework diagram of the initial interleaved text-image generation model provided by the embodiment of the present invention. As Figure 4 shown, the initial interleaved text-image generation model includes an initial multi-modal large model, a text tokenizer, an image encoder, a text decoder, an image decoder, and an apparent feature extraction model. Among them, the initial multi-modal large model includes a self-attention module and a forward feedback module.
[0079] Collect a graphic and text data set, which includes web pages with interleaved graphics and text, picture book stories, instructions, recipes, etc., and preprocess the interleaved text and images. The preprocessing includes image size adjustment and normalization, text tokenization and encoding, etc. Finally, obtain the respective interleaved sample texts and sample images. For each sample text in the interleaved sample texts and sample images, input the sample text into a text tokenizer to obtain a second sample text token group corresponding to the sample text output by the text tokenizer. For each sample image in the interleaved sample texts and sample images, input the sample image into an image encoder to obtain a second sample image token group corresponding to the sample image output by the image encoder. Alternately combine the second sample text token group and the second sample image token group to obtain a sample token sequence. Input the sample token sequence into an initial multi-modal large model to obtain at least one interleaved first sample text token group and first sample image token group output by the initial multi-modal large model.
[0080] Step 302: Input the current first sample image token group into the image decoder in the initial interleaved graphic and text generation model to obtain the current sample intermediate feature map.
[0081] Exemplarily, traverse all interleaved first sample text token groups and first sample image token groups, input the current first sample image token group into the image decoder in the initial interleaved graphic and text generation model, and output the current sample intermediate feature map through the intermediate layer of the image decoder.
[0082] Step 303: Extract the first feature of the current sample intermediate feature map through the apparent feature extraction model in the initial interleaved graphic and text generation model, determine the sample similarity between the first feature and the second features of the previous sample intermediate feature maps, and determine the consistency loss based on each of the sample similarities.
[0083] Exemplarily, extract the features of the current sample intermediate feature map output by the intermediate layer of the image decoder through the apparent feature extraction model to obtain the first feature, calculate the sample similarity between the first feature and the second features of the previous sample intermediate feature maps, and then determine the consistency loss based on each sample similarity through the following formula (1) :
[0084]
[0085] where represents the similarity function, represents the image decoder, which is used to restore the discrete image tokens into an image, represents the th sample image token group. Indicates the th sample image token group, represents an apparent feature extraction model for extracting features from the current sample intermediate feature map output by the intermediate layer of the image decoder.
[0086] It should be noted that the present invention does not limit the specific structures of the apparent feature extraction model and the image decoder. For example, the apparent feature extraction model can be dinov2, and the image decoder can be a diffusion model.
[0087] Step 304: Determine a classification loss based on the first sample text token group and the first sample image token group.
[0088] Exemplarily, when obtaining the first sample text token group and the first sample image token group, determine the classification loss based on the difference between the first sample text token group and the second sample text token group of the corresponding sample text at the time of input, and the difference between the first sample image token group and the second sample image token group of the corresponding sample image at the time of input , and this classification loss can be a cross-entropy loss.
[0089] Step 305: Iteratively optimize the parameters of the initial interleaved text-image generation model based on the consistency loss and the classification loss to obtain the interleaved text-image generation model.
[0090] Exemplarily, when obtaining the consistency loss and the classification loss, determine the complete loss based on the following formula (2) , and based on the complete loss iteratively optimize the parameters of the initial interleaved text-image generation model to obtain the interleaved text-image generation model. Specifically, it can be to iteratively optimize the parameters of the self-attention module and the forward feedback module in the initial interleaved text-image generation model until the performance of the interleaved text-image generation model on the validation set reaches the expectation, and finally obtain the interleaved text-image generation model.
[0091]
[0092] Among them, is a loss coefficient used to modulate the balance between loss functions.
[0093] It should be noted that the performance of the interleaved text-image generation model can also be evaluated based on a test set to ensure the generalization ability of the interleaved text-image generation model, and the parameters of the finally trained interleaved text-image generation model are saved for subsequent inference use.
[0094] In this embodiment, an apparent feature extraction model is adopted, and the calculation of the sample similarity between the first feature and the second features of the intermediate feature maps of each previous sample is introduced. The consistency loss is determined based on each sample similarity, and the initial interleaved text-image generation model is constrained based on the consistency loss, so as to improve the consistency between the images in the interleaved text-images generated by the trained interleaved text-image generation model.
[0095] In one embodiment, step 301 inputs the interleaved sample text and the sample image into the initial multi-modal large model in the initial interleaved text-image generation model, and obtains at least one interleaved first sample text token group and first sample image token group output by the initial multi-modal large model. Specifically, it can be implemented through the following steps:
[0096] Based on the interleaved sample text and the sample image, a sample token sequence is determined. The sample token sequence includes an interleaved second sample text token group and second sample image token group; the interleaved second sample text token group and second sample image token group are input into the self-attention module in the first group of modules of the initial multi-modal large model, and the self-attention feature output by the self-attention module in an autoregressive manner is obtained; the self-attention feature is input into the forward feedback module in the first group of modules, and the hidden layer feature output by the forward feedback module in an autoregressive manner is obtained; the hidden layer feature is input into the second group of modules to obtain the hidden layer feature output by the second group of modules until the hidden layer feature output by the last group of modules is obtained; based on the hidden layer feature output by the last group of modules, the interleaved first sample text token group and the first sample image token group are obtained.
[0097] Exemplarily, for each sample text in the interleaved sample text and sample image, the sample text is input into a text tokenizer, and the second sample text token group corresponding to the sample text output by the text tokenizer is obtained. For each sample image in the interleaved sample text and sample image, the sample image is input into an image encoder, and the second sample image token group corresponding to the sample image output by the image encoder is obtained. The second sample text token group and the second sample image token group are alternately combined to obtain a sample token sequence 401 as shown in Figure 4 In the sample token sequence 401, the parts filled with slashes are all second sample text token groups, and the parts filled with white are all second sample image token groups. The sample token sequence is input into the self-attention module in the first group of modules of the initial multi-modal large model, and the self-attention feature output by the self-attention module in an autoregressive manner is obtained. The self-attention feature can be represented by the following formula (3):
[0098]
[0099] Among them, represents the self-attention feature, represents the query matrix, represents the key matrix, represents the value matrix, represents the dimension of K, which is used to scale the dot product result to prevent the dot product value from being too large and causing the softmax function to saturate, represents the memory matrix, represents the transpose of represents the transpose of
[0100] Input the self-attention feature into the forward feedback module in the first group of modules, and obtain the hidden layer feature output by the forward feedback module in an autoregressive manner. The hidden layer feature can be specifically represented by the following formula (4):
[0101]
[0102] Among them, represents the hidden layer feature corresponding to the current second sample text token group and the second sample image token group, represents the forward feedback module, which usually includes two linear layers and a non-linear activation function, represents the layer normalization operation, which is used to stabilize and accelerate the training process, represents the previous second sample text token group and the second sample image token group, represents the memory matrix corresponding to the previous second sample text token group and the second sample image token group, represents the current query matrix, represents the current key matrix, represents the current value matrix. The forward feedback module usually involves adjusting the model output to improve the accuracy and consistency of the generated content.
[0103] Input the hidden layer feature output by the first group of modules into the second group of modules to obtain the hidden layer feature output by the second group of modules. Then input the hidden layer feature output by the second group of modules into the third group of modules until the hidden layer feature output by the last group of modules is obtained. Then input the hidden layer feature output by the last group of modules into the output layer to obtain the interleaved first sample text token group and the first sample image token group output by the output layer. The interleaved first sample text token group and the first sample image token group are as Figure 4The shown sequence 402, the slanted-line filled parts in sequence 402 are all the first sample text token groups, and the first sample text token groups correspond to the second sample text token groups in the sample token sequence 401. The white-filled parts in sequence 402 are all the first sample image token groups, and the first sample image token groups correspond to the second sample image token groups in the sample token sequence 401.
[0104] In this embodiment, a combined representation of multiple self-attention modules and a forward feedback module is adopted to represent an adaptive alternating generation mechanism, enabling text and images to be naturally interspersed and alternated during the generation process. Different from traditional single outputs, this method flexibly switches between text and images according to user needs when generating content, achieving continuous and smooth interleaving of text and images. This approach not only improves the coherence of the content but also makes the information conveyance more vivid. Moreover, an autoregressive manner is adopted, and text and image generation are carried out in the same context, using the previous text and previous image information in each generation step, allowing the subsequently generated images and text to fully refer to the previous content, enabling the images and text to achieve higher consistency in terms of vision and semantics.
[0105] In one embodiment, inputting the interleaved second sample text token groups and second sample image token groups into the self-attention module in the first group of modules of the initial multi-modal large model to obtain the self-attention features output by the self-attention module in an autoregressive manner can be specifically implemented by the following method:
[0106] Input the interleaved second sample text token groups and the second sample image token groups into the self-attention module in the first group of modules to obtain the text intermediate hidden layer features corresponding to each of the second sample text token groups and the image intermediate hidden layer features corresponding to each of the second sample image token groups; determine the memory matrix based on each of the text intermediate hidden layer features, each of the image intermediate hidden layer features, the text update intensity, and the image update intensity in an autoregressive manner; determine the self-attention features based on the memory matrix.
[0107] Exemplarily, the self-attention module adopts a memory enhancement mechanism to ensure that the previous information is not lost by maintaining and updating the external memory matrix M, enhancing the ability of the self-attention module to process long sequences. Since there are two modalities, text and image, the memory matrix is updated in segments, and the update of the memory matrix can be specifically represented by the following formula (5):
[0108]
[0109] Where, Represents the memory matrix corresponding to the previous second sample text token group and the second sample image token group, Represents the memory matrix corresponding to the current second sample text token group and the second sample image token group, Represents the text intermediate hidden layer features and the image intermediate hidden layer features corresponding to the previous second sample text token group and the second sample image token group, Represents the text intermediate hidden layer features and the image intermediate hidden layer features corresponding to the current second sample text token group and the second sample image token group, When encoding text, it represents the text update intensity. When encoding an image, it represents the image update intensity. When updating the memory matrix, it pays attention to the type of generated tokens, and the memory matrix update weights of different types of tokens are different, that is, the text update intensity and the image update intensity are different; Represents the update memory function, including but not limited to moving average, GRU (Gated Recurrent Unit), attention mechanism, conditional update, non-linear activation function, etc.
[0110] In this embodiment, the initial multi-modal large model autoregressively generates interleaved first sample text token groups and first sample image token groups, and iteratively updates the memory matrix to improve the consistency between images in the subsequent interleaved text and images generated based on the multi-modal large model.
[0111] In one embodiment, after the above step 102, the method for generating the interleaved text and image further includes the following steps:
[0112] Receive a modification instruction for the interleaved text and image, where the modification instruction includes user requirement modification information; input the user requirement modification information, the user requirement information, and the interleaved text and image into the interleaved text and image generation model to obtain the modified interleaved text and image output by the interleaved text and image generation model.
[0113] Exemplarily, Figure 5 is the third flowchart of the method for generating interleaved text and image provided by the embodiment of the present invention, as Figure 5As shown, when the user wants to generate interlaced graphics, the user demand information can be input into the application loaded with the generation method of interlaced graphics, and the generation control is clicked, so that the electronic device installed with the application receives the interlaced graphics generation instruction including the user demand information, and the interlaced graphics generation instruction is parsed to obtain the user demand information, and the user demand information is input into the pre-trained interlaced graphics generation model to obtain the interlaced graphics corresponding to the user demand information output by the interlaced graphics generation model. After the interlaced graphics are generated, if the user receives the modification instruction of the interlaced graphics, the user demand modification information included in the modification instruction is input into the intent classifier, and the user's modification requirements are classified by the intent classifier to modify the interlaced graphics more accurately. When the user's modification requirement is to modify the content of the intermediate node, the process of modifying the intermediate node is entered. If the generation of the intermediate image does not meet expectations, the user demand modification information representing the modification of the intermediate image, the user demand information input for the first time, and the generated interlaced graphics are all input into the interlaced graphics generation model, and the interlaced graphics generation model regenerates the image of the current node. Figure 6 is one of the example diagrams of interlaced graphics and text and interlaced graphics and text modification provided by an embodiment of the present invention, such as Figure 6 As shown in the figure, the user's first input user demand information is "generate a recipe for scrambled eggs with tomatoes", and the first generated interlaced graphics and texts include Figure 6 The five generation steps shown on the left, when generating a recipe, the ingredients, steps described in the text and the generated images can be accurately matched, achieving true text and picture combination. Each generation step represents a node, and each generation step includes text and images. If the user's requirement modification information is "the image of generation step 1 is not clear enough, the image of generation step 5 is better", the modified interlaced text and images include Figure 6 The five generation steps shown on the right modify the image in relation to the user's desired information.
[0114] If subsequent nodes need to be adjusted, the user demand modification information representing the need for adjustment of subsequent nodes, the user demand information input for the first time, and the generated interlaced graphics and text are all input into the interlaced graphics and text generation model. The interlaced graphics and text generation model modifies the subsequent text and images according to the user demand modification information based on the text and image before the current node. Figure 7 FIG. 2 is a second example of interlaced graphics and text and interlaced graphics and text modification provided by an embodiment of the present invention. Figure 7 As shown in the figure, the user's first input user demand information is "generate a recipe for scrambled eggs with tomatoes", and the first generated interlaced graphics and texts include Figure 7 The five generation steps shown on the left each represent a node, and each generation step includes text and images. If the user needs to modify the information to "no salt, add some pepper", the modified interlaced text and images include Figure 7The five generation steps shown on the right modify the text and images related to the user demand information.
[0115] If it is necessary to refine a specific node, the user demand modification information representing the refinement of a specific node, the user demand information input for the first time, and the generated interleaved text and images are all input into the interleaved text and image generation model. The interleaved text and image generation model supplements and generates refined text and images according to the text and images before the current node in accordance with the user demand modification information. Figure 8 It is the third example diagram of the interleaved text and images and the modification of the interleaved text and images provided by the embodiment of the present invention. As Figure 8 shown, the user demand information input for the first time is "generate the instruction manual of an intelligent coffee machine", and the interleaved text and images generated for the first time include Figure 8 the seven parts shown on the left, each part represents a node, and each part includes text and images. The user demand modification information is "I want to know the specific process of brewing coffee", then the modified interleaved text and images include Figure 8 the five parts about the specific process of brewing coffee shown on the right.
[0116] Ask the user if there are any new demands. If the user still has new demands, then make corresponding adjustments to the interleaved text and images again according to the new demands. If the user has no new demands, then end the current text and image generation task and output the final interleaved text and images.
[0117] In this embodiment, in order to increase the controllability and flexibility of the interleaved text and image generation model, the interleaved text and image generation model can modify the interleaved text and images based on the user demand modification information input by the user, realizing the editing function of the interleaved text and images, making the finally generated interleaved text and images more in line with the personalized needs of the user.
[0118] Next, the generation device of the interleaved text and images provided by the present invention will be described. The generation device of the interleaved text and images described below can be correspondingly referred to the generation method of the interleaved text and images described above.
[0119] Figure 9 It is the structural schematic diagram of the generation device of the interleaved text and images provided by the embodiment of the present invention. As Figure 9 shown, the generation device 900 of the interleaved text and images includes an acquisition unit 901 and a first generation unit 902; wherein:
[0120] The acquisition unit 901 is used to acquire an interleaved text and image generation instruction, and the interleaved text and image generation instruction includes user demand information;
[0121] The first generation unit 902 is used to input the user demand information into the interleaved text and image generation model to obtain the interleaved text and images output by the interleaved text and image generation model;
[0122] Among them, the interleaved graphic and text generation model is obtained by training an initial interleaved graphic and text generation model based on a consistency loss. The consistency loss is determined based on the sample similarity between the first feature of the sample intermediate feature map corresponding to the current first sample image token group and the second features of the previous sample intermediate feature maps. The current first sample image token group is obtained by inputting the interleaved sample text and the sample image into the initial multimodal large model in the initial interleaved graphic and text generation model.
[0123] The interleaved graphic and text generation device provided by the present invention inputs the user demand information included in the interleaved graphic and text generation instruction into the trained interleaved graphic and text generation model, and obtains the interleaved graphic and text output by the interleaved graphic and text generation model. Among them, the interleaved graphic and text generation model is obtained by training an initial interleaved graphic and text generation model based on a consistency loss. The consistency loss is determined based on the sample similarity between the first feature of the sample intermediate feature map corresponding to the current first sample image token group and the second features of the previous sample intermediate feature maps. It can be seen that the present invention determines the consistency loss based on the sample similarity between the first feature of the sample intermediate feature map corresponding to the current first sample image token group and the second features of the previous sample intermediate feature maps to train the interleaved graphic and text generation model, taking into account the similarity between the sample intermediate feature maps, thereby being able to improve the consistency between the images in the interleaved graphic and text generated by the interleaved graphic and text generation model.
[0124] Based on any of the above embodiments, the first generation unit 902 is specifically configured to:
[0125] Input the user demand information into the multimodal large model in the interleaved graphic and text generation model, and obtain the current text token group and at least one current image token group output by the multimodal large model;
[0126] For each of the current image token groups, determine the similarity between the first feature map corresponding to the current image token group and the second feature maps corresponding to the previous image token groups;
[0127] Based on each of the similarities, determine a target current image token group from all the current image token groups, generate a current target image based on the target current image token group, and generate a current target text based on the current text token group;
[0128] Generate the interleaved graphic and text based on each of the target images and each of the target texts.
[0129] Based on any of the above embodiments, the first generation unit 902 is further specifically configured to:
[0130] Input the user demand information into the multi-modal large model, and obtain the current text token group and the at least one current image token group output by the multi-modal large model through an autoregressive manner.
[0131] Based on any of the above embodiments, the interleaved text and image generation model is trained in the following manner:
[0132] Input the interleaved sample text and the sample image into the initial multi-modal large model, and obtain at least one interleaved first sample text token group and first sample image token group output by the initial multi-modal large model;
[0133] Input the current first sample image token group into the image decoder in the initial interleaved text and image generation model to obtain the current sample intermediate feature map;
[0134] Extract the first feature of the current sample intermediate feature map through the apparent feature extraction model in the initial interleaved text and image generation model, determine the sample similarity between the first feature and the second features of the previous sample intermediate feature maps, and determine the consistency loss based on each of the sample similarities;
[0135] Determine the classification loss based on the first sample text token group and the first sample image token group;
[0136] Iteratively optimize the parameters of the initial interleaved text and image generation model based on the consistency loss and the classification loss to obtain the interleaved text and image generation model.
[0137] Based on any of the above embodiments, inputting the interleaved sample text and the sample image into the initial multi-modal large model in the initial interleaved text and image generation model to obtain at least one interleaved first sample text token group and first sample image token group output by the initial multi-modal large model includes:
[0138] Determine a sample token sequence based on the interleaved sample text and the sample image, where the sample token sequence includes an interleaved second sample text token group and second sample image token group;
[0139] Input the interleaved second sample text token group and second sample image token group into the self-attention module in the first group of modules of the initial multi-modal large model to obtain the self-attention feature output by the self-attention module through an autoregressive manner;
[0140] Input the self-attention feature into the forward feedback module in the first group of modules to obtain the hidden layer feature output by the forward feedback module through an autoregressive manner;
[0141] Input the hidden layer features into the second set of modules to obtain the hidden layer features output by the second set of modules, until the hidden layer features output by the last set of modules are obtained;
[0142] Based on the hidden layer features output by the last set of modules, obtain the interleaved first sample text token group and the first sample image token group.
[0143] Based on any of the above embodiments, inputting the interleaved second sample text token group and the second sample image token group into the self-attention module of the first set of modules of the initial multi-modal large model to obtain the self-attention features output by the self-attention module in an autoregressive manner includes:
[0144] Input the interleaved second sample text token group and the second sample image token group into the self-attention module of the first set of modules to obtain the text intermediate hidden layer features corresponding to each second sample text token group and the image intermediate hidden layer features corresponding to each second sample image token group;
[0145] Determine the memory matrix in an autoregressive manner based on each of the text intermediate hidden layer features, each image intermediate hidden layer feature, the text update intensity, and the image update intensity;
[0146] Determine the self-attention features based on the memory matrix.
[0147] Based on any of the above embodiments, the generation device 900 for interleaved text and images further includes:
[0148] A receiving unit, configured to receive a modification instruction for the interleaved text and images, where the modification instruction includes user requirement modification information;
[0149] A second generation unit, configured to input the user requirement modification information, the user requirement information, and the interleaved text and images into the interleaved text and image generation model to obtain the modified interleaved text and images output by the interleaved text and image generation model.
[0150] Figure 10 It is a schematic diagram of the physical structure of the electronic device provided by the embodiments of the present invention, as Figure 10As shown in the figure, the electronic device may include: a processor 1010, a communications interface 1020, a memory 1030, and a communication bus 1040. Among them, the processor 1010, the communications interface 1020, and the memory 1030 communicate with each other through the communication bus 1040. The processor 1010 can call the logical instructions in the memory 1030 to execute the method for generating interleaved graphics and texts, and the method includes: obtaining an interleaved graphics and text generation instruction, where the interleaved graphics and text generation instruction includes user requirement information;
[0151] Input the user requirement information into an interleaved graphics and text generation model to obtain the interleaved graphics and texts output by the interleaved graphics and text generation model;
[0152] Among them, the interleaved graphics and text generation model is obtained by training an initial interleaved graphics and text generation model based on a consistency loss. The consistency loss is determined based on the sample similarity between the first feature of the sample intermediate feature map corresponding to the current first sample image token group and the second features of the previous sample intermediate feature maps. The current first sample image token group is obtained by inputting the interleaved sample texts and sample images into the initial multi-modal large model in the initial interleaved graphics and text generation model.
[0153] In addition, when the logical instructions in the above-mentioned memory 1030 can be implemented in the form of software functional units and sold or used as an independent product, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk, or an optical disc that can store program codes.
[0154] On the other hand, the present invention also provides a computer program product. The computer program product includes a computer program. The computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the method for generating interleaved graphics and texts provided by the above-mentioned various methods. The method includes: obtaining an interleaved graphics and text generation instruction, where the interleaved graphics and text generation instruction includes user requirement information;
[0155] Input the user requirement information into the interleaved graphic and text generation model to obtain the interleaved graphics and text output by the interleaved graphic and text generation model;
[0156] Among them, the interleaved graphic and text generation model is obtained by training an initial interleaved graphic and text generation model based on a consistency loss, and the consistency loss is determined based on the sample similarity between the first feature of the sample intermediate feature map corresponding to the current first sample image token group and the second features of the previous sample intermediate feature maps. The current first sample image token group is obtained by inputting the interleaved sample text and sample image into the initial multi-modal large model in the initial interleaved graphic and text generation model.
[0157] In another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it is configured to execute the method for generating interleaved graphics and text provided by the above-mentioned various methods. The method includes: obtaining an interleaved graphic and text generation instruction, where the interleaved graphic and text generation instruction includes user requirement information;
[0158] Input the user requirement information into the interleaved graphic and text generation model to obtain the interleaved graphics and text output by the interleaved graphic and text generation model;
[0159] Among them, the interleaved graphic and text generation model is obtained by training an initial interleaved graphic and text generation model based on a consistency loss, and the consistency loss is determined based on the sample similarity between the first feature of the sample intermediate feature map corresponding to the current first sample image token group and the second features of the previous sample intermediate feature maps. The current first sample image token group is obtained by inputting the interleaved sample text and sample image into the initial multi-modal large model in the initial interleaved graphic and text generation model.
[0160] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative efforts.
[0161] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0162] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for generating interleaved graphics and texts, characterized in that, Including: Obtain an interleaved graphic text generation instruction, where the interleaved graphic text generation instruction includes user requirement information; Input the user requirement information into an interleaved graphic text generation model to obtain interleaved graphic text output by the interleaved graphic text generation model; Among them, the interleaved graphic text generation model is obtained by training an initial interleaved graphic text generation model based on a consistency loss, and the consistency loss is determined based on the sample similarity between the first feature of the sample intermediate feature map corresponding to the current first sample image token group and the second features of each previous sample intermediate feature map. The current first sample image token group is obtained by inputting the interleaved sample text and sample image into the initial multimodal large model in the initial interleaved graphic text generation model; The step of inputting the user requirement information into the interleaved graphic text generation model to obtain the interleaved graphic text output by the interleaved graphic text generation model includes: Input the user requirement information into the multimodal large model in the interleaved graphic text generation model to obtain the current text token group and at least one current image token group output by the multimodal large model; For each of the current image token groups, determine the similarity between the first feature map corresponding to the current image token group and the second feature maps corresponding to the previous image token groups; Arrange all the similarities in descending order, take the top k similarities, use the top k similarities as target similarities, and determine the current image token groups corresponding to the target similarities as target current image token groups; Generate a current target image based on the target current image token groups, and generate a current target text based on the current text token group; Generate the interleaved graphic text based on each of the target images and each of the target texts.
2. The method for generating interleaved graphics and texts according to claim 1, wherein The step of inputting the user requirement information into the multimodal large model in the interleaved graphic text generation model to obtain the current text token group and at least one current image token group output by the multimodal large model includes: Input the user requirement information into the multimodal large model, and obtain the current text token group and the at least one current image token group output by the multimodal large model through an autoregressive manner.
3. The method for generating interleaved graphics and texts according to claim 1, wherein The interleaved graphic text generation model is trained based on the following method: Input the interleaved sample text and sample image into the initial multimodal large model to obtain at least one interleaved first sample text token group and first sample image token group output by the initial multimodal large model; Input the current first sample image token group into the image decoder in the initial interleaved graphic text generation model to obtain the current sample intermediate feature map; Extract the first feature of the current sample intermediate feature map through the apparent feature extraction model in the initial interleaved graphic text generation model, determine the sample similarity between the first feature and the second features of each previous sample intermediate feature map, and determine the consistency loss based on each of the sample similarities; Determine a classification loss based on the first sample text token group and the first sample image token group; Iteratively optimize the parameters of the initial interleaved text-image generation model based on the consistency loss and the classification loss to obtain the interleaved text-image generation model.
4. The method for generating interleaved graphics and texts according to claim 3, wherein The step of inputting the interleaved sample text and the sample image into the initial multi-modal large model in the initial interleaved text-image generation model to obtain at least one interleaved first sample text token group and first sample image token group output by the initial multi-modal large model includes: Determine a sample token sequence based on the interleaved sample text and the sample image, where the sample token sequence includes an interleaved second sample text token group and second sample image token group; Input the interleaved second sample text token group and second sample image token group into the self-attention module in the first group of modules of the initial multi-modal large model to obtain self-attention features output by the self-attention module in an autoregressive manner; Input the self-attention features into the feed-forward module in the first group of modules to obtain hidden layer features output by the feed-forward module in an autoregressive manner; Input the hidden layer features into the second group of modules to obtain hidden layer features output by the second group of modules until hidden layer features output by the last group of modules are obtained; Based on the hidden layer features output by the last group of modules, obtain the interleaved first sample text token group and first sample image token group.
5. The method for generating interleaved graphics and texts according to claim 4, wherein The step of inputting the interleaved second sample text token group and second sample image token group into the self-attention module in the first group of modules of the initial multi-modal large model to obtain self-attention features output by the self-attention module in an autoregressive manner includes: Input the interleaved second sample text token group and second sample image token group into the self-attention module in the first group of modules to obtain text intermediate hidden layer features corresponding to each second sample text token group and image intermediate hidden layer features corresponding to each second sample image token group; Determine a memory matrix in an autoregressive manner based on each text intermediate hidden layer feature, each image intermediate hidden layer feature, a text update intensity, and an image update intensity; Determine the self-attention features based on the memory matrix.
6. The method for generating interleaved graphics and texts according to any one of claims 1-5, characterized in that The method further includes: Receiving a modification instruction for the interleaved text-image, where the modification instruction includes user requirement modification information; Inputting the user requirement modification information, the user requirement information, and the interleaved text-image into the interleaved text-image generation model to obtain a modified interleaved text-image output by the interleaved text-image generation model.
7. An apparatus for generating interleaved graphics and texts, characterized in that, including: An acquisition unit for acquiring an interleaved text-image generation instruction, where the interleaved text-image generation instruction includes user requirement information; A first generation unit for inputting the user requirement information into an interleaved text-image generation model to obtain an interleaved text-image output by the interleaved text-image generation model; Among them, the interleaved graphic and text generation model is obtained by training an initial interleaved graphic and text generation model based on a consistency loss, and the consistency loss is determined based on the sample similarity between the first feature of the sample intermediate feature map corresponding to the current first sample image token group and the second features of the previous sample intermediate feature maps. The current first sample image token group is obtained by inputting the interleaved sample text and sample image into the initial multi-modal large model in the initial interleaved graphic and text generation model; The first generation unit is specifically configured to: Input the user requirement information into the multi-modal large model in the interleaved graphic and text generation model to obtain the current text token group and at least one current image token group output by the multi-modal large model; For each of the current image token groups, determine the similarity between the first feature map corresponding to the current image token group and the second feature maps corresponding to the previous image token groups respectively; Arrange all the similarities in descending order, take the top k similarities, use the top k similarities as the target similarities, and determine the current image token groups corresponding to the target similarities as the target current image token groups; Generate the current target image based on the target current image token group, and generate the current target text based on the current text token group; Generate the interleaved graphic and text based on each of the target images and each of the target texts.
8. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the method for generating the interleaved graphic and text according to any one of claims 1 to 6.
9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the method for generating the interleaved graphic and text according to any one of claims 1 to 6.
Citation Information
Patent Citations
Multi-modal image-text interlaced generation model based on dynamic characteristic synchronizer
CN118364433A