Artificial intelligence text creation method based on multimodal information input

By constructing multimodal data samples and designing text creation models, the problem of combining multimodal information in the AI ​​creation field is solved, and the text generation of multimodal information in the AI ​​creation field is realized, which is more in line with the process of human creation.

CN115309886BActive Publication Date: 2025-05-13RENMIN UNIVERSITY OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202210932040.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-04
Publication Date
2025-05-13
Estimated Expiration
2042-08-04

AI Technical Summary

Technical Problem

The prior art is difficult to use multimodal information of images and text to generate text simultaneously in the field of AI creation, and there is a lack of methods to simulate human creative processes.

Method used

Multimodal data construction method is adopted to construct multimodal data samples by crawling lyrics and combining them with graphic data sets. Design a text creation model, including a multi-channel sequence processor, customized module captures the impact of input on output, attention mechanism between modes and decoder, to achieve the fusion of images and text.

Benefits of technology

It realizes that text related to input images and text is generated through multimodal information under a given topic, expanding the multimodal input capability of AI creation, which is more in line with the process of human creation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115309886B_ABST
    Figure CN115309886B_ABST
Patent Text Reader

Abstract

The present invention discloses an artificial intelligence text creation method based on multimodal information input, which includes two parts: multimodal data construction and text creation model. The present invention can simultaneously process multimodal image and text sequence information as input, generate text under the condition of given keywords, and expand the work of AI creation field from single modality to text generation to multiple modalities to text generation, which is more in line with the process of human creation. In addition, the model structure and training method of the present invention are more reasonable in terms of method, and the experimental results are reliable. At the same time, the effectiveness of the method of the present invention is also confirmed, and the method is also easier to expand, migrate and re-create in the future.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of artificial intelligence, deep learning, and natural language generation technology, and in particular to an artificial intelligence text creation method based on multimodal information input. Background Art

[0002] Lyric generation and poetry writing are two typical AI creative tasks, where the generated text needs to follow some format and rhythm. Early lyric generation works were mostly based on constraint or retrieval-based methods, trying to generate by matching the best relevant next sentence with the previous one. Later studies used neural networks such as long short-term memory (LSTM) or autoencoders to handle this task, or added hierarchical attention mechanisms in the decoder. Recently, pre-trained language models can provide better conditional-based results and consider more rhythm and rhythm. In the task of poetry generation, early models focused on keyword expansion and modeling the poet's intention until the emergence of large pre-trained language models like GPT became a milestone. In addition to text information, other works have also tried to inspire poetry generation with images. These studies use visual input to simulate the human scene perception process. Basically, these methods generate poetry from a single image input. The existing Images2Poem generates Chinese classical poetry from an image stream by selecting representative images from the image stream and decoding them with an adaptive self-attention mechanism, which is similar to the work of this application.

[0003] Another related field is multimodal summarization technology that generates text summaries by adopting multimodal data. However, the generated summary is highly dependent on the source text, which is different from the multimodal creation task limited by the subject of this application. Other related tasks are visual narratives, which take multiple consecutive images as input and aim to generate a coherent story. To solve this problem, many works use CNN to encode image streams and use RNN-like modules to generate story sentences, or use hierarchical structures as well as some specially designed attention mechanisms. There are also other works that give the model the ability to adapt to the topic or combine videos for visual narratives.

[0004] Although the above AI creation-related works are either based on text or image-based text generation, none of them simultaneously use the multimodal information of images and text and combine them with keywords as input or conditions for creation. Although there are many promising results in the work of writing poems based on images, most of them identify keywords from images, such as objects or emotions in pictures, and use keywords as input to influence the poetry generation process. At the same time, the Images2Poem method that only inputs multiple pictures is similar to the work of this application, but the constructed images (about 20 images per poem) are mainly objects mentioned in a poem, which is very different from the model that attempts to capture sequential semantics from a series of images and their corresponding texts in this application. In order to simulate the embodied experience of humans in the creative process, and not all experiences (such as feelings) can be well visualized and represented, this application constructs a specific data set to adapt to the settings and tasks of this application. The goal of this application is to simulate the embodied experience of humans under a given theme, given multiple sets of image-text pairs with sequential relationships, and to generate text that is closely related to the input image and the corresponding text, so as to fill the gap in the field of artificial intelligence creation to adapt to various multimodal inputs for text generation.

[0005] For multimodal summary generation and visual narrative tasks, although there are works based on multimodal information generation, few works use both topics and paired image-text inputs for freer text creation like the setting of this application, which is a more realistic simulation of human past experiences.

[0006] The information disclosed in this background technology section is only intended to deepen the understanding of the overall background technology of the present invention, and should not be regarded as acknowledging or suggesting in any form that the information constitutes the prior art already known to those skilled in the art. Summary of the invention

[0007] The purpose of the present invention is to provide an artificial intelligence text creation method based on multimodal information input to solve the problems existing in the prior art.

[0008] In order to achieve the above object, the present invention adopts the following technical solutions:

[0009] The present invention provides an artificial intelligence text creation method based on multimodal information input, wherein the text creation method includes two parts: multimodal data construction and text creation model; wherein:

[0010] The specific method of constructing the multimodal data is as follows: first, a large number of lyrics are crawled from the Internet, and they are split into different paragraphs in accordance with a specific pattern, and these paragraphs are composed of different sentences; the song title is used as the theme information required in the task, and each sentence is used as the key information of the query on a large-scale movie synopsis image and text dataset GraphMovie. The CADM model is used to retrieve multiple image-text pair candidates; a part of the image-text pair candidates is manually annotated and a refined ranking model is trained with this part containing the annotated information to improve the quality of the image-text pair candidates; at the same time, the ranking information of different relevance will help to construct positive and negative samples of different qualities for subsequent model training; thus, for each lyric paragraph, a sequence of image-text pair candidates with different relevance qualities can be obtained to constitute a data sample, thereby forming a data set under a specific task;

[0011] The text creation model consists of four parts; the first three parts constitute the encoder, specifically, the original image and text are first processed by a multi-channel sequence processor to generate their semantic embeddings; then, the embeddings at each step are divided into different parts to affect the final output; finally, different modalities are fused together with an attention network; the last part is the decoder, which aims to predict the final output sentence.

[0012] As a further technical solution, the first part of the text creation model is specifically as follows: the formats and semantics of the original images and texts are presented in different spaces; in order to adapt to them, a multi-channel sequence processor is designed, and different modal sequences are first mapped to the same high-dimensional space through the multimodal pre-training model WenLan, and then input into these encoder neural networks; these encoders can be recurrent neural networks or Transformers, and ultimately the specific modules to be adopted can be selected by weighing effectiveness and efficiency; the output is an implicit embedding sequence; both the input image and text sequences are processed in this way.

[0013] As a further technical solution, the second part of the text creation model is specifically as follows: the text creation model is a sequence-to-sequence architecture; however, unlike traditional machine translation tasks, each input word strictly corresponds to an output word. In the problem of this application, images or texts may affect the span of the output sequence; in order to model these restrictive features, a customized module is designed to capture the impact of input on output; specifically, the hidden embeddings derived in the previous section are allowed to specifically affect the output sequence; for these hidden embeddings, an inter-modal attention mechanism is designed within each channel to capture the impact of different time steps on other time steps, so as to obtain the inclusion A comprehensive hidden embedding of a certain time step that is different from the information of other time steps; In order to encode intuition into the customized module, a regularizer is further introduced to constrain the learning of attention weights; Formally, the distance between the attention weights and a predefined distribution is minimized, thereby defining a KL loss function between the two for optimization and learning; By minimizing the KL loss, the attention weights are regularized using a priori, which encodes the intuition that a larger input-output distance should lead to a lower impact, thereby allowing the model to have good sensitivity to the order of the inputs; Prior knowledge about the distribution of attention weights is used to narrow the exploration space to bring better convergence rate and optimization solution.

[0014] As a further technical solution, the third part of the text creation model is specifically: based on the partial hidden embeddings outputted above, different modalities are fused to derive the output of the encoder; specifically, the output of the encoder consists of L embeddings, each of which comprehensively encodes the topic, visual and textual information; the total output embedding is calculated by iterating the influence of hidden embeddings from different steps on the kth step; for each pair of steps, different modalities are weighted and combined together in a specific attention manner; intuitively, for the same output sentence, different modalities may play different roles; therefore, an inter-modal attention mechanism is adopted when combining them; if the above two attention mechanisms are compared, it may be found that the former is deployed in different steps of the same modality, while the latter aims to capture the contributions of different modalities in the same step; such a design actually forms a 2D attention mechanism, thereby modeling the influence of different positions and modalities in a more fine-grained manner.

[0015] As a further technical solution, the fourth part of the text creation model is specifically as follows: for the embedding generation output based on the output of the above module, different embeddings are merged as prompts, and all generated sentences are directly summarized and output; however, this strategy may not be optimal for retaining the sequential semantics of the input, because the ordered information may be weakened by the merging operation; in order to solve the above problem, each empirical embedding is allowed to affect the output sentence separately; formally, at each step, the output embedding of that step and the word embedding are added, and the subject word is used as a prompt, and then the whole is input into the decoder for generation; this method can maximize the retention of the influence of different time steps on different parts of the generated sentence.

[0016] As a further technical solution, in order to maximize the probability of generating the target output from the positive sample input and minimize the probability of generating the target output from the negative sample input, the text creation model is trained through curriculum learning. The specific training method is: first learn the most negative sample to better initialize the model optimization; once the model has learned enough patterns to handle the most negative patterns, more difficult samples are gradually introduced near the positive and negative boundaries; more specifically, evaluate the correlation between the input image / text and the output, and construct 5 levels of samples; Level-5 represents the most relevant input, and Level-1 represents the least relevant input and output; during the training process, first train the model with Level-5 and Level-1 samples, and then include Level-4 and Level-2 in the positive and negative sample sets respectively, and guide the learning of the model in a gradually increasing manner from easy to difficult.

[0017] By adopting the above technical solution, the present invention has the following beneficial effects:

[0018] 1. It can simultaneously process multi-modal image and text sequence information as input, generate text under the condition of given keywords, and expand the work of AI creation from single modality to text generation to multiple modalities to be more in line with the human creation process.

[0019] 2. In terms of method, the model structure and training method of the present invention are more reasonable, the experimental results are reliable, and the effectiveness of the method of the present invention is also confirmed. The method is also easier to expand, migrate and re-create in the future. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] In order to more clearly illustrate the specific implementation methods of the present invention or the technical solutions in the prior art, the drawings required for use in the specific implementation methods or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some implementation methods of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0021] Figure 1 It is a schematic diagram of a model of the present invention;

[0022] Figure 2 A schematic diagram of a specific example of multimodal data construction provided by an embodiment of the present invention;

[0023] Figure 3 A schematic diagram of a specific example of generating text by the text creation model provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0024] The technical solution of the present invention will be described clearly and completely below in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0025] The specific implementation of the present invention is described in detail below in conjunction with the accompanying drawings. It should be understood that the specific implementation described here is only used to illustrate and explain the present invention, and is not used to limit the present invention.

[0026] The present embodiment provides an artificial intelligence text creation method based on multimodal information input. The method of the present application first includes the construction of specific data. Since there is no publicly available data set for the task of the present application, the present application first crawls a large number of lyrics from the Internet and splits them into different paragraphs in accordance with a specific pattern. These paragraphs are composed of different sentences. The present application uses the song title as the subject information needed in the task of the present application. The present application uses the CADM model on a large-scale image and text data set GraphMovie of movie synopsis, retrieves each sentence as the key information of the query and obtains multiple image and text pair candidates. The present application manually annotates a part of these image and text pair candidates and trains a fine ranking model with this part containing the annotation information to improve the quality of the image and text pair candidates. At the same time, the ranking information of different relevance will help the present application to construct positive and negative samples of different qualities for the training of subsequent models. Thus, for each lyric paragraph, the present application can obtain image and text pair candidate sequences of different relevance qualities to constitute the data sample of the present application, thereby forming a data set under the specific task of the present application.

[0027] The model architecture of this application is as follows Figure 1 As shown, specifically, the original image and text are first processed by a multi-channel sequence processor to generate their semantic embeddings. Then, the embeddings of each step are divided into different parts to affect the final output. Finally, the different modalities are fused together with the attention network. The last part is the decoder, which aims to predict the final output sentence. Below, this application will explain the model of this application in more detail.

[0028] Model part 1: The formats and semantics of the original images and texts are presented in different spaces. In order to adapt to them, this application designs a multi-channel sequence processor, which first maps different modal sequences to the same high-dimensional space through the multimodal pre-trained model WenLan, and then inputs them into these encoder neural networks. These encoders can be recurrent neural networks or Transformers, and this application can ultimately select the specific modules to be adopted by weighing effectiveness and efficiency. The output is an implicit embedding sequence. Both the input image and text sequences are processed in this way.

[0029] Model Part 2: Roughly speaking, the model of this application is a sequence-to-sequence architecture. However, unlike traditional tasks such as machine translation, where each input word usually strictly corresponds to one output word, in the problem of this application, images or text may affect the span of the output sequence. In order to model these limiting features, this application designs a customized module to capture the impact of input on output. Specifically, this application allows the hidden embedding derived in the previous section to specifically affect the output sequence. For these hidden embeddings, this application designs an inter-modal attention mechanism in each channel to capture the impact of different time steps on other time steps, so as to obtain a comprehensive hidden embedding of a time step containing information from different other time steps. However, this application believes that the impact of input on output should also follow some intuitive patterns. For example, if the distance between the input and output time steps is large, then the impact should be small. In order to encode these intuitions into the model of this application, this application further introduces a regularizer to constrain the learning of attention weights. Formally, this application minimizes the distance between the attention weight and a predefined distribution, thereby defining a KL loss function between the two for optimization and learning. By minimizing the KL loss, this application uses a priori regularization of the attention weights, which encodes the intuition that larger input-output distances should lead to lower influence, allowing the model to be more sensitive to the order of the inputs. Using prior knowledge about the distribution of attention weights to narrow the exploration space can lead to better convergence rates and optimized solutions.

[0030] Model part three: Based on the partial hidden embeddings output above, this application fuses different modalities to derive the output of the encoder. Specifically, the output of the encoder consists of L embeddings, each of which comprehensively encodes the topic, visual, and textual information. The total output embedding is calculated by iterating the influence of hidden embeddings from different steps on the kth step. For each pair of steps, different modalities are weighted and combined together in a specific attention manner. Intuitively, different modalities may play different roles for the same output sentence. Therefore, this application adopts an inter-modal attention mechanism when combining them. If you compare the above two attention mechanisms, you may find that the former is deployed in different steps of the same modality, while the latter aims to capture the contributions of different modalities in the same step. Such a design actually forms a 2D attention mechanism, which is expected to model the influence of different positions and modalities in a more fine-grained manner.

[0031] Model Part 4: For the embedding generation output based on the output of the above module, it is straightforward to merge different embeddings as prompts and directly summarize and output all generated sentences. However, this strategy may not be optimal for retaining the sequential semantics of the input, because the ordered information may be weakened by the merging operation. In order to solve the above problems, the present application allows each empirical embedding to affect the output sentence separately. Formally, the present application adds the output embedding of that step to the word embedding at each step, uses the subject word as a prompt, and then inputs the whole into the decoder for generation. This approach can maximize the retention of the impact of different time steps on different parts of the generated sentence.

[0032] Following the strategies of some previous works, this application maximizes the probability of generating the target output from the positive sample input while minimizing the probability of generating the target output from the negative sample input. In the task of this application, the input is a sequence. As the sequence becomes longer, the negative sample space expands exponentially, and it is impossible to select all negative samples. In order to better learn the model of this application, this application selects negative samples in the form of course learning. The overall idea of ​​this application is to first learn the most negative samples in order to better initialize model optimization. Once the model has learned enough patterns to handle the most negative patterns, this application will gradually introduce more difficult samples near the positive and negative boundaries. More specifically, this application evaluates the correlation between the input image / text and the output and constructs 5 levels of samples. Level-5 represents the most relevant input, and Level-1 represents the least relevant input and output. During the training process, this application first trains the model with Level-5 and Level-1 samples, and then includes Level-4 and Level-2 in the positive and negative sample sets respectively, guiding the learning of the model in a gradually increasing manner from easy to difficult.

[0033] In order to further illustrate the present invention in more detail, Figure 2 and Figure 3 The process of generating text by the text creation model of the present invention and the process of generating text by the text creation model are respectively provided. Figure 2 and Figure 3 It can be seen that the present invention can simultaneously process multimodal image and text sequence information as input, generate text under the condition of given keywords, and expand the work of AI creation field from single modality to text generation to multiple modalities to text generation, which is more in line with the human creation process.

[0034] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. An artificial intelligence text creation method based on multimodal information input, characterized in that: The text creation method includes two parts: multimodal data construction and text creation model; wherein, The specific method of constructing the multimodal data is as follows: first, a large number of lyrics are crawled from the Internet, and they are split into different paragraphs in accordance with a specific pattern, and these paragraphs are composed of different sentences; the song title is used as the theme information required in the task, and each sentence is used as the key information of the query on a large-scale movie synopsis image and text dataset GraphMovie. The CADM model is used to retrieve multiple image-text pair candidates; a part of the image-text pair candidates is manually annotated and a refined ranking model is trained with this part containing the annotated information to improve the quality of the image-text pair candidates; at the same time, the ranking information of different relevance will help to construct positive and negative samples of different qualities for subsequent model training; thus, for each lyric paragraph, a sequence of image-text pair candidates with different relevance qualities can be obtained to constitute a data sample, thereby forming a data set under a specific task; The text creation model consists of four parts; the first three parts constitute the encoder, specifically, the original image and text are first processed by a multi-channel sequence processor to generate their semantic embeddings; then, the embeddings at each step are divided into different parts to affect the final output; finally, different modalities are fused together with an attention network; the last part is the decoder, which aims to predict the final output sentence.

2. The artificial intelligence text creation method based on multimodal information input according to claim 1 is characterized in that: The first part of the text creation model is specifically as follows: the formats and semantics of the original images and texts are presented in different spaces; in order to adapt to them, a multi-channel sequence processor is designed, and the different modal sequences are first mapped to the same high-dimensional space through the multimodal pre-training model WenLan, and then input into these encoder neural networks; these encoders can be recurrent neural networks or Transformers, and ultimately the specific modules to be adopted can be selected by weighing effectiveness and efficiency; the output is an implicit embedding sequence; both the input image and text sequences are processed in this way.

3. The artificial intelligence text creation method based on multimodal information input according to claim 1 is characterized in that: The second part of the text creation model is as follows: the text creation model is a sequence-to-sequence architecture; however, unlike traditional machine translation tasks, each input word strictly corresponds to an output word. In the problem of this application, images or text may affect the span of the output sequence; in order to model these restrictions, a customized module is designed to capture the impact of input on output; specifically, the hidden embeddings derived in the previous section are made to specifically affect the output sequence; for these hidden embeddings, an inter-modal attention mechanism is designed within each channel to capture the impact of different time steps on other time steps, so as to obtain the span containing different other time steps. The comprehensive hidden embedding of a certain time step of the step information; in order to encode intuition into the customized module, a regularizer is further introduced to constrain the learning of attention weights; formally, the distance between the attention weights and a predefined distribution is minimized, thereby defining a KL loss function between the two for optimization and learning; by minimizing the KL loss, the attention weights are regularized using a priori, which encodes the intuition that a larger input-output distance should lead to a lower impact, thereby allowing the model to have good sensitivity to the order of the inputs; the prior knowledge about the distribution of attention weights is used to narrow the exploration space to bring better convergence rate and optimization solution.

4. The artificial intelligence text creation method based on multimodal information input according to claim 1 is characterized in that: The third part of the text creation model is specifically as follows: based on the partial hidden embeddings outputted above, different modalities are fused to derive the output of the encoder; specifically, the output of the encoder consists of L embeddings, each of which comprehensively encodes the topic, visual and textual information; the total output embedding is calculated by iterating the influence of hidden embeddings from different steps on the kth step; for each pair of steps, different modalities are weighted and combined together in a specific attention manner; intuitively, for the same output sentence, different modalities may play different roles; therefore, an inter-modal attention mechanism is adopted when combining them; if the above two attention mechanisms are compared, it may be found that the former is deployed in different steps of the same modality, while the latter aims to capture the contributions of different modalities in the same step; such a design actually forms a 2D attention mechanism, thereby modeling the influence of different positions and modalities in a more fine-grained manner.

5. The artificial intelligence text creation method based on multimodal information input according to claim 1 is characterized in that: The fourth part of the text creation model is specifically as follows: for the embedding generation output based on the output of the above module, different embeddings are merged as prompts, and all generated sentences are directly summarized and output; however, this strategy may not be optimal for retaining the sequential semantics of the input, because the ordered information may be weakened by the merging operation; in order to solve the above problem, each empirical embedding is allowed to affect the output sentence separately; formally, at each step, the output embedding of that step and the word embedding are added, and the topic word is used as a prompt, and then the whole is input into the decoder for generation; this method can maximize the retention of the influence of different time steps on different parts of the generated sentence.

6. The artificial intelligence text creation method based on multimodal information input according to claim 1 is characterized in that: In order to maximize the probability of generating the target output from the positive sample input and minimize the probability of generating the target output from the negative sample input, the text creation model is trained through curriculum learning. The specific training method is as follows: first learn the most negative sample to better initialize the model optimization; once the model has learned enough patterns to handle the most negative patterns, more difficult samples are gradually introduced near the positive and negative boundaries; more specifically, evaluate the correlation between the input image / text and the output, and construct 5 levels of samples; Level-5 represents the most relevant input, and Level-1 represents the least relevant input and output; during the training process, first train the model with Level-5 and Level-1 samples, and then include Level-4 and Level-2 in the positive and negative sample sets respectively, guiding the learning of the model in a gradually increasing manner from easy to difficult.