Visual language alignment-based visual narrative generation method and device, electronic equipment and storage medium
Through the multimodal big model of visual language alignment, the problems of narrative consistency and relevance of visual generative models are solved, and high-quality visual narrative texts are generated with a coherent plot, improving the creative effect of visual stories.
Patent Information
- Application Number
- CN202510309756.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2025-01-07
- Filing Date
- 2025-03-17
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-03-17
AI Technical Summary
When generating image sequence narratives, existing visual generation models are difficult to maintain visual relevance and narrative consistency, resulting in a lack of emotional resonance and in-depth plot portrayal of the generated content.
A multimodal large model based on visual language alignment is adopted. Through the combination of visual encoder, multimodal mapper and large language model, image encoding, semantic alignment and logical reasoning are performed to generate visual narrative text, combining multi-stage training and human preference optimization to improve the alignment of visual information and language output.
The narrative quality and consistency of visual story creation have been improved. The generated visual narrative text has high visual correlation and plot coherence with the image sequence, which is in line with the story content of human preferences.
Smart Images

Figure CN120279301A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the technical field of image understanding, and particularly relates to a visual narrative generation method, device, electronic device, and storage medium based on visual-language alignment. Background Art
[0002] Humans can easily recognize objects and targets in a given image, understand the logical relationships between them, and tell the content to someone who has never seen the image. They can even use common sense and external knowledge for supplementation. In this case, enabling a computer to have similar capabilities is a very challenging task.
[0003] Visual narrative is a task of generating a text story related to the information in the image or visual concepts with non-text data such as images as input. With the development of multi-modal large language models, visual narrative has become an important direction in content creation. Multi-modal text generation not only requires the system to understand the image content, but also to parse the text syntactically and grammatically, and effectively fuse the two modalities to achieve multi-modal reasoning. The difficulty of the text generation task is high.
[0004] Currently, vision generation models based on the Transformer architecture and large multi-modal models (Large Language and Vision Assistant, LLaVA) can be applied to generate narrative content by combining vision and language. However, when generating narrative for an image sequence, it often contains details unrelated to the image, resulting in a lack of visual relevance in the generated content, making it difficult to maintain the coherence of the plot, and causing the narrative content to be inconsistent with the timeline of the image sequence. Therefore, the generated narrative is difficult to achieve the richness and attractiveness of human creation, and the story lacks emotional resonance and in-depth plot portrayal.
[0005] In view of this, there is an urgent need for a new visual narrative generation method that can effectively align visual information with language output and improve the narrative quality and consistency of visual story creation. Summary of the Invention
[0006] This application aims to solve at least one of the technical problems existing in the prior art. For this purpose, this application provides a visual narrative generation method, device, electronic device, and storage medium based on visual-language alignment, which can effectively align visual information with language output and improve the visual relevance of multi-image narrative.
[0007] In a first aspect, this application provides a visual narrative generation method based on visual-language alignment, and the method includes:
[0008] Obtain a sequence of images to be processed and an instruction text;
[0009] Input the to-be-processed image sequence and the instruction text into a multi-modal large model to obtain the visual narrative text output by the multi-modal large model;
[0010] The multi-modal large model includes a visual encoder, a multi-modal mapper, and a large language model connected in sequence, and also includes a text encoder, and the output end of the text encoder is connected to the input end of the large language model;
[0011] The step of inputting the to-be-processed image sequence and the instruction text into the multi-modal large model to obtain the visual narrative text output by the multi-modal large model includes:
[0012] Encode the images in the to-be-processed image sequence sequentially through the visual encoder to generate visual features;
[0013] Project the visual features sequentially into the word embedding space of the large language model through the multi-modal mapper for semantic alignment to obtain language embedding vectors;
[0014] Convert the instruction text into an instruction embedding vector through the text encoder;
[0015] Generate the visual narrative text through the large language model according to the language embedding vector and the instruction embedding vector.
[0016] According to an embodiment of the present application, before inputting the to-be-processed image sequence and the instruction text into the multi-modal large model, the method further includes:
[0017] Obtain multiple image sequence samples from a sample set, and the initial storylines of the multiple image sequence samples, and the image sequence samples include multiple sample images;
[0018] Generate an overall description for each of the sample images in the image sequence sample;
[0019] Perform local region division on each of the sample images and generate corresponding local descriptions respectively;
[0020] Obtain the image description of the sample image according to the overall description and the local description;
[0021] Determine the image theme of the image sequence sample according to the initial storyline of the image sequence sample and the image description of each sample image in the image sequence sample;
[0022] Fuse the initial storyline, the image description, and the image theme to generate the story narrative label of the image sequence sample;
[0023] Input the combination of the image sequence sample, the corresponding story narrative label, and the sample instruction text as a training sample into the multi-modal large model to train the multi-modal large model.
[0024] According to an embodiment of the present application, determining the image theme of the image sequence sample based on the initial story line of the image sequence sample and the image description of each sample image in the image sequence sample includes:
[0025] Generate a first theme for each of the image sequence samples according to the image description and the initial story line;
[0026] Based on the data source of the image sequence sample, divide the first theme into at least one theme set;
[0027] Deduplicate and screen the first theme in the theme set to obtain at least one second theme;
[0028] Based on the initial story line, determine the image theme of each image sequence sample by performing theme matching between the second theme and the image sequence sample corresponding to the data source.
[0029] According to an embodiment of the present application, based on the initial story line, determining the image theme of each image sequence sample by performing theme matching between the second theme and the image sequence sample corresponding to the data source includes:
[0030] Put the multiple image sequence samples back into the sample set;
[0031] Divide the image sequences in the sample set according to the data source;
[0032] Perform theme matching on the initial story line of the image sequence sample through the second theme corresponding to the data source to determine the image theme of the image sequence sample;
[0033] Generate the image theme according to the initial story line and the image description of the unmatched image sequence sample.
[0034] According to an embodiment of the present application, before inputting the combination of the image sequence sample, the corresponding story narrative label, and the sample instruction text as a training sample into the multi-modal large model to train the multi-modal large model, the method further includes:
[0035] Obtain a plurality of image-text pairs, where the image-text pair includes an image file and the corresponding image description text;
[0036] Input the image-text pair into the multi-modal large model to train the modality mapper in the multi-modal large model.
[0037] According to an embodiment of the present application, a combination of the image sequence sample, the corresponding story narrative label, and the sample instruction text is used as a training sample and input into a multi-modal large model to train the multi-modal large model, including:
[0038] Determine the probability of the multi-modal large model generating a story narrative through the training sample:
[0039]
[0040] where s k is the output of the multi-modal large model according to the k-th training sample; in the sample instruction text, X Instrution is the specification of the output content by the sample instruction text, K is the style type of the generated s k ; r i is the word count requirement for the generated s k ; is the input image sequence sample, and the image sequence sample includes m sample images, and X′ vj is the j-th sample image in the image sequence sample;
[0041] Fine-tune the parameters of the large language model in the multi-modal large model according to the probability.
[0042] According to an embodiment of the present application, after using the combination of the image sequence sample, the corresponding story narrative label, and the sample instruction text as a training sample and inputting it into a multi-modal large model to train the multi-modal large model, the method further includes:
[0043] Input the training sample into the multi-modal large model in a loop to obtain multiple narrative outputs output by the multi-modal large model;
[0044] Reorder the sample images in the training sample and input them into the multi-modal large model to obtain multiple narrative outputs output by the multi-modal large model;
[0045] Score and sort the multiple narrative outputs, use the narrative output with the highest ranking as the positive example of the image sequence sample in the training sample, and use the narrative output with the lowest ranking as the negative example of the image sequence sample in the training sample;
[0046] Optimize the multi-modal large model through the positive example and the negative example corresponding to the training sample.
[0047] In a second aspect, the present application provides a visual narrative generation device based on visual language alignment, and the device includes:
[0048] An acquisition module for acquiring an image sequence to be processed and an instruction text;
[0049] A processing module for inputting the image sequence to be processed and the instruction text into a multi-modal large model to obtain a visual narrative text output by the multi-modal large model;
[0050] The multi-modal large model includes a visual encoder, a multi-modal mapper, and a large language model connected in sequence, and also includes a text encoder, and an output end of the text encoder is connected to an input end of the large language model;
[0051] The processing module is further configured to:
[0052] Encode the images in the image sequence to be processed in sequence through the visual encoder to generate visual features;
[0053] Project the visual features into the word embedding space of the large language model in sequence through the multi-modal mapper for semantic alignment to obtain language embedding vectors;
[0054] Convert the instruction text into an instruction embedding vector through the text encoder;
[0055] Perform logical reasoning according to the language embedding vector and the instruction embedding vector through the large language model to generate the visual narrative text.
[0056] In a third aspect, the present application provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the computer program, it implements the visual narrative generation method based on visual language alignment as described in the first aspect above.
[0057] In a fourth aspect, the present application provides a non-transitory computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the visual narrative generation method based on visual language alignment as described in the first aspect above.
[0058] In a fifth aspect, the present application provides a chip, the chip includes a processor and a communication interface, the communication interface is coupled to the processor, and the processor is configured to run a program or an instruction to implement the visual narrative generation method based on visual language alignment as described in the first aspect.
[0059] In a sixth aspect, the present application provides a computer program product, including a computer program, and when the computer program is executed by a processor, it implements the visual narrative generation method based on visual language alignment as described in the first aspect above.
[0060] Additional aspects and advantages of the present application will be given in part in the following description, become apparent in part from the following description, or be learned through the practice of the present application.
[0061] A visual narrative generation method, device, electronic device and storage medium based on visual language alignment provided by the present application have the following beneficial effects compared with the prior art:
[0062] (1) Through the modality mapper, visual features can be sequentially projected into the word embedding space of the large language model, which can adapt to the visual narrative generation task, effectively align visual information with language output, improve the narrative quality and consistency of visual story creation. The visual narrative text output by the large language model has visual relevance with the input image sequence to be processed and instruction text, can maintain the coherence of the plot, and the narrative content is consistent with the time axis of the image sequence, realizing high-quality visual narrative, and can be widely applied to the fields of intelligent content creation, virtual narrative generation and multi-modal artificial intelligence.
[0063] (2) Through rich story data optimization strategies such as image description generation, theme generation and story generation by the multi-modal large model, deep visual semantics and logical story evolution are extracted to ensure the clear theme and logical coherence of the narrative content. Through three-stage training of visual language alignment pre-training, story data fine-tuning and human preference optimization, the coherence, visual relevance and human preference adaptability of the visual narrative text generated by the multi-modal large model are significantly improved, and the generation from visual information to high-quality text narrative is completed, making the generated visual narrative significantly improved in terms of theme coherence, detail richness and human aesthetic adaptability, generating more attractive story content that conforms to human preferences, and improving the quality of visual narrative generation. BRIEF DESCRIPTION OF THE DRAWINGS
[0064] The above and / or additional aspects and advantages of the present application will become apparent and be readily understood from the description of the embodiments in conjunction with the following drawings, where:
[0065] Figure 1 is one of the flow schematic diagrams of the visual narrative generation method based on visual language alignment provided by the embodiment of the present application;
[0066] Figure 2 is the second flow schematic diagram of the visual narrative generation method based on visual language alignment provided by the embodiment of the present application;
[0067] Figure 3 is a schematic diagram of the visual narrative provided by the related art;
[0068] Figure 4 is the model framework diagram of the multi-modal large model provided by the embodiment of the present application;
[0069] Figure 5It is a flowchart of the story data enrichment method provided by an embodiment of the present application;
[0070] Figure 6 It is a schematic flowchart of the three-stage training provided by an embodiment of the present application;
[0071] Figure 7 It is a schematic diagram of the generation result of the multi-modal large model provided by an embodiment of the present application;
[0072] Figure 8 It is a schematic structural diagram of the visual narrative generation device based on visual language alignment provided by an embodiment of the present application;
[0073] Figure 9 It is a schematic structural diagram of the electronic device provided by an embodiment of the present application. Detailed implementation manners
[0074] Next, the technical solutions in the embodiments of the present application will be clearly described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are some, rather than all, of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application belong to the scope of protection of the present application.
[0075] The terms "first", "second", etc. in the specification and claims of the present application are used to distinguish similar objects, rather than to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present application can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first", "second", etc. are usually of the same category, and the number of objects is not limited. For example, the first object can be one or multiple. In addition, "and / or" in the specification and claims means at least one of the connected objects, and the character " / " generally indicates an "or" relationship between the associated objects before and after.
[0076] Next, in conjunction with the accompanying drawings, the visual narrative generation method based on visual language alignment, the visual narrative generation device based on visual language alignment, the electronic device, and the readable storage medium provided by the embodiments of the present application will be described in detail through specific embodiments and their application scenarios.
[0077] Among them, the visual narrative generation method based on visual language alignment can be applied to a terminal, and can be specifically executed by hardware or software in the terminal.
[0078] The terminal includes, but is not limited to, portable communication devices such as mobile phones or tablet computers having a touch-sensitive surface (e.g., a touch screen display and / or a touchpad). It should also be understood that, in some embodiments, the terminal may not be a portable communication device, but a desktop computer having a touch-sensitive surface (e.g., a touch screen display and / or a touchpad).
[0079] In the following various embodiments, a terminal including a display and a touch-sensitive surface is described. However, it should be understood that the terminal may include one or more other physical user interface devices such as a physical keyboard, a mouse, and a joystick.
[0080] The method for generating visual narratives based on visual-language alignment provided in the embodiments of the present application. The execution subject of the method for generating visual narratives based on visual-language alignment may be an electronic device or a functional module or functional entity in the electronic device that can implement the method for generating visual narratives based on visual-language alignment. The electronic devices mentioned in the embodiments of the present application include, but are not limited to, mobile phones, tablet computers, computers, cameras, and wearable devices, etc. Hereinafter, taking the electronic device as the execution subject as an example, the method for generating visual narratives based on visual-language alignment provided in the embodiments of the present application will be described.
[0081] As Figure 1 shown, the method for generating visual narratives based on visual-language alignment includes:
[0082] Step 110, obtain a sequence of images to be processed and an instruction text;
[0083] Step 120, input the sequence of images to be processed and the instruction text into a multimodal large model to obtain a visual narrative text output by the multimodal large model;
[0084] The multimodal large model includes a visual encoder, a multimodal mapper, and a large language model connected in sequence, and also includes a text encoder. The output end of the text encoder is connected to the input end of the large language model;
[0085] As Figure 2 shown, step 120, input the sequence of images to be processed and the instruction text into a multimodal large model to obtain a visual narrative text output by the multimodal large model, includes:
[0086] Step 121a, sequentially encode the images in the sequence of images to be processed through the visual encoder to generate visual features;
[0087] Step 122, project the visual features sequentially into the word embedding space of the large language model through the multimodal mapper for semantic alignment to obtain language embedding vectors;
[0088] Step 121b: Convert the instruction text into an instruction embedding vector through the text encoder.
[0089] Step 123: Through the large language model, perform logical reasoning based on the language embedding vector and the instruction embedding vector to generate the visual narrative text.
[0090] As Figure 3 shown, given an image, humans can easily identify the objects and targets in it, understand the logical relationships between them, and tell the content to someone who has never seen the image. They can even use common sense and external knowledge for supplementation to achieve visual narrative (Visual Embedding). In the embodiments of this application, a similar effect is achieved by setting up a multimodal large model.
[0091] As Figure 4 shown, the multimodal large model can use the pre-trained SigLIP-so400m as the visual encoder to process the input sequence of images to be processed at a high resolution of 384×384, and generate a feature map from each image as the visual feature, so that the visual information is converted into a representation compatible with the language model.
[0092] Through a two-layer Multilayer Perceptron (MLP) as a multimodal mapper (Projection), project the visual features into the word embedding space of the language model for semantic alignment, and generate a language embedding vector with the same spatial dimension as the word embedding space, so that the visual information can be fused with the text information.
[0093] The Phi-3 Mini-128K model adopted by the large language model can support up to 128K tokens to ensure the effective processing of multi-image input and reduce the consumption of computing resources.
[0094] In actual execution, the sequence of images to be processed can be a set of n continuously input images is the i-th image, i ∈ n. Use the visual encoder to encode each image to obtain the corresponding visual feature is the visual feature of the i-th image.
[0095] Through a trainable two-layer MLP as a multimodal mapper, convert these visual features into a language embedding vector with the same dimension as the word embedding space of the large language model:
[0096]
[0097] Among them, is the language embedding vector of the i-th image.
[0098] Through the text encoder, the user's instruction text (Language Instruction) is vectorized and encoded. The instruction text is preprocessed through the tokenizer&embedding technology to obtain information such as instructions (Instruction), image topics (Topic Inf), and story requirements (Story Requirement). These information are vectorized to obtain the instruction embedding vector H c .
[0099] Visual narrative (Visual Embedding) is realized through the large language model. According to the language embedding vector and the instruction embedding vector H c , the visual narrative text X corresponding to the image is inferred S , which can be expressed as:
[0100]
[0101] In the multimodal large model, the multimodal mapper can adapt to the visual narrative generation task by projecting visual features into the word embedding space of the large language model, ensuring the efficient integration of visual information and language information.
[0102] According to the visual narrative generation method based on visual language alignment provided by the embodiments of the present application, through the modality mapper, visual features can be successively projected into the word embedding space of the large language model, which can adapt to the visual narrative generation task, effectively align visual information with language output, improve the narrative quality and consistency of visual story creation. The visual narrative text output by the large language model has visual relevance with the input image sequence to be processed and the instruction text, can maintain the coherence of the plot, and the narrative content is consistent with the time axis of the image sequence, realizing high-quality visual narrative, and can be widely applied to the fields of intelligent content creation, virtual narrative generation, and multimodal artificial intelligence.
[0103] In some embodiments, before inputting the image sequence to be processed and the instruction text into the multimodal large model, the method further includes:
[0104] Obtain a plurality of image sequence samples and the initial storylines of the plurality of image sequence samples from the sample set. The image sequence samples include a plurality of sample images;
[0105] Generate an overall description for each of the sample images in the image sequence sample;
[0106] Perform local region division on each of the sample images and generate corresponding local descriptions respectively;
[0107] Obtain the image description of the sample image according to the overall description and the partial description;
[0108] Determine the image theme of the image sequence sample according to the initial storyline of the image sequence sample and the image description of each sample image in the image sequence sample;
[0109] Fuse the initial storyline, the image description, and the image theme to generate the story narrative label of the image sequence sample;
[0110] Use the combination of the image sequence sample, the corresponding story narrative label, and the sample instruction text as a training sample to input into a multi-modal large model for training the multi-modal large model.
[0111] Among them, the sample instruction text can be determined according to the relationship between the image sequence sample and the story narrative label.
[0112] In some embodiments, obtaining a plurality of image sequence samples from a sample set includes:
[0113] Obtain a plurality of image sequences from the VIST dataset, VWP dataset, PororoSV dataset, and FlintstonesSV dataset as image sequence samples to construct a sample set.
[0114] Extract the plurality of image sequence samples from the sample set based on the data source.
[0115] According to the scale of the original dataset, image sequence samples are selected from different datasets. For example, 20k, 10k, 5k, and 5k image sequence samples are respectively selected from datasets such as VIST, VWP, PororoSV, and FlintstonesSV for low-rank adaptation (Lora) fine-tuning, and the data is optimized. Finally, a total of 40k image sequence samples for fine-tuning are generated.
[0116] During the process of generating the narrative, automatically extract the image theme of the image sequence sample, and optimize it in combination with the GPT-4 model to ensure that the generated narrative content has a clear theme and avoid redundant and irrelevant details. Generate multiple versions of the story, and rank these versions using human preference data, and select the narrative that best meets human aesthetics to optimize the quality of the content generated by the model.
[0117] In actual execution, each image sequence sample is processed to obtain the initial storyline corresponding to the image sequence sample. For example, an image sequence sample includes 5 sample images.
[0118] Such as Figure 5As shown, based on caption generation technology and combined with local captions, rich and high-quality caption data can significantly enhance the capabilities of multimodal models.
[0119] In the overall caption stage, using a vision foundation model such as Florence-2, for each sample image in the image sequence sample, an overall description is generated separately, and at the same time, through the Segment Anything Model (SAM), each sample image is divided into local regions, and detailed local descriptions are generated for each region in the sample image.
[0120] To make these descriptions closer to human language expressions, a phrase-level generation strategy can also be adopted to refine the overall description and local description of each sample image, and finally obtain a refined and structured image description for each sample image, providing comprehensive information support for the next step of topic generation.
[0121] In some embodiments, determining the image topic of the image sequence sample according to the initial storyline of the image sequence sample and the image description of each sample image in the image sequence sample includes:
[0122] Generating a first topic for each image sequence sample according to the image description and the initial storyline;
[0123] Based on the data source of the image sequence sample, dividing the first topic into at least one topic set;
[0124] Removing duplicates and screening the first topics in the topic set to obtain at least one second topic;
[0125] Based on the initial storyline, determining the image topic of each image sequence sample by matching the second topic with the corresponding image sequence sample of the data source.
[0126] In the process of story topic generation, the TopicGPT model is used to generate a first topic related to the initial storyline according to the image description. Since different image sequence samples are obtained from different data sets and have different data sources, the first topic can be divided into different topic sets according to different data sources, and each topic set corresponds to a data source. After post-processing the first topics in the topic set by removing duplicates, merging duplicate topics, and eliminating rare topics, the second topic corresponding to each data source is obtained, ensuring quality and coverage, and making the topics more accurate and diverse.
[0127] In some embodiments, based on the initial storyline, the image theme of each image sequence sample is determined by matching the second theme with the image sequence samples of the corresponding data source, including:
[0128] Put the multiple image sequence samples back into the sample set;
[0129] Divide the image sequences in the sample set according to the data source;
[0130] Match the initial storyline of the image sequence sample with the second theme corresponding to the data source to determine the image theme of the image sequence sample;
[0131] Generate the image theme (Topic Generation) according to the initial storyline and the image description of the unmatched image sequence samples.
[0132] Put the extracted multiple image sequence samples back into the sample set, divide the image sequence samples in the sample set according to the data source, and perform topic matching (Topic Assignment) on the image sequence samples of the corresponding data source according to the initial storyline (Storyline) through the second theme according to the assignment prompt (Assignment Prompt). Some image sequence samples can (Success) match the second theme, and the second theme is used as the image theme (Story Topic) of the image sequence sample.
[0133] For the image sequence samples without a matching theme, use the predefined examples (Generation Prompt) of the TopicGPT model to generate a new image theme according to the initial storyline and the image description to ensure that the initial storyline of each image sequence sample has a clear direction. When it is necessary to reduce the computing cost, an open-source model such as Mistral-7B can also be used instead of GPT-4 for topic assignment to achieve local computing.
[0134] Fuse the initial storyline, image description, and image theme of each image sequence sample, and use the GPT model to focus on the logical coherence and theme consistency of the story narration to generate a high-quality complete story narration that combines visual information and theme, avoiding false details unrelated to the image content or theme direction. By integrating visual details and theme orientation, the generated story narration not only accurately reflects the information of the input image but also is more attractive emotionally and logically, and outputs a high-quality story S:
[0135] S = GPT(y, C, T, Prompt)
[0136] Among them, y is the initial story line, C is the image description, T is the image theme, and the prompt will guide GPT-4 to fuse these inputs into a coherent story S.
[0137] Taking the story S as the story narrative label for the corresponding image sequence sample, and taking the combination of the image sequence sample and the corresponding story narrative label as a training sample, multiple training samples can be obtained. Then, input the multiple training samples into the multi-modal large model for training in sequence.
[0138] In this embodiment, by enriching the story data, the story can be automatically rewritten and enriched, fusing deeper visual semantics and logical story evolution; by generating image descriptions, it provides detailed visual information support for subsequent theme generation and story generation, extracts or generates a set of content-related themes from the image descriptions, provides directional support for story narration, and generates logical and attractive story narrative labels based on the image descriptions, theme sets, and initial story lines, providing a basis for the training of the multi-modal large model.
[0139] Such as Figure 6 As shown, high-quality story generation is achieved through multi-stage training and inference. In multi-stage training, through three stages of training: image-text pre-training, story data fine-tuning training, and preference optimization training, the multi-modal large model is gradually optimized in each stage to improve the quality of visual narrative generation, and finally high-quality visual narrative content is generated.
[0140] In some embodiments, before taking the combination of the image sequence sample, the corresponding story narrative label, and the sample instruction text as a training sample and inputting it into the multi-modal large model for training the multi-modal large model, the method further includes:
[0141] Obtain a plurality of image-text pairs, where the image-text pairs include image files and corresponding image description texts;
[0142] Input the image-text pairs into the multi-modal large model to train the modality mapper in the multi-modal large model.
[0143] In this embodiment, through pre-training of the multi-modal large model, the modality mapper learns to align visual features into the word embedding space of the large language model, and the large language model (LLM) can understand the output of the modality mapper.
[0144] In some embodiments, taking the combination of the image sequence sample, the corresponding story narrative label, and the sample instruction text as a training sample and inputting it into the multi-modal large model for training the multi-modal large model includes:
[0145] Determine the probability of the multi-modal large model generating a story narrative through the training samples:
[0146]
[0147] where s k is the narrative output of the multi-modal large model according to the k-th training sample; in the sample instruction text, X Insrrution is the requirement of the sample instruction text for the output content, K is the style type of the generated s k ; r i is the word count requirement for the generated s k ; is the input image sequence sample, and the image sequence sample includes m sample images, X' vj is the j-th sample image in the image sequence sample;
[0148] Fine-tune the parameters of the large language model in the multi-modal large model according to the probability.
[0149] The multi-modal large model performs reinforcement learning from human feedback, effectively guiding the large language model to generate more honest, useful, and harmless content, enabling the multi-modal large model to not only learn to match the image order but also learn to reasonably adjust the story plot in different image arrangements and improve the overall fluency of the story.
[0150] In this embodiment, by fine-tuning the parameters of the large language model in a supervised manner to maximize the probability of generating a story narrative, the multi-modal large model enhances its understanding of images and generates more coherent stories on the basis of visual-language alignment.
[0151] Through fine-tuning, the length of the stories generated by the multi-modal large model is increased, but the hallucination of the multi-modal large model increases significantly, and further optimization of the multi-modal large model is required.
[0152] In some embodiments, after using the combination of the image sequence sample, the corresponding story narrative label, and the sample instruction text as training samples to input into the multi-modal large model for training the multi-modal large model, the method further includes:
[0153] Circularly input the training samples into the multi-modal large model to obtain multiple narrative outputs output by the multi-modal large model;
[0154] Re-order the sample images in the training samples and input them into the multi-modal large model to obtain multiple narrative outputs output by the multi-modal large model;
[0155] Score and rank the multiple narrative outputs, and use the narrative output with the highest rank as the positive example of the image sequence sample in the training sample, and use the narrative output with the lowest rank as the negative example of the image sequence sample in the training sample;
[0156] Optimize the multi-modal large model through the training sample corresponding to the positive example and the negative example.
[0157] After training is completed, evaluate the 6 narrative outputs generated by sampling each training sample by the multi-modal large model, then swap the order of the sample images in each training sample, continue to sample and generate narrative outputs, score and rank all the narrative outputs according to the preset quality criteria, select the narrative output with the best score and the highest rank as the positive example, and select the narrative output with the lowest score and the lowest rank as the negative example.
[0158] Use the combination of the image sequence sample (image sequence), positive example (chosen story), and negative example (rejected story) in the training sample as preference data. A total of 6k pieces of preference data are constructed, and each piece of data contains <imagesequence, chosen story, rejected story> samples, combined with the corresponding sample instruction text, for use in direct preference optimization (DPO) training to optimize the multi-modal large model.
[0159] In some embodiments, the training objective is defined as:
[0160]
[0161] where, π θ is the multi-modal large model in DPO training; π ref is the baseline reference model after fine-tuning, I is the image sequence, S w is the chosen story, S l is the rejected story, E(I, S w , S l ) represents the expected value on the dataset of the image sequence and story pair, σ is the logistic function, and β is a parameter set to 0.1.
[0162] In this embodiment, using the preference data, optimize the model through the DPO method to make the stories generated by the multi-modal large model more in line with human preferences. After adjusting the model parameters of the multi-modal large model, the generated stories can achieve the best in terms of coherence, non-redundancy, and fluency, achieving a high-quality narrative effect, and are significantly superior to the prior art in terms of the coherence, fluency, and non-redundancy of visual narrative.
[0163] The existing models with similar structures and the multi-modal large model provided by the embodiments of this application were verified on the VIST dataset, and the experimental results are as follows. As Figure 7 shown, it qualitatively demonstrates the results generated by sampling test set images of the multi-modal large model of this application on the cartoon dataset. The experimental results show that the multi-modal large model of this application has achieved optimal performance in narrative text generation.
[0164] Table 1. Reference-free metrics on the VIST dataset
[0165] Model Name Intra-Repetition GROOVIST RoVIST-C Coherence RoVIST-NR AREL 23.5 0.584 0.577 0.833 MCSM + BART 2.8 0.852 0.666 0.865 TAPM 6.8 0.734 0.671 0.903 LLaVA 6.5 0.541 0.809 0.851 Multimodal Large Model 0.5 0.764 0.833 0.905
[0166] Among the reference-free metrics, the smaller the Intra-Repetition value of sentence-level repetition, the higher the quality of the story.
[0167] GROOViST for visual relevance is a metric used to evaluate the alignment between vision and text in visual narratives. It extracts noun phrases in the story and compares them with visual objects in the image to calculate the visual alignment score. Compared with existing metrics, it is more in line with human intuition about visual associations and is an evaluation tool with strong interpretability and scalability. The larger the GROOViST for visual relevance, the higher the quality of the story.
[0168] RoViST-C is used to evaluate the inter-sentence coherence of the story. The model captures the logical fluency of the story by judging whether adjacent sentences are coherent, calculates the coherence probability of each pair of adjacent sentences, and finally generates an overall coherence score. It effectively captures the natural transition between sentences and is suitable for analyzing the structural integrity of the generated story. The greater the RoVIST-C coherence, the higher the quality of the story.
[0169] RoViST-NR is used to evaluate the redundancy of the story, including inter-sentence and intra-sentence repetition. By calculating the Jaccard similarity of words between sentences and within sentences, the degree of repetition is measured. The higher the score, the less redundant the story and the higher the quality of the story.
[0170] As shown in Table 1, in the results of applying the reference-free automatic evaluation metrics on the VIST test set, which uses five consecutive images without real label annotations, the multi-modal large model of this application is continuously superior to the AREL model, MCSM+BART model, TAPM model, and LlaVA model in terms of coherence, non-redundancy, and fluency.
[0171] Table 2. Reference-based metrics on the VIST dataset
[0172] Model Name METEOR ROUGE-L(R) CIDEr(C) AREL 35.2 29.3 9.1 MCSM + BART 36.1 30.7 11.0 TAPM 33.1 37.2 13.8 Multimodal Large Model 29.9 33.7 14.5
[0173] Among the reference metrics, the larger the METEOR (Metric for Evaluation of Translation with Explicit Ordering), the higher the quality of the story; the larger the ROUGE-L (Recall-Oriented Understudy for Gisting Evaluation - Longest common subsequence), the higher the quality of the story; the larger the CIDEr (Cosensus-based Image Description Evaluation), the higher the quality of the story.
[0174] As shown in Table 2, in the results of applying reference-based automatic evaluation metrics on the VIST test set, the multi-modal large model of this application performs excellently on traditional word-overlap-based evaluation metrics.
[0175] In this embodiment, through rich story data optimization strategies such as image description generation, theme generation, and story generation by the multi-modal large model, deep visual semantics and logical story evolution are extracted to ensure clear themes and logical coherence of the narrative content. Through three-stage training of visual-language alignment pre-training, story data fine-tuning, and human preference optimization, the coherence, visual relevance, and human preference adaptability of the multi-modal large model in generating visual narrative texts are significantly improved, completing the generation from visual information to high-quality text narrative, and significantly improving the generated visual narrative in terms of theme coherence, detail richness, and human aesthetic adaptability, generating more engaging story content that conforms to human preferences and improving the quality of visual narrative generation.
[0176] In the method for visual narrative generation based on visual-language alignment provided by the embodiments of this application, the execution subject can be a device for visual narrative generation based on visual-language alignment. In the embodiments of this application, taking the device for visual narrative generation based on visual-language alignment as an example to execute the method for visual narrative generation based on visual-language alignment, the device for visual narrative generation based on visual-language alignment provided by the embodiments of this application is described.
[0177] The embodiments of this application also provide a device for visual narrative generation based on visual-language alignment.
[0178] As Figure 8 shown, the device for visual narrative generation based on visual-language alignment includes:
[0179] An acquisition module 810, configured to acquire an image sequence to be processed and an instruction text;
[0180] A processing module 820, configured to input the image sequence to be processed and the instruction text into a multi-modal large model, and obtain a visual narrative text output by the multi-modal large model;
[0181] The multimodal large model includes a visual encoder, a multimodal mapper, and a large language model connected in sequence, and also includes a text encoder, and an output end of the text encoder is connected to an input end of the large language model;
[0182] The processing module 820 is further configured to:
[0183] Encode the images in the to-be-processed image sequence in sequence through the visual encoder to generate visual features;
[0184] Project the visual features into the word embedding space of the large language model in sequence through the multimodal mapper for semantic alignment to obtain language embedding vectors;
[0185] Convert the instruction text into an instruction embedding vector through the text encoder;
[0186] Perform logical reasoning according to the language embedding vector and the instruction embedding vector through the large language model to generate the visual narrative text.
[0187] According to the visual narrative generation device based on visual-language alignment provided by the embodiments of the present application, through the modal mapper, the visual features can be projected into the word embedding space of the large language model in sequence, which can adapt to the visual narrative generation task, effectively align visual information with language output, improve the narrative quality and consistency of visual story creation, and there is a visual correlation between the visual narrative text output by the large language model and the input to-be-processed image sequence and instruction text, and the coherence of the plot can be maintained, and the narrative content is consistent with the time axis of the image sequence, realizing high-quality visual narrative, and can be widely applied to the fields of intelligent content creation, virtual narrative generation, and multimodal artificial intelligence.
[0188] In addition, the device is further provided with a user interface for receiving the input to-be-processed image sequence and outputting the generated visual narrative text.
[0189] The visual narrative generation device based on visual language alignment in the embodiments of this application can be an electronic device or a component in an electronic device, such as an integrated circuit or a chip. The electronic device can be a terminal or other devices other than terminals. Exemplarily, the electronic device can be a mobile phone, a tablet computer, a laptop computer, a handheld computer, an in-vehicle electronic device, a Mobile Internet Device (MID), an Augmented Reality (AR) / Virtual Reality (VR) device, a robot, a wearable device, an Ultra-Mobile Personal Computer (UMPC), a netbook, or a Personal Digital Assistant (PDA), etc., and can also be a server, a Network Attached Storage (NAS), a Personal Computer (PC), a Television (TV), a teller machine, or a self-service machine, etc. The embodiments of this application do not make specific limitations.
[0190] The visual narrative generation device based on visual language alignment in the embodiments of this application can be a device with an operating system. The operating system can be an Android operating system, an iOS operating system, or other possible operating systems. The embodiments of this application do not make specific limitations.
[0191] The visual narrative generation device based on visual language alignment provided in the embodiments of this application can implement each process realized by the embodiments of the visual narrative generation method based on visual language alignment as described above. To avoid repetition, it will not be elaborated here.
[0192] In some embodiments, as Figure 9 shown, the embodiments of this application also provide an electronic device 900, including a processor 901, a memory 902, and a computer program stored on the memory 902 and executable on the processor 901. When the program is executed by the processor 901, it implements each process of the embodiments of the above-mentioned visual narrative generation method based on visual language alignment and can achieve the same technical effects. To avoid repetition, it will not be elaborated here.
[0193] It should be noted that the electronic devices in the embodiments of this application include the above-mentioned mobile electronic devices and non-mobile electronic devices.
[0194] An embodiment of the present application also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements each process of the above-mentioned embodiment of the visual narrative generation method based on vision-language alignment and can achieve the same technical effects. To avoid repetition, it will not be elaborated here.
[0195] Among them, the processor is the processor in the electronic device in the above-mentioned embodiment. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disk or optical disc, etc.
[0196] An embodiment of the present application also provides a computer program product, including a computer program, which implements the above-mentioned visual narrative generation method based on vision-language alignment when executed by a processor.
[0197] Among them, the processor is the processor in the electronic device in the above-mentioned embodiment. The readable storage medium includes computer-readable storage media, such as computer read-only memory ROM, random access memory RAM, magnetic disk or optical disc, etc.
[0198] Another embodiment of the present application provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is used to run programs or instructions to implement each process of the above-mentioned embodiment of the visual narrative generation method based on vision-language alignment and can achieve the same technical effects. To avoid repetition, it will not be elaborated here.
[0199] It should be understood that the chip mentioned in the embodiment of the present application may also be referred to as a system-on-chip, system chip, chip system, or system-on-chip, etc.
[0200] It should be noted that in this text, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements not only includes those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the phrase "comprising a..." does not exclude the presence of additional identical elements in the process, method, article or device comprising that element. In addition, it should be pointed out that the scope of the methods and devices in the embodiments of the present application is not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in the reverse order according to the functions involved. For example, the described methods may be performed in an order different from that described, and various steps may be added, omitted, or combined. Additionally, the features described with reference to certain examples may be combined in other examples.
[0201] Through the description of the above embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a computer software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions for causing a terminal (which can be a mobile phone, computer, server, or network device, etc.) to execute the visual narrative generation method based on visual language alignment in various embodiments of the present application.
[0202] In the description of the present application, "the first feature", "the second feature" may include one or more of such features.
[0203] In the description of the present application, the meaning of "a plurality of" is two or more.
[0204] The embodiments of the present application have been described above in conjunction with the accompanying drawings. However, the present application is not limited to the above specific embodiments. The above specific embodiments are merely illustrative and not restrictive. Those of ordinary skill in the art, under the inspiration of the present application and without departing from the purpose of the present application and the scope protected by the claims, can still make many forms, all of which fall within the protection scope of the present application.
[0205] In the description of this specification, the description referring to terms such as "one embodiment", "some embodiments", "schematic embodiments", "examples", "specific examples", or "some examples", etc. means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of this application. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples.
[0206] Although the embodiments of this application have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and purposes of this application, and the scope of this application is defined by the claims and their equivalents.
Claims
1. A visual narrative generation method based on visual-language alignment, characterized in that Including: Obtain the image sequence to be processed and the instruction text; Input the image sequence to be processed and the instruction text into a multimodal large model to obtain the visual narrative text output by the multimodal large model; The multimodal large model includes a visual encoder, a multimodal mapper, and a large language model connected in sequence, and also includes a text encoder, and the output end of the text encoder is connected to the input end of the large language model; The step of inputting the image sequence to be processed and the instruction text into the multimodal large model to obtain the visual narrative text output by the multimodal large model includes: Through the visual encoder, encode the images in the image sequence to be processed in sequence to generate visual features; Through the multimodal mapper, project the visual features into the word embedding space of the large language model in sequence for semantic alignment to obtain language embedding vectors; Through the text encoder, convert the instruction text into an instruction embedding vector; Through the large language model, perform logical reasoning based on the language embedding vector and the instruction embedding vector to generate the visual narrative text.
2. The visual narrative generation method based on visual-language alignment according to claim 1, wherein Before inputting the image sequence to be processed and the instruction text into the multimodal large model, the method further includes: Obtain multiple image sequence samples from the sample set, and the initial storylines of the multiple image sequence samples, and the image sequence samples include multiple sample images; Generate an overall description for each of the sample images in the image sequence sample; Perform local region division on each of the sample images and generate corresponding local descriptions respectively; According to the overall description and the local description, obtain the image description of the sample image; According to the initial storyline of the image sequence sample and the image description of each sample image in the image sequence sample, determine the image theme of the image sequence sample; Fuse the initial storyline, the image description, and the image theme to generate the story narrative label of the image sequence sample; Input the combination of the image sequence sample, the corresponding story narrative label, and the sample instruction text into the multimodal large model as a training sample to train the multimodal large model.
3. The visual narrative generation method based on visual-language alignment according to claim 2, wherein According to the initial storyline of the image sequence sample and the image description of each sample image in the image sequence sample, determining the image theme of the image sequence sample includes: Generate a first theme for each of the image sequence samples according to the image description and the initial storyline; Based on the data source of the image sequence sample, divide the first theme into at least one theme set; Deduplicate and screen the first themes in the theme set to obtain at least one second theme; Based on the initial storyline, perform theme matching between the second theme and the image sequence sample of the corresponding data source to determine the image theme of each image sequence sample.
4. The visual narrative generation method based on visual-language alignment according to claim 3, wherein, Based on the initial storyline, performing theme matching between the second theme and the image sequence sample of the corresponding data source to determine the image theme of each image sequence sample includes: Put the multiple image sequence samples back into the sample set; Divide the image sequences in the sample set according to the data source; Perform theme matching on the initial storyline of the image sequence samples corresponding to the data source to determine the image themes of the image sequence samples; Generate the image themes according to the initial storyline and the image descriptions of the unmatched image sequence samples.
5. The visual narrative generation method based on visual-language alignment according to claim 2, wherein Before using the combination of the image sequence samples, the corresponding story narrative tags, and the sample instruction text as training samples to input into the multi-modal large model for training the multi-modal large model, the method further includes: Obtain a plurality of image-text pairs, where the image-text pairs include image files and corresponding image description texts; Input the image-text pairs into the multi-modal large model to train the modality mapper in the multi-modal large model.
6. The method for generating visual narrative based on visual-language alignment according to claim 2, wherein Using the combination of the image sequence samples, the corresponding story narrative tags, and the sample instruction text as training samples to input into the multi-modal large model for training the multi-modal large model includes: Determine the probability of the multi-modal large model generating a story narrative through the training samples; Among them, s k is the output of the multi-modal large model according to the k-th training sample; in the sample instruction text, X Instrution is the specification of the output content by the sample instruction text, K is the style type of the generated s k ; r i is the word count requirement for the generated s k ; is the input image sequence sample, and the image sequence sample includes m sample images, X′ vj is the j-th sample image in the image sequence sample; Fine-tune the parameters of the large language model in the multi-modal large model according to the probability.
7. The visual narrative generation method based on visual-language alignment according to claim 2, wherein After using the combination of the image sequence samples, the corresponding story narrative tags, and the sample instruction text as training samples to input into the multi-modal large model for training the multi-modal large model, the method further includes: Loop the training samples into the multi-modal large model to obtain multiple narrative outputs output by the multi-modal large model; Reorder the sample images in the training samples and input them into the multi-modal large model to obtain multiple narrative outputs output by the multi-modal large model; Score and sort the multiple narrative outputs, regard the narrative output with the highest ranking as the positive example of the image sequence sample in the training sample, and regard the narrative output with the lowest ranking as the negative example of the image sequence sample in the training sample; Optimize the multi-modal large model through the positive example and the negative example corresponding to the training samples.
8. A visual narrative generation device based on visual-language alignment, characterized in that, Includes: An acquisition module for acquiring an image sequence to be processed and an instruction text; A processing module for inputting the image sequence to be processed and the instruction text into the multi-modal large model to obtain the visual narrative text output by the multi-modal large model; The multi-modal large model includes a visual encoder, a multi-modal mapper, and a large language model connected in sequence, and further includes a text encoder, and the output end of the text encoder is connected to the input end of the large language model; The processing module is further configured to: Sequentially encode the images in the image sequence to be processed through the visual encoder to generate visual features; Sequentially project the visual features into the word embedding space of the large language model through the multi-modal mapper for semantic alignment to obtain language embedding vectors; Convert the instruction text into an instruction embedding vector through the text encoder; Generate the visual narrative text through the large language model according to the language embedding vector and the instruction embedding vector.
9. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the method for generating visual narratives based on visual-language alignment according to any one of claims 1-7.
10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method for generating visual narratives based on visual-language alignment according to any one of claims 1-7.
Citation Information
Patent Citations
Few-sample visual story narrating method based on theme adaptation and prototype coding
CN111708904A
Model training method and device, perception reasoning method, electronic equipment and medium
CN117290459A
Sports data news generation system based on big language model
CN118395948A
Model training method and device, electronic equipment and storage medium
CN118824282A
Building a customized story
US20120246562A1
Cited By
Multi-modal space-time trajectory fusion method and system
CN120724396A
A multi-modal spatio-temporal trajectory fusion method and system
CN120724396B
Method, device and equipment for generating multi-modal visual language model
CN121094125A
Role playing data set construction method based on image enhancement
CN121416006A
Image processing method and device, computer equipment and storage medium
CN121527596A