Low-cost multi-modal article generation method
Through the low-cost multimodal article generation method, Lora's fine-tuned StableDiffusion-XL model and High-class corpus generation model solve the problems of strong dependence on LLM and limited input mode compatibility of existing systems, and achieve high-quality multimodal article generation.
Patent Information
- Application Number
- CN202510222630.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-27
- Publication Date
- 2025-06-13
AI Technical Summary
The existing article generation system has strong dependence on large-scale pre-trained language model (LLM), and the compatibility of input modes is limited, making it difficult to effectively integrate visual modal information to enhance the richness of generated content.
The low-cost multimodal article generation method is adopted to generate image data by obtaining user input data (target text data or target text data and target image data), using the StableDiffusion-XL model fine-tuned by Lora, and fusing it with image data through the High-class corpus generation model to generate multimodal articles.
The local small parameter LLM is implemented to expand the amount of information in the image mode, output high-quality predictions, reduce the dependence on online LLM, and support flexible combination of multimodal inputs, improving the richness and quality of article generation.
Smart Images

Figure CN120144739A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of article generation, and more specifically, to a low-cost multi-modal article generation method. Background Art
[0002] As a cutting-edge direction in the field of natural language processing, the core challenge of the article generation task lies in constructing a text generation system with long-term semantic relevance and discourse coherence. Traditional methods mainly build an encoder-decoder architecture based on recurrent neural networks (RNNs) and attention mechanisms, but there are generally semantic discontinuity problems when dealing with long texts exceeding a thousand characters. With the introduction of the Transformer architecture, article generation methods based on large-scale pre-trained language models (LLMs) have significantly improved the text quality, but still face the following two key problems:
[0003] (1) Existing systems have a strong dependence on LLMs: Article generation is achieved by invoking online LLMs through prompt engineering, with a high dependence on external APIs. The knowledge contained in local small-parameter LLMs is insufficient to support the creative requirements of the article generation task. A feasible method is to fine-tune local small-parameter LLMs to inject knowledge, but academic institutions and small and medium-sized enterprises can hardly afford the training costs;
[0004] (2) There are significant limitations in the compatibility of input modalities: The current mainstream architectures can be divided into single-text models and text-image dual-modal models; among them, single-text models only accept text-only prompting as input and cannot effectively integrate visual modality information to enhance the richness of the generated content; while text-image dual-modal models require text-image pairs as input. For existing text-image dual-modal models, if only text data is input and no image data is input, the performance of the text-image dual-modal models will drop sharply. This either-or design paradigm limits the actual application scenarios and is difficult to meet the needs of users for flexible combination of multi-modal inputs.
[0005] Therefore, how to enable local small-parameter LLMs to expand the information volume of the image modality and output high-quality text is an urgent problem to be solved by those skilled in the art. Summary of the Invention
[0006] In view of the above problems, the present invention provides a low-cost multi-modal article generation method to at least solve some of the technical problems mentioned in the above background art.
[0007] To achieve the above object, the present invention adopts the following technical solutions:
[0008] The present invention provides a low-cost multi-modal article generation method, including the following steps:
[0009] Obtain user input data, where the user input data is target text data, or target text data and target image data;
[0010] Input the target text data in the user input data into the StableDiffusion-XL model fine-tuned by Lora to generate m pieces of image data;
[0011] Fuse the m pieces of image data, or fuse the target image data with the m pieces of image data to generate a fused feature;
[0012] Input the target text data, the fused feature, and the m pieces of image data respectively into the trained and optimized High-class corpus generation model to output m corresponding target corpora;
[0013] Input the m target corpora and the target text data into the local Writer LLM module to generate a multimodal article.
[0014] Furthermore, the training process of the High-class corpus generation model includes:
[0015] Obtain a large amount of input text data and corresponding input image data;
[0016] Input the input text data into the Stable Diffusion-XL model fine-tuned by Lora to generate multiple pieces of image data;
[0017] Fuse the multiple pieces of image data to generate a first fused feature;
[0018] Fuse the multiple pieces of image data and the input image data to generate a second fused feature;
[0019] Use the input text data, the multiple pieces of image data, and the first fused feature as the first training data;
[0020] Use the input text data, the multiple pieces of image data, and the second fused feature as the second training data;
[0021] Based on the first training data and the second training data, train the High-class corpus generation model respectively to obtain the trained High-class corpus generation model.
[0022] Furthermore, the optimization process of the High-class corpus generation model includes:
[0023] Optimize the trained High-class corpus generation model through the Eval LLM module using DPO to obtain the optimized High-class corpus generation model.
[0024] Furthermore, the optimization of the trained High-class corpus generation model through the Eval LLM module specifically includes:
[0025] Determine the prompting strategy; the prompting strategy includes pointwise-to-pairwise prompting and reference-to-reference-free prompting;
[0026] Based on the prompting strategy, design two corresponding evaluation paths to evaluate the output corpus of the trained High-class corpus generation model, obtaining two corresponding evaluation results;
[0027] Perform cross-validation on the two evaluation results. If the two evaluation results are consistent, complete the evaluation and use the evaluation result at this time as the final evaluation result; if the two evaluation results are inconsistent, re-evaluate through the Eval LLM module;
[0028] Generate a better text according to the final evaluation result;
[0029] Use the better text as the chosen of DPO and the output corpus of the trained High-class corpus generation model as the reject of DPO to optimize the trained High-class corpus generation model.
[0030] Furthermore, the two evaluation paths include the first evaluation path and the second evaluation path; expressed as:
[0031] The first evaluation path is expressed as:
[0032] The second evaluation path is expressed as:
[0033] Where D represents the path; point represents pointwise evaluation; pair represents pairwise evaluation; r represents reference evaluation; rf represents reference-free evaluation; f p2p represents pointwise-to-pairwise prompting; f R2RF represents reference-to-reference-free prompting.
[0034] Furthermore, the High-class corpus generation model includes a Fusion Block module and a T5 module;
[0035] The FusionBlock module includes a language feature extraction layer, a visual feature extraction layer, a single-head attention layer, and a gated fusion layer;
[0036] The output of the gated fusion layer is used as the input of the T5 module, and the corpus is output through the T5 module.
[0037] Furthermore, in the Fusion Block module:
[0038] (1) The target text data is subjected to feature extraction through the language feature extraction layer to obtain the corresponding language representation; specifically:
[0039] H language = LanguageEncoder(T user )
[0040] where H language is the language representation, and where n is the length of the language input; d is the hidden dimension; LanguageEncoder(·) is the language feature extraction layer; T user is the target text data;
[0041] (2) The m image data and the fusion feature are subjected to feature extraction through the visual feature extraction layer to obtain the corresponding patch-level visual representation, and the visual representation is converted into a language representation by using a projection matrix; specifically:
[0042] H vision = W h ·VisionExtractor(I g , F I )
[0043] where H vision is the patch-level visual representation, and where m is the number of patches; VisionExtractor(·) is the visual feature extraction layer; I g is the m image data; F I is the fusion feature; W h is the projection matrix;
[0044] (3) The language representation and the patch-level visual representation are associated through the single-head attention layer; among them, the query, key, and value in the single-head attention layer are H language , H vision and H vision respectively. Based on this, the output representation of the single-head attention layer is:
[0045]
[0046] Among them, is the output of the single-head attention layer; Q represents the query; K represents the key; V represents the value; d k represents the dimension of K;
[0047] (4) The gated fusion layer fuses the language representation and the patch-level visual representation, and inputs the output of the gated fusion layer into the T5 decoder; among them, the output of the gated fusion layer is expressed as:
[0048]
[0049] Among them, H fuse is the output of the gated fusion layer; W l and W v are both learnable parameters; λ is a trainable parameter, representing the ratio used to control the fusion of the language representation and the patch-level visual representation.
[0050] Furthermore, the language feature extraction layer is implemented based on the Transformer encoder. Specifically, the hidden state of the last layer in the Transformer encoder is used as the language representation.
[0051] Furthermore, it also includes: combining the multi-modal article generated by the Writer LLM module with m pieces of image data to obtain the final article with attached drawings.
[0052] It can be seen from the above technical solutions that, compared with the prior art, the present invention discloses a low-cost multi-modal article generation method, which has the following beneficial effects:
[0053] 1. The present invention innovatively expands the information volume of the image modality through its own image generation model (the Stable Diffusion-XL model fine-tuned by Lora), realizes expanding the information volume of the image modality through the local small-parameter LLM, and outputs the technical effect of high-quality prediction.
[0054] 2. The present invention performs DPO optimization on the T5 model after autoregressive training through the Eval LLM module to make the model output closer to high-quality corpus.
[0055] Next, through the drawings and embodiments, the technical solutions of the present invention will be further described in detail. Description of the Drawings
[0056] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the accompanying drawings required in the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can also be obtained based on the provided drawings.
[0057] Figure 1 Schematic flowchart of a low-cost multi-modal article generation method provided by an embodiment of the present invention.
[0058] Figure 2 Schematic diagram of a framework for generating a multi-modal article when the user input data is target text data and target image data provided by an embodiment of the present invention. Detailed implementation manners
[0059] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0060] An embodiment of the present invention discloses a low-cost multi-modal article generation method. Refer to Figure 1 as shown, which includes the following steps:
[0061] S1. Obtain user input data, where the user input data is target text data, or target text data and target image data;
[0062] S2. Input the target text data in the user input data into the StableDiffusion-XL model fine-tuned by Lora to generate m pieces of image data;
[0063] S3. Fuse the m pieces of image data, or fuse the target image data with the m pieces of image data to generate a fused feature;
[0064] S4. Input the target text data, the fused feature, and the m pieces of image data into the trained and optimized High-class corpus generation model respectively to output corresponding m pieces of target corpus;
[0065] S5. Input the m pieces of target corpus and the target text data into the local Writer LLM module to generate a multi-modal article;
[0066] S6. Combine the multi-modal article generated by the Writer LLM module with the m pieces of image data to obtain the final article with attached drawings.
[0067] Next, each of the above steps will be described separately.
[0068] In the above step S1, user input data is obtained, and the user input data is target text data T user , or, is target text data T user and target image data I user ; among them, the target image data Among them, is the nth target image in the target image data;
[0069] For current existing text-image dual-modal models, good performance can only be demonstrated when both text data and image data are input simultaneously; if only text data is input into the text-image dual-modal model, the performance of the text-image dual-modal model will drop sharply. In the embodiments of the present invention, whether both text data and image data are input simultaneously or only text data is input, ideal results can be obtained.
[0070] In the above step S2, the target text data T in the user input data user is input into the Stable Diffusion-XL model fine-tuned by LoRA to generate m pieces of image data; expressed as:
[0071] I g = SDXL(T user )
[0072] Among them, I g is an image data set including m pieces of image data, denoted as is the mth image generated based on the target text data;
[0073] Among them, the Stable Diffusion-XL model fine-tuned by LoRA refers to optimizing or adjusting the Stable Diffusion XL model using a technique called LoRA (Low-Rank Adaptation). StableDiffusion is a deep learning model mainly used for generating images, based on a text-to-image generation method. It can generate corresponding images according to the given text description.
[0074] LoRA is a model fine-tuning technique aimed at achieving efficient, low-rank adaptation to a pre-trained model by adding or adjusting a small number of parameters in the model, thereby improving the performance of the model on specific tasks or data sets while keeping the computational cost and storage requirements low. This method is particularly suitable for customizing large models in situations with limited resources.
[0075] In the above step S3, m pieces of image data are fused, or the target image data is fused with m pieces of image data to generate a fused feature; specifically:
[0076] (1) Fusing m pieces of image data is expressed as:
[0077]
[0078] (2) Fusing the target image data with m pieces of image data is expressed as:
[0079]
[0080] Among them, F I is the fused feature; λ m is a hyperparameter representing the weight.
[0081] In the above step S4, the target text data T user , the fused feature F I are respectively input into the trained and optimized High-class corpus generation model together with m pieces of image data, that is, each time the target text data T user , the fused feature F I are respectively input into the trained and optimized High-class corpus generation model together with a single piece of image data, and input m times to output the corresponding m pieces of target corpus; among them:
[0082] 1. The training process of the High-class corpus generation model includes:
[0083] Obtain a large amount of input text data and corresponding input image data;
[0084] Input the input text data into the Stable Diffusion-XL model fine-tuned by Lora to generate multiple pieces of image data;
[0085] Fuse multiple pieces of image data to generate a first fused feature;
[0086] Fuse multiple pieces of image data and the input image data to generate a second fused feature;
[0087] Take the input text data, multiple pieces of image data and the first fused feature as the first training data;
[0088] Take the input text data, multiple pieces of image data and the second fused feature as the second training data;
[0089] Based on the first training data and the second training data, train the High-class corpus generation model respectively to obtain a trained High-class corpus generation model.
[0090] 2. The optimization process of the High-class corpus generation model includes: directly performing preference optimization (DPO) on the trained High-class corpus generation model through the Eval LLM module to obtain an optimized High-class corpus generation model; specifically including:
[0091] (1) Determine the prompting strategy; the prompting strategy includes pointwise-to-pair prompting and reference-to-reference-free prompting; where:
[0092] Pointwise-to-pair prompting (f p2p ): This prompting strategy injects pointwise scoring comments of the generated text into pairwise comparison comments, thus enriching them with information about the respective text quality. At the same time, it requires self-reflection on the pointwise comments generated by the LLM before obtaining the final pairwise comparison result.
[0093] Reference-to-reference-free prompting (f R2RF ): This prompting strategy aims to eliminate direct comparison with references while retaining the information content in the references. It also requires the LLM to self-reflect on whether the evaluation results (including scores / labels and revised explanations) are consistent and modify the results if necessary.
[0094] (2) Based on the prompting strategy, design two corresponding evaluation paths to evaluate the output corpus of the trained High-class corpus generation model to obtain two corresponding evaluation results; where the two evaluation paths include the first evaluation path and the second evaluation path; expressed as:
[0095] The first evaluation path is expressed as:
[0096] The second evaluation path is expressed as:
[0097] where D represents the path; point represents pointwise evaluation; pair represents pairwise evaluation; r represents reference evaluation; rf represents reference-free evaluation; f p2p represents pointwise-to-pair prompting; f R2RF represents reference-to-reference-free prompting.
[0098] (3) Cross-validate the two evaluation results. If the two evaluation results are consistent, the evaluation is completed, and the evaluation result at this time is used as the final evaluation result. If the two evaluation results are inconsistent, re-evaluate through the Eval LLM module;
[0099] (4) Generate a better text according to the final evaluation result;
[0100] (5) Use the better text as the chosen of DPO, and use the output corpus of the trained High-class corpus generation model as the reject of DPO to optimize the trained High-class corpus generation model.
[0101] 3. High-class corpus generation model:
[0102] The above High-class corpus generation model includes a FusionBlock module and a T5 module; the Fusion Block module includes a language feature extraction layer, a visual feature extraction layer, a single-head attention layer, and a gated fusion layer; the output of the gated fusion layer is used as the input of the T5 module, and the corpus is output through the T5 module. Among them, the language feature extraction layer is implemented based on the Transformer encoder. Specifically, the hidden state of the last layer in the Transformer encoder is used as the language representation.
[0103] In the above Fusion Block module:
[0104] (1) Extract features from the target text data through the language feature extraction layer to obtain the corresponding language representation; specifically:
[0105] H language =LanguageEncoder(T user )
[0106] where H language is the language representation, and where n is the length of the language input; d is the hidden dimension; LanguageEncoder(·) is the language feature extraction layer; T user is the target text data;
[0107] (2) Extract features from the m image data and the fusion features through the visual feature extraction layer to obtain the corresponding patch-level visual representation, and use the projection matrix to convert the visual representation into a language representation; specifically:
[0108] H vision =W h ·VisionExtractor(Ig , F I )
[0109] Among them, H vision is the patch-level visual representation, and where m is the number of patches; VisionExtractor(·) is the visual feature extraction layer; I g is the m image data; F I is the fused feature; W h is the projection matrix;
[0110] Among them, VisionExtractor(·) is used to vectorize the input image into visual features. We extract features by freezing the visual extraction model. After obtaining the patch-level visual features, we apply the learnable projection matrix W h to convert the shape of VisionExtractor(I g , F I ) to the shape of H language ;
[0111] (3) Associate the language representation with the patch-level visual representation through a single-head attention layer; among them, the query, key, and value in the single-head attention layer are H language , H vision , and H vision , respectively. Based on this, the output representation of the single-head attention layer is:
[0112]
[0113] Among them, is the output of the single-head attention layer; Q represents the query; K represents the key; V represents the value; d k represents the dimension of K; since single-head attention is used, d k is the same as the dimension of H language .
[0114] (4) Fuse the language representation and the patch-level visual representation through a gated fusion layer, and input the output of the gated fusion layer into the T5 decoder; among them, the output of the gated fusion layer is represented as:
[0115]
[0116] Among them, H fuse is the output of the gated fusion layer; W l and W v are both learnable parameters; λ is a trainable parameter, representing the ratio used to control the fusion of the language representation and the patch-level visual representation.
[0117] In the above step S5, m pieces of target corpus and target text data are input into the local Writer LLM module to generate a multimodal article. Figure 2 It is a schematic diagram of the framework for generating a multimodal article when the user input data is target text data and target image data.
[0118] In the above step S6, the multimodal article generated by the Writer LLM module is combined with m pieces of image data to obtain the final article with attached drawings.
[0119] In summary, the existing methods mainly generate articles by calling an online LLM or fine-tuning a local LLM. However, a low-cost multimodal article generation method provided by the embodiments of the present invention only needs to fine-tune the T5 model in the High-class corpus generation module, and then a local LLM that does not require fine-tuning can be used to achieve it.
[0120] The various embodiments in this specification are described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. The same or similar parts among the various embodiments can be referred to each other.
[0121] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be obvious to those skilled in the art. The general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown herein, but will be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A low-cost multimodal article generation method, characterized in that: The steps include: Acquire user input data, wherein the user input data is target text data, or target text data and target image data; Input the target text data in the user input data into the Stable Diffusion-XL model fine-tuned by Lora to generate m images; Fusing m image data, or fusing target image data with m image data to generate fusion features; Input the target text data, fusion features and m image data into the trained and optimized High-classcorpus generation model, and output the corresponding m target corpus; Input m target corpora and target text data into the local Writer LLM module to generate multimodal articles.
2. A low-cost multimodal article generation method according to claim 1, characterized in that: The training process of the High-class corpus generation model includes: Obtaining a large amount of input text data and corresponding input image data; Input text data into the Stable Diffusion-XL model fine-tuned by Lora to generate multiple image data; Fusing multiple image data to generate a first fusion feature; Fusing the plurality of image data with the input image data to generate a second fusion feature; Using the input text data, the plurality of image data and the first fusion feature as first training data; Using the input text data, the plurality of image data and the second fusion feature as the second training data; Based on the first training data and the second training data, the High-class corpus generation model is trained respectively to obtain a trained High-class corpus generation model.
3. A low-cost multimodal article generation method according to claim 2, characterized in that: The optimization process of the High-class corpus generation model includes: The trained High-class corpus generation model is optimized by DPO through the Eval LLM module to obtain the optimized High-class corpus generation model.
4. A low-cost multimodal article generation method according to claim 3, characterized in that: The DPO optimization of the trained High-class corpus generation model through the Eval LLM module specifically includes: Determine the cueing strategy; cueing strategies range from point-by-point to paired cueing, and from reference to no-reference cueing; Based on the prompt strategy, two corresponding evaluation paths are designed to evaluate the output corpus of the trained High-class corpus generation model, and two corresponding evaluation results are obtained; The two evaluation results are cross-validated. If the two evaluation results are consistent, the evaluation is completed and the evaluation result at this time is used as the final evaluation result. If the two evaluation results are inconsistent, re-evaluation is performed through the Eval LLM module. Generate better text based on the final evaluation results; The better text is used as the chosen one of DPO, and the output corpus of the trained High-class corpus generation model is used as the rejected one of DPO, so as to optimize the trained High-class corpus generation model.
5. A low-cost multimodal article generation method according to claim 4, characterized in that: The two evaluation paths include a first evaluation path and a second evaluation path; expressed as: The first evaluation path is expressed as: The second evaluation path is expressed as: Where D represents the path; point represents point-by-point evaluation; pair represents paired evaluation; r represents reference evaluation; rf represents no-reference evaluation; f p2p Indicates point-by-point to paired prompts; f R2RF Indicates reference to no reference prompt.
6. A low-cost multimodal article generation method according to claim 1, characterized in that: The High-class corpus generation model includes a FusionBlock module and a T5 module; The FusionBlock module includes a language feature extraction layer, a visual feature extraction layer, a single-head attention layer and a gated fusion layer; The output of the gated fusion layer is used as the input of the T5 module, and the corpus is output through the T5 module.
7. A low-cost multimodal article generation method according to claim 6, characterized in that: In the Fusion Block module: (1) extracting features from the target text data through a language feature extraction layer to obtain corresponding language representation; specifically: H language =LanguageEncoder(T user ) Among them, H language is a language representation, and Where n is the length of the language input; d is the hidden dimension; LanguageEncoder(·) is the language feature extraction layer; T user is the target text data; (2) extracting features from the m image data and fusion features through a visual feature extraction layer to obtain corresponding patch-level visual representations, and converting the visual representations into language representations using a projection matrix; specifically: H vision =W h ·VisionExtractor(I g ,F I ) Among them, H vision is a patch-level visual representation, and Where m is the number of patches; VisionExtractor(·) is the visual feature extraction layer; I g is m images data; F I is the fusion feature; W h is the projection matrix; (3) The language representation and patch-level visual representation are associated through a single-head attention layer; the query, keyword, and value in the single-head attention layer are H language , H vision and H vision , based on this, the output of the single-head attention layer is expressed as: in, is the output of the single-head attention layer; Q represents the query; K represents the keyword; V represents the value; d k represents the dimension of K; (4) The language representation and the patch-level visual representation are fused through the gated fusion layer, and the output of the gated fusion layer is input into the T5 decoder; the output of the gated fusion layer is expressed as: Among them, H fuse is the output of the gated fusion layer; W l and W v are all learnable parameters; λ is a trainable parameter, which is used to control the ratio of the fusion of language representation and patch-level visual representation.
8. A low-cost multimodal article generation method according to claim 6, characterized in that: The language feature extraction layer is implemented based on the Transformer encoder. Specifically, the hidden state of the last layer in the Transformer encoder is used as the language representation.
9. A low-cost multimodal article generation method according to claim 1, characterized in that: Also includes: Combine the multimodal article generated by the Writer LLM module with the m image data to obtain the final article with illustrations.