Multimodal controllable text generation method and device, electronic equipment, storage medium and program product
Through multiple iterations of text generation reasoning and visual information matching, the problems of high training cost and unsmooth text semantics in existing technologies are solved, and the matching of text and visual information and high-level multimodal control are achieved.
Patent Information
- Application Number
- CN202411416813.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-11
- Publication Date
- 2025-10-10
- Estimated Expiration
- 2044-10-11
AI Technical Summary
The controllable text generation methods in the existing technology have high training costs, the generated text context is not semantically fluent, and it is impossible to match text with visual information, and it is impossible to achieve a higher level of multimodal control.
By determining the text control elements and visual control elements, at least two text generation reasonings are performed to obtain the intermediate results of text generation. The text generation is iterated through the thought chain prompts until the results converge, and the large language model and contrastive language image model are used to match the text and visual information.
The model training cost is reduced, the generated text context is semantically fluent, and the matching of text and visual information is achieved, achieving a higher level of multimodal control.
Smart Images

Figure CN119514501B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the technical field of natural language generation, and in particular to a multimodal controllable text generation method, device, electronic device, storage medium, and program product. Background Art
[0002] This section is intended to provide a background or context to the embodiments of the present disclosure that are recited in the claims. No statement herein is admitted to be prior art by virtue of its inclusion in this section.
[0003] Controllable text generation refers to the process of controlling the specific attributes or content of the generated text by introducing certain constraints or guiding factors during the text generation process. This technology allows users to specify the style, theme, emotion and other characteristics of the generated text, so that the generated text better meets the expected needs.
[0004] However, in related technologies, controllable text generation methods have high training costs and are expensive, the generated text context is not semantically fluent, and the text cannot be matched with visual information, thus failing to achieve a higher level of multimodal control. Summary of the Invention
[0005] In view of this, the purpose of the present disclosure is to propose a multimodal controllable text generation method, device, electronic device, storage medium and program product, which at least to a certain extent solve one of the technical problems in the related art.
[0006] Based on the above objectives, a first aspect of an exemplary embodiment of the present disclosure provides a multimodal controllable text generation method, which is applied to a server. The method includes:
[0007] Identify textual and visual control elements;
[0008] Based on the text control element and the visual control element, perform at least two text generation reasonings to obtain at least two text generation intermediate results;
[0009] Splicing the intermediate results generated from the at least two texts to obtain a thought chain prompt;
[0010] Based on the text control elements, the thinking chain prompts and the visual control elements, text generation reasoning is iteratively performed until the text generation intermediate results converge, and the converged text generation intermediate results are used as the text generation results, wherein the text generation intermediate results output from this round of iteration are spliced into the thinking chain prompts input from this round of iteration as the thinking chain prompts input from the next round of iteration.
[0011] Based on the same inventive concept, a second aspect of the exemplary embodiments of the present disclosure provides a multimodal controllable text generation device, comprising:
[0012] an element determination module configured to determine text control elements and visual control elements;
[0013] an intermediate result determination module, configured to perform at least two text generation inferences based on the text control element and the visual control element to obtain at least two text generation intermediate results;
[0014] a thought chain prompt determining module, configured to splice intermediate results generated from the at least two texts to obtain a thought chain prompt;
[0015] The iteration module is configured to iteratively perform text generation reasoning based on the text control elements, the thinking chain prompts and the visual control elements until the text generation intermediate results converge, and use the converged text generation intermediate results as the text generation results, wherein the text generation intermediate results output by this round of iteration are spliced into the thinking chain prompts input by this round of iteration as the thinking chain prompts input by the next round of iteration.
[0016] Based on the same inventive concept, the third aspect of the exemplary embodiment of the present disclosure provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the method described in the first aspect is implemented.
[0017] Based on the same inventive concept, a fourth aspect of the exemplary embodiment of the present disclosure provides a non-transitory computer-readable storage medium, wherein the non-transitory computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a computer to execute the method described in the first aspect.
[0018] Based on the same inventive concept, a fifth aspect of the exemplary embodiments of the present disclosure provides a computer program product, including computer program instructions. When the computer program instructions are executed on a computer, the computer is caused to execute the method described in the first aspect.
[0019] From the above, it can be seen that the embodiments of the present disclosure provide a multimodal controllable text generation method, device, electronic device, storage medium and program product, the method comprising: determining a text control element and a visual control element, performing at least two text generation reasonings based on the text control element and the visual control element, obtaining at least two text generation intermediate results, splicing the at least two text generation intermediate results to obtain a thinking chain prompt, iteratively performing text generation reasoning based on the text control element, the thinking chain prompt and the visual control element until the text generation intermediate result converges, and using the converged text generation intermediate result as the text generation result, wherein the text generation intermediate result output by this round of iteration is spliced into the thinking chain prompt input by this round of iteration as the thinking chain prompt input by the next round of iteration. The present disclosure has low training cost, the generated text context semantics are fluent, and the text can be matched with visual information, thereby achieving a higher level of multimodal control. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] In order to more clearly illustrate the technical solutions in the present disclosure or related technologies, the following briefly introduces the drawings required for use in the embodiments or related technical descriptions. Obviously, the drawings described below are only embodiments of the present disclosure. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0021] Figure 1 A schematic diagram of an application scenario of the multimodal controllable text generation method provided by an exemplary embodiment of the present disclosure;
[0022] Figure 2 A schematic diagram of a process for generating multimodal controllable text provided by an exemplary embodiment of the present disclosure;
[0023] Figure 3 A schematic structural diagram of a multimodal controllable text generation device provided by an exemplary embodiment of the present disclosure;
[0024] Figure 4 A schematic diagram of the hardware structure of an electronic device according to an embodiment of the present disclosure;
[0025] Figure 5 A schematic diagram of a graphical user interface for generating multimodal controllable text provided by an exemplary embodiment of the present disclosure;
[0026] Figure 6 A schematic diagram of a graphical user interface in simplified output mode for generating multimodal controllable text provided by an exemplary embodiment of the present disclosure;
[0027] Figure 7A schematic diagram of a graphical user interface in detailed output mode for generating multimodal controllable text provided by an exemplary embodiment of the present disclosure. DETAILED DESCRIPTION
[0028] It is understandable that before using the technical solutions disclosed in the embodiments of this application, the type, scope of use, usage scenarios, etc. of the personal information involved in this application should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.
[0029] For example, in response to receiving a user's active request, a prompt message is sent to the user to clearly inform the user that the operation requested will require the acquisition and use of the user's personal information. Thus, based on the prompt message, the user can independently choose whether to provide personal information to the electronic device, application, server, storage medium, or other software or hardware that performs the operation of the technical solution of this application.
[0030] As an optional but non-limiting implementation, in response to receiving a user's active request, the prompt information may be sent to the user in the form of a pop-up window, in which the prompt information may be presented in text form. Furthermore, the pop-up window may also contain a selection control for the user to select "agree" or "disagree" to provide personal information to the electronic device.
[0031] It is understandable that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this application. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this application.
[0032] It is understandable that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) must comply with the requirements of relevant laws, regulations and relevant provisions.
[0033] To make the objectives, technical solutions, and advantages of the present disclosure more clearly understood, the principles and spirit of the present disclosure will be described below with reference to several exemplary embodiments. It should be understood that these embodiments are provided solely to enable those skilled in the art to better understand and implement the present disclosure, and are not intended to limit the scope of the present disclosure in any way. On the contrary, these embodiments are provided to make the present disclosure more thorough and complete, and to fully convey the scope of the present disclosure to those skilled in the art.
[0034] It should be understood herein that any number of elements in the drawings is for illustration only and not for limitation, and any naming is only for distinction and does not have any limiting meaning.
[0035] It should be noted that, unless otherwise defined, the technical terms or scientific terms used in the embodiments of the present disclosure should have the usual meanings understood by people with ordinary skills in the field to which the present disclosure belongs. The "first", "second" and similar words used in the embodiments of the present disclosure do not indicate any order, quantity or importance, but are only used to distinguish different components. "Include" or "comprise" and similar words mean that the elements or objects appearing before the word include the elements or objects listed after the word and their equivalents, without excluding other elements or objects. "Connect" or "connected" and similar words are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. "Up", "down", "left", "right" and the like are only used to indicate relative positional relationships. When the absolute position of the described object changes, the relative positional relationship may also change accordingly. The article "one" or "an" before an element does not exclude the presence of multiple such elements.
[0036] The principles and spirit of the present disclosure are explained in detail below with reference to several representative embodiments of the present disclosure.
[0037] As described in the background technology, in the related art, the training cost of multimodal controllable text generation method is high and expensive, the generated text context semantics are not smooth, and it is impossible to match the text with visual information to achieve a higher level of multimodal control.
[0038] Specifically, in the related art, a controllable text generation method based on a large language model, first, constructs an opinion database based on opinion data. Then, the first hot event of the first media platform and the first hot content of the first hot event are input into the hot content summary model to obtain the first summary content of the first hot event. Then, based on the opinion database, the first post of the first hot event and the first summary content of the first hot event, the post filtering model is used to filter the first post of the first hot event that is inconsistent with the opinion database. Finally, the filtered first post of the first hot event, the first summary content of the first hot event and the target user group information are input into the text generation large model to obtain controllable text.
[0039] However, in order to generate more controllable text, this method may need to combine multiple models and algorithms. Although it improves the accuracy and efficiency of the generated text, it requires a lot of computing resources.
[0040] In related technologies, a zero-shot multimodal controllable text generation method uses ZeroGen. During the inference phase, for text control, ZeroGen first calculates the cosine similarity between each candidate token and the text control word and integrates this into the original probability distribution to achieve text control. Then, for visual control, it uses CLIP to evaluate the similarity between k candidate tokens and k candidate sentences and images concatenated from the generated text, and further integrates this probability distribution into the original probability distribution.
[0041] However, this method, CLIP, calculates similarity at the image feature level rather than directly assessing the association between text and image, which can lead to less precise text control. Furthermore, CLIP trains the model by maximizing the similarity of positive samples and minimizing the similarity of negative samples, but this approach may not be fully adaptable to all types of text and image data. These factors can affect ZeroGen's performance and accuracy in multimodal controllable text generation, resulting in semantically unsmooth generated text contexts.
[0042] In related technology, a controllable text generation method based on a pretrained language model is proposed. The core of this method is to train a discriminator model for topics, sentiment, and writing style. This model then uses a Bayesian probabilistic decomposition to combine the output probabilities of the pretrained language model and the discriminator model. Compared to models trained solely to meet constraints, this method does not require any modifications to the pretrained language model itself. Instead, it uses an attribute discriminator during the model inference phase to guide the model in generating content that meets the constraints.
[0043] However, while this approach saves computational resources for training large-scale pre-trained language models, it fails to match text with visual information to achieve a higher level of multimodal control.
[0044] To solve the above problems, the present disclosure provides a multimodal controllable text generation method, device, electronic device, storage medium, and program product solution, specifically including:
[0045] Determine text control elements and visual control elements, perform at least two text generation reasonings based on the text control elements and the visual control elements, obtain at least two text generation intermediate results, splice the at least two text generation intermediate results to obtain a thinking chain prompt, iterate and perform text generation reasoning based on the text control elements, the thinking chain prompt and the visual control elements until the text generation intermediate results converge, and use the converged text generation intermediate results as the text generation results, wherein the text generation intermediate results output from this round of iteration are spliced into the thinking chain prompt input from this round of iteration as the thinking chain prompt input for the next round of iteration. This solution makes full use of the pre-training capabilities of the large language model and does not require additional large-scale data training, thereby reducing the cost and time of model training. Secondly, this solution uses the powerful natural language generation capabilities of the large language model to guide text generation through thought chain prompts to ensure that the generated text is fluent and consistent with the contextual semantics. In addition, this solution combines the image-text matching model in visual control and generates chain thought prompts through multiple iterations to gradually improve the matching and consistency between text and visual information. Referring to Table 1, when the chain thought prompts are removed, the BLEU-1, BLEU-4, ROUGE-L and CIDEr scores decrease slightly. BLEU-1, BLEU-4, ROUGE-L and CIDEr are four commonly used evaluation indicators in the field of natural language processing (NLP), which are used for different tasks and application scenarios respectively. This shows that the model has a decreased accuracy in generating text that is consistent with the given control content. Secondly, the cyclic iterative generation is removed. This means that the generation process has only one round. At the same time, the thought chain prompts are also removed because it is impossible to construct a thought chain prompt when only one generation is performed. After removing the cyclic iterative generation, all indicators have dropped significantly, which shows that iterative generation is crucial to improving the accuracy and quality of the output. Therefore, by generating chain thought prompts through multiple iterations, a higher level of multimodal control can be achieved.
[0046] Table 1 Natural language processing index evaluation table
[0047]
[0048] After introducing the basic principles of the present disclosure, various non-limiting embodiments of the present disclosure are described in detail below.
[0049] refer to Figure 1 , which is a schematic diagram of an application scenario of the multimodal controllable text generation method provided by an exemplary embodiment of the present disclosure.
[0050] This application scenario includes a terminal device 101 and a server 102. The terminal device 101 and the server 102 can be connected via a wired or wireless communication network to achieve data interaction.
[0051] The terminal device 101 can be an electronic device close to the user side with data transmission, multimedia input / output functions, including but not limited to a desktop computer, a mobile phone, a mobile computer, a tablet computer, a media player, a smart wearable device, a personal digital assistant (PDA), or other electronic devices capable of realizing the above functions, etc. The electronic device can include a processor and a display screen with touch input function, the display screen is used to present a graphical user interface, the graphical user interface can display an application interface, the processor is used to process application data, generate a graphical user interface, and control the display of the graphical user interface on the display screen.
[0052] The server 102 can be a standalone physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN (Content Delivery Network), and basic cloud computing services such as big data and artificial intelligence platforms.
[0053] In some example embodiments, the multi-modal controllable text generation method can run on the terminal device 101 or the server 102.
[0054] When the multi-modal controllable text generation method runs on the server 102, the server 102 is used to provide a multi-modal controllable text generation service to the user of the terminal device 101.
[0055] The terminal device 101 determines the text control element and the visual control element and sends them to the server 102;
[0056] The server 102 accepts the text control element and the visual control element sent by the terminal device, and the server 102 performs at least two text generation inferences based on the text control element and the visual control element, to obtain at least two text generation intermediate results;
[0057] The server 102 splices the at least two text generation intermediate results to obtain a thought chain prompt;
[0058] The server 102 iteratively performs text generation reasoning based on the text control elements, the thinking chain prompts and the visual control elements, and sends several output results of the generated reasoning to the terminal device 101. The server 102 continues until the text generation intermediate results converge, and then uses the converged text generation intermediate results as the text generation results and sends the text generation results to the terminal device 101. The server 102 splices the text generation intermediate results output from this round of iteration into the thinking chain prompts input from this round of iteration as the thinking chain prompts input for the next round of iteration.
[0059] refer to Figure 2 , a multimodal controllable text generation method is applied to a server, and the method comprises the following steps:
[0060] Step S210: Determine text control elements and visual control elements.
[0061] The following describes how to determine text control elements and visual control elements:
[0062] In this exemplary embodiment, a graphical user interface is provided by a terminal, and the content displayed by the graphical user interface includes an input interface and an output interface, and the input interface includes a text control control and a visual control control;
[0063] Identify textual and visual control elements, including:
[0064] In response to an input operation on the text control control, the input text control element is acquired; in response to an input operation on the visual control control, the input visual control element is acquired.
[0065] During specific implementation, determine the text control elements and visual control elements, including:
[0066] As a specific example, refer to Figure 5 The content displayed by the graphical user interface includes an input interface and an output interface. The input interface includes a text control control and a visual control control. The user can input natural language through the text control control or click on the text control control to select preset text options, such as emotional text options (such as positive, negative, etc.) or thematic text options (such as scenery, animals or plants, etc.); the user can click on the visual control control to select files such as photos or videos from the local computer and upload them to the input interface.
[0067] Step S220: Based on the text control element and the visual control element, perform at least two text generation reasonings to obtain at least two text generation intermediate results.
[0068] The following describes how to obtain each text generation inference in at least two intermediate text generation results:
[0069] In this exemplary embodiment, at least two text generation inferences are performed based on the text control element and the visual control element to obtain each of at least two text generation intermediate results, including:
[0070] A large language model is used to make inferences based on the text control elements to obtain several candidate texts, the several candidate texts are matched with the visual control elements based on a contrastive language image model, and one of the candidate texts is used as the text to generate an intermediate result.
[0071] In specific implementation, the comparison of language image models includes:
[0072] The Contrastive Language Image Processing (CLIP) is a multimodal pre-trained neural network model designed to associate images and text through contrastive learning. Its core idea is to pre-train the model using a large amount of unlabeled image and text data, thereby achieving mutual understanding and alignment between images and text.
[0073] In the above exemplary embodiment, a method for obtaining each text generation inference in at least two text generation intermediate results is introduced. The following describes a method for obtaining several candidate texts:
[0074] In this exemplary embodiment, a large language model is used to perform inference based on the text control elements to obtain several candidate texts, including:
[0075] Determine several natural language instructions of the text control element, use the several natural language instructions as observation parameters of one of the several candidate texts based on the large language model, update the probability distribution of the observation parameters, obtain the posterior probability of one of the several candidate texts, perform probability calculation on the posterior probability, and obtain one of the several candidate texts.
[0076] In specific implementation, the large language model includes:
[0077] Large Language Models (LLMs) are deep learning-based AI models designed to understand and generate human language. Trained on large amounts of text data, they are capable of performing a wide range of natural language processing (NLP) tasks, such as translating text, answering questions conversationally, and classifying and generating words based on knowledge gained from diverse datasets.
[0078] In a specific implementation, determining a plurality of natural language instructions of the text control element, and using the plurality of natural language instructions as observation parameters of one of the plurality of candidate texts based on the large language model, includes:
[0079] Through the natural language instructions input by the user on the terminal device, the server can understand the natural language instructions based on the large language model, representing the natural language instructions of text control. 1:n Used to directly represent text controls.
[0080] In a specific implementation, updating the probability distribution of the observation parameters to obtain the posterior probability of one of the candidate texts includes:
[0081] Due to the powerful ability of large language models in understanding natural language instructions, in order to maximize P(X 1:n |C T ) to search for X 1:n , which can be regarded as maximizing P(X 1:n |I 1:n ), which can be proved by the following calculation formula:
[0082] According to Bayes' theorem:
[0083]
[0084] Thanks to the powerful ability of large language models in understanding natural language instructions, these models can effectively and accurately generate text containing the required text control content. Therefore, P(I|C T ) and P(C T |I) are close to 1, so they can be regarded as 1. Therefore,
[0085]
[0086] As a specific example, suppose that we want to prompt the model to generate positive text. By inputting the sentence "Generate some positive text" into the large language model, it can effectively and accurately generate the required positive text.
[0087] In a specific implementation, the posterior probability is calculated to obtain a candidate text from the plurality of candidate texts, including:
[0088] By comparing the posterior probabilities of the candidate texts, the candidate text with the largest posterior probability is selected as the output.
[0089] In the above exemplary embodiment, a method of obtaining several candidate texts is introduced. The following describes a method of obtaining texts and generating intermediate results:
[0090] In this exemplary embodiment, matching the candidate texts with the visual control element based on the contrastive language image model and using one of the candidate texts as the text to generate an intermediate result includes:
[0091] The matching value between the visual control element and each candidate text in the plurality of candidate texts is determined respectively; and the candidate text with the largest matching value is used as the text to generate an intermediate result.
[0092] In a specific implementation, determining the matching value between the visual control element and each candidate text in the plurality of candidate texts includes:
[0093] Several candidate texts and visual control elements are input into a comparative language image model (CLIP) model, and matching scores between the image and text are output to obtain scores of matching between the candidate texts and the visual control elements.
[0094] In a specific implementation, the candidate text with the largest matching value is used as the text to generate an intermediate result, including:
[0095] Select the text with the highest score from several candidate texts and use it as the current output.
[0096] Step S230: splicing the intermediate results generated by the at least two texts to obtain a thought chain prompt.
[0097] Here are some tips on thinking chain:
[0098] When implementing it, the thought chain prompts include:
[0099] The thought chaining prompt is an improved prompting strategy for improving the performance of large language models (LLMs) in complex reasoning tasks. It achieves this goal by guiding the model to generate a series of intermediate reasoning steps, thereby helping the model better understand and solve complex arithmetic, common sense, and symbolic reasoning problems.
[0100] In the above exemplary embodiment, the thought chain prompt is introduced. The following describes how to obtain the thought chain prompt:
[0101] In a specific implementation, the step of generating intermediate results from at least two texts and splicing them together to obtain a thought chain prompt includes:
[0102] According to the generation order of the at least two text generation intermediate results, the at least two text generation intermediate results are spliced to obtain the thought chain prompt.
[0103] Step S240: Based on the text control elements, the thinking chain prompts and the visual control elements, iteratively perform text generation reasoning until the text generation intermediate results converge, and use the converged text generation intermediate results as the text generation results, wherein the text generation intermediate results output by this round of iteration are spliced into the thinking chain prompts input by this round of iteration as the thinking chain prompts input by the next round of iteration.
[0104] In this exemplary embodiment, step S240 specifically includes:
[0105] Inputting the text control elements and the first thought chain prompt into the large language model, and using the large language model to make inferences based on the text control elements and the first thought chain prompt to obtain a number of candidate texts;
[0106] Inputting the visual control element and several candidate texts into a contrastive language image model, matching the several candidate texts with the visual control element based on the contrastive language image model, and taking one of the candidate texts as a first text to generate an intermediate result;
[0107] Splicing the intermediate result generated by the first text and the first thought chain prompt to obtain the second thought chain prompt;
[0108] Inputting the text control elements and the second thought chain prompts into the large language model, and using the large language model to make inferences based on the text control elements and the second thought chain prompts to obtain several candidate texts;
[0109] Inputting the visual control element and several candidate texts into a contrastive language image model, matching the several candidate texts with the visual control element based on the contrastive language image model, and taking one of the candidate texts as a second text to generate an intermediate result;
[0110] The intermediate result generated by the first text is spliced with the second thinking chain prompt to obtain the third thinking chain prompt;
[0111] Inputting the text control elements and the third thought chain prompts into the large language model, and using the large language model to make inferences based on the text control elements and the third thought chain prompts to obtain several candidate texts;
[0112] Inputting the visual control element and several candidate texts into a contrastive language image model, matching the several candidate texts with the visual control element based on the contrastive language image model, and taking one of the candidate texts as a third text to generate an intermediate result;
[0113] …
[0114] And so on, the loop is iterated until the intermediate result of text generation converges, and the converged intermediate result of text generation is used as the text generation result.
[0115] In implementation, whether the model converges is determined by comparing the similarity of the model output at different time points. If the model gives almost the same output for the same input after multiple training, it can be considered that the model has converged.
[0116] Next, the visualization method in the process of obtaining the text generation result is introduced.
[0117] In the example embodiment, based on the text control element, the thought chain prompt and the visual control element, the text generation inference is iteratively performed until the text generation intermediate result converges, and the converged text generation intermediate result is taken as the text generation result, including:
[0118] The text generation intermediate result generated in each iteration is sequentially displayed on the output interface, and the text generation result is displayed on the output interface.
[0119] In implementation, the text generation intermediate result generated in each iteration is sequentially displayed on the output interface, and the text generation result is displayed on the output interface, including:
[0120] In the embodiment of step S210, refer to Figure 5 When the user inputs the text control element through the text control control and inputs the visual control element through the visual control control in the input interface displayed by the graphical user interface, the user can click the text generation control in the input interface to display the output result of each iteration and the final generation result generated after the model converges in the output interface displayed by the graphical user interface. The user can select between the brief mode and the detailed mode of the output interface through the mode switching control in the output interface. When the text output is in the brief mode, refer to Figure 6 , the output interface only displays the output result of each iteration and the final generation result. When the text output is in the detailed mode, refer to Figure 7 , the output interface displays the input thought chain prompt and the generated candidate of each iteration in addition to the output result of each iteration and the final generation result.
[0121] It should be noted that the method of the embodiment of the present disclosure can be executed by a single device, such as a computer or a server. The method of the embodiment can also be applied to a distributed scenario and completed by multiple devices cooperating with each other. In this distributed scenario, one of the multiple devices can only execute one or more steps in the method of the embodiment of the present disclosure, and the multiple devices can interact with each other to complete the method.
[0122] It should be noted that the above description is limited to some embodiments of the present disclosure. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in an order different from that described in the above embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order or sequential order shown to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0123] Based on the same inventive concept, corresponding to any of the above-mentioned embodiment methods, the present disclosure also provides a multimodal controllable text generation device.
[0124] refer to Figure 3 , the multimodal controllable text generation device comprises:
[0125] An element determination module 310 is configured to determine text control elements and visual control elements;
[0126] The intermediate result determination module 320 is configured to perform at least two text generation inferences based on the text control element and the visual control element to obtain at least two text generation intermediate results;
[0127] The thought chain prompt determining module 330 is configured to splice the intermediate results generated by the at least two texts to obtain a thought chain prompt;
[0128] The iteration module 340 is configured to iteratively perform text generation reasoning based on the text control elements, the thinking chain prompts and the visual control elements until the text generation intermediate results converge, and use the converged text generation intermediate results as the text generation results, wherein the text generation intermediate results output from this round of iteration are spliced into the thinking chain prompts input from this round of iteration as the thinking chain prompts input from the next round of iteration.
[0129] In this exemplary embodiment, the element determination module 310 is specifically configured to:
[0130] In response to an input operation on the text control control, the input text control element is acquired; in response to an input operation on the visual control control, the input visual control element is acquired.
[0131] In this exemplary embodiment, the intermediate result determination module 320 is specifically configured to:
[0132] Determine several natural language instructions of the text control element, use the several natural language instructions as observation parameters of one of the several candidate texts based on the large language model, update the probability distribution of the observation parameters, obtain the posterior probability of one of the several candidate texts, perform probability calculation on the posterior probability to obtain one of the several candidate texts, determine the matching value between the visual control element and each of the several candidate texts, and use the candidate text with the largest matching value as the text to generate an intermediate result.
[0133] In this exemplary embodiment, the thought chain prompt determination module 330 is specifically configured to:
[0134] According to the generation order of the at least two text generation intermediate results, the at least two text generation intermediate results are spliced to obtain the thought chain prompt.
[0135] In this exemplary embodiment, the iteration module 340 is specifically configured to:
[0136] Based on the text control elements, the thinking chain prompts and the visual control elements, text generation reasoning is performed iteratively, and the text generation intermediate results generated in each iteration are displayed in sequence on the output interface until the text generation intermediate results converge, and the converged text generation intermediate results are used as text generation results, and the text generation results are displayed on the output interface, wherein the text generation intermediate results outputted in this round of iteration are spliced into the thinking chain prompts inputted in this round of iteration as the thinking chain prompts inputted in the next round of iteration.
[0137] For the convenience of description, the above devices are described as being functionally divided into various modules. Of course, when implementing the present disclosure, the functions of each module can be implemented in the same or multiple software and / or hardware.
[0138] The device of the above embodiment is used to implement the corresponding multimodal controllable text generation method in any of the above embodiments, and has the beneficial effects of the corresponding method embodiment, which will not be repeated here.
[0139] Based on the same inventive concept, corresponding to any of the above-mentioned embodiments and methods, the present disclosure also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and runnable on the processor, wherein when the processor executes the program, the multimodal controllable text generation method described in any of the above embodiments is implemented.
[0140] Figure 410 is a schematic diagram showing a more specific hardware structure of an electronic device provided in this embodiment. The device may include: a processor 1010, a memory 1020, an input / output interface 1030, a communication interface 1040, and a bus 1050. The processor 1010, the memory 1020, the input / output interface 1030, and the communication interface 1040 are communicatively connected to each other within the device via the bus 1050.
[0141] The processor 1010 can be implemented using a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this specification.
[0142] The memory 1020 can be implemented in the form of ROM (Read Only Memory), RAM (Random Access Memory), static storage devices, dynamic storage devices, etc. The memory 1020 can store an operating system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 1020 and is called and executed by the processor 1010.
[0143] The input / output interface 1030 is used to connect an input / output module to implement information input and output. The input / output module can be configured as a component in the device (not shown in the figure) or can be externally connected to the device to provide corresponding functions. Input devices may include a keyboard, mouse, touch screen, microphone, various sensors, etc., and output devices may include a display, speaker, vibrator, indicator light, etc.
[0144] The communication interface 1040 is used to connect to a communication module (not shown) to enable communication between the device and other devices. The communication module can communicate via a wired method (such as USB, network cable, etc.) or a wireless method (such as mobile network, WiFi, Bluetooth, etc.).
[0145] The bus 1050 comprises a path for transmitting information between the various components of the device (eg, the processor 1010 , the memory 1020 , the input / output interface 1030 , and the communication interface 1040 ).
[0146] It should be noted that although the above device only shows the processor 1010, the memory 1020, the input / output interface 1030, the communication interface 1040, and the bus 1050, in a specific implementation, the device may also include other components necessary for normal operation. In addition, it will be understood by those skilled in the art that the above device may only include the components necessary to implement the embodiments of this specification, and does not necessarily include all the components shown in the figure.
[0147] The electronic device of the above embodiment is used to implement the corresponding multimodal controllable text generation method in any of the above embodiments, and has the beneficial effects of the corresponding method embodiment, which will not be repeated here.
[0148] Based on the same inventive concept, corresponding to any of the above-mentioned embodiment methods, the present disclosure also provides a non-transitory computer-readable storage medium, which stores computer instructions, and the computer instructions are used to enable the computer to execute the multimodal controllable text generation method described in any of the above embodiments.
[0149] The computer-readable media of this embodiment include permanent and non-permanent, removable and non-removable media that can be used to store information by any method or technology. The information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, read-only compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device.
[0150] The above-mentioned non-transitory computer-readable storage medium can be any available medium or data storage device that can be accessed by a computer, including but not limited to magnetic storage (such as floppy disks, hard disks, magnetic tapes, magneto-optical disks (MO)), optical storage (such as CDs, DVDs, BDs, HVDs, etc.), and semiconductor storage (such as ROMs, EPROMs, EEPROMs, non-volatile memories (NAND FLASH), solid-state drives (SSDs)), etc.
[0151] The computer instructions stored in the storage medium of the above embodiment are used to enable the computer to execute the multimodal controllable text generation method described in any embodiment in the above exemplary method part, and have the beneficial effects of the corresponding method embodiment, which will not be repeated here.
[0152] Based on the same inventive concept, corresponding to the multimodal controllable text generation method described in any of the above embodiments, the present disclosure also provides a computer program product, which includes computer program instructions. In some embodiments, the computer program instructions can be executed by one or more processors of a computer so that the computer and / or the processor executes the multimodal controllable text generation method. Corresponding to the execution subject corresponding to each step in each embodiment of the multimodal controllable text generation method, the processor that executes the corresponding step can belong to the corresponding execution subject.
[0153] The computer program product of the above embodiment is used to enable the computer and / or the processor to execute the multimodal controllable text generation method described in any of the above embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be repeated here.
[0154] Those skilled in the art will appreciate that embodiments of the present disclosure may be implemented as a system, method, or computer program product. Therefore, the present disclosure may be implemented in the following forms: entirely in hardware, entirely in software (including firmware, resident software, microcode, etc.), or in a combination of hardware and software, generally referred to herein as a "circuit," "module," or "system." Furthermore, in some embodiments, the present disclosure may also be implemented in the form of a computer program product in one or more computer-readable media containing computer-readable program code.
[0155] Any combination of one or more computer-readable media can be used. A computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or component, or any combination thereof. More specific examples (non-exhaustive examples) of computer-readable storage media can include, for example: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In this document, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in combination with an instruction execution system, device or device.
[0156] A computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device.
[0157] Program code embodied on a computer readable medium may be transmitted using any appropriate medium, including but not limited to wireless, wireline, optical fiber cable, RF, etc., or any suitable combination of the foregoing.
[0158] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages, or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).
[0159] It should be understood that each block in the flowchart and / or block diagram, and combinations of blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device to produce a machine. These computer program instructions are executed by the computer or other programmable data processing device to produce a device that implements the functions / operations specified in the blocks in the flowchart and / or block diagram.
[0160] These computer program instructions can also be stored in a computer-readable medium that enables a computer or other programmable data processing device to operate in a specific manner. In this way, the instructions stored in the computer-readable medium produce a product that includes an instruction device that implements the functions / operations specified in the blocks in the flowchart and / or block diagram.
[0161] Computer program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device so that a series of operational steps are performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, such that the instructions executed on the computer or other programmable apparatus provide a process that implements the functions / operations specified in the blocks in the flowchart and / or block diagram.
[0162] Furthermore, although the operations of the disclosed method are described in a particular order in the accompanying drawings, this does not require or imply that the operations must be performed in that particular order, or that all of the operations shown must be performed to achieve the desired results. Rather, the steps depicted in the flowcharts may be performed in a different order. Additionally or alternatively, certain steps may be omitted, multiple steps may be combined into a single step, and / or a single step may be broken down into multiple steps.
[0163] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present application. Among them, each box in the flowchart or block diagram can represent a module, program segment, or part of the code, and the above-mentioned module, program segment, or part of the code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or flowchart, and the combination of boxes in the block diagram or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0164] It should be noted that, although several modules or units of the device for action execution are mentioned in the above detailed description, this division is not mandatory. In fact, according to the embodiment of the application, the features and functions of two or more modules or units described above can be concretized in one module or unit. On the contrary, the features and functions of one module or unit described above can be further divided into multiple modules or units to be concretized.
[0165] Those skilled in the art should understand that the above discussion of any embodiment is merely exemplary and is not intended to be limiting of the scope of the application (including the claims) which is intended to be limited only by the claims. The above embodiments or technical features among different embodiments can also be combined, steps can be implemented in any order, and there are many other variations of the aspects of the embodiments of the application as described above, which are not provided in detail in order to be brief. The embodiments of the application are not limited in scope by the sum of the embodiments disclosed because the embodiments of the application include any combination of the embodiments.
[0166] In addition, to simplify the description and discussion, and so as not to make the embodiments of the application difficult to understand, the well-known power / ground connections to integrated circuit (IC) chips and other components can or can not be shown in the provided drawings. Furthermore, devices can be shown in block diagram form in order to avoid making the embodiments of the application difficult to understand, and this also takes into account the fact that the details regarding the implementation of these block diagram devices are highly dependent on the platform on which the embodiments of the application are to be implemented (i.e., these details should be well within the understanding of those skilled in the art). Where specific details (e.g., circuitry) are set forth in order to describe an illustrative embodiment of the application, it should be apparent to those skilled in the art that the embodiment of the application can be practiced without these specific details or with an equivalent arrangement. Therefore, the description should not be construed as limiting, but merely as illustrative.
[0167] Although the application has been described in conjunction with specific embodiments thereof, numerous alternatives, modifications, and variations will be readily apparent to those skilled in the art. For example, other memory architectures (e.g., dynamic RAM (DRAM)) can use the embodiments discussed.
[0168] The embodiments of the application are intended to cover all such alternatives, modifications, and variations which fall within the broad scope of the appended claims. Accordingly, any one or more of the omitted, modified, equivalently replaced, improved, etc., should be included within the scope of the application.
[0169] While the principles of the disclosure have been described above in connection with specific embodiments, it is to be understood that this disclosure is not limited to the disclosed embodiments, but is intended to cover various arrangements and modifications which fall within the spirit and scope of the appended claims. The scope of the appended claims is to be accorded the broadest interpretation so as to include all such modifications and equivalent arrangements and equivalents thereof.
Claims
1. A multimodal controllable text generation method, characterized in that: include: Identify textual and visual control elements; Determining a plurality of natural language instructions of the text control element, and using the plurality of natural language instructions as observation parameters of one of the plurality of candidate texts based on a large language model; updating a probability distribution of the observation parameters to obtain a posterior probability of the one of the plurality of candidate texts; performing a probability calculation on the posterior probability to obtain one of the plurality of candidate texts; matching the plurality of candidate texts with the visual control element based on a contrastive language image model, and using the one of the plurality of candidate texts as a text to generate an intermediate result; Splicing the intermediate results generated from at least two of the texts to obtain a thought chain prompt; Based on the text control elements, the thinking chain prompts and the visual control elements, text generation reasoning is iteratively performed until the text generation intermediate results converge, and the converged text generation intermediate results are used as the text generation results, wherein the text generation intermediate results output from this round of iteration are spliced into the thinking chain prompts input from this round of iteration as the thinking chain prompts input from the next round of iteration.
2. The method according to claim 1, characterized in that The matching of the plurality of candidate texts with the visual control element based on the contrastive language image model and taking one of the plurality of candidate texts as text to generate an intermediate result includes: determining a matching value between the visual control element and each candidate text in the plurality of candidate texts respectively; The candidate text with the largest matching value is used as the text to generate an intermediate result.
3. The method according to claim 1, characterized in that The step of splicing the intermediate results generated from the at least two texts to obtain a thought chain prompt includes: According to the generation order of the at least two text generation intermediate results, the at least two text generation intermediate results are spliced to obtain the thought chain prompt.
4. The method according to claim 1, wherein Providing a graphical user interface through a terminal, wherein the content displayed by the graphical user interface includes an input interface and an output interface, wherein the input interface includes a text control control and a visual control control; The determining of the text control element and the visual control element includes: In response to an input operation on the text control component, obtaining the input text control element; In response to an input operation on the visual control component, obtaining the input visual control element; The iterative text generation reasoning based on the text control element, the thought chain prompt, and the visual control element until the text generation intermediate result converges, and the converged text generation intermediate result is used as the text generation result, including: The text generated in each iteration is sequentially displayed on the output interface; The text generation result is displayed on the output interface.
5. A multimodal controllable text generation device, characterized in that: include: an element determination module configured to determine text control elements and visual control elements; An intermediate result determination module is configured to determine a plurality of natural language instructions of the text control element, use the plurality of natural language instructions as observation parameters of one of the plurality of candidate texts based on a large language model; update a probability distribution of the observation parameters to obtain a posterior probability of the one of the plurality of candidate texts; perform a probability calculation on the posterior probability to obtain one of the plurality of candidate texts; match the plurality of candidate texts with the visual control element based on a contrastive language image model, and use the one of the plurality of candidate texts as the text to generate an intermediate result; a thought chain prompt determining module, configured to splice intermediate results generated from at least two of the texts to obtain a thought chain prompt; The iteration module is configured to iteratively perform text generation reasoning based on the text control elements, the thinking chain prompts and the visual control elements until the text generation intermediate results converge, and use the converged text generation intermediate results as the text generation results, wherein the text generation intermediate results output by this round of iteration are spliced into the thinking chain prompts input by this round of iteration as the thinking chain prompts input by the next round of iteration.
6. An electronic device, characterized in that: The method comprises a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the method according to any one of claims 1 to 4 is implemented.
7. A non-transitory computer-readable storage medium, characterized in that The non-transitory computer-readable storage medium stores computer instructions, and the computer instructions are used to cause a computer to execute the method according to any one of claims 1 to 4.
8. A computer program product, characterized in that The method comprises computer program instructions, which, when executed on a computer, cause the computer to execute the method according to any one of claims 1 to 4.
Citation Information
Patent Citations
Generative artificial intelligence causal thinking chain generation method, device and equipment
CN116862000A
Generative visual common sense reasoning and explaining method based on scene graph generation
CN116955672A