Information processing method, device and system for image generation, and related apparatuses
By processing the image generation requirements information through the Prompt engineering system and extracting the image generation parameters adapted to the image generation system, the shortcomings of image generation technology in terms of quality and user experience are solved, and higher quality and more personalized image generation is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-18
- Publication Date
- 2026-04-07
AI Technical Summary
Existing image generation technologies are inadequate in terms of image quality and user experience, making it difficult to meet personalized needs and security requirements.
The Prompt engineering system is introduced, which processes the image generation requirement information through multiple neural network models and extracts the image generation parameters suitable for the image generation system, including risk information, image generation intent information, model type, rewriting prompts and response information, thereby improving the adaptability and accuracy of the image generation system.
It improves the quality of image generation and user experience, enhances the ability to identify and filter malicious requests, meets personalized needs, and improves the practical application of the image generation system.
Smart Images

Figure CN119810234B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of artificial intelligence technology, particularly to the fields of computer vision, deep learning, and large models, and can be applied to AIGC (Artificial Intelligence Generative Content) and other scenarios based on artificial intelligence content generation. Background Technology
[0002] Image generation technology is an important research area in the field of artificial intelligence. With the development of AI technology, image generation technology is also constantly iterating and updating. Image generation technology is diverse and practical in various industries. For example, it can assist in artistic creation and design, and in education, it can help in understanding abstract concepts.
[0003] With the continuous advancement of technology, the application areas of image generation technology will be further expanded. Summary of the Invention
[0004] This disclosure provides an information processing method, apparatus, system, and related equipment for image generation.
[0005] According to one aspect of this disclosure, an information processing method for image generation is provided, comprising:
[0006] Obtain the raw image requirements of the target object for generating the image;
[0007] The raw image requirement information is input into the prompting engineering system, so that the raw image requirement information is converted into raw image parameters for use by the image generation system.
[0008] According to another aspect of this disclosure, an information processing apparatus for image generation is provided, comprising:
[0009] The acquisition unit is used to acquire the raw image requirement information of the target object for generating images;
[0010] The processing unit is used to input the raw image requirement information into the prompting engineering system, so as to convert the raw image requirement information into raw image parameters for use by the image generation system based on the prompting engineering system.
[0011] According to another aspect of this disclosure, an information processing system for image generation is provided, comprising:
[0012] The prompting engineering system is used to process the image generation requirement information of the target object, so as to convert the image generation requirement information into image generation parameters for use by the image generation system based on the prompting engineering system;
[0013] An image generation system for generating a target image based on the generated image parameters.
[0014] According to another aspect of this disclosure, an electronic device is provided, comprising:
[0015] At least one processor; and
[0016] The memory is communicatively connected to the at least one processor; wherein,
[0017] The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform any of the methods described in the present disclosure.
[0018] According to another aspect of this disclosure, a non-transitory computer-readable storage medium is provided storing computer instructions, wherein the computer instructions are used to cause the computer to perform any of the methods according to embodiments of this disclosure.
[0019] According to another aspect of this disclosure, a computer program product is provided, including a computer program that, when executed by a processor, implements any of the methods according to embodiments of this disclosure.
[0020] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0021] The accompanying drawings are provided to better understand this solution and do not constitute a limitation of this disclosure. Wherein:
[0022] Figure 1 This is a schematic flowchart of an information processing method for image generation according to an embodiment of the present disclosure;
[0023] Figure 2 This is a schematic diagram of a scenario for an information processing method for image generation according to an embodiment of the present disclosure;
[0024] Figure 3 This is a schematic diagram illustrating the training of a first intent recognition sub-model according to an embodiment of the present disclosure;
[0025] Figure 4 This is another schematic flowchart of an information processing method for image generation according to an embodiment of the present disclosure;
[0026] Figure 5 This is a schematic diagram illustrating the intent of identifying raw image requirement information according to an embodiment of the present disclosure;
[0027] Figure 6This is an architecture diagram of a prompting engineering system according to an embodiment of the present disclosure;
[0028] Figure 7 It is a model type of the target image model according to an embodiment of the present disclosure;
[0029] Figure 8 This is a schematic diagram of a prompting engineering system according to an embodiment of the present disclosure;
[0030] Figure 9 This is a schematic diagram of a multi-turn dialogue prompting engineering system according to an embodiment of the present disclosure;
[0031] Figure 10 This is a schematic diagram of a security management strategy according to an embodiment of the present disclosure;
[0032] Figure 11 This is an architecture diagram of an information processing system for image generation according to an embodiment of the present disclosure;
[0033] Figure 12 This is a schematic diagram of an information processing apparatus for image generation according to an embodiment of the present disclosure;
[0034] Figure 13 This is a block diagram of an electronic device for implementing the information processing method for image generation according to embodiments of the present disclosure. Detailed Implementation
[0035] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.
[0036] The terms “first,” “second,” etc., used in this disclosure are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover non-exclusive inclusion, such as including a series of steps or units. A method, system, product, or apparatus is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to these processes, methods, products, or apparatuses.
[0037] To better improve the practical application of image generation systems in real-world products, this disclosure proposes a Prompt engineering system to provide key parameters for the image generation system, thereby facilitating its implementation and improving the quality of images generated by the system.
[0038] like Figure 1 The diagram shown is a flowchart illustrating the information processing method provided in this embodiment of the disclosure, including the following:
[0039] S101, Obtain the raw image requirement information of the target object for generating the image.
[0040] The raw image requirement information refers to the descriptive information of the target object for at least one dimension of the image to be generated, such as subject, style, and content description. This can be a text description or a combination of text and images.
[0041] When the requirement for generating an image is in the form of voice, the voice can be converted into text, which is then used to construct the requirement information for generating the image.
[0042] S102, input the raw image requirement information into the prompting engineering system, so that the raw image requirement information is converted into raw image parameters for use by the image generation system based on the prompting engineering system.
[0043] In this embodiment of the disclosure, multiple neural network models can be set up within the engineering system. Different neural network models process all or part of the information in the image generation requirement information based on different task requirements, in order to extract image generation parameters suitable for the image generation system from the image generation requirement information. These image generation parameters are more adaptable to the image generation system than the image generation requirement information. For example, they can describe the requirements for generating images using expressions that the image generation system can accurately understand. They can also more accurately describe the services that the image generation system should provide. For example, the image generation system provides multiple models for generating images, and different models meet different requirements. The image generation parameters can better utilize these models to meet the personalized needs of the target object.
[0044] In this embodiment of the disclosure, the content of the image generation requirement information is processed based on the prompting engineering system. By utilizing the prompting engineering system, the image generation requirement information can be systematically understood based on the characteristics of the image generation system, so that the image generation requirements of the target object are clearer and the image generation system can be used better. This can effectively improve the quality of the generated image and enhance the user experience.
[0045] In some embodiments, the image parameters include at least one of the following:
[0046] 1) Risk information corresponding to the raw image request. This risk information indicates whether the raw image request contains malicious requests. Malicious raw image requests need to be filtered or rewritten as normal requests.
[0047] 2) Image generation intent information corresponding to image generation requirement information. In this embodiment of the disclosure, image generation intent information can describe the expected image content and / or the method of image generation from multiple dimensions or perspectives.
[0048] The raw image intent information includes, but is not limited to, the raw image type (such as first-round text-to-raw image, first-round image-to-raw image, text-to-raw image in multiple rounds, etc.), image size, image style, reference image information, and the type of target subject to be generated.
[0049] 3) The model type of the target image generation model that matches the image generation requirements. That is, the image generation system provides a variety of image generation models. Different types of models have their own characteristics. By predicting the model type, a suitable model can be selected as the target image generation model to generate images, which can improve the implementation effect of the image generation system.
[0050] 4) Rewrite the target prompts obtained from the raw image requirement information. By rewriting the raw image requirement information through prompts, the engineering system can automatically adapt to the characteristics of the image generation system, thereby improving the quality of the images generated by the image generation system. Furthermore, it can support the target object to flexibly express its raw image requirements using different descriptive methods, improving the user experience.
[0051] 5) Response information based on raw image requests. This response information is used to provide personalized replies to the target audience's raw image requests, facilitating interaction and improving user experience.
[0052] 6) Image generation requirements: Layout information of the target subject type in the generated image.
[0053] The target subject types include, but are not limited to, modern personal names, ancient personal names, fictional personal names, animal names, plant names, product models, brand names, location landmarks, film and television culture, and popular internet slang.
[0054] The image generation requirement information includes the layout information of the target subject type in the generated image, which indicates the position of the subject corresponding to the target subject type in the generated image. For example, the image can be divided into multiple regions, and this layout information can express the image region where the subject corresponding to the target subject type is located.
[0055] During implementation, when the generated image parameters include risk information, generated image intent information, model type, rewritten target prompts, response information, and layout information of the target subject type, such as... Figure 2As shown, inputting the raw image requirement information 201 into the prompting engineering system 202 yields the aforementioned six raw image parameters, namely, the risk information 2021 corresponding to the raw image requirement information, the raw image intent information 2022 corresponding to the raw image requirement information, the model type 2023 of the target raw image model adapted to the raw image requirement information, the prompt words 2024 obtained by rewriting the raw image requirement information, the response information 2025 for the raw image requirement information, and the layout information 2026 of the target subject type in the image to be generated by the raw image requirement information. These parameters are then transmitted to the downstream network application service 203 (i.e., the image generation system, exemplarily an iRAG (image based Retrieval-Augmented Generation) system) to obtain the target image 204 generated by the downstream network application service.
[0056] In this embodiment of the disclosure, multi-dimensional raw image parameters are used for the image generation system, which can better utilize the image generation system and improve the quality of the generated images and the user experience.
[0057] The above-mentioned parameters for generating images will be explained in detail below:
[0058] 1) Risk information corresponding to raw image demand information
[0059] In some embodiments, risk information is determined by inputting the requirement text in the raw image requirement information into a security risk identification model to identify the risk information in the raw image requirement information; the risk information includes risk type and / or risk level.
[0060] Among them, the security risk identification model can be obtained by training a pre-trained language model on a scale of 3B (3 billion).
[0061] Among these, pre-trained language models can be large language models (LLMs). Large language models refer to a specific type of large-scale model specifically designed for processing text data. These models are neural network-based natural language processing models that can be used to generate, understand, and process text data. Large language models can have hundreds of billions of parameters, generate high-quality text, and can be used for various natural language processing tasks, such as question answering, text generation, and dialogue systems.
[0062] Large language models possess excellent reasoning and few-shot learning capabilities. This is because large language models are comprehensible in that they can be trained on a large number of samples. Based on such large models, accurate semantic understanding can be achieved.
[0063] Here, risk type refers to the types within the risk type set. This set can define different risk types with finer granularity, allowing the security risk identification model to perform classification and prediction. In implementation, the finer the granularity of the risk type division within the risk type set, the more accurately specific risk types can be identified, and corresponding subsequent processing can be performed, such as deciding whether to filter or rewrite based on the situation.
[0064] A risk level is used to indicate the degree of risk involved. For example, risk levels can be divided into three categories: high risk, medium risk, and low risk. In practice, the risk level can be determined based on the types of risks involved and / or the number of risk types.
[0065] If the risk level is determined to be high, the generation requirement for the target object may not be supported, and the target object may be prompted to modify the raw image requirement information.
[0066] If the risk level is determined to be low and the risk type is a preset type, content with risk in the image request information can be deleted. For example, if the image content of this preset type is not the core content of the image (such as background or decoration), then the corresponding image content of this preset type will not be generated.
[0067] In addition, if there is a replaceable safe text, the corresponding risk text can be replaced with the safe text. After the replacement is completed, the safe text is input into the safety risk identification model to identify the risk information of the replaced raw image requirement information. If there is no risk, the subsequent processing flow can continue.
[0068] In this embodiment of the disclosure, risk detection is performed on the raw image requirement information to reduce risk factors in the raw image requirement information and improve the quality of the generated image.
[0069] In some embodiments, a supervised training method can be used to train the security risk identification model. Specifically, this can be implemented as follows: A first sample set is obtained, comprising multiple first samples, labeled risks corresponding to each first sample (e.g., including risk type and / or risk level), and a first task instruction for each first sample; the first sample set is input into an initial security risk identification model to obtain the predicted risk output by the model; the predicted risk is compared with the labeled risk to obtain a loss value; the model parameters of the initial security risk identification model are adjusted based on the loss value; and the security risk identification model is trained under convergence conditions; wherein the first task instruction indicates that the initial security risk identification model needs to perform a risk detection task.
[0070] 2) The intent information of the raw image corresponding to the raw image requirement information
[0071] In some embodiments, the raw image intent information is determined by identifying raw image demand information based on a multimodal intent understanding model to obtain the raw image intent information.
[0072] A multimodal intent understanding model is a model capable of understanding information from multiple modalities to predict the intent of a generated image. This model can be a large multimodal language model (MLLM) or constructed from multiple sub-models. Different sub-models perform different intent understanding tasks. For example, a multimodal intent understanding model may include a first intent recognition sub-model, a second intent recognition sub-model, and the large multimodal language model.
[0073] In this embodiment of the disclosure, the image generation requirement information is identified based on a multimodal intent understanding model. On the one hand, through multimodal intent understanding, the target object can use information from multiple modalities to express its image generation intention, improving the flexibility and convenience of expressing image generation requirements. On the other hand, the multimodal intent understanding model can accurately understand the target object's image generation requirements, thereby improving the quality of the generated image.
[0074] In some embodiments, the process of identifying raw image requirement information based on a multimodal intent understanding model to obtain raw image intent information can be implemented as follows: constructing text understanding prompts based on the requirement text and target prompt template in the raw image requirement information; inputting the text understanding prompts into the first intent recognition sub-model in the multimodal intent understanding model to obtain at least one of the following information in the raw image intent information: raw image type, image size, and image style; wherein, the raw image type is a type in a preset type set.
[0075] In some embodiments, the preset type set includes at least one of the following types: first-round text-to-image type, first-round image-to-image type, text-to-image multiple-round successor type, image-to-image multiple-round successor type, and image editing type.
[0076] Among them, the first round of text-to-image type refers to the image requirement information obtained in the first round of dialogue in an image generation task. This image requirement information is the text content provided by the target object.
[0077] The first-round image type refers to the image requirement information obtained in the first round of dialogue in an image generation task, which is expressed using an image.
[0078] The text-to-image multi-turn successor type indicates that after the initial dialogue uses text to express the image generation requirement, the target audience continues to use multiple rounds of dialogue to improve the image generation requirement in order to continue generating images;
[0079] The "Image-to-Image Multi-Turn Successor Type" indicates that after the initial dialogue uses an image to express the image requirement, the target audience continues to use multiple turns of dialogue to improve the image requirement and continue generating images.
[0080] Image editing type indicates that there are local contents in a known image that need to be added, modified, or deleted. The known image can be an image generated from a text-based image or an image-based image, or it can be an image provided by the target object.
[0081] In this embodiment of the disclosure, the raw image type is divided into at least one of the aforementioned types, which allows for a more refined understanding of the raw image intent, improves the accuracy of understanding the raw image requirements, and further enhances the quality of the generated image.
[0082] The image size is used to indicate the required size of the generated image. For example, if the image generation requirement is for a poster, the image needs to be adjusted to the size corresponding to a poster. Furthermore, in a multi-round image generation task, for example, the first round might require a poster size, while the second round might require a book cover size; the image size is predicted based on different usage requirements.
[0083] Image style refers to the style of the image to be generated, such as abstract, realistic, or anime styles. Furthermore, in multi-round image generation tasks, the style can vary from round to round.
[0084] In this embodiment of the disclosure, the first intent recognition sub-model is used to input the demand text. The first intent recognition sub-model focuses on recognizing the type, size, and style of the raw image. Based on this, the accuracy of obtaining the intent information of the raw image can be improved, and the quality of the generated image can be improved.
[0085] In some embodiments, the first intent recognition sub-model is obtained through training in the following manner, which can be specifically implemented as follows:
[0086] Step A1: Distill training sample sets of different raw image intent labels based on the large language model.
[0087] In some embodiments, an initial sample set can be obtained through random sampling. The samples are then distilled from the large language model to train the first intent recognition sub-model. Distilling the samples can be implemented as follows:
[0088] Step A11: Obtain the initial sample set.
[0089] The initial sample set can be randomly sampled data. This sample data includes the text content provided online when implementing the text-to-image task.
[0090] Step A12: Based on the sequence labeling task, multiple intermediate samples are selected from the initial sample set to obtain an intermediate sample set.
[0091] Sequence labeling is a fundamental task in Natural Language Processing (NLP), which involves assigning a label or category to each word or subword in a text.
[0092] Sequence labeling can be viewed as a generalization of classification problems, or a simplified form of more complex structural prediction problems. In sequence labeling, the input is an observation sequence, and the output is a label sequence or state sequence. The goal is to provide a label sequence as a prediction for the observation sequence.
[0093] Since the quality of the initial sample set obtained by random sampling varies and may even contain noisy data, the sequence labeling task can be used to filter the samples and obtain intermediate samples of better quality.
[0094] The intermediate sample set contains multiple user queries. In order to train the first intent recognition sub-model, the large language model needs to distill the training samples.
[0095] Step A13: Use a large language model to generate the raw image intent labels for each sample in the intermediate sample set to obtain the training sample set.
[0096] That is, the large language model processes each user query in the intermediate sample set and labels it according to the task requirements of the first intent recognition sub-model.
[0097] Among them, the raw image intent tags include, but are not limited to, raw image type (including first-round text-based raw image intent, first-round image-based raw image intent, text-based raw image with multiple rounds of subsequent intent, image-based raw image with multiple rounds of subsequent intent, and image editing intent), image size, and image style.
[0098] In this embodiment of the disclosure, the initial sample set is initially screened based on the sequence labeling task to optimize the sample quality. The samples are distilled through the large language model, which enables automated labeling. This allows the knowledge of the large model to be transferred to the first intent recognition sub-model through training samples, thereby improving the intent recognition accuracy of the first intent recognition sub-model.
[0099] Step A2: Classify the training sample set based on the raw image intent labels to obtain sample subsets corresponding to each raw image intent label.
[0100] For example, if training sample A is {sample A - raw image intent label 1}, training sample B is {sample B - raw image intent label 2}, training sample C is {sample C - raw image intent label 1}, and training sample D is {sample D - raw image intent label 1}, then the generated sample subset includes two types: one is raw image intent label 1 {training sample A; training sample C; training sample D}, and the other is raw image intent label 2 {training sample B}.
[0101] Step A3: Based on the sample subsets corresponding to each raw image intent label, supervised training is performed on the pre-trained initial model to obtain the first intent recognition sub-model.
[0102] During implementation, the initial model is trained using samples from a general task. If the convergence condition is met, a pre-trained initial model is obtained. Then, steps A1-A3 are executed to obtain the first intent recognition sub-model.
[0103] During implementation, a subset of samples of each raw image intent label is sampled according to a preset ratio to obtain multiple samples in a batch. These multiple samples are then input into a pre-trained initial model for supervised training, enabling the pre-trained initial model to learn the recognition patterns of each intent. Under the condition of meeting the convergence condition, the first intent recognition sub-model is obtained.
[0104] In this embodiment of the disclosure, a training sample set is distilled based on a large language model, and the pre-trained initial model is trained in a supervised manner based on this sample set. This allows the initial model to learn how to recognize the intent information of raw images through supervised training, based on a certain amount of knowledge. As a result, the first intent recognition sub-model can more accurately recognize the intent of raw images.
[0105] In some embodiments, to improve the recognition accuracy of the first intent recognition sub-model, the reflective optimization mechanism of the large language model can be relied upon to optimize the target prompt template adopted by the first intent recognition sub-model, which can be implemented as follows:
[0106] Step B1: Using the prompt template to be adjusted, predict the initial intent label of the target sample set based on the first intent recognition sub-model.
[0107] During implementation, the powerful semantic understanding capabilities of large models are utilized to predict the initial intent labels of the target sample set.
[0108] The target sample set and the training sample set can be the same sample set or different sample sets.
[0109] Step B2 uses a large language model to determine whether the initial intent label meets the requirements of multiple preset rules in the rule set.
[0110] Among them, the preset rules are used to restrict the initial intent tag. For example, it can be used to determine whether the initial intent tag meets the preset rule requirements. The preset rule requirements can be that the initial intent tag cannot contain metaphors, cannot be ambiguous, etc. These preset rules can set the descriptions or content that the initial intent tag should not contain.
[0111] Furthermore, the aforementioned raw image types can be further subdivided into subtypes based on different operational intentions. Preset rules can also be used to distinguish between similar raw image subtypes; specific subtypes can be referred to later. Figure 5 The description will not be detailed here. For ease of understanding, for example, the raw image requirement information is "remove glasses from a person's face," which means the raw image subtype is local deletion processing. The raw image requirement information is "remove blemishes from a person's face," which means the raw image subtype is local modification processing. The two are similar in language, but the fine-grained intents are different. Therefore, the first intent recognition sub-model can be trained to distinguish between similar raw image subtypes by using the constraints required by the preset rules.
[0112] Step B3: If any preset rule requirement is not met, optimize the prompt template to be adjusted based on the large language model and the preset rule requirement as a benchmark to obtain the target prompt template.
[0113] In this embodiment, the prompt template to be adjusted is optimized based on preset rules to avoid the model outputting irrelevant content or creating illusions. At the same time, the use of a standardized target prompt template further improves the accuracy of intent recognition by the first intent recognition sub-model, thereby improving the quality of the raw image parameters and the quality of the generated image.
[0114] In some embodiments, the first intent recognition sub-model is trained using a 10B pre-trained language model as a base model. For example... Figure 3The diagram shows the training framework: an initial sample set 301 is obtained by randomly collecting samples online. Based on the sequence labeling task, multiple intermediate samples are selected from the initial sample set 301 to obtain an intermediate sample set 302. The intermediate sample set 302 is input into a large-size LLM / MLLM 303 for distillation. The large model 303 becomes a large-size LLM / MLLM, obtaining the raw image intent labels for each sample. This results in a synthetic sample 304 constructed from the user query and its corresponding raw image intent labels. This synthetic sample 304 consists of sample pairs including the user query and its corresponding raw image intent labels. To improve sample quality, the synthetic sample 304 can be modified and optimized through manual review 305. New sample pairs can be automatically synthesized, ultimately obtaining the training sample set 306. The training sample set can be divided into different sample subsets 307 based on each raw image intent label. Training samples from the same batch are selected from different sample subsets 307 according to a set ratio to train a first intent recognition sub-model 308, which is a small-sized LLM / MLLM. Then, a larger model 303 uses multiple preset rules from a rule set to judge the intent labels predicted by the first intent recognition sub-model 308, i.e., the evaluation samples 309. Multidimensional evaluation is performed using the larger model 303. If any preset rule requirement is not met, the larger model 303 reflects on and corrects the prompt template used by the first intent recognition sub-model 308.
[0115] Furthermore, the first intent recognition sub-model may include multiple model modules, each specifically designed to identify the image type, image size, and image style. Each model module may employ... Figure 3 Training is conducted in this manner.
[0116] In some embodiments, the target object may also provide a first reference image for the image generation system to generate an image. In this case, the image generation intent expressed by the target object based on the first reference image can be identified, which can be implemented as follows: when the image generation requirement information includes the first reference image provided by the target object, the reference image information of the first reference image is obtained based on the multimodal large model in the multimodal intent understanding model; the image generation intent information includes the reference image information. The reference image information includes at least one of the following: reference subject type, number of reference subjects, and textual description information of the first reference image.
[0117] In practice, when a first reference image is included, a multimodal large model can be used to understand the first reference image. When the first reference image includes multiple reference subjects, the number of each reference subject can be identified, and the operation corresponding to each reference subject may be different.
[0118] For example, the first reference image may contain subject A and subject B, and its image requirement information is "adjust subject A to a realistic style and delete subject B". Therefore, the intention of "changing the color of subject A" is a local modification, while the intention of "deleting subject B" is a local deletion of the image. The two intentions are not the same, so the corresponding operations are different.
[0119] Examples of the first reference subject types may include modern personal names, ancient personal names, fictional personal names, animal names, plant names, product models, brand names, location landmarks, film and television culture, popular internet slang, etc.
[0120] The textual description information of the first reference graph can be obtained by using the powerful reasoning and content understanding capabilities of the multimodal large model to perform content understanding on the first reference graph.
[0121] The specific information included in the reference figures can be set based on actual circumstances, and this disclosure does not limit which information is included.
[0122] In this embodiment of the disclosure, when using the first reference image, the multimodal intent understanding model can accurately understand the content expressed by the first reference image, thereby assisting the model's intent recognition task and helping the image generation system accurately understand the user's purpose in using the first reference image, thus improving the quality of the generated image.
[0123] In some embodiments, not only can the reference subject in the reference diagram be identified, but also the subject expected to be generated in the text description can be understood. Thus, the process of identifying the raw image requirement information based on the multimodal intent understanding model to obtain raw image intent information can be implemented as follows: inputting the requirement text in the raw image requirement information into the second intent recognition sub-model in the multimodal intent understanding model to obtain the target subject type indicated by the requirement text; the raw image intent information includes the target subject type.
[0124] The target subject type covers the same range of types as the first reference subject type, as described above, and will not be repeated here in this disclosure.
[0125] In this embodiment of the disclosure, the target subject type is identified using the requirement text in the raw image requirement information, so as to help the image generation system generate an image around the subject and improve the quality of the generated image.
[0126] 3) Model type of the target image model that adapts to the image generation requirements.
[0127] In some embodiments, the model type of the target image generation model is determined in the following manner, specifically by: processing the image generation intent information based on the distribution strategy model to determine the model type of the target image generation model in the image generation system that generates images for the image generation requirement information.
[0128] In this embodiment of the disclosure, a suitable raw image model type can be selected for the target object based on the distribution strategy model, and the model type can be provided as a raw image parameter to the image generation system. This allows the system to generate images for the target object using a suitable model, thereby improving the user experience.
[0129] In some embodiments, the model type of the target image model includes at least one of the following:
[0130] The base model supports general image generation tasks; the base model can be an SD model (stabledefusion, a general text-to-image generation model) for general image generation tasks; the base model can support various image generation tasks, focusing on supporting a wide range of task types, but the accuracy may be somewhat limited.
[0131] A precise model is used to generate high-fidelity images; the precise model can be used to generate images that are required to maintain a high degree of consistency with the actual scene, for example, specific models of vehicles, trademarks, etc.
[0132] A high-generalization model is used to generate images that meet preset fidelity requirements and possess artistic effects. Compared to a precise model, a high-generalization model allows for some modifications or optimizations, and allows for the addition of ideas and creative elements. For example, a portrait image, while still recognizable as a human figure, may be given a certain artistic style.
[0133] An editing model is used to perform editing operations on a known image. This model is used to make local adjustments to a known image, such as removing glasses from a face, modifying the position of a mole on a face, or changing the shape and color of glasses.
[0134] In this embodiment of the disclosure, based on the aforementioned multiple model adaptations and corresponding intentions, most raw image tasks can be processed as much as possible to improve the image generation system's ability to adapt to multiple tasks and improve image generation quality.
[0135] In some embodiments, when the target subject type in the image to be generated by the raw image requirement information is a preset type, the target model type for processing the raw image requirement information is a security model; wherein, the security model is used to generate images that avoid public opinion risks.
[0136] For example, preset types could be iconic buildings of each region, iconic buildings of tourist attractions, etc. During implementation, the main text and image corresponding to each preset type can be associated and stored as key-value pairs. When a target subject is detected as being in the list of key-value pairs, it is determined to be a preset type, and a security model is used for generation.
[0137] In this embodiment of the disclosure, a security model is used to process key target subject types that require a high degree of consistency in order to avoid public opinion risks.
[0138] 4) Rewrite the target prompts obtained from the raw image requirement information, and the response information based on the raw image requirement information.
[0139] In some embodiments, the target prompts and response information for raw image demand information are determined in the following ways: Figure 4 As shown, it can be implemented as follows:
[0140] S401, based on the distribution strategy model, processes the target object's image requirement information to obtain the processing result.
[0141] The processing results include those requiring search enhancement and those not requiring search enhancement. In the case where search enhancement is not required, the raw image requirement information is rewritten based on the prompt word optimization model to obtain the target prompt words and the response information. In the case where search enhancement is required, S402 is executed.
[0142] S402, if the processing result indicates that retrieval enhancement is needed, obtain the retrieval statement generated by the distribution strategy model from the processing result.
[0143] The search query can be the description information of the key subject obtained from the raw image requirement information of the target object for generating the image. The key subject can be the target subject in the raw image requirement information.
[0144] In some embodiments, the training method for the distribution strategy model may be implemented as follows: Obtaining a second sample set; wherein the second sample set includes multiple second samples, corresponding labels for each second sample, and second task instructions. The labels indicate the applicable model type for the second sample, whether retrieval enhancement is performed, and may even include the target subject type in the second sample. The labeled model types include, but are not limited to, base model, accurate model, high generalization model, and edit model. The second sample set is input into the initial distribution strategy model to obtain prediction results. The prediction results include whether to enhance the binary classification result, the predicted subject, and the prediction type, where the prediction type indicates the model type to be used for the predicted second sample. Based on the comparison between the prediction results and the labels, and if the convergence condition is met, a distribution strategy model is obtained. The second task instructions are used to instruct the initial distribution strategy model on the tasks to be performed, including a classification task (whether to enhance), a target subject type prediction task, and a distribution strategy task.
[0145] S403, based on the search query, retrieve the second reference map for the birth map requirement information.
[0146] The second reference image is used to enhance the image creation intent of the image creation requirement information.
[0147] S404: Based on the prompt word optimization model, the second reference image and the raw image requirement information are rewritten to obtain the target prompt words of the target raw image model in the adapted image generation system, and a response message is generated.
[0148] In this embodiment of the disclosure, it can be automatically determined whether the raw image requirement information needs to be enhanced. If enhancement is required, a second reference image is generated to assist the model in understanding, making the intent of the raw image requirement information clearer and further improving the image generation system's ability to understand raw image requirements.
[0149] In some embodiments, the second reference image and the raw image requirement information are rewritten based on the prompt word optimization model to obtain the target prompt word for the target raw image model adapted to the image generation system. This can be implemented as follows: obtaining the target dialogue text of the current round of dialogue of the target object from the raw image requirement information; constructing the information to be rewritten based on the target dialogue text, the historical dialogue information in the raw image requirement information, and the second reference image; and inputting the information to be rewritten into the prompt word optimization model to obtain the target prompt word.
[0150] In this embodiment of the disclosure, the drawing requirements can be expressed more clearly based on the second reference drawing, thereby improving the drawing quality.
[0151] Specifically, if the current round does not require retrieval enhancement, the information to be rewritten is constructed based on the target dialogue text and historical dialogue information of the current round. This information is then input into the prompt word optimization model to obtain the target prompt words. In particular, if the current round of dialogue is the first round, the target dialogue text and target characters are concatenated into structured information to be rewritten according to a preset format.
[0152] During implementation, the default format can be set to a format such as <target dialogue text, target character>. In the case of the first round of dialogue, this format should meet the following requirements:
[0153] 1) The target dialogue text is represented by a target field, and the target dialogue text is the value of the target field. The start and end of the target dialogue text are represented by a dialogue start symbol and a dialogue end symbol, so that the text-generated graph model can identify and verify the integrity of the target dialogue text.
[0154] 2) Use target characters to represent historical dialogue records. If the target character is empty, it means that the current round of dialogue is the first round of dialogue and there are no historical dialogue records.
[0155] In this embodiment of the disclosure, the target dialogue text and target characters are concatenated according to a preset format, which can ensure that the information to be rewritten has regularity and logic, so that the large language model can understand and sort out the current round of dialogue's requirements for image generation, as well as the specific content of historical situations. This can reduce the understanding error of the large model, improve the accuracy of target prompts, and thus improve the quality of the generated image.
[0156] When the current dialogue is the latest in a multi-turn dialogue, the information to be rewritten can be constructed based on the dialogue text and historical dialogue information through the following steps:
[0157] Step D1: Obtain the historical dialogue text provided by the target object in the historical dialogue record, as well as the historical prompt words generated by the prompt word optimization model according to the requirements of the historical dialogue text;
[0158] Step D2: Concatenate the target dialogue text and the historical dialogue record according to a preset format to obtain the information to be rewritten; wherein, in the information to be rewritten, the historical dialogue text and the historical prompts of the historical dialogue text are grouped into information tuples according to their association relationship, and the information tuples are the values of the target characters.
[0159] Based on the preceding explanation, the default format can be set to a format such as <target dialogue text, target character>. When the current round of dialogue is the latest round in a multi-turn dialogue, and the current round is the second round of dialogue, this format must meet the following requirements:
[0160] 1) A target field is used to describe the target dialogue text of the current round of dialogue, and the target dialogue text is the value of the target field. A dialogue start symbol and a dialogue end symbol are used to indicate the beginning and end of the target dialogue text, so that the text-generated graph model can identify and verify the integrity of the target dialogue text.
[0161] 2) The target character is used to represent the historical dialogue record. The target character contains two key fields, namely the first subfield and the second subfield.
[0162] 3) The first subfield represents the historical dialogue text of the first round of dialogue, and this historical dialogue text is the value of the first subfield. Furthermore, the historical dialogue text uses a first start character and a first end character to indicate the beginning and end of the historical dialogue text, used to identify and verify the integrity of the historical dialogue text. Similarly, the second subfield represents the historical prompt words of the first round of dialogue, and this historical prompt word is the value of the second subfield. Furthermore, the historical prompt words use a second start character and a second end character to indicate the beginning and end of the historical prompt words, used to identify and verify the integrity of the historical prompt words.
[0163] 4) Use the target start symbol and target end symbol from the historical dialogue record to indicate the start and end positions of the target character value, so as to facilitate the identification and verification of the integrity of the target character value.
[0164] In this embodiment of the disclosure, in a multi-turn dialogue scenario, the dialogue text and historical dialogue records are concatenated according to a preset format. This allows the collection of all image generation requests expressed in the historical dialogue records, which are then combined with the latest dialogue text to obtain the information to be rewritten. In this rewritten information, the formatted representation of information tuples facilitates the understanding of drawing requirements and preferences by the larger model, thereby improving the quality of the generated target prompts.
[0165] Correspondingly, during the creation process, the target audience may refer to the generated images and change their image requirements. This may lead to multiple rounds of dialogue iteratively refining the drawing requirements to ultimately obtain the desired image. When the historical dialogue record includes multiple rounds of dialogue, there will be multiple sets of information tuples within the historical dialogue record. Specifically, the historical dialogue text of each round and the corresponding historical prompts constitute the corresponding information tuple.
[0166] Taking a multi-round historical dialogue as an example, the format of the prompt message to be rewritten should meet the following requirements:
[0167] 1) A target field is used to describe the target dialogue text of the current round of dialogue, and the target dialogue text is the value of the target field. A dialogue start symbol and a dialogue end symbol are used to indicate the beginning and end of the target dialogue text, so that the text-generated graph model can identify and verify the integrity of the target dialogue text.
[0168] 2) The target character is used to represent the historical dialogue record. The target character contains two key fields, namely the first subfield and the second subfield.
[0169] 3) In cases involving multiple rounds of historical dialogue, each round includes a corresponding pair of first and second subfields. The first and second subfields of the same round of historical dialogue are represented by a third start symbol and a third end symbol. This ensures that the text-based graph model can accurately understand the historical dialogue text and historical prompts within the same round of historical dialogue.
[0170] 4) In each round of historical dialogue, the corresponding first subfield represents the historical dialogue text for that round, and this historical dialogue text is the value of the first subfield. Furthermore, the historical dialogue text uses a first start symbol and a first end symbol to indicate the beginning and end of the historical dialogue text, used to identify and verify the integrity of the historical dialogue text. Similarly, the second subfield represents the historical prompt words for that round of historical dialogue, and this historical prompt word is the value of the second subfield. Furthermore, the historical prompt words use a second start symbol and a second end symbol to indicate the beginning and end of the historical prompt words, used to identify and verify the integrity of the historical prompt words.
[0171] 5) Use the target start symbol and target end symbol from the historical dialogue record to indicate the start and end positions of the target character value, so as to facilitate the identification and verification of the integrity of the target character value.
[0172] In this embodiment of the disclosure, the historical dialogue text and corresponding historical prompt words of each round of historical dialogue are constructed into corresponding information tuples, which can provide a structured storage method for dialogue information. By arranging these information tuples according to the rounds, the evolution process of the target object's image generation needs can be clearly expressed. This enables the large language model to accurately capture the target object's image generation needs and preferences, clarify the rewriting direction and purpose of the generated historical prompt words, and improve the quality of the rewritten target prompt words.
[0173] In this embodiment of the disclosure, the target object can generate one or more images using a piece of text. Each image has its own characteristics and may also have its own corresponding core content. Therefore, the prompt words used to generate each image are different and independent. Based on this, when multiple images are generated in the first target round dialogue in the historical dialogue record, the historical prompt words of the first target round dialogue include the sub-prompt words corresponding to each of the multiple images.
[0174] That is, the multiple sub-cues of the multiple images and the historical dialogue text of the first target round dialogue are constructed into the information tuple of the first target round dialogue.
[0175] Taking any round of historical dialogue as an example, assuming that multiple images are generated in any round of historical dialogue, in addition to meeting the above requirements, the expression of the historical prompts in any round of historical dialogue must meet the following requirements:
[0176] 1) The history prompt words in the second subfield include multiple sub-prompt words, which are separated by a delimiter to facilitate the identification of sub-prompt words for each image.
[0177] 2) It can express the correspondence between each sub-prompt word and its corresponding image. In implementation, the same sorting order as the generated multiple images can be used to sort the sub-prompt words, thereby establishing the correspondence between the sub-prompt words and images through the sorting position. In addition, special descriptors can be added to the second sub-field to express the correspondence between the sub-prompt words and images. For example, a first special symbol and a second special symbol can be used to express that these two special symbols contain a sub-prompt word and a corresponding image identifier. It is understood that it is sufficient to be able to identify the correspondence between sub-prompt words and images, and this embodiment of the disclosure does not limit this.
[0178] Thus, a set of information tuples is constructed from multiple sub-cues from multiple images and the historical dialogue text of the first target round dialogue.
[0179] In some embodiments, the target object uploads a first reference image and expects to generate an image with a similar style or content based on the first reference image. In this case, if the target object provides the first reference image in the current round of dialogue, the text description of the first reference image can be used as the content of the historical prompts corresponding to the previous round of dialogue in the current round of dialogue, in order to construct the information to be rewritten.
[0180] Similarly, when retrieval enhancement is required, a second reference image can be automatically mined based on the aforementioned retrieval enhancement methods. The text description of the second reference image can be used as the content of the historical prompt words corresponding to the previous round of dialogue in the current round of dialogue, in order to construct the information to be rewritten.
[0181] In this embodiment, the second reference image obtained through retrieval enhancement is transformed into a text description. This text description is then used as the content of historical prompts corresponding to the previous round of dialogue in the current round of dialogue. This is used to construct the image generation requirement information, allowing the information contained in the second reference image to be integrated with the previous dialogue information. Employing a single information tuple representation allows the prompt optimization model to process and understand the rewriting task, thereby ensuring that the generated target prompts meet expectations and improving the quality of the generated images.
[0182] In other words, regardless of whether it's the first or second reference image, if a reference image is available, a corresponding text description can be generated for this reference image using a combination of image recognition and natural language generation techniques. The generated text description can be determined based on the target dialogue text of the current round of dialogue. For example, if the target object specifies generating a similar subject based on a certain subject in the reference image, then image recognition technology can be used to identify that subject and generate a corresponding description. If the target object specifies generating an image based on the overall layout style of the reference image, then image recognition technology can be used to classify the layout style of the reference image to obtain the corresponding layout style category. In implementation, a multimodal large model within a multimodal intent understanding model can be used to extract reference information from the reference image for constructing the information to be rewritten.
[0183] In multi-turn dialogues, the time and resource costs of processing and analyzing historical dialogue data increase accordingly with the number of dialogue turns. Therefore, in this embodiment of the disclosure, when the number of dialogue turns in the historical dialogue record is greater than m, the historical dialogue text of m turns is selected from the historical dialogue record; where m is a positive integer greater than 1.
[0184] During implementation, the most recent m rounds of historical dialogue can be selected and retained. This ensures that the generated image is more focused on the current drawing needs of the target object, thus making the generated image more in line with the target object's expectations. Furthermore, selecting m rounds of historical dialogue from a large number of historical dialogues can effectively reduce the memory resources required to store and process a large number of historical dialogues, thereby improving the efficiency of image generation.
[0185] In this embodiment of the disclosure, when the current round of dialogue is the first round of dialogue, the target prompt word is the prompt word obtained by the prompt word optimization model rewriting the dialogue text. Based on the rewriting capability of the prompt word optimization model, the dialogue text of the target object can be transformed into standardized and professional instructions that are more in line with the image generation technology requirements of the image generation system, which helps to better understand the raw image requirements of the target object.
[0186] When the current round of dialogue is the latest round in a multi-round dialogue, the target cue word is the cue word obtained by rewriting the specified cue word by the cue word optimization model; wherein, the specified cue word is the historical cue word of the second target round of dialogue indicated by the dialogue text.
[0187] Taking the second round of dialogue as the current round as an example, assuming the current round is the latest round in a multi-round dialogue, if the dialogue text input by the target in the current round is: "The silver sword should be longer," the corresponding "prompts" would be: ["The scene depicts a fantasy anime-style setting. A man wearing leather armor is in a fantastical forest filled with monsters. He carries an iron sword on his back and holds a longer silver sword in his right hand. His eyes are focused, gazing intently at a monster in front of him, giving a sense of fantastical mystery."]. Therefore, it can be seen that the historical prompts in the second round of historical dialogue are rewritten based on the historical prompts in the first round of historical dialogue combined with the target dialogue text of the second round.
[0188] The prompt word obtained by rewriting the specified prompt word through the prompt word optimization model can be: "The screen displays a fantasy anime-style scene. A man wearing leather armor is in a fantasy forest full of monsters. He has an iron sword on his back and a long silver sword in his right hand. His eyes are focused, staring at a monster in front of him, giving people a sense of fantasy and mystery."
[0189] In other words, when generating an image, if the target dialogue text in the current round of dialogue does not explicitly specify which prompt word needs to be rewritten, the historical prompt word from the previous round of dialogue will be used as the specified prompt word by default. Similarly, when generating an image, if the target dialogue text in the current round of dialogue explicitly refers to a historical prompt word, that historical prompt word will be used as the specified prompt word.
[0190] When multiple images are generated, if the target dialogue text in the current round does not explicitly specify which prompt word to rewrite, nor does it specify which image's prompt word to rewrite, the sub-prompt word of a randomly selected image from the previous round will be used as the designated prompt word by default. If the target dialogue text in the current round explicitly specifies which image's prompt word from that round, the sub-prompt word of that explicitly specified round will be used as the designated prompt word.
[0191] In this embodiment of the disclosure, during multi-turn dialogue, the prompt words obtained by rewriting the specified prompt words through the prompt word optimization model can incorporate new image generation requirements based on historical prompt words. The rewritten target prompt words can provide more comprehensive and accurate instructions for the image generation system.
[0192] Based on the prompt word optimization model, the information on the requirements for raw images is fully rewritten and optimized so that various differentiated raw image requirements can be aligned with the expression of the training data of the target raw image model. In other words, it mainly aligns with the expression of the training data, which can also be understood as aligning with the expression that the target raw image model can adapt to.
[0193] In addition, during the alignment optimization process, an art tag system is constructed for the text-to-image scenario. This system comprehensively improves the richness and completeness of the prompt words from various dimensions, including core image content, image style description, image subject and limitation description, image detail description, image background modification description, special effects, composition, color tone, clarity description, and quality description, thus ensuring the artistry of the text-to-image result.
[0194] in:
[0195] The description of an image's style refers to its artistic style, showcasing its representative artistic characteristics, such as anime, illustration, Chinese style, ink painting, cartoon, animation, sketch, line drawing, landscape painting, etc. Descriptions that are merely descriptive, such as photographs, realistic, or lifelike, are not considered artistic styles.
[0196] The subject and limiting description of the image refer to the main subject presented in the image, which may include people, animals, plants, buildings, landscapes, etc., as well as limiting descriptions, such as an 18-year-old girl, a blue sky, and white clouds.
[0197] Detailed description of the image refers to the detailed description of the subject, further refining and enriching the subject. For example, an 18-year-old girl with delicate features, wearing comfortable clothes with beautiful patterns.
[0198] Background embellishment refers to adding background details to the subject to enhance the overall aesthetics of the image and prevent it from appearing simple and empty. For example, a background might be a green meadow with flowers and trees. Furthermore, a blank background, a transparent background, or a solid color background can also make the image more visually appealing.
[0199] Special effects refer to enhancing and enriching visual content, making it more exquisite, and increasing its impact and appeal. Terms include visual effects, lighting, and focus; examples include: thick painting, perfect lighting and shadow, texture, ray tracing, Tyndall effect, blurring, sharp focus, light and shadow, reflection, CG (Computer Graphics) rendering, 3D (Three Dimensions) rendering, multiple exposure, and double exposure.
[0200] Composition refers to depicting the position and state of the subject, that is, appropriately organizing the subject to be represented according to the requirements of the subject to form a harmonious and complete picture. It involves terms such as shot, perspective, and depth of field, including: long shot, close-up, center shot, extreme wide-angle, frontal view, and overhead view.
[0201] Hue refers to the relative brightness and darkness of an image, describing the color of the subject and the background color, such as: white, gold, color, rich color, colorful, rich in color, and multicolored.
[0202] Sharpness description refers to the clarity of the overall image quality, such as: high definition, ultra-high definition, 8K high definition.
[0203] Quality description refers to depicting the overall texture and subjective feeling of the image, such as: masterpiece, highest quality, work of art, award-winning work, immersive, simple, dilapidated.
[0204] When rewriting, the earlier the rewrite order, the more important it is. The content included in the rewrite order depends on the dialogue text (user query) in the image requirement information. For less important content, if it is not mentioned in the dialogue text, it can be omitted from the rewritten prompts. If the dialogue text requests image quality, then the rewritten prompts will include a description of image quality. This applies to both the training and inference phases.
[0205] When the target image model is an accurate model, the layout information of the main content can be automatically planned and generated based on the search query. The layout information, the second reference image, and the image requirement information are rewritten based on the prompt keyword optimization model to obtain prompt keywords adapted to the accurate model, and a response is generated. The layout information refers to the placement of the main content, exemplified by top left, bottom left, top right, and bottom right.
[0206] With the continuous development of artificial intelligence, in some scenarios, although related technologies can provide responses along with the generated images, these responses are difficult to personalize and cannot effectively improve the user experience and satisfaction of the target audience.
[0207] In view of this, in this embodiment of the disclosure, the prompt word optimization model can be trained by fine-tuning on the basis of a large language model to have the ability to rewrite and generate prompt words. At the same time, the prompt word optimization model can generate personalized responses based on the fine-tuning training. Thus, the response text generated by the prompt word optimization model based on the image requirement information can be obtained; the response text, together with the generated image, is fed back to the target object.
[0208] In other words, the prompt word optimization model generates response text tailored to the image requirement information. This text is then paired with the generated image and presented to the target audience to provide a more intuitive and richer feedback experience.
[0209] In this embodiment, for the prompt word optimization model, after constructing the training corpus, the base model can be fine-tuned using SFT (Supervised Fine-Tuning) and DPO (Direct Preference Optimization) training methods to obtain the optimized prompt word model. This optimized prompt word model can provide an interface for providing inference and prediction services. The main function of this service is to rewrite text, obtain high-quality prompt words and personalized responses, and provide them to downstream image generation systems. Furthermore, during the online application process, prompt word examples and response examples can be collected and added to the training corpus to iteratively optimize the prompt word optimization model. For example, for a new scenario, only the corresponding training corpus needs to be constructed to optimize the prompt word optimization model.
[0210] In some embodiments, if the image requirement information includes historical dialogue records, and if the current dialogue is a successor to a multi-turn dialogue, text with potential risks in the historical dialogue records can also be deleted.
[0211] Among them, texts with potential risks can be the risky content identified by the aforementioned security risk identification model.
[0212] In this embodiment of the disclosure, text containing potential risks in historical dialogue records is deleted to avoid public opinion risks and improve the content quality of the generated images.
[0213] In some embodiments, the raw image types described above can be subjected to fine-grained identification, the fine-grained content of which is as follows: Figure 5 As shown, based on the first-round reference image, the target dialogue text for this round, and the historical dialogue, the intent is determined to be a first-round generated image 511 and multiple-round successors 512. The first-round generated image 511 can include text-generated images 521 and image-generated images 522. Multiple-round successors 512 can include text-generated image successors 523 and image-generated image successors 524. Text-generated image successors 523 further include requirement design class 541, redraw class 542, and editing class 543. Requirement design class 541 further includes changing the requirement type 551, which means changing the type of the generated image; for example, changing a poster to a wide image. Editing class 543 includes modifying aspect ratio 552, adding elements 553, deleting elements 554, replacing elements 555, and style transformation 556, etc.
[0214] Furthermore, the subsequent image 524 is based on at least one of the first reference image, the second reference image, and the reference image generated in the previous round in reference image 531. The resulting intents include requirement design type 544, redraw type 545, editing type 546, and creative similarity type 547. Requirement design type, redraw type, and editing type are similar to those described above. Creative similarity type 547 includes reference image subject similarity 567, reference image background similarity 568, and reference image style similarity 569.
[0215] In summary, the architecture diagram of the prompting engineering system provided in this embodiment is as follows: Figure 6 As shown, the entire system architecture is divided into multiple small models and three layers of cognitive capabilities. These three layers of cognitive capabilities can include cognitive correction, thinking planning, and basic knowledge. Basic knowledge is used to obtain the results of intent recognition. Basic knowledge includes the ability to determine security intent 611 based on the security risk recognition model 610, where security intent is used to identify risk information corresponding to the raw image requirement information. Basic knowledge also includes various intents obtained from multimodal intent understanding 620, including reference image intent 621, multi-round intent 622 (including first-round text-to-image type, first-round image-to-image type, text-to-image multi-round successor type, and image-to-image multi-round successor type in raw image types), size intent 623 (i.e., identifying the image size in the raw image intent information), type intent 624 (i.e., identifying the model type), style intent 625 (i.e., image style), subject intent 626 (i.e., target subject type), and editing intent 627 (i.e., image editing type in the raw image type). Among them, the reference graph intent 621 (i.e., the reference graph information in the birth graph intent information) can be used to perform content understanding on the first reference graph provided by the target object, extract the types of reference subjects included, and the number of reference subjects. The multi-turn intent 622 is used in multi-turn dialogue scenarios to delete risky content in historical dialogues based on the risks identified by the security intent, and then identify the birth graph type.
[0216] The intention recognition results are used for planning, which mainly includes retrieval enhancement planning 630 and prompt information alignment enhancement planning 640. Retrieval enhancement planning 630 includes whether to enhance retrieval 631, retrieval query planning 632, raw image model planning 633, and subject layout planning 634. Whether to enhance retrieval 631 determines whether the intention recognition result needs retrieval enhancement; retrieval query planning 632 generates retrieval statements based on raw image requirement information; raw image model planning 633 retrieves a second reference image based on the retrieval statement; and subject layout planning 634 determines the layout information of the target subject type. Prompt information alignment enhancement planning 640 includes prompt information structure alignment 641, prompt information subject semantic attribute enhancement 642, prompt information visual art enhancement 643, and personalized response generation 644. Prompt information structure alignment 641 aligns to an expression mode that the target raw image model can adapt to; prompt information subject semantic attribute enhancement 642 semantically enhances the subject in the second reference image and / or raw image requirement information; and prompt information visual art enhancement 643 enhances content such as art tags. Personalized response generation 644 is used to generate personalized responses based on raw image requirements and the generated images.
[0217] Cognitive correction includes an online system intervention mechanism 650 and an offline evaluation and reflection optimization mechanism 660. The online system intervention mechanism 650 includes a recall mechanism 651 and a ranking mechanism 652, which are used to optimize the recall and ranking mechanisms of the online system, such as quickly avoiding risks that arise in real time. The offline evaluation and reflection optimization mechanism 660 includes a multi-dimensional judgment-based evaluation task 661 and an evaluation feedback thought chain optimization 662. The multi-dimensional judgment-based evaluation task 661 and the evaluation feedback thought chain optimization 662 are used to reflect on and optimize the model's recognition results, so as to output more accurate intent labels.
[0218] In some embodiments, the prompt word optimization model in this disclosure may include a prompt information rewriting model based on a second reference graph and a general prompt information rewriting model. Based on this, the prompt engineering system is described from the perspective of a neural network model, such as... Figure 7As shown: For the raw image requirement text 71 in the raw image requirement information, the security risk identification model 72 is invoked to perform security identification to obtain risk information. If the risk information indicates that there is no risk, then text requirement information is constructed, and the first intent identification sub-model 731 in the multimodal intent understanding model 73 is used to analyze the raw image intent information to obtain information such as raw image type and image size. If the risk information indicates that there is a risk, it is adjusted. After adjustment, it is input into the first intent identification sub-model 731 in the multimodal intent understanding model 73 to obtain information such as raw image type and image size in its raw image intent information. Where the target object provides a first reference image 78, the first reference image 78 can be processed by the image-based risk identification model 79. If it is determined that the first reference image 78 does not pose a security risk, the multimodal large model 710 is used to identify the reference information of the first reference image as a supplement to the aforementioned raw image intent information. Furthermore, when the target object provides a first reference image 78 and a raw image requirement text 71, and it is determined that neither of them poses a security risk, a multimodal large model 710 is used to identify the reference information of the first reference image, and a second intent recognition sub-model 732 is used to identify the intent information in the raw image requirement text. Based on the intent information in the raw image requirement text and the reference information of the first reference image, the raw image intent information is obtained.
[0219] Following the aforementioned raw image intent information, the distribution strategy model 74 identifies the target subject type in the raw image demand information and determines whether the raw image intent information needs enhancement. If enhancement is required, a second reference image is retrieved. The second reference image and the text demand information are then provided to the prompt information rewriting model 77 based on the second reference image for rewriting, resulting in target prompt words. If no retrieval enhancement is needed, the general prompt information rewriting model 76 is used to rewrite the text demand information to obtain target prompt words.
[0220] In addition, such as Figure 7 As shown, the distribution strategy model 74 can determine the model type of the target image generation model used to generate images in this round based on the image intent information and text requirement information. The selectable model types include base model 741, accurate model 742, high generalization model 743, and editing model 744.
[0221] The system can use a layout planning model to generate layout information 75 for the target subject type of the image generation requirement text 71. This layout information 75 can then be used by the precise model 742 in the image generation system.
[0222] This disclosure also provides a prompting engineering system, such as... Figure 8 As shown, it includes at least one of the following services:
[0223] Service 801 is used to process the raw image requirement information of the target object for generating images and to determine the risk information corresponding to the raw image requirement information.
[0224] In addition to determining risk information, the first service can also determine the raw image intent information corresponding to the raw image demand information.
[0225] The second service 802 is used to determine the model type of the target raw image model that is adapted to the raw image requirement information;
[0226] The third service, 803, is used to rewrite the target prompts obtained from the raw image requirement information.
[0227] In this embodiment of the disclosure, multiple services are used to apply to different text-to-image scenarios. The multimodal text-to-image requirement information and multi-turn historical dialogue information of the target object are used as initial input content and provided to the prompting engineering system. The system will output risk information, model type of the target text-to-image model, prompt words, etc. Based on the optimized information, the generated image is closer to the user's expectations, thus improving the user experience.
[0228] In some embodiments, the second service is specifically used for:
[0229] Determine the intent information of the generated image corresponding to the generated image requirement information;
[0230] Based on the raw image intent information, determine the model type of the target raw image model that is adapted to the raw image requirement information.
[0231] The target image model includes at least one of the following model types:
[0232] The base model supports general image generation tasks;
[0233] Accurate models are used to generate high-fidelity images.
[0234] High generalization model: The high generalization model is used to generate images that meet the preset fidelity requirements and have artistic effects;
[0235] Editing models are used to perform editing operations on known images.
[0236] In this embodiment of the disclosure, based on the second service, the intention information of the raw image corresponding to the raw image requirement information can be clearly understood, so as to adapt the model type of the target raw image model of the raw image requirement information, so that the model can accurately generate the image.
[0237] In some embodiments, the third service is also used to generate response information in response to the raw image request information.
[0238] The response information is personalized and can adapt to various scenarios.
[0239] In this embodiment, a third service is used to generate response text adapted to the image requirement information. This text is then paired with the generated image and fed back to the target audience, providing them with a more intuitive and richer feedback experience.
[0240] In some embodiments, a fourth service is also included for determining the layout information of the target subject type in the image to be generated based on the image generation requirements.
[0241] During implementation, since the raw image requirement information may have requirements for the placement of the target subject, the target subject can be placed based on the requirements in the raw image requirement information. If the raw image requirement information does not have layout information, and it is determined that search enhancement is needed, the layout information of the target subject type can be determined based on the second reference image.
[0242] In this embodiment of the disclosure, using the layout information of the target subject type can make the generated image more in line with user expectations and improve the user experience.
[0243] For example, in a multi-turn dialogue, such as Figure 9 As shown, Image901 (the first reference image in the raw image requirement information), Query902 (the target dialogue text in the current round), and History903 (historical dialogue information) can be input into the prompting engineering system904 to realize the input of raw image requirement information into the unified multi-round raw image type understanding and security understanding service905 (i.e., the first service).
[0244] The historical dialogue information includes at least one of the following:
[0245] Prompt indicates a prompt message; including a rewritten Prompt generated from the dialogue history.
[0246] sub_task_id represents the ID of the raw image task;
[0247] refer_image_caption represents the descriptive text of the second reference image;
[0248] risk_type indicates the risk level;
[0249] is_rag_trigger indicates whether to trigger search enhancement;
[0250] is_user_refer_image indicates whether the target object has uploaded the first reference image;
[0251] search_query_list represents the rewriting requirements after search enhancement;
[0252] query_type_list represents the subject type for intent understanding;
[0253] "Styles" refers to the visual style, such as realistic or artistic.
[0254] Specifically, the first service 905 calls the security risk identification model (i.e., the security intent operator 909) to process the text in the raw image requirement information, identify risks, and filter out or rewrite risky content. The first service then calls the multimodal intent understanding model (i.e., the multi-turn intent operator 910) to obtain the raw image type, which can include raw image type 1 and raw image type 2, wherein:
[0255] imggen_type1 indicates that the raw image type is 1;
[0256] imggen_type2 represents raw image type 2; raw image type 2 represents the fine-grained raw image intent of raw image type 1, such as the specific operations required in the image editing type, such as deleting local elements or modifying elements;
[0257] refer_dialogue_id represents the reference round ID, and the first round of dialogue is used as the reference round ID for a single image task;
[0258] edit_elements represents the edit elements, that is, which main elements in the reference image are edited when a reference image exists;
[0259] risk_level indicates the risk level;
[0260] risk_cls_type indicates the risk type.
[0261] The information obtained above is input into the unified multi-round iRAG trigger check understanding service 906 (i.e., the aforementioned second service). The second service depends on the iRAG trigger operator 911 to obtain the following:
[0262] is_user_refer_imag: Whether a first reference graph is present;
[0263] search_query_list: A list of enhanced search queries;
[0264] query_type_list is intended to understand the target subject type;
[0265] "Styles" refers to the visual style.
[0266] Then, the general multi-round Prompt rewriting generation service 907 (the third service) is input. The third service calls the general multi-round Prompt rewriting operator 912 (i.e., the prompt word optimization model) to realize the target prompt words obtained by rewriting the raw image requirement information and the response information for the raw image requirement information.
[0267] In addition, before the unified multi-round iRAG trigger check understanding service 906 and the general multi-round Prompt rewrite generation service 907 call the corresponding operators, History 903 can also be processed based on History preprocessing strategy 908, including, for example, deleting risky content in History 903.
[0268] This disclosure also provides a security management strategy, such as... Figure 10 As shown, the process can be divided into three stages: training, generation, and propagation. The training stage includes risk element cleaning of the training dataset (1001), used to collect the training dataset and clean up any risk elements within it. The generation stage includes input text / image review strategies (1002), generated image review strategies (1003), prompt information security optimization strategies (1004), image retrieval enhancement strategies (1005), and image-text intervention strategies based on knowledge strategies and text semantics (1006). Specifically, the input text / image review strategy (1002) reviews the input text / images from the target object, and the generated image review strategy (1003) reviews the generated images. The prompt information security optimization strategy (1004) uses a rewriting model to generate prompt information after identifying risks based on the security risk identification model. The image retrieval enhancement strategy (1005) determines whether the intent information of the raw image needs retrieval enhancement. The image-text intervention strategy based on knowledge strategies and text semantics (1006) supplements the security risk identification capability. Because the security risk identification model is trained based on training samples, its recognition ability may need to be strengthened as new vocabulary emerges. Optimizing the model by accumulating training samples is time-consuming, and using coding for risk identification also increases development costs. Therefore, in this embodiment, the model's ability to understand images and text can be utilized. The model can be configured with the descriptions of information to be filtered, allowing for real-time risk identification through its understanding capabilities. In the event of an emergency, the intervention system can intervene directly without coding, avoiding large-scale public opinion problems. A manual review mechanism (1007) can be introduced during the dissemination phase to ensure that image content meets relevant requirements. In summary, based on multiple protection mechanisms, the security of the entire alert system is improved.
[0269] Based on the same technical concept, this disclosure also provides an information processing system 1100 for image generation, such as... Figure 11 As shown, it includes:
[0270] The prompting engineering system 1101 is used to process the image generation requirement information of the target object, so as to convert the image generation requirement information into image generation parameters for use by the image generation system based on the prompting engineering system;
[0271] The generated image parameters include at least one of the following:
[0272] The risk information corresponding to the raw image requirement information;
[0273] The raw image requirement information corresponds to the raw image intent information;
[0274] The model type of the target image model that adapts to the image generation requirements;
[0275] The target prompt words are obtained by rewriting the raw image requirement information;
[0276] Response information regarding the requested raw image;
[0277] The image generation requirement information refers to the layout information of the target subject type in the generated image.
[0278] Image generation system 1102 is used to generate a target image based on the image parameters.
[0279] When the raw image parameters include risk information corresponding to the raw image requirement information, the image generation system determines whether the raw image requirement information is risky based on the risk information corresponding to the raw image requirement information. If there is no risk, subsequent raw image processing is performed; if there is risk, risk removal processing is performed. The specific methods for removing risk information have been explained above and will not be repeated here.
[0280] When the image generation parameters include the image generation intent information corresponding to the image generation requirement information, the image generation system can obtain the explicit image generation intent of the image generation requirement information.
[0281] For example, if the target image model's model type is included in the image generation parameters as part of the image generation requirement information, the image generation system calls the corresponding model type as the target image model to generate the target image.
[0282] When the target prompt word is included in the raw image parameters, the image generation system can input the target prompt word into the target raw image model so that the target raw image model can generate the target image based on the target prompt word.
[0283] When the raw image information includes response information, the image generation system needs to output the response information along with the target image generated by the target raw image model to the target object.
[0284] When the raw image parameters include layout information of the target subject type, the image generation system needs to provide this layout information to the target raw image model in order to limit the layout of the generated target image.
[0285] In this embodiment of the disclosure, the use of the prompting engineering system can make the raw image parameters more explicit, thereby improving the accuracy of the target image obtained by the image generation system based on the raw image parameters and making it closer to the user's expectations.
[0286] Based on the same technical concept, this disclosure also provides an information processing apparatus 1200 for image generation, such as... Figure 12 As shown, it includes:
[0287] The acquisition unit 1201 is used to acquire the raw image requirement information of the target object for generating the image;
[0288] The processing unit 1202 is used to input the raw image requirement information into the prompting engineering system, so as to convert the raw image requirement information into raw image parameters for use by the image generation system based on the prompting engineering system.
[0289] In some embodiments, the image parameters include at least one of the following:
[0290] The risk information corresponding to the raw image requirement information;
[0291] The raw image requirement information corresponds to the raw image intent information;
[0292] The model type of the target image model that adapts to the image generation requirements;
[0293] The target prompt words are obtained by rewriting the raw image requirement information;
[0294] Response information regarding the requested raw image;
[0295] The image generation requirement information refers to the layout information of the target subject type in the generated image.
[0296] In some embodiments, the processing unit includes:
[0297] The risk identification subunit is used to input the requirement text in the raw image requirement information into the security risk identification model to identify the risk information corresponding to the raw image requirement information.
[0298] The risk information includes the risk type and / or risk level.
[0299] In some embodiments, the processing unit includes:
[0300] The intent acquisition subunit is used to identify the raw image requirement information based on a multimodal intent understanding model, and obtain the raw image intent information.
[0301] In some embodiments, the intent to obtain the subunit is specifically used for:
[0302] Based on the requirement text and target prompt template in the raw image requirement information, construct text understanding prompt words;
[0303] The text understanding prompts are input into the first intent recognition sub-model in the multimodal intent understanding model to obtain at least one of the following information in the raw image intent information: raw image type, image size, and image style;
[0304] The raw image type is a type from a preset type set.
[0305] In some embodiments, the preset type set includes at least one of the following types:
[0306] First-round text-to-image type, first-round image-to-image type, text-to-image multiple-round successor type, image-to-image multiple-round successor type, image editing type.
[0307] In some embodiments, a first training unit is further included, for:
[0308] Training sample sets with different raw image intent labels were distilled based on a large language model and target cue templates;
[0309] The training sample set is classified based on the raw image intent labels to obtain sample subsets corresponding to each raw image intent label;
[0310] Based on the sample subsets corresponding to each of the raw image intent labels, supervised training is performed on the pre-trained initial model to obtain the first intent recognition sub-model.
[0311] In some embodiments, the first training unit is specifically used to: obtain an initial sample set;
[0312] Based on the sequence labeling task, multiple intermediate samples are selected from the initial sample set to obtain an intermediate sample set;
[0313] Construct sample prompt words based on the target prompt template and the intermediate sample set;
[0314] A large language model is used to generate the raw image intent labels for each sample in the intermediate sample set to obtain the training sample set.
[0315] In some embodiments, an adjustment unit is further included, for:
[0316] When using the prompt template to be adjusted, the initial intent label of the target sample set is predicted based on the first intent recognition sub-model;
[0317] The large language model is used to determine whether the initial intent label meets the requirements of multiple preset rules in the rule set;
[0318] If any preset rule requirement is not met, the prompt template to be adjusted is optimized based on the large language model and the preset rule requirement as a benchmark to obtain the target prompt template.
[0319] In some embodiments, the intent to obtain the subunit is further configured to:
[0320] When the image generation requirement information includes a first reference image provided by the target object, reference image information of the first reference image is obtained based on the multimodal large model in the multimodal intent understanding model; the image generation intent information includes the reference image information.
[0321] The reference image information includes at least one of the following: the type of the first reference subject, the number of the first reference subjects, and the text description information of the first reference image.
[0322] In some embodiments, the intent to obtain the subunit is further configured to:
[0323] The requirement text in the raw image requirement information is input into the second intent recognition sub-model in the multimodal intent understanding model to obtain the target subject type indicated by the requirement text; the raw image intent information includes the target subject type.
[0324] In some embodiments, the processing unit includes:
[0325] The model determination subunit is used to process the raw image intent information based on the distribution strategy model and determine the model type of the target raw image model in the image generation system that generates images for the raw image demand information.
[0326] In some embodiments, the model type of the target raw image model includes at least one of the following:
[0327] A base model that supports general image generation tasks;
[0328] A precise model used to generate high-fidelity images;
[0329] A high generalization model is used to generate images that meet preset fidelity requirements and have artistic effects.
[0330] An editing model is used to perform editing operations on a known image.
[0331] In some embodiments, a determining unit is further included, configured to:
[0332] If the target subject type in the image to be generated according to the image generation requirements is a preset type, the model type of the target image generation model is a secure model.
[0333] The security model is used to generate images that mitigate public opinion risks.
[0334] In some embodiments, the processing unit includes:
[0335] The processing subunit is used to process the image generation requirement information of the target object based on the distribution strategy model to obtain the processing result;
[0336] A generation subunit is used to obtain the retrieval statement generated by the distribution strategy model from the processing result when the processing result indicates that retrieval enhancement is needed;
[0337] The retrieval subunit is used to retrieve the second reference image of the image requirement information based on the retrieval statement;
[0338] The rewriting subunit is used to rewrite the second reference image and the image generation requirement information based on the prompt word optimization model to obtain the target prompt words that are adapted to the target image generation model in the image generation system, and to generate the response information.
[0339] In some embodiments, the rewriting subunit is specifically used for:
[0340] Obtain the target dialogue text of the current round of dialogue of the target object from the raw image requirement information;
[0341] Based on the target dialogue text, the historical dialogue information in the raw image requirement information, and the second reference image, construct the information to be rewritten;
[0342] The information to be rewritten is input into the prompt word optimization model to obtain the target prompt word.
[0343] In some embodiments, where the raw image request information includes historical conversation records, a deletion unit is further included, for:
[0344] If the current dialogue is a successor to a multi-turn dialogue in the historical dialogue record, text with potential risks will be deleted.
[0345] The specific functions and examples of each module and submodule of the apparatus in this disclosure can be found in the relevant descriptions of the corresponding steps in the above method embodiments, and will not be repeated here.
[0346] The acquisition, storage, and application of user personal information involved in the technical solution disclosed herein comply with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0347] According to embodiments of this disclosure, this disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0348] Figure 13 A schematic block diagram of an example electronic device 1300 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0349] like Figure 13 As shown, device 1300 includes a computing unit 1301, which can perform various appropriate actions and processes according to a computer program stored in read-only memory (ROM) 1302 or a computer program loaded from storage unit 1308 into random access memory (RAM) 1303. The RAM 1303 may also store various programs and data required for the operation of device 1300. The computing unit 1301, ROM 1302, and RAM 1303 are interconnected via bus 1304. Input / output (I / O) interface 1305 is also connected to bus 1304.
[0350] Multiple components in device 1300 are connected to I / O interface 1305, including: input unit 1306, such as keyboard, mouse, etc.; output unit 1307, such as various types of monitors, speakers, etc.; storage unit 1308, such as disk, optical disk, etc.; and communication unit 1309, such as network card, modem, wireless transceiver, etc. Communication unit 1309 allows device 1300 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0351] The computing unit 1301 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 1301 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 1301 performs the various methods and processes described above, such as information processing methods for image generation. For example, in some embodiments, the information processing methods for image generation can be implemented as computer software programs tangibly contained in a machine-readable medium, such as storage unit 1308. In some embodiments, part or all of the computer program can be loaded and / or installed on device 1300 via ROM 1302 and / or communication unit 1309. When the computer program is loaded into RAM 1303 and executed by the computing unit 1301, one or more steps of the information processing methods for image generation described above can be performed. Alternatively, in other embodiments, the computing unit 1301 may be configured to perform an information processing method for image generation by any other suitable means (e.g., by means of firmware).
[0352] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0353] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0354] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. Machine-readable media can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0355] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0356] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with embodiments of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.
[0357] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, servers in distributed systems, or servers incorporating blockchain technology.
[0358] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.
[0359] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the principles of this disclosure should be included within the scope of protection of this disclosure.
Claims
1. An information processing method for image generation, comprising: Obtain the raw image requirements of the target object for generating the image; The image generation system inputs the image generation requirement information into a prompting engineering system, and converts the image generation requirement information into image generation parameters for use by the image generation system based on the prompting engineering system. This includes: inputting the image generation intent information of the image generation requirement information into a distribution strategy model to obtain the model type of the target image generation model that is adapted to the image generation requirement information. The raw image demand information is processed based on a distribution strategy model to obtain a processing result; the processing result indicates whether retrieval enhancement of the raw image demand information is needed. Based on the processing results, target prompts are generated using a prompt word optimization model. In multi-turn dialogues, the target prompts are obtained by rewriting specified prompts using the prompt word optimization model, so as to incorporate new image requirements into historical prompts. The parameters of the generated image include the model type and the target prompt words.
2. The method according to claim 1, wherein, The raw image parameters also include at least one of the following: The risk information corresponding to the raw image requirement information; The raw image requirement information corresponds to the raw image intent information; Response information regarding the requested raw image; The image generation requirement information refers to the layout information of the target subject type in the generated image.
3. The method according to claim 2, wherein, The risk information is determined in the following ways, including: Input the requirement text in the raw image requirement information into the security risk identification model to identify the risk information corresponding to the raw image requirement information; The risk information includes the risk type and / or risk level.
4. The method according to claim 2, wherein, The raw image intent information is determined in the following ways, including: The raw image requirement information is identified based on a multimodal intent understanding model to obtain the raw image intent information.
5. The method according to claim 4, wherein, The method of identifying the raw image requirement information based on the multimodal intent understanding model to obtain the raw image intent information includes: Based on the requirement text and target prompt template in the raw image requirement information, construct text understanding prompt words; The text understanding prompts are input into the first intent recognition sub-model in the multimodal intent understanding model to obtain at least one of the following information in the raw image intent information: raw image type, image size, and image style; The raw image type is a type from a preset type set.
6. The method according to claim 5, wherein the preset type set includes at least one of the following types: First-round text-to-image type, first-round image-to-image type, text-to-image multiple-round successor type, image-to-image multiple-round successor type, image editing type.
7. The method according to claim 5, wherein, The first intent recognition sub-model is obtained through the following training methods: Training sample sets with different raw image intent labels were distilled based on a large language model; The training sample set is classified based on the raw image intent labels to obtain sample subsets corresponding to each raw image intent label; Based on the sample subsets corresponding to each of the raw image intent labels, supervised training is performed on the pre-trained initial model to obtain the first intent recognition sub-model.
8. The method according to claim 7, wherein, The training sample set based on the distillation of different raw image intent labels using a large language model includes: Obtain the initial sample set; Based on the sequence labeling task, multiple intermediate samples are selected from the initial sample set to obtain an intermediate sample set; A large language model is used to generate the raw image intent labels for each sample in the intermediate sample set to obtain the training sample set.
9. The method according to claim 7, further comprising: When using the prompt template to be adjusted, the initial intent label of the target sample set is predicted based on the first intent recognition sub-model; The large language model is used to determine whether the initial intent label meets the requirements of multiple preset rules in the rule set; If any preset rule requirement is not met, the prompt template to be adjusted is optimized based on the large language model and the preset rule requirement as a benchmark to obtain the target prompt template.
10. The method according to claim 4, wherein, The method of identifying the raw image requirement information based on the multimodal intent understanding model to obtain the raw image intent information includes: When the image generation requirement information includes a first reference image provided by the target object, reference image information of the first reference image is obtained based on the multimodal large model in the multimodal intent understanding model; the image generation intent information includes the reference image information. The reference image information includes at least one of the following: the type of the first reference subject, the number of the first reference subjects, and the text description information of the first reference image.
11. The method according to claim 4, wherein, The method of identifying the raw image requirement information based on the multimodal intent understanding model to obtain the raw image intent information includes: The requirement text in the raw image requirement information is input into the second intent recognition sub-model in the multimodal intent understanding model to obtain the target subject type indicated by the requirement text; the raw image intent information includes the target subject type.
12. The method according to claim 1, wherein, The target image model includes at least one of the following model types: A base model that supports general image generation tasks; A precise model used to generate high-fidelity images; A high generalization model is used to generate images that meet preset fidelity requirements and have artistic effects. An editing model is used to perform editing operations on a known image.
13. The method of claim 12, further comprising: If the target subject type in the image to be generated according to the image generation requirements is a preset type, the model type of the target image generation model is a secure model. The security model is used to generate images that mitigate public opinion risks.
14. The method according to claim 2, wherein, The target prompts and response information for the raw image request information are determined in the following ways: The target object's image generation requirements information is processed based on the distribution strategy model to obtain the processing result; If the processing result indicates that retrieval enhancement is needed, the retrieval statement generated by the distribution strategy model is obtained from the processing result; Based on the search query, a second reference image for the raw image requirement information is retrieved; The second reference image and the raw image requirement information are rewritten based on the prompt word optimization model to obtain the target prompt words that are adapted to the target raw image model in the image generation system, and the response information is generated.
15. The method according to claim 14, wherein, The step of rewriting the second reference image and the raw image requirement information based on the prompt word optimization model to obtain target prompt words that adapt to the target raw image model in the image generation system includes: Obtain the target dialogue text of the current round of dialogue of the target object from the raw image requirement information; Based on the target dialogue text, the historical dialogue information in the raw image requirement information, and the second reference image, construct the information to be rewritten; The information to be rewritten is input into the prompt word optimization model to obtain the target prompt word.
16. The method according to claim 15, wherein, If the image request information includes historical conversation records, it also includes: If the current dialogue is a follow-up to a multi-turn dialogue, text with potential risks in the historical dialogue record will be deleted.
17. An information processing apparatus for image generation, comprising: The acquisition unit is used to acquire the raw image requirement information of the target object for generating images; The processing unit is used to input the image generation requirement information into the prompting engineering system, so as to convert the image generation requirement information into image generation parameters for use by the image generation system based on the prompting engineering system, including: a model determination subunit, used to input the image generation intent information of the image generation requirement information into the distribution strategy model to obtain the model type of the target image generation model that is adapted to the image generation requirement information; The raw image demand information is processed based on a distribution strategy model to obtain a processing result; the processing result indicates whether retrieval enhancement of the raw image demand information is needed. The processing unit is further configured to generate target prompt words based on the prompt word optimization model based on the processing result; wherein, in multi-turn dialogue, the target prompt words are obtained by rewriting the specified prompt words through the prompt word optimization model, so as to incorporate new raw image requirements into the historical prompt words; The parameters of the generated image include the model type and the target prompt words.
18. The apparatus according to claim 17, wherein, The raw image parameters also include at least one of the following: The risk information corresponding to the raw image requirement information; The raw image requirement information corresponds to the raw image intent information; Response information regarding the requested raw image; The image generation requirement information refers to the layout information of the target subject type in the generated image.
19. The apparatus according to claim 18, wherein, The processing unit includes: The risk identification subunit is used to input the requirement text in the raw image requirement information into the security risk identification model to identify the risk information corresponding to the raw image requirement information. The risk information includes the risk type and / or risk level.
20. The apparatus according to claim 18, wherein, The processing unit includes: The intent acquisition subunit is used to identify the raw image requirement information based on a multimodal intent understanding model, and obtain the raw image intent information.
21. The apparatus according to claim 20, wherein, The intent acquisition subunit is specifically used for: Based on the requirement text and target prompt template in the raw image requirement information, construct text understanding prompt words; The text understanding prompts are input into the first intent recognition sub-model in the multimodal intent understanding model to obtain at least one of the following information in the raw image intent information: raw image type, image size, and image style; The raw image type is a type from a preset type set.
22. The apparatus according to claim 21, wherein the preset type set includes at least one of the following types: First-round text-to-image type, first-round image-to-image type, text-to-image multiple-round successor type, image-to-image multiple-round successor type, image editing type.
23. The apparatus of claim 21, further comprising a first training unit, configured to: Training sample sets with different raw image intent labels were distilled based on a large language model; The training sample set is classified based on the raw image intent labels to obtain sample subsets corresponding to each raw image intent label; Based on the sample subsets corresponding to each of the raw image intent labels, supervised training is performed on the pre-trained initial model to obtain the first intent recognition sub-model.
24. The apparatus according to claim 23, wherein, The first training unit is specifically used for: Obtain the initial sample set; Based on the sequence labeling task, multiple intermediate samples are selected from the initial sample set to obtain an intermediate sample set; Construct sample prompt words based on the target prompt template and the intermediate sample set; A large language model is used to generate the raw image intent labels for each sample in the intermediate sample set to obtain the training sample set.
25. The apparatus of claim 23, further comprising an adjustment unit for: When using the prompt template to be adjusted, the initial intent label of the target sample set is predicted based on the first intent recognition sub-model; The large language model is used to determine whether the initial intent label meets the requirements of multiple preset rules in the rule set; If any preset rule requirement is not met, the prompt template to be adjusted is optimized based on the large language model and the preset rule requirement as a benchmark to obtain the target prompt template.
26. The apparatus according to claim 20, wherein, The intent acquisition subunit is also used for: When the image generation requirement information includes a first reference image provided by the target object, the reference image information of the first reference image is obtained based on the multimodal large model in the multimodal intent understanding model. The raw image intent information includes the reference image information; The reference image information includes at least one of the following: the type of the first reference subject, the number of the first reference subjects, and the text description information of the first reference image.
27. The apparatus according to claim 20, wherein, The intent acquisition subunit is also used for: The requirement text in the raw image requirement information is input into the second intent recognition sub-model in the multimodal intent understanding model to obtain the target subject type indicated by the requirement text; the raw image intent information includes the target subject type.
28. The apparatus according to claim 17, wherein, The target image model includes at least one of the following model types: A base model that supports general image generation tasks; A precise model used to generate high-fidelity images; A high generalization model is used to generate images that meet preset fidelity requirements and have artistic effects. An editing model is used to perform editing operations on a known image.
29. The apparatus of claim 28, further comprising a determining unit, configured to: If the target subject type in the image to be generated according to the image generation requirements is a preset type, the model type of the target image generation model is a secure model. in, The security model is used to generate images that mitigate public opinion risks.
30. The apparatus according to claim 18, wherein, The processing unit includes: The processing subunit is used to process the image generation requirement information of the target object based on the distribution strategy model to obtain the processing result; A generation subunit is used to obtain the retrieval statement generated by the distribution strategy model from the processing result when the processing result indicates that retrieval enhancement is needed; The retrieval subunit is used to retrieve the second reference image of the image requirement information based on the retrieval statement; The rewriting subunit is used to rewrite the second reference image and the image generation requirement information based on the prompt word optimization model to obtain the target prompt words that are adapted to the target image generation model in the image generation system, and to generate the response information.
31. The apparatus according to claim 30, wherein, The rewriting subunit is specifically used for: Obtain the target dialogue text of the current round of dialogue of the target object from the raw image requirement information; Based on the target dialogue text, the historical dialogue information in the raw image requirement information, and the second reference image, construct the information to be rewritten; The information to be rewritten is input into the prompt word optimization model to obtain the target prompt word.
32. The apparatus according to claim 31, wherein, If the image request information includes historical conversation records, a deletion unit is also included, used for: If the current dialogue is a successor to a multi-turn dialogue in the historical dialogue record, text with potential risks will be deleted.
33. An information processing system for image generation, comprising: A prompting engineering system is used to process the image generation requirement information of a target object, and to convert the image generation requirement information into image generation parameters for use by an image generation system. The system includes: inputting the image generation intent information of the image generation requirement information into a distribution strategy model to obtain a model type for a target image generation model adapted to the image generation requirement information; processing the image generation requirement information based on the distribution strategy model to obtain a processing result; the processing result indicating whether retrieval enhancement of the image generation requirement information is needed; and generating target prompt words based on a prompt word optimization model based on the processing result. In multi-turn dialogues, the target prompt words are obtained by rewriting specified prompt words through the prompt word optimization model to incorporate new image generation requirements into historical prompt words. The image generation parameters include the model type and the target prompt words. An image generation system for generating a target image based on the generated image parameters.
34. An electronic device, comprising: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-16.
35. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-16.
36. A computer program product comprising a computer program that, when executed by a processor, implements the method according to any one of claims 1-16.
Citation Information
Patent Citations
Model training method and device, image classification method and device, electronic equipment and medium
CN115482390A
Image automatic generation method and device based on AIGC, equipment and medium
CN117496302A
Prompt word attack detection method and device for large language model
CN118445815A