Image generation method, apparatus, intelligent agent, intelligent agent system, and storage medium
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- BEIJING BAIDU NETCOM SCI & TECH CO LTD
- Filing Date
- 2025-06-17
- Publication Date
- 2026-08-04
AI Technical Summary
【0013】 本開示によって提供される画像生成方法、装置、インテリジェントエージェント、インテリジェントエージェントシステム及び記憶媒体には、以下のような有益な効果が存在する。
Smart Images

Figure 0007900567000001 
Figure 0007900567000002 
Figure 0007900567000003
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technology, specifically to the fields of technologies such as computer vision, deep learning, and large-scale models, and can be applied to scenarios of Artificial Intelligence Generated Content (AIGC) based on the content of artificial intelligence. In particular, it relates to an image generation method, apparatus, intelligent agent, intelligent agent system, and storage medium.
Background Art
[0002] The core of AI image generation technology based on AIGC is to use an AI image generation model to realize the conversion from text to image or from image to image. However, the image generation effect of AI image generation technology highly depends on the quality and timeliness of model training data. Specifically, after the training of the AI image generation model is completed, the generated image content is often restricted by the timeliness of the model training data. Before the model training data is updated, the AI image generation model cannot obtain the latest content, and there may be a certain delay in the generated image.
Summary of the Invention
[0003] The present disclosure provides an image generation method, apparatus, intelligent agent, intelligent agent system, and storage medium.
[0004] According to a first aspect of the present disclosure, an image generation method is provided. The method includes steps of obtaining image generation needs information, determining a corresponding target image generation method based on the image generation needs information, querying and obtaining a first reference image based on the image generation needs information, and generating a target image using the target image generation method based on the image generation needs information and the first reference image.
[0005] A second aspect of the present disclosure provides an image generation method, the method comprising: acquiring image generation needs information; determining a corresponding target image generation method based on the image generation needs information; determining whether the image generation needs information requires a reference image query; if the image generation needs information requires a reference image query, querying and acquiring a first reference image based on the image generation needs information; and generating a target image using the target image generation method based on the image generation needs information and the first reference image.
[0006] A third aspect of the present disclosure provides an image generation apparatus, the apparatus comprising: an acquisition module for acquiring image generation needs information; a determination module for determining a corresponding target image generation method based on the image generation needs information; a query module for querying and acquiring a first reference image based on the image generation needs information; and a generation module for generating a target image using the target image generation method based on the image generation needs information and the first reference image.
[0007] A fourth aspect of the present disclosure provides an image generation apparatus, the apparatus comprising: an acquisition module for acquiring image generation needs information; a decision module for determining a corresponding target image generation method based on the image generation needs information; a determination module for determining whether the image generation needs information requires a reference image query; a query module for obtaining a first reference image by querying based on the image generation needs information if the image generation needs information requires a reference image query; and a generation module for generating a target image using the target image generation method based on the image generation needs information and the first reference image.
[0008] A fifth aspect of the present disclosure provides an intelligent agent which includes an input module for acquiring image generation needs information, a processing module for determining a corresponding target image generation method based on the image generation needs information, querying and acquiring a first reference image based on the image generation needs information, and generating a target image using the target image generation method based on the image generation needs information and the first reference image, and an output module for outputting the target image.
[0009] A sixth aspect of the present disclosure provides an intelligent agent, the intelligent agent including: an input module for acquiring image generation needs information; a processing module for determining a corresponding target image generation method based on the image generation needs information, determining whether the image generation needs information requires a reference image query, and if the image generation needs information requires a reference image query, querying and acquiring a first reference image based on the image generation needs information, and generating a target image using the target image generation method based on the image generation needs information and the first reference image; and an output module for outputting the target image.
[0010] According to a seventh aspect of this disclosure, an intelligent agent system is provided, At least one processor, Includes memory that is communicably connected to at least one processor, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can perform the image generation method according to the first embodiment or the image generation method according to the second embodiment.
[0011] According to an eighth aspect of the present disclosure, a non-temporary computer-readable storage medium is provided which stores computer instructions, the computer instructions causing the computer to execute the image generation method described in the first aspect or the image generation method described in the second aspect.
[0012] According to a ninth aspect of this disclosure, a computer program is provided which includes computer instructions, when executed by a processor, realizes a step of the image generation method described in the first aspect or a step of the image generation method described in the second aspect.
[0013] The image generation method, apparatus, intelligent agent, intelligent agent system, and storage medium provided by this disclosure have the following beneficial effects:
[0014] This disclosure obtains image generation needs information, determines a corresponding target image generation method based on the image generation needs information, queries for a first reference image based on the image generation needs information, and generates a target image using the target image generation method based on the image generation needs information and the first reference image. This disclosure generates a target image based on the queried first reference image, that is, uses image Retrieval-Augmented Generation (iRAG) technology to align the timeliness of image generation with the timeliness of image retrieval, thereby avoiding the problem of potential delays in the target image. Furthermore, since image retrieval covers almost all known knowledge, this disclosure does not have a memory capacity bottleneck, and the problem of limited memory capacity in AI image generation models is also solved.
[0015] It should be understood that the content described in this section is not intended to identify any essential or important features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will be readily apparent through the following description. [Brief explanation of the drawing]
[0016] The drawings are provided for the purpose of better understanding this technical proposal and do not limit the scope of this disclosure. [Figure 1] This is a schematic flowchart of an image generation method provided by one embodiment of the present disclosure. [Figure 2] This is a schematic flowchart of an image generation method provided by another embodiment of the present disclosure. [Figure 3] This is a schematic flowchart of an image generation method provided by another embodiment of the present disclosure. [Figure 4] This is a schematic diagram of the alternating image and text features provided by one embodiment of the present disclosure. [Figure 5] This is a schematic flowchart of an image generation method provided by another embodiment of the present disclosure. [Figure 6] This is a schematic flowchart of an image generation method provided by one embodiment of the present disclosure. [Figure 7] This is a schematic flowchart of an image generation method provided by another embodiment of the present disclosure. [Figure 8] This is a schematic diagram of an image generation device provided by one embodiment of the disclosure. [Figure 9] This is a schematic diagram of an image generation device provided by one embodiment of the disclosure. [Figure 10] This is a schematic diagram of an intelligent agent provided by one embodiment of the present disclosure. [Figure 11] This is a block diagram of an intelligent agent system for realizing the image generation method of the embodiment of this disclosure. [Modes for carrying out the invention]
[0017] Hereinafter, exemplary embodiments of the present disclosure will be described in combination with the drawings. For the ease of understanding, various details of the embodiments of the present disclosure are included therein, and they should be regarded as merely exemplary. Therefore, those skilled in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, the following description omits the description of well-known functions and structures.
[0018] In the technical solution of the present disclosure, any processing such as collection, storage, use, processing, transmission, provision, and disclosure of relevant user personal information is carried out with the consent of the user, in compliance with the provisions of relevant laws and regulations, and does not violate public order and good customs.
[0019] Hereinafter, with reference to the drawings, an image generation method, apparatus, intelligent agent, intelligent agent system, and storage medium according to embodiments of the present disclosure will be described.
[0020] It should be noted that the execution subject of the image generation method in this embodiment is an image generation apparatus, and the image generation apparatus can be realized by software and / or hardware and can be set in an intelligent agent.
[0021] FIG. 1 is a schematic flowchart of an image generation method provided by an embodiment of the present disclosure.
[0022] As shown in FIG. 1, this image generation method includes the following steps 101 to 104.
[0023] Step 101, obtain image generation requirement information.
[0024] The image generation requirement information can be used to indicate information such as the image style, image subject, and image size of the image to be generated.
[0025] As an example, the image generation requirement information can include an image generation requirement prompt input by the user.
[0026] Another example is that image generation needs information can include a prompt and a reference image entered by the user.
[0027] In step 102, the corresponding target image generation method is determined based on the image generation needs information.
[0028] Image generation methods should be tailored to specific image generation needs; different image generation needs require different image generation methods. For example, different AI image generation models may be compatible with different image generation methods.
[0029] For example, based on the correspondence between image generation needs and image generation methods, a target image generation method corresponding to image generation needs information can be determined.
[0030] For example, a target image generation method can be determined based on a large-scale model to address image generation needs information.
[0031] In step 103, a query is performed to obtain the first reference image based on the image generation needs information.
[0032] Based on the image generation needs information, an image query can be performed within the configured image library to obtain the first reference image. The first reference image may consist of one or more images.
[0033] For example, a vector feature corresponding to the image generation needs information can be obtained, and a first reference image can be obtained by performing a similarity match between the vector feature and the descriptive features of each image in the image library.
[0034] In step 104, a target image is generated using a target image generation method based on the image generation needs information and the first reference image.
[0035] Image generation needs information and a first reference image are inputs to the target image generation method. After obtaining the image generation needs information and the first reference image, the target image can be obtained by processing the image generation needs information and the first reference image according to the target image generation method.
[0036] Furthermore, if the image generation needs information includes a reference image, the system can obtain the user's editing intent targeting the reference image based on the image generation needs information, and then generate a target image using the first reference image, the editing intent, and the target image generation method based on the image generation needs information.
[0037] In the embodiments of this disclosure, image generation needs information is acquired, a corresponding target image generation method is determined based on the image generation needs information, a first reference image is obtained by querying based on the image generation needs information, and a target image is generated using the target image generation method based on the image generation needs information and the first reference image. This disclosure generates a target image based on the queried first reference image, that is, by using image Retrieval-Augmented Generation (iRAG) technology, the timeliness of image generation is aligned with the timeliness of image retrieval, thereby avoiding the problem of potential delays in the target image. Furthermore, since image retrieval covers almost all known knowledge, this disclosure does not have a memory capacity bottleneck, and the problem of limited memory capacity in AI image generation models is also solved.
[0038] Figure 2 is a schematic flowchart of an image generation method provided by another embodiment of the present disclosure.
[0039] As shown in Figure 2, this image generation method includes the following steps 201 to 204.
[0040] In step 201, image generation needs information is obtained.
[0041] In step 202, the corresponding target image generation method is determined based on the image generation needs information.
[0042] In step 203, the image subject information included in the image generation needs information is obtained, and based on the image subject information, an image query is performed within the configured image library to obtain the first reference image.
[0043] Image generation needs information includes image subject information, which can refer to information such as the type of subject and the name of the subject.
[0044] For example, if the image generation needs information is "draw a red vehicle with a modern design and a futuristic feel," the image subject information could be a red vehicle. Furthermore, if the image generation needs information also includes the vehicle model, the image subject information could be both the red vehicle and the vehicle model.
[0045] For example, based on a large-scale model, image subject information can be extracted from image generation needs information, and then an image search can be performed based on the image subject information to obtain a first reference image.
[0046] By performing image queries based on image subject information, it is possible to accurately filter reference images from the image library that match the image generation needs, thereby further improving the effectiveness of image generation.
[0047] The quality of the first reference image affects the effectiveness of image generation. If the quality of the first reference image is low, the quality of the final generated image is likely to be low as well. Therefore, to ensure the effectiveness of image generation and further improve the quality of the generated image, one example is to perform an image query in a configured image library based on the image subject information to obtain candidate reference images, obtain the image quality of the candidate reference images, obtain the configured quality requirements corresponding to the image generation needs information, and then filter the candidate reference images based on the image quality and configured quality requirements to obtain the first reference image.
[0048] The corresponding quality requirements will differ depending on the image generation needs. For example, if the image generation need is to generate images of people, the corresponding quality requirement may be that the resolution of faces in the image is greater than a set resolution threshold. If the image generation need is to generate poster images, the corresponding quality requirement may be that the image does not contain a watermark.
[0049] One possible implementation is to determine whether the quality of a candidate reference image meets the set quality requirements, thereby designating a candidate reference image whose image quality meets the set quality requirements as the first reference image.
[0050] Another possible implementation is to sort the candidate reference images in descending order of image quality based on the image quality and the set quality requirements, and then select the first N candidate reference images as the first reference image, where N is a positive integer.
[0051] In step 204, a target image is generated using a target image generation method based on the image generation needs information and the first reference image.
[0052] The explanations for steps 201, 202, and 204 can be found in the relevant explanations in any of the embodiments of this disclosure, and are therefore omitted here.
[0053] In the embodiments of this disclosure, image generation needs information is acquired, a corresponding target image generation method is determined based on the image generation needs information, image subject information included in the image generation needs information is acquired, an image query is performed within a set image library based on the image subject information to acquire a first reference image, and a target image is generated using the target image generation method based on the image generation needs information and the first reference image. By performing an image query based on the image subject information, reference images that match the image generation needs can be precisely filtered from the image library, further improving the effectiveness of image generation.
[0054] Figure 3 is a schematic flowchart of an image generation method provided by another embodiment of the present disclosure.
[0055] As shown in Figure 3, this image generation method includes the following steps 301 to 306.
[0056] In step 301, image generation needs information is obtained.
[0057] In step 302, the needs text in the image generation needs information is obtained.
[0058] The needs text can refer to the prompt in the finger image generation needs information.
[0059] In step 303, the needs text is input into the first large-scale model to detect the subject modification intent and obtain the target subject modification intent corresponding to the image generation needs information. The target subject modification intent is used to indicate whether or not the subject in the first reference image needs to be modified.
[0060] The first large-scale model can determine whether or not the subject in the first reference image needs to be modified when generating an image based on the needs text. If the target subject modification intention indicates that the subject in the first reference image needs to be modified, the image generation needs information corresponds to a generalization needs scenario at a high cognitive level. If the target subject modification intention indicates that the subject in the first reference image does not need to be modified, the image generation needs information corresponds to a pixel-level fidelity needs scenario.
[0061] For example, assuming the image generation needs information is "generate an image of vehicle A driving in a desert" or "generate an image of vehicle A driving in a forest," the target subject modification intention corresponding to the image generation needs information is that there is no need to modify the subject in the first reference image. If the image generation needs information is "generate an image of person B wearing clothing style C" or "generate a figure image of person B," then the target subject modification intention corresponding to the image generation needs information is that there is a need to modify the subject in the first reference image.
[0062] Furthermore, if the image generation needs information includes a second reference image entered by the user, the subject modification intent obtained through subject modification intent detection can also be used to indicate whether or not the subject in the second reference image needs to be modified.
[0063] In step 304, a target image generation method corresponding to the target subject modification intention is obtained based on the mapping relationship between the subject modification intention and the image generation method.
[0064] Different subject modification intentions require different corresponding image generation methods. For example, a mapping relationship between subject modification intentions and image generation methods can be pre-defined. After the target subject modification intention is determined, the corresponding target image generation method can be determined based on the pre-defined mapping relationship.
[0065] Different subjective modification intentions require different attention during image generation. Therefore, by selecting an image generation method that matches the subjective modification intention, the effectiveness and efficiency of image generation can be significantly improved.
[0066] In step 305, a query is performed to obtain the first reference image based on the image generation needs information.
[0067] In step 306, a target image is generated using a target image generation method based on the image generation needs information and the first reference image.
[0068] To extract key features from the extracted needs text, the needs text can be encoded to obtain text features.
[0069] The following describes the corresponding image generation methods and processes that combine two types of subjective modification intentions.
[0070] For example, in response to the target subject modification intention indicating that there is no need to modify the subject in the first reference image, feature extraction is performed on the first reference image to obtain the first image features, and a subject segmentation image in the first reference image is obtained. The first image features and text features are input into the first image generation model to obtain a background image and subject layout information. Based on the subject layout information, the subject segmentation image and background image are merged to obtain the target image.
[0071] The image generation method corresponding to this example can be called a precise image generation method, where the first image feature refers to a vector feature obtained using a small-scale feature extractor, the subject segmentation image refers to an image of the subject obtained by performing semantic segmentation on the first reference image, the background image refers to an image of the background part other than the subject, and the subject layout information is used to indicate the layout of the subject in the background image, for example, the subject layout information includes information such as the position and size of the subject in the background image.
[0072] In this disclosure, the content of the first reference image, i.e., the subject segmentation image, is obtained at the pixel level. The subject segmentation image retains the essential details and features of the subject in the first reference image, achieving high fidelity to the subject and ensuring the integrity of the subject in the target image. Furthermore, by performing image fusion based on the subject layout information, the subject can be naturally blended into the background image, maintaining harmony and unity with other parts of the background image.
[0073] Another example is when, in response to the target subject modification intention requiring modification of the subject in the first reference image, feature extraction is performed on the first reference image to obtain second image features, and based on the text features and the descriptions corresponding to each sub-feature in the second image features, the text features and the second image features are spliced alternately to obtain alternating image and text features, and these alternating image and text features are input into the second image generation model to obtain the target image.
[0074] The image generation method corresponding to this example can be called a highly generalized image generation method, where the second image feature is a vector feature, the second image feature and the first image feature may be different, text features include text sub-features, image features include image sub-features, text sub-features and image sub-features each correspond to a description object, and alternating splicing refers to combining image features and text features in a specific manner (for example, splicing text sub-features and image sub-features corresponding to the same description object) to form a new encoded feature, i.e., an alternating image and text feature. Referring to Figure 4, which is a schematic diagram of an alternating image and text feature, this method combines information from two different modalities, image features and text features, in a specific manner, and realizes associations between multimodal features.
[0075] For example, assuming the image generation needs information is "generate figure images of person B and person D," the corresponding description includes person B and person D, with person B having corresponding text sub-features and image sub-features, and person D having corresponding word sub-features and image sub-features. When splicing them alternately, the image sub-features corresponding to person B can be inserted into the text sub-features corresponding to person B in the text features, and then the image sub-features corresponding to person D can be inserted into the text sub-features corresponding to person D in the text features.
[0076] By alternately splicing text features and second image features, precise alignment of both is achieved in the description target. This alternating splicing method not only preserves essential information in each feature but also promotes deep fusion between text and image features. When generating a target image based on alternating image and text features, it is possible to fully consider the characteristics of the subject in the first reference image, while simultaneously maintaining fidelity to the subject and enabling generalized derivative works, resulting in a good image generation effect.
[0077] Image generation needs information may include a second reference image entered by the user. In this case, the method for obtaining alternating image and text features is as follows: Feature extraction is performed on the second reference image to obtain a third image feature. Based on the descriptive targets corresponding to the sub-features in the text feature, the second image feature, and the third image feature, the text feature, the second image feature, and the third image feature are spliced alternately to obtain alternating image and text features.
[0078] For example, assuming the image generation needs information is a second reference image containing person B and "to generate a photographic image of a figure group of person B and person D," when generating the target image, it is necessary to alternately splice the text sub-features corresponding to person B with the image sub-features of person B in the third image feature, and alternately splice the text sub-features corresponding to person D with the image sub-features of person D in the second image feature.
[0079] If the image generation needs information includes a second reference image entered by the user, the system can perform alternating feature splicing based on the needs text, the first reference image, and the second reference image entered by the user, and then generate a target image. This not only satisfies the user's individualized image generation needs and enables flexible, personalized image generation, but also ensures the effectiveness of the image generation.
[0080] Furthermore, for knowledge that has not been disclosed, this disclosure avoids the problem that the image generation model does not remember this content by allowing users to provide reference images themselves.
[0081] The explanations regarding steps 301, 302, and 305 can be found in the relevant explanations in any of the embodiments of this disclosure, and are therefore omitted here.
[0082] In the embodiments of this disclosure, image generation needs information is acquired, needs text is acquired from the image generation needs information, the needs text is input into a first large-scale model to detect subject modification intent, and a target subject modification intent corresponding to the image generation needs information is acquired. The target subject modification intent is used to indicate whether or not the subject in the first reference image needs to be modified. Based on the mapping relationship between the subject modification intent and the image generation method, a target image generation method corresponding to the target subject modification intent is acquired. Based on the image generation needs information, a first reference image is obtained by querying, and a target image is generated using the target image generation method based on the image generation needs information and the first reference image. Different subject modification intents require different attention during image generation. Therefore, by selecting an image generation method that matches the subject modification intent, the effectiveness and efficiency of image generation can be significantly improved.
[0083] Figure 5 is a schematic flowchart of an image generation method provided by another embodiment of the present disclosure.
[0084] As shown in Figure 5, the method for generating this image includes the following steps 501 to 504.
[0085] In step 501, image generation needs information is obtained.
[0086] In step 502, the corresponding target image generation method is determined based on the image generation needs information.
[0087] In step 503, a query is performed to obtain the first reference image based on the image generation needs information.
[0088] In step 504, the needs text in the image generation needs information is obtained, the needs text is revised and expanded to obtain the target needs text, and the target image is generated using the target image generation method based on the target needs text and the first reference image.
[0089] Compared to needs text, target needs text has a more structured format and more comprehensive, detailed descriptions. For example, needs text can be input into a large-scale model, which then revises and expands upon the needs text based on elements such as core screen content, screen style descriptions, screen subject and constraint descriptions, screen detail descriptions, screen background decoration descriptions, special effects, composition, color scheme, resolution descriptions, and quality descriptions.
[0090] There is a possibility that the needs text in the image generation needs information may be abstract and ambiguous in its expression. Since directly generating images based on the needs text often does not yield optimal results, this disclosure shows that the effectiveness of image generation can be improved by revising and expanding the needs text.
[0091] Another example involves obtaining contextual information from image generation needs information, and if the image generation needs information includes a second reference image entered by the user, obtaining descriptive information from the second reference image, and inputting at least one of the contextual information and descriptive information, along with the needs text, into a second large-scale model for modification and expansion to obtain the target needs text.
[0092] The second large-scale model and the first large-scale model may be the same large-scale model or they may be different large-scale models.
[0093] In the image generation process, the image generation needs information provided by the user may be incomplete, potentially resulting in the generated image not meeting the user's expectations. To better understand user needs and improve the accuracy and effectiveness of image generation, this disclosure also allows for the acquisition of further contextual information related to the image generation needs information, and, if the user inputs a second reference image, the acquisition of descriptive information for that second reference image. By combining this information, this disclosure can gain a more comprehensive understanding of the user's true image generation intentions, further improving the effectiveness of image generation and enhancing the user experience.
[0094] For explanations regarding steps 501 to 503, please refer to the relevant explanations in any of the embodiments of this disclosure, and the explanations are omitted here.
[0095] In the embodiments of this disclosure, the needs text in the image generation needs information is obtained, the needs text is revised and expanded to obtain the target needs text, and the target image is generated using the target image generation method based on the target needs text and the first reference image. There is a possibility that the needs text in the image generation needs information may be abstract and have ambiguous expressions, and since directly generating an image based on the needs text often does not yield optimal results, this disclosure can improve the effectiveness of image generation by revising and expanding the needs text.
[0096] Figure 6 is a schematic flowchart of an image generation method provided by one embodiment of the present disclosure.
[0097] As shown in Figure 6, this image generation method includes the following steps 601 to 605.
[0098] In step 601, image generation needs information is obtained.
[0099] In step 602, the corresponding target image generation method is determined based on the image generation needs information.
[0100] In step 603, it is determined whether the image generation needs information requires a query for a reference image.
[0101] When generating images, not all image generation needs require enhanced efficiency through search augmentation. For example, if the image generation need information is "draw a small rabbit," this can be handled by the model itself that generates images from common text, and does not require triggering search augmentation. Based on this, to avoid unnecessary complex calculations and wasted resources, this disclosure performs a query determination of reference images before generating images.
[0102] For example, image generation needs information can be input into a third large-scale model, and the third large-scale model can determine whether the image generation needs information requires a query for a reference image. The third large-scale model, the second large-scale model, and the first large-scale model may be the same large-scale model, or they may refer to different large-scale models.
[0103] In step 604, if the image generation needs information requires a reference image query, the system queries based on the image generation needs information to obtain a first reference image.
[0104] In step 605, a target image is generated using a target image generation method based on the image generation needs information and the first reference image.
[0105] The explanations for steps 601, 602, 604, and 605 can be found in the relevant explanations in any of the embodiments of this disclosure, and are therefore omitted here.
[0106] In the embodiments of this disclosure, image generation needs information is acquired, a corresponding target image generation method is determined based on the image generation needs information, it is determined whether the image generation needs information requires a reference image query, and if the image generation needs information requires a reference image query, a first reference image is obtained by querying based on the image generation needs information, and a target image is generated using the target image generation method based on the image generation needs information and the first reference image. This disclosure determines whether the image generation needs information requires a reference image query, and if the image generation needs information requires a reference image query, it uses iRAG technology to generate the image. This aligns the timeliness of image generation with the timeliness of image retrieval, avoiding the problem of potential delays in the target image, avoiding unnecessary complex calculations and wasted resources when image augmentation is not required, and improving overall image generation efficiency. Furthermore, since image retrieval covers almost all known knowledge, this disclosure does not have a memory capacity bottleneck, and the problem of limited memory capacity in AI image generation models is also solved.
[0107] Figure 7 is a schematic flowchart of an image generation method provided by another embodiment of the present disclosure.
[0108] As shown in Figure 7, this image generation method includes the following steps 701 to 708.
[0109] In step 701, image generation needs information is obtained.
[0110] In step 702, the corresponding target image generation method is determined based on the image generation needs information.
[0111] In addition to precise and highly generalizable image generation methods, this disclosure further provides complementary methods for generating images from common text and generating images from images. For image generation needs information, a matching target image generation method can be selected from the aforementioned methods.
[0112] For example, if we assume that the image generation need information is "draw a small rabbit," the corresponding target image generation method may be a method that generates images from common text. If we assume that the image generation need information is "draw a Flemish giant rabbit," the corresponding target image generation method may be a precise image generation method.
[0113] For example, the system obtains the needs text from the image generation needs information, inputs the needs text into a first large-scale model to detect the subject modification intent, obtains the target subject modification intent corresponding to the image generation needs information, uses the target subject modification intent to indicate whether or not the subject in the first reference image needs to be modified, and obtains the target image generation method corresponding to the target subject modification intent based on the mapping relationship between the subject modification intent and the image generation method.
[0114] In step 703, the needs text in the image generation needs information is obtained.
[0115] In step 704, needs understanding is performed on the needs text, and the type of needs corresponding to the image generation needs information is obtained.
[0116] The type of need can be used to indicate the type of knowledge related to image generation needs information, for example, the type of need may include types such as basic knowledge, specialized knowledge, long-tail knowledge, and timeliness knowledge.
[0117] For example, needs text can be input into a fourth large-scale model to understand the needs and obtain the type of needs corresponding to the image generation needs information. The fourth large-scale model, the third large-scale model, the second large-scale model, and the first large-scale model may be large-scale models, or they may refer to different large-scale models.
[0118] In step 705, based on the type of need, it is determined whether the image generation needs information requires a reference image query.
[0119] There is a correspondence between the type of need and whether or not a reference image query is required. For example, basic knowledge does not require a reference image query, meaning there is no need to obtain a first reference image, while specialized knowledge, long-tail knowledge, and timeliness knowledge require a reference image query, meaning there is a need to obtain a first reference image.
[0120] When handling image generation needs, this disclosure employs an intelligent strategy, namely, a flexible decision on whether or not to perform reference image queries based on the type of need, recognizing that not all needs need to be improved through search augmentation. By precisely determining needs, this disclosure can accurately perform search augmentation, thereby not only avoiding excessive search augmentation load but also meeting user needs, achieving optimal resource allocation, and avoiding unnecessary resource waste.
[0121] In step 706, if the image generation needs information requires a reference image query, the system queries based on the image generation needs information to obtain a first reference image.
[0122] For example, image subject information included in image generation needs information is obtained, and based on the image subject information, an image query is performed within the configured image library to obtain a first reference image.
[0123] Another example involves obtaining image subject information included in image generation needs information, performing image queries in a configured image library based on the image subject information to obtain candidate reference images, obtaining the image quality of the candidate reference images, obtaining the configured quality requirements corresponding to the image generation needs information, and filtering the candidate reference images based on the image quality and configured quality requirements to obtain a first reference image.
[0124] In step 707, a target image is generated using a target image generation method based on the image generation needs information and the first reference image.
[0125] For example, text features correspond to the needs text, and in response to the target subject modification intention indicating that there is no need to modify the subject in the first reference image, feature extraction is performed on the first reference image to obtain the first image features, and a subject segmentation image in the first reference image is obtained. The first image features and text features are input into the first image generation model to obtain the background image and subject layout information, and based on the subject layout information, the subject segmentation image and background image are merged to obtain the target image.
[0126] Another example is where a needs text corresponds to text features, and in response to the target subject modification intention requiring modification of the subject in the first reference image, feature extraction is performed on the first reference image to obtain second image features. Based on the descriptive targets corresponding to the sub-features in the text features and the second image features, the text features and the second image features are spliced alternately to obtain alternating image and text features. These alternating image and text features are then input into a second image generation model to obtain the target image.
[0127] One possible implementation of the embodiments of the present disclosure includes the steps of obtaining alternating image and text features by alternately splicing a text feature and a second image feature based on descriptive objects corresponding to the respective sub-features of the text feature and the second image feature, the image generation needs information includes a second reference image input by the user, and the steps of obtaining a third image feature by performing feature extraction on the second reference image, and the steps of obtaining alternating image and text features by alternately splicing a text feature, a second image feature and a third image feature based on descriptive objects corresponding to the respective sub-features of the text feature, the second image feature and the third image feature.
[0128] For example, the process involves obtaining the needs text from the image generation needs information, revising and expanding the needs text to obtain the target needs text, and then generating the target image using the target image generation method based on the target needs text and the first reference image.
[0129] For example, contextual information of image generation needs information is obtained, and if the image generation needs information includes a second reference image entered by the user, descriptive information of the second reference image is obtained, and at least one of the contextual information and descriptive information, along with the needs text, is input into a second large-scale model for modification and expansion to obtain the target needs text.
[0130] In step 708, if the image generation needs information does not require a reference image query, the target image is generated using the target image generation method based on the image generation needs information.
[0131] Even when reference image queries are not required, this disclosure can expand the application scenarios of this disclosure by meeting diverse image generation needs by efficiently and accurately completing image generation tasks based on image generation needs information.
[0132] For example, if the detected subject modification intention indicates that there is no need to modify the subject in the second reference image, the target image generation method may be a precise image generation method. Specifically, feature extraction is performed on the second reference image to obtain a fourth image feature, and a subject segmentation image is obtained in the second reference image. The fourth image feature and text feature are input into the first image generation model to obtain a background image and subject layout information. Based on the subject layout information, the subject segmentation image and background image are merged to obtain the target image.
[0133] As another example, if the detected subject modification intent indicates that the subject in the second reference image needs to be modified, the target image generation method may be a highly generalized image generation method. Specifically, feature extraction is performed on the second reference image to obtain a fifth image feature, and the text feature and the fifth image feature are spliced alternately based on the descriptive targets corresponding to the respective sub-features in the text feature and the fifth image feature to obtain alternating image and text features. These alternating image and text features are then input into the second image generation model to obtain the target image.
[0134] The explanations for steps 701, 702, 703, 706, and 707 can be found in the relevant explanations in any of the embodiments of this disclosure, and are therefore omitted here.
[0135] In the embodiments of this disclosure, the need text in the image generation needs information is obtained, the need understanding is performed on the need text, the type of need corresponding to the image generation needs information is obtained, and based on the type of need, it is determined whether or not the image generation needs information requires a reference image query. If the image generation needs information does not require a reference image query, a target image is generated using the target image generation method based on the image generation needs information. If the image generation needs information requires a reference image query, a first reference image is obtained by querying based on the image generation needs information, and a target image is generated using the target image generation method based on the image generation needs information and the first reference image. Since not all needs require improvement through search enhancement, this disclosure flexibly determines whether to obtain a first reference image based on the type of need. By precisely determining the needs, this disclosure can accurately perform search enhancement, thereby not only avoiding excessive load on the search enhancement system but also meeting user needs, achieving optimal resource allocation, and avoiding unnecessary resource waste.
[0136] Figure 8 is a schematic diagram of an image generation apparatus provided by one embodiment of the disclosure.
[0137] As shown in Figure 8, this image generation device includes an acquisition module 801, a decision module 802, a query module 803, and a generation module 804.
[0138] The acquisition module 801 acquires image generation needs information. The decision module 802 determines the corresponding target image generation method based on the image generation needs information. The query module 803 queries based on the image generation needs information to acquire a first reference image. The generation module 804 generates a target image using the target image generation method based on the image generation needs information and the first reference image.
[0139] One possible implementation of the embodiments of this disclosure is that the query module 803 obtains image subject information included in the image generation needs information, and based on the image subject information, performs an image query within a configured image library to obtain a first reference image.
[0140] One possible implementation of the embodiment of this disclosure is that the query module 803 performs an image query in a configured image library based on image subject information to obtain candidate reference images, obtain the image quality of the candidate reference images, obtain configured quality requirements corresponding to image generation needs information, and filter the candidate reference images based on the image quality and configured quality requirements to obtain a first reference image.
[0141] In one possible implementation of the embodiments of this disclosure, the decision module 802 acquires the needs text in the image generation needs information, inputs the needs text into a first large model to perform subject modification intent detection, acquires the target subject modification intent corresponding to the image generation needs information, the target subject modification intent is used to indicate whether or not the subject in the first reference image needs to be modified, and acquires the target image generation method corresponding to the target subject modification intent based on the mapping relationship between the subject modification intent and the image generation method.
[0142] One possible implementation of the embodiments of this disclosure is as follows: Needs text corresponds to text features, and in response to the target subject modification intent indicating that there is no need to modify the subject in the first reference image, the generation module 804 performs feature extraction on the first reference image to obtain first image features and subject segmentation images in the first reference image, inputs the first image features and text features into a first image generation model to obtain background image and subject layout information, and merges the subject segmentation images and background image based on the subject layout information to obtain a target image.
[0143] One possible implementation of the embodiments of this disclosure is as follows: Needs text corresponds to text features, and the generation module 804, in response to the target subject modification intention requiring modification of a subject in a first reference image, performs feature extraction on the first reference image to obtain second image features, alternately splices the text features and the second image features based on the descriptions corresponding to the respective sub-features in the text features and the second image features to obtain alternating image and text features, inputs the alternating image and text features into a second image generation model to obtain a target image.
[0144] One possible implementation of the embodiments of this disclosure is that the image generation needs information includes a second reference image input by the user, the generation module 804 performs feature extraction on the second reference image to obtain a third image feature, and then alternately splices the text feature, the second image feature and the third image feature based on the description targets corresponding to the sub-features in the text feature, the second image feature and the third image feature to obtain alternating image and text features.
[0145] One possible implementation of the embodiment of this disclosure is that the generation module 804 obtains the needs text in the image generation needs information, modifies and expands the needs text to obtain the target needs text, and generates a target image using a target image generation method based on the target needs text and the first reference image.
[0146] One possible implementation of the embodiments of this disclosure is that the generation module 804 acquires context information of image generation needs information, and if the image generation needs information includes a second reference image entered by the user, acquires description information of the second reference image, and inputs at least one of the context information and description information, along with the needs text, into a second large model to revise and expand it to obtain target needs text.
[0147] Figure 9 is a schematic diagram of an image generation apparatus provided by one embodiment of the disclosure.
[0148] As shown in Figure 9, this image generation device includes an acquisition module 901, a decision module 902, a judgment module 903, a query module 904, and a generation module 905. The acquisition module 901 acquires image generation needs information, the decision module 902 determines the corresponding target image generation method based on the image generation needs information, the judgment module 903 determines whether the image generation needs information requires a reference image query, the query module 904, if the image generation needs information requires a reference image query, queries based on the image generation needs information to acquire a first reference image, and the generation module 905 generates a target image using the target image generation method based on the image generation needs information and the first reference image.
[0149] One possible implementation of the embodiments of this disclosure is that the decision module 903 obtains the needs text in the image generation needs information, performs a needs understanding on the needs text, obtains the type of needs corresponding to the image generation needs information, and determines whether the image generation needs information requires a reference image query based on the type of needs.
[0150] One possible implementation of the embodiments of this disclosure is that, if the image generation needs information does not require a reference image query, the generation module 905 generates a target image using a target image generation method based on the image generation needs information.
[0151] One possible implementation of the embodiments of this disclosure is that the query module 904 obtains image subject information included in the image generation needs information, and based on the image subject information, performs an image query within the configured image library to obtain a first reference image.
[0152] One possible implementation of the embodiment of this disclosure is that the query module 904 performs an image query in a configured image library based on image subject information to obtain candidate reference images, obtain the image quality of the candidate reference images, obtain configured quality requirements corresponding to image generation needs information, and filter the candidate reference images based on the image quality and configured quality requirements to obtain a first reference image.
[0153] In one possible implementation of the embodiments of this disclosure, the decision module 902 acquires the needs text in the image generation needs information, inputs the needs text into a first large-scale model to perform subject modification intent detection, acquires the target subject modification intent corresponding to the image generation needs information, the target subject modification intent is used to indicate whether or not the subject in the first reference image needs to be modified, and acquires the target image generation method corresponding to the target subject modification intent based on the mapping relationship between the subject modification intent and the image generation method.
[0154] One possible implementation of the embodiments of this disclosure is as follows: Needs text corresponds to text features, and in response to the target subject modification intent indicating that there is no need to modify the subject in the first reference image, the generation module 905 performs feature extraction on the first reference image to obtain first image features and subject segmentation images in the first reference image, inputs the first image features and text features into a first image generation model to obtain background image and subject layout information, and merges the subject segmentation images and background image based on the subject layout information to obtain a target image.
[0155] In one possible implementation of the embodiments of this disclosure, the needs text corresponds to text features, and the generation module 905, in response to the target subject modification intention requiring modification of a subject in a first reference image, performs feature extraction on the first reference image to obtain second image features, alternately splices the text features and the second image features based on the descriptions corresponding to the respective sub-features in the text features and the second image features to obtain alternating image and text features, inputs the alternating image and text features into a second image generation model to obtain a target image.
[0156] One possible implementation of the embodiments of this disclosure is that the image generation needs information includes a second reference image input by the user, the generation module 905 performs feature extraction on the second reference image to obtain a third image feature, and then alternately splices the text feature, the second image feature and the third image feature based on the description targets corresponding to the sub-features in the text feature, the second image feature and the third image feature to obtain alternating image and text features.
[0157] One possible implementation of the embodiment of this disclosure is that the generation module 905 obtains the needs text in the image generation needs information, modifies and expands the needs text to obtain the target needs text, and generates a target image using the target image generation method based on the target needs text and the first reference image.
[0158] One possible implementation of the embodiments of this disclosure is that the generation module 905 acquires context information of image generation needs information, and if the image generation needs information includes a second reference image entered by the user, acquires description information of the second reference image, and inputs at least one of the context information and description information, along with the needs text, into a second large model to revise and expand it to obtain target needs text.
[0159] The above-mentioned explanation of the image generation method also applies to the image generation device of this embodiment, so the explanation is omitted here.
[0160] Figure 10 is a schematic diagram of an intelligent agent provided by one embodiment of the present disclosure.
[0161] As shown in Figure 10, this intelligent agent 1000 may include an input module 1001, a processing module 1002, and an output module 1003.
[0162] For example, input module 1001 acquires image generation needs information, processing module 1002 determines a corresponding target image generation method based on the image generation needs information, queries for a first reference image based on the image generation needs information, generates a target image using the target image generation method based on the image generation needs information and the first reference image, and output module 1003 outputs the target image.
[0163] In another example, the input module 1001 acquires image generation needs information, the processing module 1002 determines the corresponding target image generation method based on the image generation needs information, determines whether the image generation needs information requires a reference image query, and if the image generation needs information requires a reference image query, it queries based on the image generation needs information to acquire a first reference image, generates a target image using the target image generation method based on the image generation needs information and the first reference image, and the output module 1003 outputs the target image.
[0164] According to embodiments of the present disclosure, the present disclosure further provides an intelligent agent system, a readable storage medium, and a computer program.
[0165] Figure 11 is a schematic block diagram of an exemplary intelligent agent system 1100 for performing an embodiment of the present disclosure. The intelligent agent system 1100 is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The intelligent agent system may also represent various forms of mobile devices, such as personal digital processing devices, mobile phones, smartphones, wearable devices, and other similar computing devices. The components, their connections and relationships, and their functions shown herein are merely examples and are not intended to limit the description herein and / or the implementation of the present disclosure as requested.
[0166] As shown in Figure 11, the intelligent agent system 1100 includes a computing unit 1101 that can perform various appropriate operations and processes according to a computer program stored in a ROM (Read-Only Memory) 1102 or a computer program loaded from a storage unit 1108 into a RAM (Random Access Memory) 1103. The RAM 1103 may also store various programs and data necessary for the operation of the intelligent agent system 1100. The computing unit 1101, ROM 1102, and RAM 1103 are connected to each other via a bus 1104. An I / O (Input / Output) interface 1105 is also connected to the bus 1104.
[0167] Multiple components of the intelligent agent system 1100 are connected to the I / O interface 1105, which includes input units 1106 such as a keyboard and mouse, output units 1107 such as various types of displays and speakers, storage units 1108 such as magnetic disks and optical disks, and communication units 1109 such as a network card, modem, and wireless communication transceiver. The communication units 1109 enable the intelligent agent system 1100 to exchange information / data with other devices via computer networks such as the Internet and / or various telegraph networks.
[0168] The computing unit 1101 may be a variety of general-purpose and / or dedicated processing components having processing and computing capabilities. Some examples of the computing unit 1101 include, but are not limited to, a CPU (Central Processing Unit), a GPU (Graphic Processing Unit), various dedicated AI (Artificial Intelligence) computing chips, various machine driving learning model algorithm computing units, a DSP (Digital Signal Processor), and any suitable processor, controller, microcontroller, etc. The computing unit 1101 performs each of the methods and processes described in the preceding paragraph, for example, the image generation method. For example, in some embodiments, the image generation method can be implemented as a computer software program tangibly contained in a machine-readable medium such as a storage unit 1108. In some embodiments, part or all of the computer program can be loaded and / or installed into the intelligent agent system 1100 via ROM 1102 and / or communication unit 1109. When a computer program is loaded into RAM 1103 and executed by the computing unit 1101, one or more steps of the image generation method described in the preceding paragraph may be performed. Alternatively, in other embodiments, the computing unit 1101 may be configured to perform the image generation method by any other suitable method (e.g., via firmware).
[0169] Various embodiments of the systems and technologies described herein can be implemented as digital electronic circuit systems, integrated circuit systems, FPGAs (Field Programmable Gate Arrays), ASICs (Application-Specific Integrated Circuits), ASSPs (Application Specific Standard Products), SOCs (System on Chip), CPLDs (Complex Programmable Logic Devices), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include being executed by one or more computer programs, which may be executed and / or interpreted on a programmable system including at least one programmable processor, which may be an application-specific or general-purpose programmable processor, which may receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, at least one input device, and at least one output device.
[0170] Program code for performing the methods of this disclosure can be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, a dedicated computer, or other programmable data processing device, so that when executed by the processor or controller, the functions / operations defined in the flowcharts and / or block diagrams are performed. The program code may run entirely on a machine, partially on a machine, or, as a standalone software package, partially on a machine, partially on a remote machine, or entirely on a remote machine or server.
[0171] In the context of this disclosure, a machine-readable medium may be a tangible medium that contains or can store a program for use by or in combination with an instruction execution system, apparatus, or device. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination thereof. More specific examples of machine-readable storage media include one or more line-based electrical connections, portable computer disks, hard disks, RAM, ROM, EPROM (Electrically Programmable Read-Only Memory), or flash memory, optical fibers, CD-ROM (Compact Disc Read-Only Memory), optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0172] To provide user interaction, the systems and technologies described herein can be implemented on a computer, which may have a display device for displaying information to the user (e.g., a CRT (Cathode-Ray Tube) or LCD (Liquid Crystal Display) monitor), and a keyboard and pointing device (e.g., a mouse or trackball), and the user may provide input to the computer via the keyboard and pointing device. Other types of devices may also provide user interaction, for example, the feedback provided to the user may be any form of sensing feedback (e.g., vision feedback, auditory feedback, or haptic feedback), and may receive input from the user in any form (including acoustic input and voice input or haptic input).
[0173] The systems and technologies described herein can be run on computing systems including backend components (e.g., data servers), computing systems including middleware components (e.g., application servers), computing systems including frontend components (e.g., user computers having a graphical user interface or web browser, through which users can interact with embodiments of the systems and technologies described herein), or any combination of such backend components, middleware components, and frontend components. The components of the system can be interconnected by digital data communication in any form or medium (e.g., communication networks). Examples of communication networks include LANs (Local Area Networks), WANs (Wide Area Networks), the Internet, and blockchain networks.
[0174] A computer system can include clients and servers. Clients and servers are generally geographically separated and typically interact via a communication network. The client-server relationship is generated by computer programs running on corresponding computers that have a client-server relationship with each other. A server may be a cloud server, also called a cloud computing server or cloud host, a host product in a cloud computing service system that addresses the management difficulties and limited business scalability inherent in traditional physical hosts and VPS services ("Virtual Private Server," or simply "VPS"). A server may be a server in a distributed system, or it may be a server incorporating blockchain technology.
[0175] Artificial intelligence (AI) is the study of enabling computers to simulate certain human thought processes and intelligent actions (learning, reasoning, thinking, planning, etc.), and it encompasses both hardware-level and software-level technologies. Artificial intelligence hardware technologies generally include sensors, dedicated AI chips, cloud computing, distributed memory, and big data processing. Artificial intelligence software technologies mainly include several areas such as computer vision technologies, speech recognition technologies, natural language processing technologies, machine learning / deep learning, big data processing technologies, and knowledge graph technologies.
[0176] It should be understood that the steps can be rearranged, added, or deleted using the various forms of flows shown above. For example, each step described in this disclosure may be performed in parallel, sequentially, or in a different order, as long as the technical proposal disclosed herein can achieve the desired results.
[0177] The specific embodiments described above do not limit the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, subcombinations, and substitutions can be made depending on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure must be within the scope of protection of this disclosure.
Claims
1. An image generation method, Steps to obtain image generation needs information, The steps include obtaining the needs text in the aforementioned image generation needs information, A step of inputting the aforementioned needs text into a first large-scale model to detect the subject modification intention and to obtain the target subject modification intention corresponding to the aforementioned image generation needs information, wherein the target subject modification intention is used to indicate whether or not the subject in the first reference image needs to be modified. A step of obtaining a target image generation method corresponding to the target subject modification intention based on the mapping relationship between the subject modification intention and the image generation method, The steps include querying and obtaining the first reference image based on the image generation needs information, A step of generating a target image using the target image generation method based on the image generation needs information and the first reference image, An image generation method that includes [a specific feature / method].
2. The step of querying and obtaining the first reference image based on the image generation needs information is as follows: The steps include: obtaining image subject information included in the aforementioned image generation needs information; The steps include: obtaining the first reference image by performing an image query in the configured image library based on the aforementioned image subject information; The image generation method according to claim 1, which includes the following:
3. The step of obtaining the first reference image by performing an image query in the configured image library based on the aforementioned image subject information is as follows: Based on the aforementioned image subject information, the step of performing an image query in the configured image library to obtain candidate reference images, The steps include obtaining the image quality of the candidate reference image and obtaining the set quality requirements corresponding to the image generation needs information, A step of obtaining the first reference image by filtering the candidate reference images based on the image quality and the set quality requirements, The image generation method according to claim 2, which includes the following:
4. Text features are associated with the aforementioned needs text. The step of generating a target image using the target image generation method based on the image generation needs information and the first reference image is as follows: In response to the target subject modification intention indicating that there is no need to modify the subject in the first reference image, the steps include: performing feature extraction on the first reference image to obtain first image features and obtaining a subject segmentation image in the first reference image; The steps include inputting the first image feature and the text feature into a first image generation model to obtain a background image and subject layout information, The steps include obtaining the target image by fusing the subject division image and the background image based on the subject layout information, The image generation method according to claim 1, which includes the following:
5. Text features are associated with the aforementioned needs text. The step of generating a target image using the target image generation method based on the image generation needs information and the first reference image is as follows: In response to the target subject modification intention indicating that the subject in the first reference image needs to be modified, the step of performing feature extraction on the first reference image to obtain a second image feature, The steps include obtaining alternating image and text features by alternately splicing the text features and the second image features based on the description targets corresponding to each sub-feature in the text features and the second image features, The steps include inputting the aforementioned image and text features into a second image generation model to obtain the target image, The image generation method according to claim 1, which includes the following:
6. The aforementioned image generation needs information includes a second reference image entered by the user. The step of obtaining alternating image and text features by alternately splicing the text features and the second image features based on the description target corresponding to each sub-feature in the text features and the second image features is as follows: The steps include: performing feature extraction on the aforementioned second reference image to obtain a third image feature; The image generation method according to claim 5, comprising the step of obtaining alternating image and text features by alternately splicing the text features, the second image features, and the third image features based on the description target corresponding to each sub-feature in the text features, the second image features, and the third image features.
7. The step of generating a target image using the target image generation method based on the image generation needs information and the first reference image is as follows: The steps include obtaining the needs text in the aforementioned image generation needs information, The steps include: revising and expanding the aforementioned needs text to obtain the target needs text; A step of generating the target image using the target image generation method based on the target needs text and the first reference image, The image generation method according to claim 1, which includes the following:
8. The step of revising and expanding the aforementioned needs text to obtain the target needs text is: The steps include obtaining contextual information for the image generation needs information, If the image generation needs information includes a second reference image entered by the user, the steps include obtaining the description information of the second reference image, The steps include: inputting at least one of the context information and the descriptive information, and the needs text into a second large-scale model, revising and expanding it, to obtain the target needs text; The image generation method according to claim 7, including the following:
9. An image generation method, Steps to obtain image generation needs information, The steps include obtaining the needs text in the aforementioned image generation needs information, A step of inputting the aforementioned needs text into a first large-scale model to detect the subject modification intention and to obtain the target subject modification intention corresponding to the aforementioned image generation needs information, wherein the target subject modification intention is used to indicate whether or not the subject in the first reference image needs to be modified. A step of obtaining a target image generation method corresponding to the target subject modification intention based on the mapping relationship between the subject modification intention and the image generation method, The steps include determining whether the aforementioned image generation needs information requires a reference image query, If the image generation needs information requires a reference image query, the step of querying based on the image generation needs information to obtain the first reference image, A step of generating a target image using the target image generation method based on the image generation needs information and the first reference image, An image generation method that includes [a specific feature / method].
10. The step of determining whether the aforementioned image generation needs information requires a reference image query is: The steps include obtaining the needs text in the aforementioned image generation needs information, The steps include: performing a needs understanding on the aforementioned needs text and obtaining the type of needs corresponding to the aforementioned image generation needs information; The steps include determining whether the image generation needs information requires a reference image query based on the type of needs, The image generation method according to claim 9, which includes the following:
11. The image generation method according to claim 9, further comprising the step of generating the target image using the target image generation method based on the image generation needs information if the image generation needs information does not require a reference image query.
12. The step of querying and obtaining a tenth reference image based on the aforementioned image generation needs information is as follows: The steps include: obtaining image subject information included in the aforementioned image generation needs information; The steps include: obtaining the first reference image by performing an image query in the configured image library based on the aforementioned image subject information; The image generation method according to claim 9, which includes the following:
13. The step of obtaining the first reference image by performing an image query in the configured image library based on the aforementioned image subject information is as follows: Based on the aforementioned image subject information, the step of performing an image query in the configured image library to obtain candidate reference images, The steps include obtaining the image quality of the candidate reference image and obtaining the set quality requirements corresponding to the image generation needs information, A step of obtaining the first reference image by filtering the candidate reference images based on the image quality and the set quality requirements, The image generation method according to claim 12, which includes the following:
14. An image generation device, A module for acquiring image generation needs information, A decision module for obtaining the needs text in the image generation needs information, inputting the needs text into a first large-scale model to detect the subject modification intention, obtaining the target subject modification intention corresponding to the image generation needs information, using the target subject modification intention to indicate whether or not the subject in the first reference image needs to be modified, and obtaining the target image generation method corresponding to the target subject modification intention based on the mapping relationship between the subject modification intention and the image generation method, A query module for querying and obtaining the first reference image based on the image generation needs information, A generation module for generating a target image using the target image generation method, based on the image generation needs information and the first reference image, An image generation device including [specific components].
15. The aforementioned query module, The image subject information included in the aforementioned image generation needs information is obtained, The image generation apparatus according to claim 14, which performs an image query in a configured image library based on the image subject information to obtain the first reference image.
16. The aforementioned query module, Based on the aforementioned image subject information, an image query is performed in the configured image library to obtain candidate reference images. The image quality of the candidate reference image is obtained, and the set quality requirements corresponding to the image generation needs information are obtained. The image generation apparatus according to claim 15, which performs filtering on the candidate reference images based on the image quality and the set quality requirements to obtain the first reference image.
17. The aforementioned needs text is associated with text features, and the generation module is, In response to the target subject modification intention indicating that there is no need to modify the subject in the first reference image, feature extraction is performed on the first reference image to obtain the first image features and obtain the subject segmentation image in the first reference image. The first image features and the text features are input into the first image generation model to obtain a background image and main body layout information. The image generation apparatus according to claim 14, which obtains the target image by fusing the subject division image and the background image based on the subject layout information.
18. The aforementioned needs text is associated with text features, and the generation module is, In response to the target subject modification intention indicating that the subject in the first reference image needs to be modified, feature extraction is performed on the first reference image to obtain a second image feature, Based on the text features and the descriptions corresponding to each sub-feature in the second image features, the text features and the second image features are alternately spliced to obtain alternating image and text features. The image generation apparatus according to claim 14, which inputs the alternating features of the image and text into a second image generation model to obtain the target image.
19. The image generation needs information includes a second reference image entered by the user, and the generation module, Feature extraction is performed on the second reference image mentioned above to obtain a third image feature. The image generation apparatus according to claim 18, which obtains alternating image and text features by alternately splicing the text features, the second image features and the third image features based on the description target corresponding to each sub-feature in the text features, the second image features and the third image features.
20. The aforementioned generation module is The needs text in the aforementioned image generation needs information is obtained, The aforementioned needs text is revised and expanded to obtain the target needs text. The image generation apparatus according to claim 14, which generates the target image using the target image generation method based on the target needs text and the first reference image.
21. The aforementioned generation module is The context information of the aforementioned image generation needs information is obtained, If the image generation needs information includes a second reference image entered by the user, the description information of the second reference image is obtained. The image generation apparatus according to claim 20, wherein at least one of the context information and the descriptive information, and the needs text are input into a second large-scale model and modified and expanded to obtain the target needs text.
22. An image generation device, A module for acquiring image generation needs information, A decision module for obtaining the needs text in the image generation needs information, inputting the needs text into a first large-scale model to detect the subject modification intention, obtaining the target subject modification intention corresponding to the image generation needs information, using the target subject modification intention to indicate whether or not the subject in the first reference image needs to be modified, and obtaining the target image generation method corresponding to the target subject modification intention based on the mapping relationship between the subject modification intention and the image generation method, A determination module for determining whether or not the aforementioned image generation needs information requires a reference image query, If the image generation needs information requires a reference image query, a query module is provided to query and obtain the first reference image based on the image generation needs information. A generation module for generating a target image using the target image generation method, based on the image generation needs information and the first reference image, An image generation device including [specific components].
23. The aforementioned determination module is The needs text in the aforementioned image generation needs information is obtained, Needs understanding is performed on the aforementioned needs text to obtain the type of needs corresponding to the image generation needs information. The image generation apparatus according to claim 22, which determines whether or not the image generation needs information needs to perform a reference image query based on the type of needs.
24. The aforementioned generation module further, The image generation apparatus according to claim 22, wherein, if the image generation needs information does not require a reference image query, the target image is generated using the target image generation method based on the image generation needs information.
25. The aforementioned query module, The image subject information included in the aforementioned image generation needs information is obtained, The image generation apparatus according to claim 22, which performs an image query in a configured image library based on the image subject information to obtain the first reference image.
26. The aforementioned query module, Based on the aforementioned image subject information, an image query is performed in the configured image library to obtain candidate reference images. The image quality of the candidate reference image is obtained, and the set quality requirements corresponding to the image generation needs information are obtained. The image generation apparatus according to claim 25, which performs filtering on the candidate reference images based on the image quality and the set quality requirements to obtain the first reference image.
27. It is an intelligent agent, An input module for obtaining image generation needs information, A processing module for generating a target image using the target image generation method, which involves obtaining the need text in the image generation needs information, inputting the need text into a first large-scale model to detect the subject modification intention, obtaining the target subject modification intention corresponding to the image generation needs information, using the target subject modification intention to indicate whether or not the subject in the first reference image needs to be modified, obtaining the target image generation method corresponding to the target subject modification intention based on the mapping relationship between the subject modification intention and the image generation method, querying and obtaining the first reference image based on the image generation needs information, and generating a target image using the target image generation method based on the image generation needs information and the first reference image. An output module for outputting the aforementioned target image, An intelligent agent that includes [this].
28. It is an intelligent agent, An input module for obtaining image generation needs information, A processing module for generating a target image by obtaining the need text in the image generation needs information, inputting the need text into a first large-scale model to detect the subject modification intention, obtaining the target subject modification intention corresponding to the image generation needs information, using the target subject modification intention to indicate whether or not the subject in the first reference image needs to be modified, obtaining the target image generation method corresponding to the target subject modification intention based on the mapping relationship between the subject modification intention and the image generation method, determining whether or not the image generation needs information needs to perform a reference image query, obtaining the first reference image by querying based on the image generation needs information, and generating a target image using the target image generation method based on the image generation needs information and the first reference image. An output module for outputting the aforementioned target image, An intelligent agent that includes [this].
29. The intelligent agent system is At least one processor, Includes a memory that is communicably connected to at least one processor, An intelligent agent system in which the memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor performs the method according to any one of claims 1 to 8 or 9 to 13.
30. A non-temporary computer-readable storage medium in which computer instructions are stored, The computer instruction is a non-temporary computer-readable storage medium that causes the computer to perform the method according to any one of claims 1 to 8 or 9 to 13.
31. A computer program that contains computer instructions, A computer program that, when the computer instruction is executed by a processor, implements the method according to any one of claims 1 to 8 or 9 to 13.