Image generation method and related equipment
By extracting prompt information from the basic images, generating similar images using the diffusion model and predicting resource allocation effects, the time consumption and quality instability caused by artificial design dependence is solved, and fast and efficient image generation and optimization are achieved.
Patent Information
- Application Number
- CN202410131077.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-30
- Publication Date
- 2025-08-01
AI Technical Summary
In the prior art, the image reproduction process relies on manual design, resulting in high time consumption and unstable content quality.
By extracting prompt information from the basic image, similar images are generated using the trained diffusion model, and resource allocation effects are predicted to determine the target image.
Quickly generate a large number of similar images under the premise of saving human resources, optimize resource allocation effects, and improve image quality stability.
Smart Images

Figure CN120411267A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of image generation technologies, and in particular, to an image generation method and related devices. Background Art
[0002] In the process of image production in the prior art, for images with good effects after resource allocation, these images need to be screened, and new images are replicated based on these images.
[0003] Traditionally, this replication process usually relies on manual design and creativity. Team members need to invest a lot of time in conceiving and producing similar high-quality images. In addition, obtaining images replicated by other members from other channels is also a new idea, but sometimes it may face the problem of unstable content quality. Summary of the Invention
[0004] In view of this, the purpose of the present disclosure is to propose an image generation method and related devices. There is no need to use manually conceived and produced replicated images of high-quality images, which can greatly reduce manual participation, reduce the consumed human resources, and also reduce the subjectivity brought by humans, effectively ensuring the quality of the replicated images.
[0005] Based on the above purpose, the present disclosure provides an image generation method, including: extracting prompt information from the obtained base image; inputting the prompt information into a trained diffusion model, and obtaining a to-be-predicted image through a diffusion process; obtaining a description feature of the to-be-predicted image, and predicting the resource allocation effect of the to-be-predicted image based on the description feature; determining a target image based on the resource allocation effect of the to-be-predicted image.
[0006] Based on the above image generation method, an embodiment of the present disclosure provides an image generation device, including:
[0007] An extraction module, configured to extract prompt information from the obtained base image;
[0008] A diffusion module, configured to input the prompt information into a trained diffusion model, and obtain a to-be-predicted image through a diffusion process;
[0009] A prediction module, configured to obtain a description feature of the to-be-predicted image, and predict the resource allocation effect of the to-be-predicted image based on the description feature;
[0010] A determination module, configured to determine a target image based on the resource allocation effect of the to-be-predicted image.
[0011] In addition, an embodiment of the present disclosure further provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the above method is implemented.
[0012] An embodiment of the present disclosure further provides a non-transitory computer-readable storage medium storing computer instructions for causing a computer to execute the above method.
[0013] An embodiment of the present disclosure further provides a computer program product including computer program instructions, which when running on a computer, cause the computer to execute the above method.
[0014] The above image generation method and related devices include: extracting prompt information from the obtained base image; inputting the prompt information into a trained diffusion model, and obtaining a to-be-predicted image through a diffusion process; obtaining a description feature of the to-be-predicted image, and predicting the resource allocation effect of the to-be-predicted image based on the description feature; and determining a target image based on the resource allocation effect of the to-be-predicted image. In the embodiments of the present disclosure, by extracting prompt information from the base image, a comprehensive interpretation of historical high-quality images is performed, and then the interpreted information is extracted to obtain prompt information. Then, a similar image (to-be-predicted image) of the base image is generated based on the prompt information. Also, since a large number of similar images can be generated at one time based on the foregoing method, the present disclosure determines the final qualified target image by predicting the resource allocation effect of the similar images. The image generation method of the embodiments of the present disclosure can quickly generate a large number of similar images on the premise of effectively saving human resources, and can screen the generated similar images by predicting the resource allocation effect in advance to obtain the final similar images, so as to optimize the actual resource allocation effect of the similar images. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] In order to more clearly illustrate the technical solutions in the present disclosure or related technologies, the following will briefly introduce the drawings required for use in the description of the embodiments or related technologies. Obviously, the drawings in the following description are only the embodiments of the present disclosure. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0016] Figure 1 Shows the implementation process of the image generation method according to some embodiments of the present disclosure;
[0017] Figure 2 Shows a schematic diagram of an image according to some embodiments of the present disclosure
[0018] Figure 3Shows a schematic diagram of an image generation device according to some embodiments of the present disclosure;
[0019] Figure 4 Shows a schematic diagram of the hardware structure of an electronic device according to an embodiment of the present disclosure. Detailed implementation manners
[0020] To make the purpose, technical solutions and advantages of the present disclosure clearer and more understandable, the following further elaborates on the present disclosure in detail with reference to specific embodiments and the accompanying drawings.
[0021] It should be noted that, unless otherwise defined, the technical terms or scientific terms used in the embodiments of the present disclosure should have the ordinary meanings understood by those with ordinary skills in the field to which the present disclosure belongs. The "first", "second" and similar terms used in the embodiments of the present disclosure do not indicate any order, quantity or importance, but are only used to distinguish different components. The terms such as "including" or "comprising" mean that the elements or objects appearing before this word cover the elements or objects listed after this word and their equivalents, without excluding other elements or objects. The terms such as "connected" or "coupled" are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. The terms such as "upper", "lower", "left" and "right" are only used to represent relative position relationships, and when the absolute position of the object being described changes, the relative position relationship may also change accordingly.
[0022] It can be understood that before using the technical solutions of the various embodiments of the present disclosure, the types, usage scopes, usage scenarios, etc. of the personal information involved will be informed to the user in an appropriate manner and the user's authorization will be obtained.
[0023] For example, when responding to receiving an active request from the user, a prompt message is sent to the user to clearly prompt the user that the operation requested by the user will require obtaining and using the user's personal information. Thus, the user can autonomously choose whether to provide personal information to software or hardware such as an electronic device, an application program, a server or a storage medium that executes the operation of the technical solution of the present disclosure according to the prompt message.
[0024] As an optional but non-limiting implementation manner, the way of sending a prompt message to the user in response to receiving an active request from the user can be, for example, in the form of a pop-up window. The prompt message can be presented in text in the pop-up window. In addition, the pop-up window can also carry a selection control for the user to choose "agree" or "disagree" to provide personal information to the electronic device.
[0025] It is understandable that the above notification and the process of obtaining user authorization are only illustrative and do not limit the implementation manner of the present disclosure. Other manners that comply with relevant laws and regulations can also be applied to the implementation manner of the present disclosure.
[0026] As described above, in the process of image production in the prior art, for images with good resource allocation effects, these images need to be screened, and new images are replicated based on these images. Traditionally, this replication process usually relies on manual design and creativity. Team members need to invest a lot of time in conceiving and producing similar high-quality images. In addition, obtaining images replicated by other members from other channels is also a new idea, but sometimes this may face the problem of unstable content quality.
[0027] To this end, some embodiments of the present disclosure provide an image generation method, which extracts prompt information from the obtained base image; inputs the prompt information into a trained diffusion model, and obtains a to-be-predicted image through a diffusion process; obtains the description features of the to-be-predicted image, and predicts the resource allocation effect of the to-be-predicted image based on the description features; determines a target image based on the resource allocation effect of the to-be-predicted image. In the embodiments of the present disclosure, by extracting prompt information from the base image, a comprehensive interpretation of historical high-quality images is performed, and then the information after interpretation is extracted to obtain prompt information. Then, a similar image (to-be-predicted image) of the base image is generated based on the prompt information. Also, because a large number of similar images can be generated at one time based on the foregoing method, the present disclosure determines the final qualified target image by predicting the resource allocation effects of the similar images. The image generation method of the embodiments of the present disclosure can quickly generate a large number of similar images on the premise of effectively saving human resources, and can screen the generated similar images by predicting the resource allocation effects in advance to obtain the final similar images, so as to optimize the actual resource allocation effect of the similar images.
[0028] Figure 1 shows the implementation process of the image generation method described in some embodiments of the present disclosure. As Figure 1 shown, the method may include the following steps:
[0029] In step 102, prompt information is extracted from the obtained base image.
[0030] In the embodiments of the present disclosure, the base image is generally an image with a good resource allocation effect. Therefore, those skilled in the art have a need to replicate these images to obtain more images with good resource allocation effects.
[0031] Step 102 can be understood as an in-depth analysis of the base image, that is, the understanding and insight into the base image. Because in order to be able to derive similar high-quality images from the base image, the embodiments of the present disclosure not only need to deeply understand the original content, but also need to accurately grasp those key elements that can ensure the resource allocation effect, and ensure that these specific elements can be retained in the newly generated images.
[0032] In the embodiments of the present disclosure, when evaluating the performance of an image, the following aspects are mainly considered:
[0033] Firstly, it is the basic characteristics of the image. The basic characteristics of the image generally may include the resolution, size, clarity, color tone, and the contrast between the theme and the background of the image.
[0034] The basic characteristics of the image can lay the overall tone of the image. Therefore, when evaluating the performance of the image, this aspect needs to be considered. It can also be said that it is necessary to first determine the basic characteristics of the image, or the basic characteristics of the image can be called a kind of constraint, so as to ensure that the subsequently generated image has a basic similarity with the base image.
[0035] Secondly, it is the main content of the image. The main content of the image involves the core elements in the image, that is, what the main body is, what activities the main body performs, and the state shown by the main body in the image.
[0036] Thirdly, it is the text information. The text information contained in the image is generally quite prominent and generally encompasses the core information of the image. Therefore, the text information contained in the image usually can play a major decision-making role in the resource allocation effect.
[0037] Finally, it is the layout design of the image. The layout design of the image refers to the positions of various elements in the image and the corresponding matching relationships. The layout design of the image needs to be added to the consideration of the image because a good layout design can significantly increase the overall expressiveness and attractiveness of the image and can effectively improve the resource allocation effect.
[0038] In some embodiments of the present disclosure, extracting the hint information from the obtained base image includes: extracting the basic information from the obtained base image; identifying the image content of the base image to obtain the content information; and extracting the hint information from the basic information and the content information.
[0039] In some embodiments of the present disclosure, the basic information includes at least one of the image size, resolution, color tone, and contrast of the base image.
[0040] Figure 2 Shows a schematic diagram of the image according to some embodiments of the present disclosure.
[0041] For reference Figure 2 , for example, Image 1 is circular, Image 2 and Image 3 are rectangular, and Image 2 and Image 3 are of the same size, and so on.
[0042] In the embodiments of the present disclosure, the foregoing basic information is the basic characteristics of the foregoing pictures. The basic characteristics can be obtained in the form of direct acquisition when obtaining, or the corresponding information can be obtained through simple computer vision methods. Computer vision is a science that studies how to make machines "see". Further, it refers to using cameras and computers to replace human eyes to perform machine vision such as target recognition, tracking, and measurement on targets, and further performing graphic processing to make the computer-processed images more suitable for human eye observation or transmission to instruments for detection.
[0043] In some embodiments of the present disclosure, the recognition of the image content of the basic image to obtain content information includes: recognizing the text content of the basic image to obtain text information; obtaining the position information of the text information to obtain text position information; recognizing the image content of the basic image to obtain image information; recognizing the image layout of the basic image to obtain layout information; and obtaining the content information based on the text information, the text position information, the image information, and the layout information.
[0044] Such as Figure 2 shown, in the embodiments of the present disclosure, the recognition of the image content of the basic image may include recognizing the text content of the text, recognizing the position of the text, recognizing the content of the image, and further including recognizing the layout of the image.
[0045] The recognition of the text content of the basic image can first be recognized through neural network technology. Specifically, it can be further divided into: first, it is necessary to detect where there is text. In the embodiments of the present disclosure, the text can be detected by means of object detection.
[0046] The task of object detection is to find all interesting targets (objects) in the image. Different from classification and regression problems, object detection also needs to determine the position of the target in the image (localization), and determine the category and position of the recognized target (classification and localization).
[0047] For example, in Figure 2 after recognizing the text, it can be obtained that the text is located in the upper left corner of the picture.
[0048] Furthermore, after recognizing the position of the text, the detected text position area can be recognized to obtain the specific content of the text.
[0049] In an embodiment of the present disclosure, the Optical Character Recognition (OCR) technology can be used to recognize the specific content of the text. Optical character recognition refers to the process of analyzing and recognizing an image file of text materials to obtain text and layout information. That is, the text in the image is recognized and returned in text form.
[0050] In an embodiment of the present disclosure, the image content of the base image can be recognized to obtain image information. Here, the image information can be very specific information. For example, it can be recognized that the image currently includes perfume of brand A, specifically how many bottles, and also includes skin care lotion of brand A, specifically how many bottles, and so on.
[0051] After that, the image layout of the base image can be recognized. The image layout here can include: multi-image splicing materials, single-image materials, text display at the top, text display in the middle, etc., or it can also refer to Figure 2 , for example, the image is a multi-image splicing material, and there is no overlap between each image. Image 1 is circular and is located on the left side of the overall image and also on the left side of Image 2. Image 2 and Image 3 are vertically aligned, Image 3 is located below Image 2, and both Image 2 and Image 3 are located on the right side of the image.
[0052] After that, based on the text information, the text position information, the image information, and the layout information, the content information is obtained. In an embodiment of the present disclosure, after obtaining the text information, the text position information, the image information, and the layout information, the foregoing information is integrated to obtain the content information of the image, that is, information that details the image content, layout, etc.
[0053] In an embodiment of the present disclosure, detailed description information corresponding to the image can also be directly generated through multi-modal technology. For example, there is a man standing on a surfboard surfing in the sea. It can directly and detailedly describe the main information of the image.
[0054] In some embodiments of the present disclosure, the information extraction of the base information and the content information to obtain the prompt information includes: obtaining the base information, the text position information, and the layout information; extracting the core information of the text information and the image information; and obtaining the prompt information based on the base information, the text information, the layout information, and the core information.
[0055] In an embodiment of the present disclosure, after obtaining the text information, text position information, image information, and layout information, detailed information corresponding to the image can be comprehensively obtained. Subsequently, when replicating the image, it is replicated based on the prompt information. However, after obtaining the aforementioned detailed information, the detailed information cannot be directly used as the prompt information, because for the embodiments of the present disclosure, it is necessary to create a similar image, and only the most core content of the base image needs to be retained. Retaining too many details is instead not conducive to the subsequent replication of the image.
[0056] Therefore, it is necessary to extract the content of the aforementioned obtained detailed information to obtain the corresponding prompt information. The specific steps can be to retain the aforementioned obtained basic information, text position information, and layout information. For example, the basic information is that image 1 is circular, image 2 is square, image 3 is square, the background color of the picture is warm-toned, and there is a contrast with the cosmetic picture.
[0057] For the text information and image information, only their core content needs to be retained. For example, referring to Figure 2 , if the text information is "A brand of cosmetics is very cost-effective", it can be extracted as "Cosmetics are cost-effective". For the image information, the aforementioned obtained image information is "The image is composed of three pictures spliced together. Image 1 is a girl's hand holding a perfume of brand A, image 2 is a girl's hand holding a skin care lotion of brand A, and image 3 is a girl's hand holding a liquid foundation of brand A".
[0058] The finally retained prompt information can be "The image is composed of three images spliced together. Each image is a human hand holding different cosmetics. The cosmetics can be skin care products or perfumes. A text description needs to be output at the upper left corner of the picture, indicating that these cosmetics are very cost-effective. The overall background color tone of the image needs to be warm-toned and there needs to be a contrast with the cosmetic picture".
[0059] After the extraction of the prompt information of the image in the aforementioned steps is completed, it is necessary to generate a similar image (image to be predicted) based on the formed prompt information.
[0060] In step 104, in some embodiments of the present disclosure, obtaining the image to be predicted through the diffusion process includes: obtaining a random noise image; gradually transforming the random noise image into a coherent image based on the prompt information to obtain the image to be predicted.
[0061] In the embodiments of the present disclosure, based on the aforementioned understanding and insight of the image, high-quality prompt information has been obtained, and then an image generation algorithm can be used to generate a corresponding similar image based on the originally obtained prompt information.
[0062] Specifically, the embodiments of the present disclosure use the Stable Diffusion algorithm to generate images.
[0063] The Stable Diffusion algorithm is a novel image generation technology aimed at generating high-fidelity images in a controllable and stable manner. This algorithm is based on the idea of image denoising and gradually transforms a random noise pattern into a coherent image structure through a diffusion process.
[0064] The core of the Stable Diffusion algorithm lies in its two main steps:
[0065] The first step is diffusion, that is, after inputting the original image, it simulates the effect of the natural diffusion process in the high-dimensional data space, continuously adds noise to the original information, and introduces randomness. The second step is reverse diffusion, that is, using a deep learning model to guide the diffusion process towards the desired image input, and finally obtaining the image to be generated.
[0066] Moreover, for the denoising process, in addition to the information after adding noise to the original image, it also includes conditional information in each denoising process. These conditional information are usually composed of some text descriptions or some constraints.
[0067] In some embodiments of the present disclosure, gradually transforming the random noise image into a coherent image based on the prompt information to obtain the image to be predicted includes: using the prompt information as the conditional information of the random noise image and denoising the random noise image to obtain the image to be predicted.
[0068] In the embodiments of the present disclosure, the conditional information can be the prompt information.
[0069] In some embodiments of the present disclosure, the trained diffusion model is obtained by the following method: obtaining an image with a preset parameter higher than a preset threshold from historical images to obtain a historical basic image;
[0070] Extracting historical prompt information according to the historical basic image; inputting the historical prompt information into an untrained diffusion model to obtain a historical image to be predicted; training the untrained diffusion model based on the historical image to be predicted and the historical basic image to obtain the trained diffusion model.
[0071] [[ID=X]]In the embodiments of the present disclosure, since the training data of the standard Stable Diffusion model are often some general pictures, which do not match the scenario of the present disclosure, the images generated by the present disclosure need to be similar to the basic images and have a better resource allocation effect. Therefore, the requirements for the images in the present disclosure are special, and the images generated by the standard Stable Diffusion model cannot be directly used.
[0072] In the embodiments of the present disclosure, there are many high-quality base images, and the resource allocation effects of these high-quality base images are all very good. Therefore, these high-quality base images can be used to retrain the stable diffusion model to improve the quality of the generated images. Of course, during training, the model can be trained from scratch or retrained on an already trained model, which can effectively save the time for training the model.
[0073] During the training process, in the process of obtaining the prompt information of the high-quality base images, the aforementioned steps for generating prompt information can be used for generation, and then an image is generated based on this prompt information, and then the relevant loss can be calculated.
[0074] Furthermore, after generating the images, since resources are limited, it is necessary to allocate resources to the highest-quality images. Therefore, the process does not end after generating the images. A model can also be set up to predict the resource allocation effect of the generated images, and the images with better resource allocation effects are used as the final target images.
[0075] In step 106, in some embodiments of the present disclosure, obtaining the description features of the image to be predicted and predicting the resource allocation effect of the image to be predicted based on the description features includes: obtaining the description features and category features of the image to be predicted; obtaining the historical resource allocation effects of the base images; obtaining the prompt information corresponding to the image to be predicted and obtaining the description features of the prompt information; predicting the resource allocation effect of the image to be predicted based on the description features and category features of the image to be predicted, the historical resource allocation effects of the base images, and the description features of the prompt information.
[0076] In some embodiments of the present disclosure, determining the target image based on the resource allocation effect of the image to be predicted includes: in response to the resource allocation effect of the image to be predicted being higher than a preset threshold, determining the image to be predicted as the target image.
[0077] In the embodiments of the present disclosure, since there have been a large number of images in history, with image information and the resource allocation effects of the corresponding images, the model can be trained based on this historical information.
[0078] After the training of the model is completed, the resource allocation effect of the aforementioned obtained image to be predicted can be predicted. In the specific prediction process, first, the description features and category features of the image to be predicted can be obtained. The description features refer to representing an object with a low-dimensional vector, which can be a word, a commodity, a movie, etc. The property of this vector is that objects corresponding to vectors with close distances have similar meanings.
[0079] Subsequently, the descriptive features of the prompt information can also be obtained. Then, due to the relationship that the image to be predicted is a replicated image of the base image, the category of the image to be predicted can also be obtained. Of course, the categories of the base image and the image to be predicted are the same.
[0080] Subsequently, dense features such as the size, resolution of the image to be predicted, and the historical resource allocation effect of the base image can also be obtained.
[0081] Subsequently, the features obtained above are input into a neural network to calculate features using a multi-task structure. Finally, each prediction target is estimated using a branch, and the underlying features of each sub-task are shared. This can effectively reduce the computational amount and improve the computational speed through parameter sharing.
[0082] In the above image generation method, prompt information is extracted from the obtained base image; the prompt information is input into a pre-trained diffusion model, and the image to be predicted is obtained through the diffusion process; the descriptive features of the image to be predicted are obtained, and the resource allocation effect of the image to be predicted is predicted based on the descriptive features; the target image is determined based on the resource allocation effect of the image to be predicted. In the embodiments of the present disclosure, prompt information is extracted from the base image to comprehensively interpret high-quality historical images, and then the interpreted information is extracted to obtain prompt information. Then, a similar image (image to be predicted) of the base image is generated based on the prompt information. Since a large number of similar images can be generated at one time based on the foregoing method, the present disclosure determines the final qualified target image by predicting the resource allocation effect of the similar images. The image generation method of the embodiments of the present disclosure can quickly generate a large number of similar images on the premise of effectively saving human resources, and can screen the generated similar images by predicting the resource allocation effect in advance to obtain the final similar images, so as to optimize the actual resource allocation effect of the similar images.
[0083] It should be noted that the method of the embodiments of the present disclosure can be executed by a single device, such as a computer or a server. The method of this embodiment can also be applied to a distributed scenario and completed by multiple devices cooperating with each other. In this case of a distributed scenario, one of the multiple devices can only execute one or more steps of the method of the embodiments of the present disclosure, and these multiple devices will interact with each other to complete the described method.
[0084] Note that the above describes some embodiments of the present disclosure. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than in the above embodiments and still achieve the desired results. Additionally, the processes depicted in the figures do not necessarily require the particular order or sequential order shown to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0085] Based on the same inventive concept, corresponding to any of the above method embodiments, the present disclosure further provides an image generation device, including:
[0086] An extraction module 302, configured to extract prompt information from the acquired base image;
[0087] A diffusion module 304, configured to input the prompt information into a trained diffusion model, and obtain a to-be-predicted image through a diffusion process;
[0088] A prediction module 306, configured to obtain a description feature of the to-be-predicted image, and predict a resource allocation effect of the to-be-predicted image based on the description feature;
[0089] A determination module 308, configured to determine a target image based on the resource allocation effect of the to-be-predicted image.
[0090] In some embodiments of the present disclosure, the extraction module 302 includes:
[0091] An extraction unit, configured to extract base information from the acquired base image;
[0092] An identification unit, configured to identify the image content of the base image to obtain content information;
[0093] An information extraction unit, configured to perform information extraction on the base information and the content information to obtain the prompt information.
[0094] In some embodiments of the present disclosure, the base information includes at least one of the image size, resolution, hue, and contrast of the base image.
[0095] In some embodiments of the present disclosure, the identification unit includes:
[0096] Identify the text content of the base image to obtain text information;
[0097] Obtain the position information of the text information to obtain text position information;
[0098] Identify the image content of the base image to obtain image information;
[0099] Identify the image layout of the base image to obtain layout information;
[0100] Based on the text information, the text position information, the image information, and the layout information, obtain the content information.
[0101] In some embodiments of the present disclosure, the extraction unit includes:
[0102] Obtain the base information, the text position information, and the layout information;
[0103] Extract the core information of the text information and the image information;
[0104] Based on the base information, the text information, the layout information, and the core information, obtain the prompt information.
[0105] In some embodiments of the present disclosure, the diffusion module 304 includes:
[0106] A first acquisition unit for acquiring a random noise image;
[0107] A conversion unit for gradually converting the random noise image into a coherent image based on the prompt information to obtain the image to be predicted.
[0108] In some embodiments of the present disclosure, the conversion unit includes:
[0109] Use the prompt information as the conditional information of the random noise image to denoise the random noise image to obtain the image to be predicted.
[0110] In some embodiments of the present disclosure, it further includes:
[0111] A historical image acquisition module for acquiring images with preset parameters higher than a preset threshold from historical images to obtain historical base images;
[0112] A historical prompt information extraction module for extracting historical prompt information according to the historical base images;
[0113] An input module for inputting the historical prompt information into an untrained diffusion model to obtain a historical image to be predicted;
[0114] A training module for training the untrained diffusion model based on the historical image to be predicted and the historical base images to obtain the trained diffusion model.
[0115] In some embodiments of the present disclosure, the prediction module 306 includes:
[0116] A second acquisition unit, configured to acquire the description features and category features of the image to be predicted;
[0117] A third acquisition unit, configured to acquire the historical resource allocation effect of the base image;
[0118] A fourth acquisition unit, configured to acquire the prompt information corresponding to the image to be predicted and acquire the description features of the prompt information;
[0119] A prediction unit, configured to predict the resource allocation effect of the image to be predicted based on the description features and category features of the image to be predicted, the historical resource allocation effect of the base image, and the description features of the prompt information.
[0120] In some embodiments of the present disclosure, the determination module 308 includes:
[0121] A determination unit, configured to determine the image to be predicted as the target image in response to the resource allocation effect of the image to be predicted being higher than a preset threshold.
[0122] For the convenience of description, the above devices are described by function as various modules respectively. Of course, when implementing the present disclosure, the functions of each module can be implemented in one or more software and / or hardware.
[0123] The devices in the above embodiments are used to implement the corresponding image generation method in any of the foregoing embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be described in detail herein.
[0124] Based on the same inventive concept, corresponding to the method in any of the above embodiments, the present disclosure further provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, where the processor implements the image generation method described in any of the above embodiments when executing the program.
[0125] Figure 4 FIG. shows a more specific schematic diagram of the hardware structure of the electronic device provided in this embodiment. The device may include: a processor 1010, a memory 1020, an input / output interface 1030, a communication interface 1040, and a bus 1050. Among them, the processor 1010, the memory 1020, the input / output interface 1030, and the communication interface 1040 are communicatively connected to each other inside the device through the bus 1050.
[0126] The processor 1010 can be implemented in the form of a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, etc., and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this specification.
[0127] The memory 1020 can be implemented in the form of a ROM (Read Only Memory), a RAM (Random Access Memory), a static storage device, a dynamic storage device, etc. The memory 1020 can store an operating system and other application programs. When implementing the technical solutions provided in the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 1020 and are called and executed by the processor 1010.
[0128] The input / output interface 1030 is used to connect to the input / output module to implement information input and output. The input / output module can be configured as a component in the device (not shown in the figure) or can be externally connected to the device to provide corresponding functions. Among them, the input device can include a keyboard, a mouse, a touch screen, a microphone, various sensors, etc., and the output device can include a display, a speaker, a vibrator, an indicator light, etc.
[0129] The communication interface 1040 is used to connect to a communication module (not shown in the figure) to implement communication interaction between this device and other devices. Among them, the communication module can implement communication through a wired method (such as USB, network cable, etc.) or can implement communication through a wireless method (such as a mobile network, WIFI, Bluetooth, etc.).
[0130] The bus 1050 includes a path for transmitting information between various components of the device (such as the processor 1010, the memory 1020, the input / output interface 1030, and the communication interface 1040).
[0131] It should be noted that although the above device only shows the processor 1010, the memory 1020, the input / output interface 1030, the communication interface 1040, and the bus 1050, in the specific implementation process, this device may also include other components necessary for normal operation. In addition, those skilled in the art can understand that the above device may also only include the components necessary to implement the solutions of the embodiments of this specification and do not necessarily include all the components shown in the figure.
[0132] The electronic device of the above embodiment is used to implement the corresponding image generation method in any of the foregoing embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be elaborated here.
[0133] Based on the same inventive concept, corresponding to the method of any of the above embodiments, the present disclosure also provides a non-transitory computer-readable storage medium storing computer instructions for causing the computer to execute the image generation method described in any of the above embodiments.
[0134] The computer-readable medium of this embodiment includes permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, magnetic tape magnetic disk storage or other magnetic storage devices, or any other non-transmission medium that can be used to store information accessible by a computing device.
[0135] The computer instructions stored in the storage medium of the above embodiment are used to cause the computer to execute the image generation method described in any of the above embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be elaborated here.
[0136] Based on the same inventive concept, corresponding to the image generation method described in any of the above embodiments, the present disclosure also provides a computer program product including computer program instructions. In some embodiments, the computer program instructions can be executed by one or more processors of the computer to cause the computer and / or the processor to execute the image generation method. Corresponding to the execution subject corresponding to each step in the embodiments of the image generation method, the processor executing the corresponding step can belong to the corresponding execution subject.
[0137] The computer program product of the above embodiment is used to cause the computer and / or the processor to execute the image generation method described in any of the above embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be elaborated here.
[0138] Those of ordinary skill in the art should understand that the discussion of any of the above embodiments is merely exemplary and is not intended to imply that the scope of the present disclosure (including the claims) is limited to these examples; under the concept of the present disclosure, the technical features in the above embodiments or different embodiments can also be combined, the steps can be implemented in any order, and there are many other variations in different aspects of the embodiments of the present disclosure as described above, and they are not provided in detail for the sake of brevity.
[0139] In addition, for simplicity of explanation and discussion, and in order not to make the embodiments of the present disclosure difficult to understand, the well-known power / ground connections to integrated circuit (IC) chips and other components may or may not be shown in the provided drawings. Further, the devices may be shown in block diagram form in order to avoid making the embodiments of the present disclosure difficult to understand, and this also takes into account the fact that the details of the implementation of these block diagram devices are highly dependent on the platform on which the embodiments of the present disclosure are to be implemented (i.e., these details should be fully within the understanding of those skilled in the art). In the case where specific details (such as circuits) are set forth to describe exemplary embodiments of the present disclosure, it will be apparent to those skilled in the art that the embodiments of the present disclosure may be implemented without these specific details or with variations of these specific details. Therefore, these descriptions should be considered illustrative rather than restrictive.
[0140] Although the present disclosure has been described in connection with specific embodiments thereof, many alternatives, modifications, and variations of these embodiments will be apparent to those of ordinary skill in the art based on the foregoing description. For example, other memory architectures (such as dynamic RAM (DRAM)) may be used with the embodiments discussed.
[0141] The embodiments of the present disclosure are intended to cover all such alternatives, modifications, and variations that fall within the broad scope of the appended claims. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the embodiments of the present disclosure shall be included within the protection scope of the present disclosure.
Claims
1. An image generation method, comprising: Extracting prompt information from the acquired base image; Inputting the prompt information into a trained diffusion model, and obtaining a to-be-predicted image through a diffusion process; Obtaining the description features of the to-be-predicted image, and predicting the resource allocation effect of the to-be-predicted image based on the description features; Determining a target image based on the resource allocation effect of the to-be-predicted image.
2. The method according to claim 1, wherein The extracting prompt information from the acquired base image includes: Extracting base information from the acquired base image; Identifying the image content of the base image to obtain content information; Performing information extraction on the base information and the content information to obtain the prompt information.
3. The method according to claim 2, wherein The base information includes at least one of the image size, resolution, hue, and contrast of the base image.
4. The method according to claim 2, wherein, The identifying the image content of the base image to obtain content information includes: Identifying the text content of the base image to obtain text information; Obtaining the position information of the text information to obtain text position information; Identifying the image content of the base image to obtain image information; Identifying the image layout of the base image to obtain layout information; Based on the text information, the text position information, the image information, and the layout information, obtaining the content information.
5. The method according to claim 4, wherein, The performing information extraction on the base information and the content information to obtain the prompt information includes: Obtaining the base information, the text position information, and the layout information; Extracting the core information of the text information and the image information; Based on the base information, the text information, the layout information, and the core information, obtaining the prompt information.
6. The method according to claim 1, wherein, The obtaining a to-be-predicted image through a diffusion process includes: Obtaining a random noise image; Based on the prompt information, gradually converting the random noise image into a coherent image to obtain the to-be-predicted image.
7. The method according to claim 6, wherein The based on the prompt information, gradually converting the random noise image into a coherent image to obtain the to-be-predicted image includes: Using the prompt information as the conditional information of the random noise image, and denoising the random noise image to obtain the to-be-predicted image.
8. The method according to claim 1, wherein, The trained diffusion model is obtained through the following method: Obtaining an image with a preset parameter higher than a preset threshold from historical images to obtain a historical base image; Extracting historical prompt information according to the historical base image; Inputting the historical prompt information into an untrained diffusion model to obtain a historical to-be-predicted image; Training the untrained diffusion model based on the historical to-be-predicted image and the historical base image to obtain the trained diffusion model.
9. The method according to claim 1, wherein The obtaining the description features of the to-be-predicted image, and predicting the resource allocation effect of the to-be-predicted image based on the description features includes: Obtaining the description features and category features of the to-be-predicted image; Obtaining the historical resource allocation effect of the base image; Obtaining the prompt information corresponding to the to-be-predicted image, and obtaining the description features of the prompt information; Predict the resource allocation effect of the image to be predicted based on the description features and category features of the image to be predicted, the historical resource allocation effect of the base image, and the description features of the prompt information.
10. The method according to claim 1, wherein Determining a target image based on the resource allocation effect of the image to be predicted includes: In response to the resource allocation effect of the image to be predicted being higher than a preset threshold, determining the image to be predicted as the target image.
11. An image generation device, comprising: An extraction module for extracting prompt information from the acquired base image; A diffusion module for inputting the prompt information into a trained diffusion model and obtaining an image to be predicted through a diffusion process; A prediction module for obtaining the description features of the image to be predicted and predicting the resource allocation effect of the image to be predicted based on the description features; A determination module for determining a target image based on the resource allocation effect of the image to be predicted.
12. The device according to claim 11, wherein, The extraction module includes: An extraction unit for extracting basic information from the acquired base image; An identification unit for identifying the image content of the base image to obtain content information; An information extraction unit for extracting information from the basic information and the content information to obtain the prompt information.
13. The device according to claim 12, wherein, The basic information includes at least one of the image size, resolution, hue, and contrast of the base image.
14. The apparatus according to claim 12, wherein, The identification unit includes: Identifying the text content of the base image to obtain text information; Obtaining the position information of the text information to obtain text position information; Identifying the image content of the base image to obtain image information; Identifying the image layout of the base image to obtain layout information; Based on the text information, the text position information, the image information, and the layout information, obtaining the content information.
15. The apparatus according to claim 14, wherein The extraction unit includes: Obtaining the basic information, the text position information, and the layout information; Extracting the core information of the text information and the image information; Based on the basic information, the text information, the layout information, and the core information, obtaining the prompt information.
16. The apparatus according to claim 11, wherein The diffusion module includes: A first acquisition unit for acquiring a random noise image; A conversion unit for gradually converting the random noise image into a coherent image based on the prompt information to obtain the image to be predicted.
17. The apparatus according to claim 16, wherein, The conversion unit includes: Using the prompt information as the conditional information of the random noise image to denoise the random noise image to obtain the image to be predicted.
18. The device according to claim 11, further comprising: A historical image acquisition module for acquiring an image with a preset parameter higher than a preset threshold from historical images to obtain a historical base image; A historical prompt information extraction module for extracting historical prompt information according to the historical base image; An input module for inputting the historical prompt information into an untrained diffusion model to obtain a historical image to be predicted; A training module, configured to train the untrained diffusion model based on the historical image to be predicted and the historical base image, so as to obtain the trained diffusion model.
19. The apparatus according to claim 11, wherein, The prediction module includes: A second acquisition unit, configured to acquire the description feature and the category feature of the image to be predicted. A third acquisition unit, configured to acquire the historical resource allocation effect of the base image. A fourth acquisition unit, configured to acquire the prompt information corresponding to the image to be predicted and acquire the description feature of the prompt information. A prediction unit, configured to predict the resource allocation effect of the image to be predicted based on the description feature and the category feature of the image to be predicted, the historical resource allocation effect of the base image, and the description feature of the prompt information.
20. The apparatus according to claim 11, wherein The determination module includes: A determination unit, configured to determine the image to be predicted as the target image in response to the resource allocation effect of the image to be predicted being higher than a preset threshold.
21. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, the method according to any one of claims 1 to 10 is implemented.
22. A non-transitory computer-readable storage medium storing computer instructions, characterized in that, The computer instructions are used to cause a computer to execute the method according to any one of claims 1 to 10.
23. A computer program product, including computer program instructions, which when running on a computer, cause the computer to execute the method according to any one of claims 1 to 10.
Citation Information
Cited By
Image generation method and related device
WO2025162038A1