Image generation method and related device

By extracting prompt information from the basic images, using the diffusion model to generate similar images and predict resource allocation effects, the instability problem caused by manual design is solved, and high-quality similar images are quickly and efficiently generated.

WO2025162038A1PCT designated stage Publication Date: 2025-08-07BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2025/073456
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-01-30
Filing Date
2025-01-20
Publication Date
2025-08-07

AI Technical Summary

Technical Problem

In the prior art, the image production process relies on manual design and creativity, resulting in unstable effects after resource allocation, and generating similar images requires a large amount of human resources.

Method used

By extracting prompt information from the basic image, a similar image is generated using the trained diffusion model, and the resource allocation effect is predicted based on the description feature, and the target image is determined.

Benefits of technology

On the premise of saving human resources, a large number of similar images are quickly generated, and the final target images that meet the conditions are filtered out by predicting resource allocation effects to optimize the resource allocation effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025073456_07082025_PF_FP_ABST
    Figure CN2025073456_07082025_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure provides an image generation method, comprising: extracting prompt information from an obtained basic image; inputting the prompt information into a trained diffusion model, and by means of a diffusion process, obtaining an image to be predicted; obtaining a description feature of the image to be predicted, and on the basis of the description feature, predicting a resource allocation effect of the image to be predicted; and determining a target image on the basis of the resource allocation effect of the image to be predicted. On the basis of the image generation method, the present disclosure further provides an image generation apparatus, an electronic device, a storage medium and a program product.
Need to check novelty before this filing date? Find Prior Art

Description

Image generation method and related equipment

[0001] This application claims priority to the Chinese invention patent application entitled “Image Generation Method and Related Equipment” and application number 2024101310776, filed on January 30, 2024, the entire contents of which are incorporated by reference into this application. Technical Field

[0002] The present disclosure relates to the field of image generation technology, and in particular to an image generation method and related equipment. Background Art

[0003] In the image production process of the prior art, for images with better effects after resource allocation, it is necessary to screen these images and reproduce new images based on these images.

[0004] Traditionally, this reproduction process usually relies on manual design and creativity, and team members need to invest a lot of time to conceive and produce similar high-quality images. In addition, obtaining images from other members to reproduce images from other channels is also a new idea, but this sometimes faces the problem of unstable content quality. Summary of the Invention

[0005] In view of this, the purpose of the present disclosure is to provide an image generation method and related devices.

[0006] The present disclosure provides an image generation method, comprising: extracting prompt information from an acquired basic image; inputting the prompt information into a trained diffusion model to obtain an image to be predicted through a diffusion process; obtaining descriptive features of the image to be predicted, and predicting a resource allocation effect of the image to be predicted based on the descriptive features; and determining a target image based on the resource allocation effect of the image to be predicted.

[0007] Based on the above-mentioned image generation method, an embodiment of the present disclosure provides an image generation device, comprising: an extraction module for extracting prompt information from an acquired basic image; a diffusion module for inputting the prompt information into a trained diffusion model to obtain an image to be predicted through a diffusion process; a prediction module for obtaining descriptive features of the image to be predicted and predicting a resource allocation effect of the image to be predicted based on the descriptive features; and a determination module for determining a target image based on the resource allocation effect of the image to be predicted.

[0008] In addition, an embodiment of the present disclosure further provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the above method when executing the program.

[0009] An embodiment of the present disclosure further provides a non-transitory computer-readable storage medium, wherein the non-transitory computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a computer to execute the above method.

[0010] An embodiment of the present disclosure further provides a computer program product, comprising computer program instructions, which, when executed on a computer, enable the computer to execute the above method.

[0011] The above-mentioned image generation method and related equipment include: extracting prompt information from the acquired basic image; inputting the prompt information into a trained diffusion model to obtain the image to be predicted through the diffusion process; obtaining the descriptive features of the image to be predicted, and predicting the resource allocation effect of the image to be predicted based on the descriptive features; and determining the target image based on the resource allocation effect of the image to be predicted. The embodiment of the present disclosure extracts prompt information from the basic image, comprehensively interprets historical high-quality images, and then extracts prompt information from the interpreted information. Then, based on the prompt information, a similar image (the image to be predicted) of the basic image is generated. Since a large number of similar images can be generated at one time based on the above-mentioned method, the present disclosure determines the target image that ultimately meets the conditions by predicting the resource allocation effect of the similar images. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] In order to more clearly illustrate the technical solutions in the present disclosure or related technologies, the following briefly introduces the drawings required for use in the embodiments or related technical descriptions. Obviously, the drawings described below are only embodiments of the present disclosure. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0013] FIG1 shows an implementation process of an image generation method according to some embodiments of the present disclosure;

[0014] FIG2 shows a schematic diagram of an image according to some embodiments of the present disclosure;

[0015] FIG3 shows a schematic diagram of an image generating device according to some embodiments of the present disclosure; and

[0016] FIG4 shows a schematic diagram of the hardware structure of an electronic device according to an embodiment of the present disclosure. DETAILED DESCRIPTION

[0017] In order to make the objectives, technical solutions and advantages of the present disclosure more clearly understood, the present disclosure is further described in detail below in conjunction with specific embodiments and with reference to the accompanying drawings.

[0018] It should be noted that, unless otherwise defined, the technical terms or scientific terms used in the embodiments of the present disclosure should have the usual meanings understood by people with ordinary skills in the field to which the present disclosure belongs. The "first", "second" and similar words used in the embodiments of the present disclosure do not indicate any order, quantity or importance, but are only used to distinguish different components. "Include" or "comprise" and similar words mean that the elements or objects appearing before the word include the elements or objects listed after the word and their equivalents, without excluding other elements or objects. "Connect" or "connected" and similar words are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. "Up", "down", "left", "right" and the like are only used to indicate relative position relationships. When the absolute position of the described object changes, the relative position relationship may also change accordingly.

[0019] It is understandable that before using the technical solutions of each embodiment of the present disclosure, the type, scope of use, usage scenarios, etc. of the personal information involved will be informed to the user in an appropriate manner, and the user's authorization will be obtained.

[0020] For example, in response to a user's active request, a prompt message is sent to the user to clearly inform the user that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the electronic device, application, server, storage medium, or other software or hardware that performs the operation of the disclosed technical solution based on the prompt message.

[0021] As an optional but non-limiting implementation, in response to a user's active request, the prompt information may be sent to the user in the form of a pop-up window, in which the prompt information may be presented in text form. Furthermore, the pop-up window may also contain a selection control for the user to select "agree" or "disagree" to provide personal information to the electronic device.

[0022] It is understandable that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of the present disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of the present disclosure.

[0023] As mentioned above, in the existing image production process, images that perform well after resource allocation need to be screened and then reproduced based on these images. Traditionally, this reproduction process often relies on manual design and creativity, requiring team members to invest a considerable amount of time to conceive and produce similar high-quality images. Alternatively, obtaining images reproduced by other team members from other sources is another novel approach, but this sometimes leads to inconsistent content quality.

[0024] To this end, some embodiments of the present disclosure provide an image generation method, which extracts prompt information from an acquired base image; inputs the prompt information into a trained diffusion model to obtain a to-be-predicted image through a diffusion process; obtains descriptive features of the to-be-predicted image, and predicts the resource allocation effect of the to-be-predicted image based on the descriptive features; and determines a target image based on the resource allocation effect of the to-be-predicted image. The present disclosure embodiment extracts prompt information from the base image, comprehensively interprets historical high-quality images, then extracts prompt information from the interpreted information, and then generates a similar image (to-be-predicted image) of the base image based on the prompt information. Because the aforementioned method can generate a large number of similar images at once, the present disclosure determines the final target image that meets the requirements by predicting the resource allocation effect of the similar images. The image generation method of the present disclosure embodiment can quickly generate a large number of similar images while effectively saving human resources, and can screen the generated similar images by pre-predicting the resource allocation effect to obtain the final similar image, thereby optimizing the actual resource allocation effect of the similar images.

[0025] FIG1 shows an implementation process of an image generation method according to some embodiments of the present disclosure. As shown in FIG1 , the method may include the following steps:

[0026] In step 102, prompt information is extracted from the obtained basic image. In the embodiments of the present disclosure, the basic image is generally an image with good resource allocation effect. Therefore, those skilled in the art need to replicate these images to obtain more images with good resource allocation effect.

[0027] Step 102 can be understood as an in-depth analysis of the base image, that is, an understanding and insight into the base image. Because in order to derive a similar high-quality image from the base image, the embodiment of the present disclosure not only needs to have an in-depth understanding of the original content, but also needs to accurately grasp those key elements that can ensure the resource allocation effect, and ensure that these characteristics can be retained in the newly generated image.

[0028] In the embodiments of the present disclosure, the performance of an image is evaluated mainly by considering the following aspects.

[0029] The first is the basic characteristics of the image, which generally include the image's resolution, size, clarity, color tonality, and the contrast between the subject and background.

[0030] The basic characteristics of an image can set the tone for the entire image, so this aspect needs to be considered when evaluating the performance of an image. In other words, the basic characteristics of the image need to be determined first, or the basic characteristics of the image can be called a constraint, so as to ensure that the subsequently generated image has a basic similarity with the basic image.

[0031] The second is the main content of the image. The main content of the image involves the core elements in the image, that is, what the subject is, what activities the subject performs, and the state of the subject in the image.

[0032] The second is text information. The text information contained in the image is generally more conspicuous and generally includes the core information of the image. Therefore, the text information contained in the image can usually play a significant role in decision-making on resource allocation effects.

[0033] Finally, there is the image layout design. The image layout design refers to the position of each element in the image and the corresponding matching relationship. The image layout design needs to be included in the image consideration because a good layout design can significantly increase the overall expressiveness and attractiveness of the image and effectively improve the effect of resource allocation.

[0034] In some embodiments of the present disclosure, the extracting of prompt information from the acquired basic image includes: extracting basic information from the acquired basic image; identifying the image content of the basic image to obtain content information; and extracting information from the basic information and the content information to obtain the prompt information.

[0035] In some embodiments of the present disclosure, the basic information includes at least one of image size, resolution, hue, and contrast of the basic image.

[0036] Figure 2 shows a schematic diagram of images according to some embodiments of the present disclosure. Referring to Figure 2 , for example, image 1 is circular, image 2 and image 3 are rectangular, image 2 and image 3 are of equal size, and so on.

[0037] In the embodiments of the present disclosure, the aforementioned basic information refers to the basic characteristics of the aforementioned image. The basic characteristics can be obtained directly or through simple computer vision methods. Computer vision is the science of how to make machines "see". More specifically, it refers to the use of cameras and computers to replace the human eye to identify, track, and measure targets, and further perform image processing to make the computer processing into an image more suitable for human observation or transmission to instruments for detection.

[0038] In some embodiments of the present disclosure, recognizing the image content of the base image to obtain content information includes: recognizing text content of the base image to obtain text information; obtaining position information of the text information to obtain text position information; recognizing the image content of the base image to obtain image information; recognizing the image layout of the base image to obtain layout information; and obtaining the content information based on the text information, the text position information, the image information, and the layout information.

[0039] As shown in FIG2 , in an embodiment of the present disclosure, recognizing the image content of a basic image may include recognizing the text content of text, recognizing the position of text, recognizing the content of the image, and also recognizing the layout of the image.

[0040] The text content of the basic image can be recognized by neural network technology, which can be further divided into the following steps: first, it is necessary to detect where the text is. In the embodiment of the present disclosure, the text can be detected by object detection.

[0041] The task of target detection is to find all targets (objects) of interest in an image. Unlike classification and regression problems, target detection also requires determining the location of the target in the image (localization), and determining the category and location of the identified target (classification and localization).

[0042] For example, after recognizing the text in FIG2 , it can be obtained that the text is located in the upper left corner of the image.

[0043] Furthermore, after the position of the text is identified, the detected text position area can be identified to obtain the specific content of the text.

[0044] In the embodiments of the present disclosure, optical character recognition (OCR) technology can be used to identify the specific content of text. Optical character recognition refers to the process of analyzing and identifying image files of text materials to obtain text and layout information. In other words, it recognizes the text in the image and returns it in the form of text.

[0045] In the embodiments of the present disclosure, the image content of the base image can be recognized to obtain image information. This image information can be very specific. For example, it can be recognized that the image currently contains brand A perfume, and the specific number of bottles, and also contains brand A skin care lotion, and the specific number of bottles, etc.

[0046] Afterwards, the image layout of the basic image can be identified. This image layout can include: multi-image mosaic material, single image material, text display at the top, text display in the middle, etc. Alternatively, refer to Figure 2. For example, the image is a multi-image mosaic material, and there is no overlap between each image. Image 1 is circular and located to the left of the entire image and to the left of Image 2. Image 2 and Image 3 are vertically aligned, Image 3 is located below Image 2, and Image 2 and Image 3 are both located to the right of the image.

[0047] Then, the content information is obtained based on the text information, the text position information, the image information, and the layout information. In the embodiment of the present disclosure, after obtaining the text information, the text position information, the image information, and the layout information, the aforementioned information is integrated to obtain the content information of the image, that is, information that describes the content, layout, etc. of the image in detail.

[0048] In the embodiments of the present disclosure, multimodal technology can also be used to directly generate detailed description information corresponding to an image. For example, a man standing on a surfboard surfing on the sea can directly and in detail describe the main information of the image.

[0049] In some embodiments of the present disclosure, extracting the basic information and the content information to obtain the prompt information includes: obtaining the basic information, the text position information, and the layout information; extracting core information of the text information and the image information; and obtaining the prompt information based on the basic information, the text information, the layout information, and the core information.

[0050] In the embodiment of the present disclosure, after obtaining the text information, text position information, image information and layout information, the detailed information corresponding to the image can be obtained comprehensively. The image is subsequently reproduced based on the prompt information. However, after obtaining the aforementioned detailed information, the detailed information cannot be directly used as prompt information, because for the embodiment of the present disclosure, it is necessary to create a similar image, and it only needs to retain the core content of the basic image. Retaining too many details is not conducive to the subsequent reproduction of the image.

[0051] Therefore, it is necessary to extract the content of the detailed information obtained above to obtain the corresponding prompt information. The specific steps may be to retain the basic information, text location information, and layout information obtained above. For example, the basic information is that image 1 is circular, image 2 is square, and image 3 is square, and the background color of the image is warm tones, which has a contrast with the cosmetics image.

[0052] For text information and image information, only the core content needs to be retained. For example, referring to Figure 2, the text information may be "Brand A cosmetics are very cost-effective", which can be extracted as "Cosmetics are cost-effective". For image information, the image information obtained above is "The image is composed of three pictures spliced ​​together. Image 1 is a girl's hand holding Brand A perfume, Image 2 is a girl's hand holding Brand A skin care water, and Image 3 is a girl's hand holding Brand A liquid foundation."

[0053] The final retained prompt information may be "The image is composed of three images. Each image shows a person holding different cosmetics. The cosmetics can be skin care products or perfumes. A text description needs to be output in the upper left corner of the image to indicate that these cosmetics are very cost-effective. The overall image background tone needs to be warm and there needs to be contrast between it and the cosmetics picture."

[0054] After the prompt information of the image is extracted in the above steps, a similar image (image to be predicted) needs to be generated based on the prompt information formed above.

[0055] In step 104, in some embodiments of the present disclosure, obtaining the image to be predicted through the diffusion process includes: acquiring a random noise image; and gradually converting the random noise image into a coherent image based on the prompt information to obtain the image to be predicted.

[0056] In the embodiments of the present disclosure, based on the aforementioned understanding and insight into the image, high-quality prompt information has been obtained, and then an image generation algorithm can be used to generate a corresponding similar image based on the originally obtained prompt information.

[0057] Specifically, the disclosed embodiments employ a stable diffusion algorithm to generate images. The stable diffusion algorithm is a novel image generation technique designed to produce high-fidelity images in a controllable and stable manner. Based on the concept of image denoising, the algorithm gradually transforms random noise patterns into a coherent image structure through a diffusion process.

[0058] The core of the Stable Diffusion algorithm lies in its two main steps:

[0059] The first step is diffusion. After inputting the original image, the algorithm simulates the effects of natural diffusion in a high-dimensional data space, continuously adding noise to the original information and introducing randomness. The second step is reverse diffusion, which uses a deep learning model to guide the diffusion process toward the desired image input, ultimately obtaining the desired generated image. Furthermore, the denoising process not only includes information about the original image after adding noise, but also includes conditional information in each denoising step. This conditional information is usually composed of text descriptions or constraints.

[0060] In some embodiments of the present disclosure, gradually converting the random noise image into a coherent image based on the prompt information to obtain the image to be predicted includes: using the prompt information as conditional information of the random noise image, denoising the random noise image, and obtaining the image to be predicted.

[0061] In the embodiment of the present disclosure, the condition information may be prompt information.

[0062] In some embodiments of the present disclosure, the trained diffusion model is obtained by training by the following method: obtaining an image with preset parameters higher than a preset threshold from a historical image to obtain a historical basic image; extracting historical prompt information based on the historical basic image; inputting the historical prompt information into an untrained diffusion model to obtain a historical image to be predicted; training the untrained diffusion model based on the historical image to be predicted and the historical basic image to obtain the trained diffusion model.

[0063] In the embodiments of the present disclosure, since the training data of the standard stable diffusion model are often some general pictures, which do not match the scenarios of the present disclosure, the images generated by the present disclosure need to be similar to the basic images and the resource allocation effect needs to be better. Therefore, the requirements of the present disclosure for images are special, and therefore the images generated by the standard stable diffusion model cannot be used directly.

[0064] Since there are many high-quality basic images in the embodiments of the present disclosure, and the resource allocation effects of these high-quality basic images are very good, these high-quality basic images can be used to retrain the stable diffusion model to improve the output image quality. Of course, during training, the model can be trained from scratch, or the trained model can be retrained, which can effectively save the time of training the model.

[0065] During the training process, in the process of obtaining high-quality basic image prompt information, the aforementioned steps of generating prompt information can be used to generate it, and then an image is generated based on the prompt information, and then the relevant loss can be calculated.

[0066] Furthermore, after the image is generated, because resources are limited, resources need to be allocated to the best quality image. Therefore, the process does not end after the image is generated. A model can also be set up to predict the resource allocation effect of the generated image, and the image with better resource allocation effect can be used as the final target image.

[0067] In step 106, in some embodiments of the present disclosure, obtaining the descriptive features of the image to be predicted and predicting the resource allocation effect of the image to be predicted based on the descriptive features include: obtaining the descriptive features and category features of the image to be predicted; obtaining the historical resource allocation effect of the basic image; obtaining prompt information corresponding to the image to be predicted, and obtaining the descriptive features of the prompt information; and predicting the resource allocation effect of the image to be predicted based on the descriptive features and category features of the image to be predicted, the historical resource allocation effect of the basic image, and the descriptive features of the prompt information.

[0068] In some embodiments of the present disclosure, determining the target image based on the resource allocation effect of the image to be predicted includes: in response to the resource allocation effect of the image to be predicted being higher than a preset threshold, determining the image to be predicted as the target image.

[0069] In the embodiment of the present disclosure, since there are a large number of images in history, image information and resource allocation effects corresponding to the images, the model can be trained based on this historical information.

[0070] After model training is complete, the resource allocation effect can be predicted for the previously acquired image to be predicted. During the prediction process, the descriptive features and categorical features of the image to be predicted are first obtained. A descriptive feature represents an object using a low-dimensional vector, such as a word, a product, or a movie. The nature of this vector ensures that objects corresponding to vectors with similar distances have similar meanings.

[0071] After that, the description features of the prompt information can be obtained. And because the image to be predicted is a replica of the base image, the category of the image to be predicted can also be obtained. Of course, the category of the base image and the image to be predicted are the same.

[0072] Afterwards, dense features such as the size and resolution of the image to be predicted and the historical resource allocation effect of the basic image can be obtained.

[0073] The features obtained above are then fed into a neural network using a multi-task architecture to calculate features. Finally, a branch is used to estimate each target, and the underlying features of each subtask are shared. This effectively reduces the amount of computation and improves computation speed through parameter sharing.

[0074] The above-mentioned image generation method obtains prompt information by extracting it from the acquired basic image; inputs the prompt information into a trained diffusion model, and obtains the image to be predicted through the diffusion process; obtains the descriptive features of the image to be predicted, and predicts the resource allocation effect of the image to be predicted based on the descriptive features; and determines the target image based on the resource allocation effect of the image to be predicted. The embodiment of the present disclosure obtains prompt information by extracting it from the basic image, comprehensively interprets the historical high-quality images, and then extracts the interpreted information to obtain prompt information, and then generates a similar image (the image to be predicted) of the basic image based on the prompt information. Since a large number of similar images can be generated at one time based on the above-mentioned method, the present disclosure determines the final target image that meets the conditions by predicting the resource allocation effect of the similar images. The image generation method of the embodiment of the present disclosure can quickly generate a large number of similar images while effectively saving human resources, and can screen the generated similar images by pre-predicting the resource allocation effect to obtain the final similar image, thereby optimizing the actual resource allocation effect of the similar images.

[0075] It should be noted that the method of the embodiments of the present disclosure can be performed by a single device, such as a computer or server. The method of the embodiments of the present disclosure can also be applied in a distributed scenario, where multiple devices cooperate to perform the method. In such a distributed scenario, one of the multiple devices may only perform one or more steps of the method of the embodiments of the present disclosure, and the multiple devices will interact with each other to complete the method.

[0076] It should be noted that the above description is limited to some embodiments of the present disclosure. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in an order different from that described in the above embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order or sequential order shown to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0077] Based on the same inventive concept, corresponding to any of the above-mentioned embodiment methods, the present disclosure further provides an image generation device, including: an extraction module 302, used to extract prompt information from the acquired basic image; a diffusion module 304, used to input the prompt information into a trained diffusion model to obtain the image to be predicted through the diffusion process; a prediction module 306, used to obtain the descriptive features of the image to be predicted, and predict the resource allocation effect of the image to be predicted based on the descriptive features; and a determination module 308, used to determine the target image based on the resource allocation effect of the image to be predicted.

[0078] In some embodiments of the present disclosure, the extraction module 302 includes: an extraction unit for extracting basic information from the acquired basic image; an identification unit for identifying the image content of the basic image to obtain content information; and an information extraction unit for extracting information from the basic information and the content information to obtain the prompt information.

[0079] In some embodiments of the present disclosure, the basic information includes at least one of image size, resolution, hue, and contrast of the basic image.

[0080] In some embodiments of the present disclosure, the recognition unit includes: recognizing the text content of the basic image to obtain text information; obtaining the position information of the text information to obtain text position information; recognizing the image content of the basic image to obtain image information; recognizing the image layout of the basic image to obtain layout information; and obtaining the content information based on the text information, the text position information, the image information and the layout information.

[0081] In some embodiments of the present disclosure, the extraction unit includes: obtaining the basic information, the text position information and the layout information; extracting the core information of the text information and the image information; and obtaining the prompt information based on the basic information, the text information, the layout information and the core information.

[0082] In some embodiments of the present disclosure, the diffusion module 304 includes: a first acquisition unit for acquiring a random noise image; and a conversion unit for gradually converting the random noise image into a coherent image based on the prompt information to obtain the image to be predicted.

[0083] In some embodiments of the present disclosure, the conversion unit includes: using the prompt information as condition information of the random noise image, denoising the random noise image, and obtaining the image to be predicted.

[0084] In some embodiments of the present disclosure, it also includes: a historical image acquisition module, which is used to obtain an image with preset parameters higher than a preset threshold from the historical image to obtain a historical basic image; a historical prompt information extraction module, which is used to extract historical prompt information based on the historical basic image; an input module, which is used to input the historical prompt information into an untrained diffusion model to obtain a historical image to be predicted; and a training module, which is used to train the untrained diffusion model based on the historical image to be predicted and the historical basic image to obtain the trained diffusion model.

[0085] In some embodiments of the present disclosure, the prediction module 306 includes: a second acquisition unit, used to obtain the descriptive features and category features of the image to be predicted; a third acquisition unit, used to obtain the historical resource allocation effect of the basic image; a fourth acquisition unit, used to obtain the prompt information corresponding to the image to be predicted, and obtain the descriptive features of the prompt information; a prediction unit, used to predict the resource allocation effect of the image to be predicted based on the descriptive features and category features of the image to be predicted, the historical resource allocation effect of the basic image, and the descriptive features of the prompt information.

[0086] In some embodiments of the present disclosure, the determination module 308 includes: a determination unit configured to determine the image to be predicted as the target image in response to a resource allocation effect of the image to be predicted being higher than a preset threshold.

[0087] For ease of description, the above devices are described separately based on their functions and modules. Of course, when implementing the present disclosure, the functions of each module can be implemented in the same software or hardware components. The devices of the above embodiments are used to implement the corresponding image generation methods of any of the aforementioned embodiments and have the beneficial effects of the corresponding method embodiments, which will not be further described here.

[0088] Based on the same inventive concept, corresponding to any of the above-mentioned embodiments and methods, the present disclosure also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and runnable on the processor, wherein when the processor executes the program, the image generation method described in any of the above embodiments is implemented.

[0089] FIG4 shows a more specific schematic diagram of the hardware structure of an electronic device provided in this embodiment. The device may include: a processor 1010, a memory 1020, an input / output interface 1030, a communication interface 1040, and a bus 1050. The processor 1010, the memory 1020, the input / output interface 1030, and the communication interface 1040 are communicatively connected to each other within the device via the bus 1050.

[0090] The processor 1010 can be implemented using a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this specification.

[0091] The memory 1020 can be implemented in the form of ROM (Read Only Memory), RAM (Random Access Memory), static storage devices, dynamic storage devices, etc. The memory 1020 can store an operating system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 1020 and is called and executed by the processor 1010.

[0092] The input / output interface 1030 is used to connect input / output modules to implement information input and output. The input / output modules can be configured as components within the device (not shown in the figure) or can be externally connected to the device to provide corresponding functions. Input devices may include a keyboard, mouse, touch screen, microphone, various sensors, etc., and output devices may include a display, speaker, vibrator, indicator light, etc.

[0093] The communication interface 1040 is used to connect to a communication module (not shown) to enable communication between the device and other devices. The communication module can communicate via a wired method (such as USB, network cable, etc.) or a wireless method (such as mobile network, WiFi, Bluetooth, etc.).

[0094] The bus 1050 comprises a path for transmitting information between the various components of the device (eg, the processor 1010 , the memory 1020 , the input / output interface 1030 , and the communication interface 1040 ).

[0095] It should be noted that although the above device only shows the processor 1010, the memory 1020, the input / output interface 1030, the communication interface 1040, and the bus 1050, in a specific implementation, the device may also include other components necessary for normal operation. In addition, it will be understood by those skilled in the art that the above device may only include the components necessary to implement the embodiments of this specification, and does not necessarily include all the components shown in the figure.

[0096] The electronic device of the above embodiment is used to implement the corresponding image generation method in any of the above embodiments, and has the beneficial effects of the corresponding method embodiment, which will not be described in detail here.

[0097] Based on the same inventive concept, corresponding to any of the above-mentioned embodiment methods, the present disclosure also provides a non-transitory computer-readable storage medium, which stores computer instructions, and the computer instructions are used to enable the computer to execute the image generation method described in any of the above embodiments.

[0098] The computer-readable media of this embodiment include permanent and non-permanent, removable and non-removable media that can be used to store information by any method or technology. The information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, read-only compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device.

[0099] The computer instructions stored in the storage medium of the above embodiment are used to enable the computer to execute the image generation method described in any of the above embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be repeated here.

[0100] Based on the same inventive concept, corresponding to the image generation method described in any of the above embodiments, the present disclosure further provides a computer program product comprising computer program instructions. In some embodiments, the computer program instructions can be executed by one or more processors of a computer to cause the computer and / or the processor to perform the image generation method. For the execution entities corresponding to the steps in each embodiment of the image generation method, the processors executing the corresponding steps can belong to the corresponding execution entities.

[0101] The computer program product of the above embodiment is used to enable the computer and / or the processor to execute the image generation method described in any of the above embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be repeated here.

[0102] Those skilled in the art should understand that the discussion of any of the above embodiments is merely illustrative and is not intended to imply that the scope of the present disclosure (including the claims) is limited to these examples. Within the scope of the present disclosure, the technical features in the above embodiments or different embodiments may be combined, the steps may be implemented in any order, and there are many other variations of the different aspects of the embodiments of the present disclosure as described above, which are not provided in detail for the sake of simplicity.

[0103] In addition, to simplify the description and discussion, and so as not to obscure the embodiments of the present disclosure, known power / ground connections to integrated circuit (IC) chips and other components may or may not be shown in the provided figures. In addition, devices may be shown in the form of block diagrams to avoid obscuring the embodiments of the present disclosure, and this also takes into account the fact that the details of the implementation of these block diagram devices are highly dependent on the platform on which the embodiments of the present disclosure are to be implemented (i.e., these details should be fully within the purview of those skilled in the art). Where specific details (e.g., circuits) are set forth to describe exemplary embodiments of the present disclosure, it will be apparent to those skilled in the art that the embodiments of the present disclosure may be implemented without these specific details or with variations in these specific details. Therefore, these descriptions should be considered illustrative rather than restrictive.

[0104] Although the present disclosure has been described in conjunction with specific embodiments thereof, many alternatives, modifications, and variations of these embodiments will be apparent to those skilled in the art based on the foregoing description. For example, other memory architectures (e.g., dynamic RAM (DRAM)) may use the embodiments discussed.

[0105] The embodiments of the present disclosure are intended to cover all such substitutions, modifications, and variations that fall within the broad scope of the appended claims. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the embodiments of the present disclosure should be included in the scope of protection of the present disclosure.

Claims

1. A method for generating an image, comprising: Extracting prompt information from the acquired basic image; Inputting the prompt information into the trained diffusion model to obtain the image to be predicted through the diffusion process; Acquiring description features of the image to be predicted, and predicting a resource allocation effect of the image to be predicted based on the description features; A target image is determined based on the resource allocation effect of the image to be predicted.

2. The method according to claim 1, wherein The step of extracting prompt information from the acquired basic image includes: Extracting basic information from the acquired basic image; Identifying the image content of the basic image to obtain content information; Information extraction is performed on the basic information and the content information to obtain the prompt information.

3. The method according to claim 2, wherein: The basic information includes at least one of image size, resolution, hue, and contrast of the basic image.

4. The method according to claim 2, wherein: The identifying the image content of the basic image to obtain content information includes: Recognizing text content of the basic image to obtain text information; Acquire the position information of the text information to obtain the text position information; Recognizing the image content of the basic image to obtain image information; Identifying the image layout of the basic image to obtain layout information; The content information is obtained based on the text information, the text position information, the image information, and the layout information.

5. The method according to claim 4, wherein The extracting the basic information and the content information to obtain the prompt information includes: Acquire the basic information, the text position information, and the layout information; extracting core information of the text information and the image information; The prompt information is obtained based on the basic information, the text information, the layout information and the core information.

6. The method according to claim 1, wherein The step of obtaining the image to be predicted through the diffusion process includes: Get a random noise image; Based on the prompt information, the random noise image is gradually converted into a coherent image to obtain the image to be predicted.

7. The method according to claim 6, wherein: The step of gradually converting the random noise image into a coherent image based on the prompt information to obtain the image to be predicted includes: The prompt information is used as condition information of the random noise image, and the random noise image is denoised to obtain the image to be predicted.

8. The method according to claim 1, wherein The trained diffusion model is obtained by training using the following method: Acquire an image with preset parameters higher than a preset threshold from the historical images to obtain a historical basic image; Extracting historical prompt information based on the historical basic image; Inputting the historical prompt information into an untrained diffusion model to obtain a historical image to be predicted; The untrained diffusion model is trained based on the historical image to be predicted and the historical basic image to obtain the trained diffusion model.

9. The method according to claim 1, wherein The acquiring description features of the image to be predicted and predicting the resource allocation effect of the image to be predicted based on the description features includes: Obtaining description features and category features of the image to be predicted; Obtaining a historical resource allocation effect of the basic image; Obtaining prompt information corresponding to the image to be predicted, and obtaining descriptive features of the prompt information; The resource allocation effect of the image to be predicted is predicted based on the descriptive features and category features of the image to be predicted, the historical resource allocation effect of the basic image, and the descriptive features of the prompt information.

10. The method according to claim 1, wherein The determining of the target image based on the resource allocation effect of the image to be predicted includes: In response to the resource allocation effect of the to-be-predicted image being higher than a preset threshold, the to-be-predicted image is determined as the target image.

11. An image generating device, comprising: An extraction module, used for extracting prompt information from the acquired basic image; A diffusion module, configured to input the prompt information into a trained diffusion model and obtain an image to be predicted through a diffusion process; A prediction module, configured to obtain descriptive features of the image to be predicted, and predict a resource allocation effect of the image to be predicted based on the descriptive features; A determination module is used to determine a target image based on the resource allocation effect of the image to be predicted.

12. The device according to claim 11, wherein The extraction module comprises: an extraction unit, configured to extract basic information from the acquired basic image; an identification unit, configured to identify the image content of the basic image to obtain content information; The information extraction unit is used to extract the basic information and the content information to obtain the prompt information.

13. The device according to claim 12, wherein The basic information includes at least one of image size, resolution, hue, and contrast of the basic image.

14. The device according to claim 12, wherein The identification unit includes: Recognizing text content of the basic image to obtain text information; Acquire the position information of the text information to obtain the text position information; Recognizing the image content of the basic image to obtain image information; Identifying the image layout of the basic image to obtain layout information; The content information is obtained based on the text information, the text position information, the image information, and the layout information.

15. The device according to claim 14, wherein The extraction unit comprises: Acquire the basic information, the text position information, and the layout information; extracting core information of the text information and the image information; The prompt information is obtained based on the basic information, the text information, the layout information and the core information.

16. The device according to claim 11, wherein The diffusion module comprises: A first acquisition unit is used to acquire a random noise image; A conversion unit is used to gradually convert the random noise image into a coherent image based on the prompt information to obtain the image to be predicted.

17. The device according to claim 16, wherein The conversion unit comprises: The prompt information is used as condition information of the random noise image, and the random noise image is denoised to obtain the image to be predicted.

18. The apparatus according to claim 11, further comprising: A historical image acquisition module is used to acquire images with preset parameters higher than a preset threshold from historical images to obtain a historical basic image; A historical prompt information extraction module is used to extract historical prompt information based on the historical basic image; An input module, configured to input the historical prompt information into an untrained diffusion model to obtain a historical image to be predicted; A training module is used to train the untrained diffusion model based on the historical image to be predicted and the historical basic image to obtain the trained diffusion model.

19. The device according to claim 11, wherein The prediction module includes: A second acquisition unit, configured to acquire descriptive features and category features of the image to be predicted; A third acquiring unit is configured to acquire a historical resource allocation effect of the basic image; a fourth acquiring unit, configured to acquire prompt information corresponding to the image to be predicted, and acquire descriptive features of the prompt information; The prediction unit is configured to predict the resource allocation effect of the image to be predicted based on the descriptive features and category features of the image to be predicted, the historical resource allocation effect of the basic image, and the descriptive features of the prompt information.

20. The device according to claim 11, wherein The determining module includes: A determining unit is configured to determine the image to be predicted as the target image in response to a resource allocation effect of the image to be predicted being higher than a preset threshold.

21. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the method according to any one of claims 1 to 10 when executing the program.

22. A non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are configured to cause a computer to execute the method according to any one of claims 1 to 10.

23. A computer program product comprising computer program instructions, which, when executed on a computer, cause the computer to perform the method according to any one of claims 1 to 10.

Citation Information

Patent Citations

  • Image generation method and related equipment

    CN120411267A

  • Training method of image editing model and image editing method and device

    CN116363261A

  • Image generation method and device and electronic equipment

    CN116363262A

  • Commodity publicity map generation method and device, computer equipment and storage medium

    CN117058275A

  • Method and device for generating picture with accurate characters according to character description and storage medium

    CN117252957A