Picture generation method and device, electronic equipment and storage medium
By generating overall layout planning information and gradually adjusting the method, the problems of complex process and difficult quality assurance in generating advertising material images were solved, and efficient and beautiful image generation effects were achieved.
Patent Information
- Application Number
- CN202510884467.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-27
- Publication Date
- 2025-10-17
AI Technical Summary
The existing technology is complex when generating advertising material images and the image quality is difficult to guarantee.
The overall layout planning information is generated by obtaining the generation requirement description information, and the initial result image is generated using the layout model. Combined with the expansion model and the redrawing model, the target result image is gradually generated, including the expansion and adjustment of the drawing area, the target migration area and the edge area. Finally, the target content information is added and the color fusion processing is performed.
The generation process is simplified, the generation efficiency and quality of the target result image are improved, and the aesthetics and overall coordination of the image are enhanced.
Smart Images

Figure CN120807680A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the field of artificial intelligence, in particular to a picture generation method and device, electronic equipment and storage medium in the fields of computer vision, image processing, deep learning and large model. BACKGROUND
[0002] At present, many platforms have developed various different new functions, in order to enable users to better understand these new functions, these new functions can be pushed out in the form of advertisement material pictures. SUMMARY
[0003] The present disclosure provides a picture generation method, device, electronic equipment and storage medium.
[0004] A picture generation method comprises:
[0005] obtaining generation requirement description information corresponding to a target result picture to be generated, and generating overall layout planning information of the target result picture according to the generation requirement description information;
[0006] generating an initial result picture according to the overall layout planning information;
[0007] obtaining target content information to be added to the target result picture, and generating the target result picture according to the initial result picture and the target content information.
[0008] A picture generation device comprises a first processing module, a second processing module and a third processing module.
[0009] The first processing module is configured to obtain generation requirement description information corresponding to a target result picture to be generated, and generate overall layout planning information of the target result picture according to the generation requirement description information;
[0010] The second processing module is configured to generate an initial result picture according to the overall layout planning information;
[0011] The third processing module is configured to obtain target content information to be added to the target result picture, and generate the target result picture according to the initial result picture and the target content information.
[0012] An electronic equipment comprises:
[0013] at least one processor; and
[0014] a memory connected with the at least one processor in communication; wherein
[0015] The memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the method as described above.
[0016] A non-transitory computer readable storage medium storing computer instructions for causing a computer to perform the method as described above.
[0017] A computer program product comprising computer programs / instructions which, when executed by a processor, implement the method as described above.
[0018] It should be understood that the contents described in this part are not intended to identify key or important features of the embodiments of the present disclosure, nor to limit the scope of the present disclosure. Other features of the present disclosure will become apparent from the following description. BRIEF DESCRIPTION OF DRAWINGS
[0019] The accompanying drawings are used to better understand the present scheme, and do not constitute a limitation on the present disclosure. Among them:
[0020] Figure 1 a flow chart of the picture generation method first embodiment described in the present disclosure;
[0021] Figure 2 a flow chart of the picture generation method second embodiment described in the present disclosure;
[0022] Figure 3 a schematic diagram of the target result picture described in the present disclosure;
[0023] Figure 4 a schematic diagram of the composition structure of the picture generation apparatus embodiment 400 described in the present disclosure;
[0024] Figure 5 A schematic block diagram of an electronic device 500 that can be used to implement embodiments of the present disclosure is shown. DETAILED DESCRIPTION
[0025] Exemplary embodiments of the present disclosure are described below with reference to the accompanying drawings, which include various details of the embodiments of the present disclosure to help in understanding, and should be considered as merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Also, in order to be clear and concise, descriptions of well-known functions and structures are omitted in the following description.
[0026] In addition, it should be understood that the term "and / or" herein is merely an associated relationship between the associated objects described, which means that there can be three relationships, for example, A and / or B can represent: A exists alone, A and B exist together, and B exists alone. In addition, the character " / " herein generally represents an "or" relationship between the associated objects before and after.
[0027] Figure 1 A flowchart of a first embodiment of the picture generation method of the present disclosure is shown in FIG. 1. As shown in FIG. 1, the embodiment includes the following specific implementation. Figure 1
[0028] In step 101, the generation requirement description information corresponding to the target result picture to be generated is obtained, and the overall layout planning information of the target result picture is generated according to the generation requirement description information.
[0029] In step 102, the initial result picture is generated according to the overall layout planning information.
[0030] In step 103, the target content information to be added to the target result picture is obtained, and the target result picture is generated according to the initial result picture and the target content information.
[0031] The target result picture can be an advertisement material picture. In the traditional way, the implementation of generating an advertisement material picture is usually complex, and the quality of the picture is difficult to guarantee.
[0032] However, by using the scheme of the above method embodiment, the overall layout planning information can be generated according to the generation requirement description information first, then the initial result picture can be generated according to the overall layout planning information, and finally the target result picture required can be generated according to the initial result picture and the target content information. The whole process is simple and convenient to realize, thereby improving the generation efficiency of the target result picture, and the quality and aesthetic degree of the generated target result picture are improved by means of the overall layout planning information.
[0033] The generation requirement description information can be input by a user, which is used to describe the user's requirements for the content presented by the target result picture, and the specific content can be determined according to actual needs, such as including: picture size, picture main body is a tiger holding a drawing board (scene picture), drawing board is used to display target content information, and the target content information includes about how many characters and / or how many pictures, etc.
[0034] After obtaining the generation requirement description information, the overall layout planning information of the target result picture can be generated according to the generation requirement description information. In some embodiments of the present disclosure, the overall layout planning information can be generated by using a layout model according to the generation requirement description information.
[0035] The layout model can be obtained by fine-tuning a predetermined large language model. The large language model refers to a deep learning model trained using a large amount of text data, which can generate natural language text or understand the meaning of natural language text, etc. The predetermined large language model can be an ERNIE Faster (Enhanced Representation through kNowledge IntEgration) model, etc.
[0036] After inputting the generated requirement description information into the layout model, overall layout planning information output by the layout model can be obtained. Accordingly, by means of the strong reasoning capability of the layout model, the accuracy of the obtained overall layout planning information can be improved, etc.
[0037] In some embodiments of the present disclosure, the overall layout planning information can include size information of the target result picture and region planning information corresponding to each region in the target result picture, and each region planning information includes position information of the corresponding region in the target result picture.
[0038] The size of the target result picture can be 500*500. The position information of any region can include the shape, size, and position coordinates of the region, etc. If necessary, some other information can be further included in the region planning information, and the specific information included can be determined according to actual needs.
[0039] By means of the overall layout planning information, the success rate of subsequent expansion generation can be improved, and the diversity of picture layout can be enriched, avoiding single and rigid, etc.
[0040] In some embodiments of the present disclosure, the regions can include a drawing region, a target migration region, and an edge region. The drawing region is a display region of a scene drawing in the target result picture, the target migration region is a display region of target content information, and the edge region is a transition region outside the drawing region and the target migration region, which can be used as an extension of the drawing region in the subsequent expansion process. This region design method is simple and convenient to implement, and can meet various information display needs of users.
[0041] In addition, in some embodiments of the present disclosure, the layout model can also generate picture content description information, and accordingly, the manner of generating the initial result picture according to the overall layout planning information can include: generating a first intermediate picture according to the overall layout planning information and the picture content description information, the drawing region in the first intermediate picture being drawn with a scene drawing, the scene drawing being drawn according to the picture content description information, and then performing expansion operations on the target migration region and the edge region in the first intermediate picture based on the drawing region to obtain a second intermediate picture, and further generating the initial result picture according to the second intermediate picture.
[0042] That is, the layout model can output not only the overall layout planning information but also the picture content description information of the drawing area, which is used to draw the scene picture according to the picture content description information. For example, the scene picture is a tiger holding a drawing board, and the picture content description information can include the specific image of the tiger, such as a pair of black headphones on the head of the tiger, the tiger is yellow, and there are black stripes on the head, the tiger's hands hold a drawing board, the size of the drawing board, and the like. In short, the picture content description information needs to include various picture detail information required to draw the scene picture.
[0043] Then, the drawing of the drawing area can be performed to obtain the first intermediate picture with the scene picture drawn, and then the first intermediate picture can be expanded to obtain the second intermediate picture, and further, the initial result picture can be generated according to the second intermediate picture. By sequentially performing different picture generation steps, the error probability is reduced, and the accuracy of the generated initial result picture is improved.
[0044] In some embodiments of the present disclosure, the drawing model can be used to generate the first intermediate picture, that is, the overall layout planning information and the picture content description information can be input into the drawing model to obtain the first intermediate picture output by the drawing model.
[0045] The size of the first intermediate picture is the size of the target result picture, which can include the drawing area, the target migration area, and the edge area, and the scene picture is drawn in the drawing area.
[0046] In addition, the drawing model can be a Flux (Fused Large-scale Unified Transformation eXtensions) model. The Flux model is an open source artificial intelligence image generation model, which has many advantages such as high detail restoration and multi-style support. Accordingly, with the help of the Flux model, the required first intermediate picture can be efficiently and accurately generated.
[0047] Based on the drawing area, the target migration area and the edge area in the first intermediate picture can also be expanded to obtain the second intermediate picture, that is, the picture content in the target migration area and the edge area can be generated based on the scene picture in the drawing area.
[0048] In some embodiments of the present disclosure, the expansion operation can be completed by using an expansion model. For example, the expansion model can be a Flux Fill model used for redrawing. In addition, in some embodiments of the present disclosure, in response to determining that the expansion operation is an expansion operation performed on the target migration area, the expansion operation can be completed by using the expansion model in combination with the obtained constraint guide information, which is used to constrain the picture content generated by the expansion.
[0049] In the expansion process, the edge area and the target migration area can be processed separately. For the edge area, the expansion operation can be completed by directly using the expansion model without the aid of constraint guide information, so that the picture content in the drawing area is naturally expanded out. For the target migration area, the expansion operation can be completed by using the expansion model in combination with the constraint guide information, which is set by a specific prompt. This is because the expansion operation can cause the generation of extra content that does not meet the requirements, which can affect the display of the target content information subsequently added to the target migration area. Accordingly, with the aid of the constraint guide information, the picture content generated in the target migration area can be constrained to ensure that as little extra content as possible that does not meet the requirements is generated, so that the background content remains simple and avoids conflicts with the target content information subsequently added to the target migration area, and reduces the visual sense of strangeness.
[0050] After obtaining the second intermediate picture, an initial result picture can be generated according to the second intermediate picture. In some embodiments of the present disclosure, the second intermediate picture can be redrawn to obtain the initial result picture.
[0051] After obtaining the second intermediate picture, the second intermediate picture can be redrawn to obtain the initial result picture.
[0052] In some embodiments of the present disclosure, the second intermediate picture can be redrawn by using a redrawing model, and a redrawing coefficient corresponding to the redrawing model is greater than or equal to a first threshold.
[0053] The redrawing model can be a Flux model, and the specific value of the first threshold can be determined according to actual needs, such as 75%. Assuming that the redrawing coefficient is set to 75%, the redrawing coefficient can be used to constrain the change range of the picture after redrawing compared with the picture before redrawing to be less than 25%, so that the overall picture structure of the initial result picture after redrawing does not change greatly compared with the second intermediate picture. In addition, in actual application, the overall layout planning information and the picture content description information can also be used as reference information for the redrawing process to further improve the content consistency of the initial result picture and the second intermediate picture.
[0054] After obtaining the initial result picture, target content information that needs to be added to the target result picture can be obtained, that is, target migration can be performed. In some embodiments of the present disclosure, the text input in the target migration area of the initial result picture can be obtained, and / or the picture pasted in the target migration area of the initial result picture can be obtained.
[0055] For example, the initial result picture can be displayed to the user, and accordingly, the user can input text and / or paste a picture in the target migration area of the initial result picture in a certain way, that is, the user can only input text, can only paste a picture, or can both input text and paste a picture, depending on actual needs. In addition, the user can set the font and size of the input text, and the user can adjust the size of the pasted picture, which is very flexible and convenient.
[0056] Assuming that the target result picture is an advertising material picture, the text input by the user can be an advertising promotion script, and the pasted picture can be an advertising promotion picture.
[0057] Further, according to the initial result picture and the target content information, the target result picture can be generated. In some embodiments of the present disclosure, the initial result picture after adding the target content information can be directly used as the target result picture, or the initial result picture after adding the target content information can be subjected to tone fusion processing to obtain the target result picture.
[0058] Preferably, the latter processing mode can be used. The input text and the pasted picture can not be consistent with the overall tone of the initial result picture, and therefore, after adding the target content information to the initial result picture, the initial result picture can be subjected to tone fusion processing to make the added target content information better fused with the surrounding tone, thereby further improving the overall picture coordination and aesthetic degree of the target result picture.
[0059] In combination with the above description, Figure 2 a flowchart of a second embodiment of the picture generation method according to the present disclosure is shown. As shown in Figure 2 the specific implementation modes include the following.
[0060] In step 201, obtain the generation requirement description information corresponding to the target result picture to be generated.
[0061] In step 202, according to the generation requirement description information, generate the following overall layout planning information using the layout model: size information of the target result picture and region planning information corresponding to each region in the target result picture, each region planning information including: position information of the corresponding region in the target result picture, the region including a drawing region, a target migration region and an edge region, and generating picture content description information using the layout model.
[0062] In step 203, generate a first intermediate picture according to the overall layout planning information and the picture content description information, the drawing region in the first intermediate picture being drawn with a scene picture, the scene picture being drawn according to the picture content description information.
[0063] For example, the first intermediate picture can be generated using a drawing model.
[0064] In step 204, based on the drawing region, expand the target migration region and the edge region in the first intermediate picture respectively to obtain a second intermediate picture.
[0065] For example, the expansion operation can be completed using an expansion model, and when the target migration region is expanded, the expansion model can be used to complete the expansion operation in combination with the obtained constraint guide information, which is used to constrain the picture content generated by the expansion.
[0066] In step 205, redraw the second intermediate picture to obtain an initial result picture.
[0067] For example, the second intermediate picture can be redrawn using a redrawing model, and the redrawing coefficient corresponding to the redrawing model is greater than or equal to a first threshold.
[0068] In step 206, obtain the target content information added in the target migration region of the initial result picture.
[0069] For example, the text input in the target migration region of the initial result picture can be obtained, and / or the picture pasted in the target migration region of the initial result picture can be obtained.
[0070] In step 207, perform tone fusion processing on the initial result picture with the added target content information to obtain a target result picture.
[0071] Figure 3 A schematic view of the target result picture of the present disclosure. As shown in the figure, Figure 3As shown, the picture content of the picture of the tiger holding the drawing board with hands is the picture content in the drawing area. Based on the drawing area, the picture content in the target migration area (the area in the drawing board) can be generated by the expansion operation. In order to prevent the display effect of the target content information from being affected, the target migration area can only include a single background color (such as light yellow), and does not include other text or patterns, etc. In addition, the picture content in the edge area, that is, the picture content outside the drawing area and the target migration area, can be generated by the expansion operation, including the picture content of the background area behind the tiger, the water cup and the desktop in front, etc. In addition, "AI second-speed writing quality is excellent, and farewell to the anxiety of manuscript" is the target content information added to the target migration area. The target content information can be written by the user, and the entire picture can be subjected to tone fusion processing after the target content information is written, so as to obtain the final target result picture required.
[0072] It should be noted that, for the foregoing method embodiments, in order to simply describe, they are all expressed as a series of action combinations, but those skilled in the art should know that the disclosure is not limited by the action sequence described, because according to the disclosure, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should know that the embodiments described in the specification all belong to preferred embodiments, and the actions and modules involved are not necessarily necessary for the disclosure. In addition, the parts not described in detail in a certain embodiment can refer to the related description in other embodiments.
[0073] The above is the introduction of the method embodiment, and the scheme described in the disclosure will be further described through the device embodiment.
[0074] Figure 4 The constituent structure schematic diagram of the picture generation device embodiment 400 of the disclosure is shown in FIG. 4. Figure 4 As shown, it can include a first processing module 401, a second processing module 402 and a third processing module 403.
[0075] The first processing module 401 is configured to obtain the generation requirement description information corresponding to the target result picture to be generated, and generate the overall layout planning information of the target result picture according to the generation requirement description information.
[0076] The second processing module 402 is configured to generate the initial result picture according to the overall layout planning information.
[0077] The third processing module 403 is configured to obtain the target content information needed to be added to the target result picture, and generate the target result picture according to the initial result picture and the target content information.
[0078] In some embodiments of the present disclosure, the first processing module 401 can generate the overall layout planning information by using a layout model according to the generated demand description information. The layout model can be obtained by fine-tuning a predetermined large language model.
[0079] In some embodiments of the present disclosure, the overall layout planning information can include size information of the target result picture and region planning information corresponding to each region in the target result picture, and each region planning information can include position information of the corresponding region in the target result picture.
[0080] In some embodiments of the present disclosure, the regions can include a drawing region, a target migration region, and an edge region. The drawing region is a display region of a scene drawing in the target result picture. The target migration region is a display region of the target content information. The edge region is a transition region outside the drawing region and the target migration region, and can be used as an extension of the drawing region in a subsequent drawing expansion process.
[0081] In addition, in some embodiments of the present disclosure, the first processing module 401 can also obtain drawing content description information generated by the layout model. Accordingly, the manner in which the second processing module 402 generates the initial result picture according to the overall layout planning information can include: generating a first intermediate picture according to the overall layout planning information and the drawing content description information. The drawing region in the first intermediate picture has a scene drawing drawn therein, and the scene drawing is drawn according to the drawing content description information. Then, the target migration region and the edge region in the first intermediate picture can be expanded based on the drawing region to obtain a second intermediate picture, and the initial result picture can be generated according to the second intermediate picture.
[0082] In some embodiments of the present disclosure, the second processing module 402 can generate the first intermediate picture by using a drawing model. That is, the overall layout planning information and the drawing content description information can be input into the drawing model to obtain the first intermediate picture output by the drawing model.
[0083] The size of the first intermediate picture is the size of the target result picture, which can include the drawing region, the target migration region, and the edge region, and the scene drawing is drawn in the drawing region.
[0084] In some embodiments of the present disclosure, the second processing module 402 can complete the expansion operation by using an expansion model. In addition, in some embodiments of the present disclosure, in response to determining that the expansion operation is an expansion operation performed on the target migration region, the second processing module 402 can complete the expansion operation by using the expansion model in combination with the obtained constraint guide information, and the constraint guide information is used to constrain the content of the picture generated by the expansion.
[0085] After the second intermediate picture is obtained, an initial result picture can be generated according to the second intermediate picture. In some embodiments of the present disclosure, the second processing module 402 can perform picture redrawing on the second intermediate picture, so as to obtain the initial result picture.
[0086] In some embodiments of the present disclosure, the second processing module 402 can perform picture redrawing on the second intermediate picture by using a redrawing model, and a redrawing coefficient corresponding to the redrawing model is greater than or equal to a first threshold.
[0087] After the initial result picture is obtained, target content information that needs to be added to the target result picture can be obtained, that is, target migration can be performed. In some embodiments of the present disclosure, the third processing module 403 can obtain the text input in the target migration area of the initial result picture, and / or obtain the picture pasted in the target migration area of the initial result picture.
[0088] Further, according to the initial result picture and the target content information, the target result picture can be generated. In some embodiments of the present disclosure, the third processing module 403 can directly take the initial result picture after adding the target content information as the target result picture, or can perform tone fusion processing on the initial result picture after adding the target content information, so as to obtain the target result picture.
[0089] Figure 4 The specific working process of the apparatus embodiment shown can refer to the related description in the foregoing method embodiments, and will not be described herein again.
[0090] The scheme described in the present disclosure can be applied to the field of artificial intelligence, and particularly relates to the fields of computer vision, image processing, deep learning, and large models. Artificial intelligence is a discipline that studies enabling computers to simulate some thinking processes and intelligent behaviors (such as learning, reasoning, thinking, planning, etc.) of humans, and includes both hardware-level technologies and software-level technologies. Artificial intelligence hardware technologies generally include technologies such as sensors, special artificial intelligence chips, cloud computing, distributed storage, big data processing, etc. Artificial intelligence software technologies mainly include computer vision technology, speech recognition technology, natural language processing technology, and machine learning / deep learning, big data processing technology, knowledge graph technology, etc.
[0091] The pictures, texts, etc. in the embodiments of the present disclosure are not directed to a specific user, and cannot reflect the personal information of a specific user. In the technical scheme of the present disclosure, the collection, storage, use, processing, transmission, provision, and disclosure of user personal information are in line with the provisions of relevant laws and regulations, and do not violate public order and good customs.
[0092] According to the embodiments of the present disclosure, the present disclosure further provides an electronic device, a readable storage medium, and a computer program product.
[0093] Figure 5 A schematic block diagram of an electronic device 500 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptops, desktops, tablets, servers, blades, mainframes, and other appropriate computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular telephones, smartphones, wearable devices, and other similar computing devices. The components shown here, their connections and relationships, and their functions, are meant to be examples only, and are not intended to limit the implementations of the present disclosure described and / or claimed in this document.
[0094] As shown in Figure 5 The electronic device 500 includes a computing unit 501 that can perform various appropriate actions and processes in accordance with a computer program stored in a read-only memory (ROM) 502 or a computer program loaded into a random access memory (RAM) 503 from a storage unit 508. Various programs and data required for the operation of the electronic device 500 can also be stored in the RAM 503. The computing unit 501, the ROM 502, and the RAM 503 are connected to each other through a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.
[0095] Various components in the electronic device 500 are connected to the I / O interface 505, including an input unit 506, such as a keyboard, a mouse, and the like; an output unit 507, such as various types of displays, speakers, and the like; a storage unit 508, such as a magnetic disk, a magneto-optical disk, and the like; and a communication unit 509, such as a network card, a modem, a wireless communication transceiver, and the like. The communication unit 509 allows the electronic device 500 to exchange information / data with other devices through a computer network, such as the Internet, and / or various telecommunication networks.
[0096] The computing unit 501 can be various general and / or special purpose processing components with processing and computing capabilities. Some examples of the computing unit 501 include, but are not limited to, a central processing unit (CPU), a graphic processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, and the like. The computing unit 501 performs various methods and processes described above, such as the methods described in the present disclosure. For example, in some embodiments, the methods described in the present disclosure can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 508. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 500 via the ROM 502 and / or the communication unit 509. When the computer program is loaded onto the RAM 503 and executed by the computing unit 501, one or more steps of the methods described in the present disclosure can be performed. Alternatively, in other embodiments, the computing unit 501 can be configured to perform the methods described in the present disclosure by any other suitable means, such as by means of firmware.
[0097] Various implementations of the systems and techniques described above can be realized in digital electronic circuitry, integrated circuitry, a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), a system on a chip (SOC), a complex programmable logic device (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.
[0098] Program code for carrying out methods of the present disclosure can be written in any combination of one or more programming languages. The program code can be provided to a processor or controller of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the program code, when executed by the processor or controller, produces the functions / operations specified in the flowcharts and / or block diagrams. The program code can be executed entirely on a machine, partially on a machine, partially on a machine as a stand-alone software package, partially on a machine and partially on a remote machine or entirely on a remote machine or server.
[0099] In the context of the present disclosure, a machine-readable medium can be a tangible medium that contains or stores a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include but is not limited to an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium would include one or more lines of electrical connections, portable computer disks, hard disk drives, random access memory, read-only memory, erasable programmable read-only memory (EPROM), flash memory, optical fibers, portable compact disc read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0100] To provide for interaction with a user, the systems and techniques described here can be implemented on a computer having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, including acoustic, speech, or tactile input.
[0101] The systems and techniques described here can be implemented in a computing system that includes a back end component (e.g., as a data server), or that includes a middleware component (e.g., an application server), or that includes a front end component (e.g., a user computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the systems and techniques described here), or any combination of such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.
[0102] The computer system can include clients and servers. This relationship can be. The servers are typically remote from the clients with the interactions between them occurring over a communication network. The relationship between a client and a server is one of client-server. The server can be a cloud server, a server of a distributed system, or a server incorporating a blockchain.
[0103] It should be understood that the steps shown in the various forms above can be reordered, added to, or deleted from. For example, the steps recited in the present disclosure can be performed in parallel, in series, or in a different order, as long as the desired results of the technical solutions of the present disclosure are achieved, and the present disclosure is not limited herein.
[0104] The specific implementation described above does not constitute a limitation on the protection scope of the present disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent replacements, and improvements within the spirit and principles of the present disclosure should be included in the protection scope of the present disclosure.
Claims
1. A method for generating an image, comprising: Obtaining generation requirement description information corresponding to the target result image to be generated, and generating overall layout planning information of the target result image according to the generation requirement description information; generating an initial result picture according to the overall layout planning information; Target content information that needs to be added to the target result picture is acquired, and the target result picture is generated according to the initial result picture and the target content information.
2. The method according to claim 1, wherein The overall layout planning information of generating the target result image according to the generation requirement description information includes: The overall layout planning information is generated using a layout model according to the generation requirement description information.
3. The method according to claim 2, wherein: The overall layout planning information includes: size information of the target result image and area planning information corresponding to each area in the target result image; The planning information of each area includes: the position information of the corresponding area in the target result image.
4. The method according to claim 3, wherein: The area includes: a drawing area, a target migration area and an edge area; the drawing area is the display area of the scene illustration in the target result image, the target migration area is the display area of the target content information, and the edge area is the transition area outside the drawing area and the target migration area.
5. The method according to claim 4, further comprising: Obtaining description information of the image content generated by the layout model; The generating of the initial result picture according to the overall layout planning information includes: generating a first intermediate image according to the overall layout planning information and the image content description information, wherein the scene image is drawn in the drawing area of the first intermediate image, and the scene image is drawn according to the image content description information; Based on the drawing area, performing an expansion operation on the target migration area and the edge area in the first intermediate image respectively to obtain a second intermediate image; The initial result image is generated according to the second intermediate image.
6. The method according to claim 5, wherein: Generating the first intermediate image includes: generating the first intermediate image using a drawing model; The expanding operation includes: using an expanding model to complete the expanding operation.
7. The method according to claim 6, wherein: The use of the map expansion model to complete the map expansion operation includes: In response to determining that the expansion operation is an expansion operation performed on the target migration area, the expansion operation is completed using the expansion model in combination with the acquired constraint guidance information, and the constraint guidance information is used to constrain the image content generated by the expansion.
8. The method according to claim 5, wherein Generating the initial result picture according to the second intermediate picture includes: Redraw the second intermediate image to obtain the initial result image.
9. The method according to claim 8, wherein Redrawing the second intermediate picture includes: The second intermediate image is redrawn using a redrawing model, and a redrawing coefficient corresponding to the redrawing model is greater than or equal to a first threshold.
10. The method according to claim 8, wherein The acquiring of target content information to be added to the target result image includes: Acquire text input in the target migration area of the initial result image, and / or acquire an image pasted in the target migration area of the initial result image.
11. The method according to claim 10, wherein: Generating the target result picture according to the initial result picture and the target content information includes: Performing tone fusion processing on the initial result image after adding the target content information to obtain the target result image.
12. A picture generating device, comprising: a first processing module, a second processing module, and a third processing module; The first processing module is used to obtain generation requirement description information corresponding to the target result image to be generated, and generate overall layout planning information of the target result image according to the generation requirement description information; The second processing module is used to generate an initial result image according to the overall layout planning information; The third processing module is configured to obtain target content information that needs to be added to the target result image, and generate the target result image according to the initial result image and the target content information.
13. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 11.
14. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to enable a computer to execute the method according to any one of claims 1 to 11.
15. A computer program product comprising a computer program / instructions, which implement the method according to any one of claims 1 to 11 when executed by a processor.