Image generation method and device, electronic equipment and storage medium
By obtaining text descriptions and selecting target models from different styles of literary and artistic graph models to generate images, the problem of single image style in the prior art is solved, and the image generation with diverse styles is realized, and the user experience is improved.
Patent Information
- Application Number
- CN202311484869.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-08
- Publication Date
- 2025-05-09
Smart Images

Figure CN119963668A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present disclosure relate to the field of computer technology, and in particular, to an image generation method, device, electronic device, and storage medium. Background Art
[0002] Nowadays, there is a widespread demand for image design. Currently, the image generation methods include obtaining a text-based image model based on deep learning, and generating images based on text through the text-based image model. The images generated in this way are often of a single style and cannot meet diverse needs. Summary of the invention
[0003] The embodiments of the present disclosure provide an image generation method, an apparatus, an electronic device, and a storage medium, which can realize image generation with diverse styles.
[0004] In a first aspect, an embodiment of the present disclosure provides an image generation method, comprising:
[0005] Get text description;
[0006] Determining a style classification label for the text description;
[0007] Determining a target model corresponding to the style classification label from each of the text-generated image models; wherein the image generation styles of each of the text-generated image models are different;
[0008] A target image corresponding to the text description is generated by the target model.
[0009] In a second aspect, the present disclosure also provides an image generating device, including:
[0010] A text acquisition module is used to obtain text descriptions;
[0011] A label determination module, used to determine a style classification label of the text description;
[0012] A model selection module, used to determine a target model corresponding to the style classification label from each text-generated image model; wherein each text-generated image model has a different image generation style;
[0013] An image generation module is used to generate a target image corresponding to the text description through the target model.
[0014] In a third aspect, an embodiment of the present disclosure further provides an electronic device, the electronic device comprising:
[0015] one or more processors;
[0016] a storage device for storing one or more programs,
[0017] When the one or more programs are executed by the one or more processors, the one or more processors implement the image generating method as described in any one of the embodiments of the present disclosure.
[0018] In a fourth aspect, the embodiments of the present disclosure further provide a storage medium comprising computer executable instructions, which, when executed by a computer processor, are used to execute the image generation method as described in any one of the embodiments of the present disclosure.
[0019] The technical solution of the embodiment of the present disclosure obtains a text description; determines a style classification label of the text description; determines a target model corresponding to the style classification label from each text-generated graph model; wherein each text-generated graph model has a different image generation style; and generates a target image corresponding to the text description through the target model. By classifying the style of the text description and selecting a target model corresponding to the classification from at least two text-generated graph models to generate an image, image generation with diverse styles can be achieved, thereby improving user experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] The above and other features, advantages and aspects of the embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. Throughout the accompanying drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and the originals and elements are not necessarily drawn to scale.
[0021] Figure 1 A flowchart of an image generation method provided by an embodiment of the present disclosure;
[0022] Figure 2 A flowchart of an image generation method provided by an embodiment of the present disclosure;
[0023] Figure 3 A flowchart of a text generation model construction in an image generation method provided in an embodiment of the present disclosure;
[0024] Figure 4 A flowchart of an image generation method provided by an embodiment of the present disclosure;
[0025] Figure 5 A schematic diagram of the structure of an image generating device provided by an embodiment of the present disclosure;
[0026] Figure 6 A schematic diagram of the structure of an electronic device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION
[0027] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as being limited to the embodiments described herein, which are instead provided for a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not intended to limit the scope of protection of the present disclosure.
[0028] It should be understood that the various steps described in the method embodiments of the present disclosure may be performed in different orders and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this respect.
[0029] The term "including" and its variations used herein are open inclusions, i.e., "including but not limited to". The term "based on" means "based at least in part on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". The relevant definitions of other terms will be given in the following description.
[0030] It should be noted that the concepts such as "first" and "second" mentioned in the present disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.
[0031] It should be noted that the modifications of "one" and "plurality" mentioned in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise clearly indicated in the context, it should be understood as "one or more".
[0032] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only used for illustrative purposes and are not used to limit the scope of these messages or information.
[0033] It is understandable that before using the technical solutions disclosed in the embodiments of the present disclosure, the types, scope of use, usage scenarios, etc. of the personal information involved in the present disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.
[0034] It is understandable that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and relevant provisions.
[0035] Figure 1The present invention is a flowchart of an image generation method provided by an embodiment of the present invention. The present invention is applicable to the case of generating an image, for example, to the case of generating an advertising image. The method can be executed by an image generation device, which can be implemented in the form of software and / or hardware, and can be configured in an electronic device, for example, in a computer.
[0036] like Figure 1 As shown, the image generation method provided in this embodiment may include:
[0037] S110: Obtain text description.
[0038] In the embodiment of the present disclosure, the text description may include characters such as words and symbols, which may be used to describe the image content of the target image to be generated. The text description may be obtained based on at least one of the following methods:
[0039] The image generation device may provide a user interface, and a text input control may be set in the user interface, and then the text description input by the user may be received through the text input control; a text recognition control may also be set in the user interface, and then the text description in the input image may be recognized through the text recognition control; a text reading control may also be set in the user interface, and then the corresponding text description under the target address may be read through the text reading control, etc. In addition, other methods for obtaining text descriptions may also be applied here, and they are not exhaustive here.
[0040] S120: Determine a style classification label for the text description.
[0041] In the embodiment of the present disclosure, at least two style classification labels can be pre-set, where the style classification labels can include, but are not limited to, realistic style, three-dimensional computer image style, cartoon style, Lego style, oil painting style, fine brushwork style, etc.
[0042] The text description can be divided into style classification labels based on existing text classification algorithms. Existing text classification algorithms may include but are not limited to classification algorithms based on classification trees (such as decision trees, random forests, etc.) and deep neural network methods (such as convolutional neural networks, recurrent neural networks), etc. In addition, other text classification methods can also be applied here, which are not exhaustive here.
[0043] S130, determining a target model corresponding to a style classification label from each text-generated image model; wherein each text-generated image model has a different image generation style.
[0044] In the disclosed embodiment, at least two existing models for generating images based on text can be used as text graph models. The image generation styles of different text graph models are different, that is, the generated images have different styles. The correspondence between each style classification label and at least two text graph models can be pre-set, and the correspondence can include a one-to-one correspondence in which one style classification label uniquely corresponds to one text graph model, and / or a many-to-one correspondence in which multiple style classification labels correspond to one text graph model, etc.
[0045] In some optional implementations, the image generation method may further include: when the style classification label fails to be determined, using a preset universal model in each text-generated image model as a target model.
[0046] In the process of classifying the text description into style classification labels, the probability value of the text description belonging to each style classification label can be determined first, and the style classification label with the largest probability value and greater than a preset value can be determined as the style classification label of the current text description. However, there is a situation where the probability value of the text description belonging to each style classification label is close to or the maximum probability value is less than the preset value. In the above case, it can be considered that the text description has no obvious style classification label corresponding to it, and it can be considered that the style classification label determination has failed. In the case of failure to determine the style classification label of the text description, the preset universal model in each text-generated image model can be used as the target model for image generation to ensure that the image can be generated normally.
[0047] In these optional implementations, by presetting a general model, when there is no obvious style classification label corresponding to the text description, the general model can be used to perform a fallback image generation to ensure normal image generation.
[0048] S140: Generate a target image corresponding to the text description through a target model.
[0049] After the target model is determined, the text description may be used as an input of the target model to generate a target image corresponding to the text description through the target model.
[0050] The technical solution of the embodiment of the present disclosure obtains a text description; determines a style classification label of the text description; determines a target model corresponding to the style classification label from each text-generated graph model; wherein each text-generated graph model has a different image generation style; and generates a target image corresponding to the text description through the target model. By classifying the style of the text description and selecting a target model corresponding to the classification from at least two text-generated graph models to generate an image, image generation with diverse styles can be achieved, thereby improving user experience.
[0051] The various optional schemes in the image generation method provided in the embodiment of the present disclosure and the above embodiment can be combined. The image generation method provided in this embodiment describes the description text acquisition process in detail. By combining the theme of text generation and the object features of the image receiving object to generate the text description, the target image generated based on the text description can be made to fit the theme, and the satisfaction of the image receiving object can be improved to a certain extent, thereby further meeting the user's needs.
[0052] Figure 2 The following is a flow chart of an image generation method provided by an embodiment of the present disclosure. Figure 2 As shown, the image generation method provided in this embodiment may include:
[0053] S210, obtaining a subject generated by the text and object features of an image receiving object.
[0054] In this embodiment, the theme generated by the text can also be considered as the theme of the subsequent image generation. The theme acquisition method may include: receiving the theme input by the user through the user interface provided by the image generation device.
[0055] The image receiving object may be considered as the image audience of the target image to be generated. The object characteristics of the image receiving object may be considered as group characteristics used to describe the image audience. The acquisition method of the object characteristics of the image receiving object may include at least one of the following:
[0056] An example of a feature text description is displayed through a user interface provided by an image generating device, and a feature text description input by a user according to the example is received, keywords are extracted based on the received feature text description, and are used as object features; multi-dimensional feature labels are provided through the user interface, and feature labels of each dimension selected by the user are received, and the received feature labels are used as object features; multiple groups of object feature sets are provided through the user interface, each group of object feature sets may include at least one object feature, a target object feature set selected by the user is received, and the object features therein are used as object features of the image receiving object.
[0057] In addition, other methods of obtaining subject and object features can also be applied here, which are not listed here exhaustively.
[0058] S220: Generate text description based on topic and object features through a text generation model.
[0059] In this embodiment, the text generation model can be an existing model for generating text based on text. The subject and object features can be input into the text generation model to generate text descriptions. Thus, text descriptions for different subjects and different object features can be obtained.
[0060] S230: Determine a style classification label for the text description.
[0061] S240, determining a target model corresponding to a style classification label from each text-generated image model; wherein each text-generated image model has a different image generation style.
[0062] S250: Generate a target image corresponding to the text description through a target model.
[0063] For example, Figure 3 This is a flowchart of the construction of a text generation model in an image generation method provided by an embodiment of the present disclosure. Figure 3 In some optional implementations, the process of constructing a text generation model may include:
[0064] First, sample object features and sample topics are obtained, wherein the sample object features and sample topics may include but are not limited to object features and topics defined by each user, as well as randomly defined pseudo object features and pseudo topics.
[0065] Next, a candidate text description is generated based on the sample object features and the sample topic through the pre-trained language model. The sample object features and the sample topic can be input into the text generation model to generate the candidate text description. The sample object features, the sample topic and the text description can be used as data pairs. Based on this method, a large number of data pairs can be generated as candidate data for building a model.
[0066] Secondly, the candidate text description is processed based on preset rules to obtain the target text description. The preset rules may include pre-set rules for text addition, deletion, modification, etc. The processing of the candidate text description based on the preset rules may include at least one of the following: screening the candidate text description based on the preset screening rules; adjusting the candidate text description based on the preset adjustment rules.
[0067] The screening of the candidate text descriptions based on the preset screening rules may include: presetting screening keywords and screening the candidate text descriptions containing the screening keywords. The preset adjustment rules may include but are not limited to at least one of word order adjustment rules, text addition rules, and text replacement rules, and the candidate text descriptions may be adjusted accordingly based on the preset adjustment rules.
[0068] By processing the candidate text descriptions based on preset rules, a target text description that meets the requirements can be obtained, and then a text generation model can be constructed based on the target description text that meets the requirements.
[0069] Thirdly, the text generation loss is determined based on the target text description and the candidate text description. The text generation loss can be obtained based on an existing text sequence loss function, such as connectionist temporal classification loss (CTC).
[0070] Finally, the pre-trained language model is adjusted according to the text generation loss to obtain a text generation model. The text generation loss can be used to reversely fine-tune the parameters of the pre-trained language model to construct a text generation model.
[0071] In these optional implementations, by generating and processing candidate text descriptions for constructing a model, a target text description for constructing a text generation model can be obtained. The text generation model can then be adjusted based on the target text description to obtain a text generation model with better performance.
[0072] The technical solution of the disclosed embodiment describes in detail the description text acquisition process. By combining the theme of text generation and the object features of the image receiving object to generate the text description, the target image generated based on the text description can be made to fit the theme, and the satisfaction of the image receiving object can be improved to a certain extent, thereby further meeting the needs of users. The image generation method provided by the disclosed embodiment and the image generation method provided by the above-mentioned embodiment belong to the same disclosed concept. The technical details not described in detail in this embodiment can be referred to the above-mentioned embodiment, and the same technical features have the same beneficial effects in this embodiment and the above-mentioned embodiment.
[0073] The various optional schemes in the image generation method provided in the embodiment of the present disclosure and the above embodiment can be combined. The image generation method provided in this embodiment describes in detail the steps of determining the style classification label. By generating text and its labels based on the constructed text generation model, a text classifier can be constructed, and then the constructed text classifier can be used to classify the text description, so as to achieve a better text classification effect.
[0074] Figure 4 The flowchart of an image generation method provided by the embodiment of the present disclosure is intended. Figure 4 As shown, the image generation method provided in this embodiment may include:
[0075] First, we obtain the subject of text generation and the object features of the image receiving object.
[0076] Next, a text description is generated based on the topic and object features through a text generation model.
[0077] Secondly, the style classification label of the text description is determined through the text classifier. For example, Figure 4 The style classification label 2 is determined as the style classification label of the text description.
[0078] The text classifier is constructed based on the sample text output by the text generation model and the sample label of the sample text. After the text generation model is constructed, the text description generated by the text generation model can be used as the sample text of the text classifier, and the style category label of each sample text (for example, the style category label of each sample text manually annotated) can be determined and used as the sample label. The sample text and the sample label can be used to construct the text classifier in a supervised manner to obtain a constructed text classifier.
[0079] By constructing a completed text classifier, the text description generated by the text generator is classified to obtain a style classification label, which can achieve a better classification effect and is conducive to selecting the target model that best matches the text description style.
[0080] Next, a target model corresponding to the style classification label is determined from each text-generated image model; wherein each text-generated image model has a different image generation style. For example, Figure 4 From the text image model 1 to the text image model m, the text image model 2 is determined as the target model corresponding to the style classification label 2.
[0081] Finally, the target model is used to generate a target image corresponding to the text description.
[0082] In some optional implementations, the target image may include an advertisement image; wherein the text description may be generated based on the advertisement theme and the object characteristics of the advertisement recipient. Figure 4 The acquired text-generated subject may include an advertisement subject, and the object characteristics of the image receiving object may include the object characteristics of the advertisement receiving object. The object characteristics of the advertisement receiving object may be considered as group characteristics of the advertisement audience. The object characteristics of the advertisement receiving object may be object characteristics customized by the advertisement publisher, or may be object characteristics acquired after authorization by the advertisement audience.
[0083] It is understandable that before using the technical solutions disclosed in the various embodiments of the present disclosure, the types of personal information obtained (such as the object characteristics of the advertising recipients), scope of use, usage scenarios, etc. involved in the present disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.
[0084] For example, in response to receiving an active request from a user, a prompt message is sent to the user to clearly prompt the user that the operation requested to be performed will require obtaining and using the user's personal information. Thus, the user can autonomously choose whether to provide personal information to software or hardware such as an electronic device, application, server, or storage medium that performs the operation of the technical solution of the present disclosure according to the prompt message.
[0085] For example, an input control or a selection control for object features may be provided on the display interface of the target object corresponding to the advertisement theme. The target object may include a physical object (e.g., an item) or a virtual object (e.g., a service). The user may actively input the object features of the advertisement receiving object in the input control, or may select the object features from preset feature tags through the selection control.
[0086] As an optional but non-limiting implementation, in response to receiving an active request from the user, the prompt information may be sent to the user in the form of a pop-up window, in which the prompt information may be presented in text form. In addition, the pop-up window may also carry a selection control for the user to choose "agree" or "disagree" to provide personal information to the electronic device.
[0087] It is understandable that the above notification and the process of obtaining user authorization are merely illustrative and do not constitute a limitation on the implementation of the present disclosure. Other methods that meet relevant laws and regulations may also be applied to the implementation of the present disclosure.
[0088] After the advertising image is generated, it can also be displayed on the display interface of the target object corresponding to the advertising theme. In some implementations, the advertising image may include at least one, and an advertising video can be generated based on at least one advertising image, and the advertising video can be played in the above display interface. Exemplarily, a sequence of advertising image frames can be generated, that is, there is a reasonable optical flow between the advertising image and the previous and next images in the sequence, and then an advertising video can be generated based on multiple advertising image frames to improve the delivery effect of the advertising video.
[0089] In these optional implementations, customized text descriptions can be generated for the advertising audience and advertising theme, and a target image can be generated based on a target generation model that best matches the text description style, which can make the target image easily accepted by the advertising audience and improve the advertising promotion effect.
[0090] In some further implementations, when the target image includes an advertisement image, the image generation method may further include:
[0091] Acquire multimedia data of a target object corresponding to the advertisement theme. The target object corresponding to the advertisement theme may include a physical object (e.g., an item) or a virtual object (e.g., a service). The multimedia data of the target object may include at least one of the following data: text data for introducing the target object, and image data, video data, etc. associated with the target object.
[0092] Accordingly, generating a target image corresponding to the text description through a target model may include: inputting the multimedia data and the text description into the target model to generate the target image through the target model.
[0093] Among them, in addition to the model capability of generating images based on text, the target model can also have the ability to process multimedia data. For example, when the multimedia data includes text data for introducing the target object, the text data can be merged with the target text description, and the merged text can be used as the control condition of the target model, so that the generated target image can simultaneously present the text data for introducing the target object and the content described in the target text. For another example, when the multimedia data includes image data associated with the target object, the image data and the target text description can be used as the control conditions of the target model, so that the generated target image can simultaneously present the content of the image data and the target text description. And, when the multimedia data includes at least two data, the at least two data and the target text description can be used as the control conditions of the target model to generate a target image that presents the content of the at least two data and the target text description.
[0094] In these further implementations, the target text description may be combined with multimedia data of the target object corresponding to the advertisement theme to generate a target image, which can make the target image more closely fit the advertisement theme to meet the needs of the advertiser.
[0095] The technical solution of the embodiment of the present disclosure describes in detail the steps of determining the style classification label. By generating text and its labels based on the constructed text generation model, a text classifier can be constructed, and then the constructed text classifier can be used to classify the text description, which can achieve a better text classification effect. The image generation method provided in the embodiment of the present disclosure and the image generation method provided in the above embodiment belong to the same public concept. The technical details not described in detail in this embodiment can be referred to the above embodiment, and the same technical features have the same beneficial effects in this embodiment and the above embodiment.
[0096] Figure 5 The image generating device provided in this embodiment is applicable to the case of generating an image, for example, the case of generating an advertisement image.
[0097] like Figure 5 As shown, the image generating device provided by the embodiment of the present disclosure may include:
[0098] A text acquisition module 510, used to acquire a text description;
[0099] A label determination module 520, used to determine a style classification label of the text description;
[0100] A model selection module 530 is used to determine a target model corresponding to a style classification label from each text-generated image model; wherein each text-generated image model has a different image generation style;
[0101] The image generation module 540 is used to generate a target image corresponding to the text description through a target model.
[0102] In some optional implementations, the text acquisition module can be used to:
[0103] Obtaining the subject of text generation and the object features of the image receiving object;
[0104] Generate text descriptions based on topic and object features through text generation models.
[0105] In some optional implementations, the image generating device may further include:
[0106] Model building module, used to build a text generation model based on the following process:
[0107] Obtain sample object features and sample topics;
[0108] Generate candidate text descriptions based on sample object features and sample topics using a pre-trained language model;
[0109] Processing the candidate text descriptions based on preset rules to obtain the target text description;
[0110] Determine the text generation loss based on the target text description and the candidate text description;
[0111] According to the text generation loss, the pre-trained language model is adjusted to obtain the text generation model.
[0112] In some optional implementations, the model building module may also be used to process the candidate text description based on at least one of the following preset rules:
[0113] Based on the preset screening rules, the candidate text descriptions are screened;
[0114] Based on preset adjustment rules, the candidate text description is adjusted.
[0115] In some optional implementations, the tag determination module may be used to:
[0116] The style classification label of the text description is determined by a text classifier; wherein the text classifier is constructed based on the sample text output by the text generation model and the sample label of the sample text.
[0117] In some optional implementations, the model selection module can also be used to:
[0118] In the case that the style classification label cannot be determined, the preset universal model in each text-generated image model is used as the target model.
[0119] In some optional implementations, the target image includes an advertisement image; wherein the text description is generated based on the advertisement theme and object features of an advertisement recipient.
[0120] In some optional implementations, the image generating device may include:
[0121] A multimedia data acquisition module, used to acquire multimedia data of a target object corresponding to an advertisement theme;
[0122] Correspondingly, the image generation module can be used to input multimedia data and text description into the target model to generate a target image through the target model.
[0123] The image generating device provided in the embodiments of the present disclosure can execute the image generating method provided in any embodiment of the present disclosure, and has the corresponding functional modules and beneficial effects of the execution method.
[0124] It is worth noting that the various units and modules included in the above-mentioned device are only divided according to functional logic, but are not limited to the above-mentioned division, as long as the corresponding functions can be achieved; in addition, the specific names of the functional units are only for the convenience of distinguishing each other, and are not used to limit the protection scope of the embodiments of the present disclosure.
[0125] Reference below Figure 6 , which shows an electronic device (eg, Figure 6 The terminal device in the embodiment of the present disclosure may include but is not limited to mobile terminals such as mobile phones, notebook computers, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 6 The electronic device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present disclosure.
[0126] like Figure 6 As shown, the electronic device 600 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 601, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 602 or a program loaded from a storage device 608 to a random access memory (RAM) 603. In the RAM 603, various programs and data required for the operation of the electronic device 600 are also stored. The processing device 601, the ROM 602, and the RAM 603 are connected to each other via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.
[0127] Typically, the following devices may be connected to the I / O interface 605: an input device 606 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 607 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 608 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 609. The communication device 609 may allow the electronic device 600 to communicate with other devices wirelessly or by wire to exchange data. Although Figure 6 The electronic device 600 is shown with various devices, but it should be understood that it is not required to implement or possess all the devices shown. More or fewer devices may be implemented or possessed instead.
[0128] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program contains program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication device 609, or installed from the storage device 608, or installed from the ROM 602. When the computer program is executed by the processing device 601, the above-mentioned functions defined in the image generation method of the embodiment of the present disclosure are executed.
[0129] The electronic device provided in the embodiment of the present disclosure and the image generation method provided in the above embodiment belong to the same disclosed concept. The technical details not fully described in this embodiment can be referred to the above embodiment, and this embodiment has the same beneficial effects as the above embodiment.
[0130] The embodiment of the present disclosure provides a computer storage medium on which a computer program is stored. When the program is executed by a processor, the image generating method provided by the above embodiment is implemented.
[0131] It should be noted that the computer-readable medium disclosed above may be a computer-readable signal medium or a computer-readable storage medium or any combination of the above two. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM) or a flash memory (FLASH), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program, which may be used by or in combination with an instruction execution system, device or device. In the present disclosure, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries a computer-readable program code. Such propagated data signals may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. Computer readable signal media may also be any computer readable medium other than computer readable storage media, which may send, propagate, or transmit programs for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer readable medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.
[0132] In some embodiments, the client and the server may communicate using any currently known or future developed network protocol such as HTTP (Hyper Text Transfer Protocol), and may be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network ("LAN"), a wide area network ("WAN"), an internet (e.g., the Internet), and a peer-to-peer network (e.g., an ad hoc peer-to-peer network), as well as any currently known or future developed network.
[0133] The computer-readable medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device.
[0134] The computer-readable medium carries one or more programs. When the one or more programs are executed by the electronic device, the electronic device:
[0135] Obtain a text description; determine a style classification label of the text description; determine a target model corresponding to the style classification label from each text-generated image model; wherein each text-generated image model has a different image generation style; and generate a target image corresponding to the text description through the target model.
[0136] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages or a combination thereof, including, but not limited to, object-oriented programming languages, such as Java, Smalltalk, C++, and conventional procedural programming languages, such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).
[0137] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, function and operation of the system, method and computer program product according to various embodiments of the present disclosure. In this regard, each square box in the flow chart or block diagram can represent a module, a program segment or a part of a code, and the module, the program segment or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some implementations as replacements, the functions marked in the square box can also occur in a sequence different from that marked in the accompanying drawings. For example, two square boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each square box in the block diagram and / or flow chart, and the combination of the square boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0138] The units involved in the embodiments described in the present disclosure may be implemented by software or hardware, wherein the names of the units and modules do not, in some cases, limit the units and modules themselves.
[0139] The functions described above herein may be performed at least in part by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that may be used include: Field Programmable Gate Array (FPGA), Application Specific Integrated Circuit (ASIC), Application Specific Standard Parts (ASSP), System on Chip (SOC), Complex Programmable Logic Device (CPLD), etc.
[0140] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or equipment. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0141] According to one or more embodiments of the present disclosure, there is provided an image generating method, the method comprising:
[0142] Get text description;
[0143] Determining a style classification label for the text description;
[0144] Determining a target model corresponding to the style classification label from each of the text-generated image models; wherein the image generation styles of each of the text-generated image models are different;
[0145] A target image corresponding to the text description is generated by the target model.
[0146] According to one or more embodiments of the present disclosure, there is provided an image generation method, further comprising:
[0147] In some optional implementations, obtaining the text description includes:
[0148] Obtaining the subject of text generation and the object features of the image receiving object;
[0149] A text description is generated based on the topic and the object features through a text generation model.
[0150] According to one or more embodiments of the present disclosure, there is provided an image generation method, further comprising:
[0151] In some optional implementations, the process of constructing the text generation model includes:
[0152] Obtain sample object features and sample topics;
[0153] Generate candidate text descriptions based on the sample object features and sample topics using a pre-trained language model;
[0154] Processing the candidate text description based on preset rules to obtain a target text description;
[0155] Determining a text generation loss according to the target text description and the candidate text description;
[0156] The pre-trained language model is adjusted according to the text generation loss to obtain a text generation model.
[0157] According to one or more embodiments of the present disclosure, there is provided an image generation method, further comprising:
[0158] In some optional implementations, the candidate text description is processed based on a preset rule, including at least one of the following:
[0159] Based on preset screening rules, screening the candidate text descriptions;
[0160] The candidate text description is adjusted based on preset adjustment rules.
[0161] According to one or more embodiments of the present disclosure, there is provided an image generation method, further comprising:
[0162] In some optional implementations, determining the style classification label of the text description includes:
[0163] The style classification label of the text description is determined by a text classifier; wherein the text classifier is constructed based on the sample text output by the text generation model and the sample label of the sample text.
[0164] According to one or more embodiments of the present disclosure, there is provided an image generation method, further comprising:
[0165] In some optional implementations, when the style classification label determination fails, a preset universal model in each of the text-generated image models is used as a target model.
[0166] According to one or more embodiments of the present disclosure, there is provided an image generation method, further comprising:
[0167] In some optional implementations, the target image includes an advertisement image; wherein the text description is generated based on an advertisement theme and object features of an advertisement recipient.
[0168] According to one or more embodiments of the present disclosure, there is provided an image generation method, further comprising:
[0169] In some optional implementations, multimedia data of a target object corresponding to the advertisement theme is obtained;
[0170] Accordingly, generating a target image corresponding to the text description through the target model includes:
[0171] The multimedia data and the text description are input into the object model to generate an object image through the object model.
[0172] According to one or more embodiments of the present disclosure, there is provided an image generating device, the device comprising:
[0173] A text acquisition module is used to obtain text descriptions;
[0174] A label determination module, used to determine a style classification label of the text description;
[0175] A model selection module, used to determine a target model corresponding to the style classification label from each text-generated image model; wherein each text-generated image model has a different image generation style;
[0176] An image generation module is used to generate a target image corresponding to the text description through the target model.
[0177] The above description is only a preferred embodiment of the present disclosure and an explanation of the technical principles used. Those skilled in the art should understand that the scope of disclosure involved in the present disclosure is not limited to the technical solutions formed by a specific combination of the above technical features, but should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosed concept. For example, the above features are replaced with the technical features with similar functions disclosed in the present disclosure (but not limited to) by each other to form a technical solution.
[0178] In addition, although each operation is described in a specific order, this should not be understood as requiring these operations to be performed in the specific order shown or in a sequential order. Under certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although some specific implementation details are included in the above discussion, these should not be interpreted as limiting the scope of the present disclosure. Some features described in the context of a separate embodiment can also be implemented in a single embodiment in combination. On the contrary, the various features described in the context of a single embodiment can also be implemented in multiple embodiments individually or in any suitable sub-combination mode.
[0179] Although the subject matter has been described in language specific to structural features and / or methodological logical actions, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. On the contrary, the specific features and actions described above are merely example forms of implementing the claims.
Claims
1. An image generation method, characterized in that: include: Get text description; Determining a style classification label for the text description; Determining a target model corresponding to the style classification label from each of the text-generated image models; wherein the image generation styles of each of the text-generated image models are different; A target image corresponding to the text description is generated by the target model.
2. The method according to claim 1, characterized in that The obtaining of the text description comprises: Obtaining the subject of text generation and the object features of the image receiving object; A text description is generated based on the topic and the object features through a text generation model.
3. The method according to claim 2, characterized in that The construction process of the text generation model includes: Obtain sample object features and sample topics; Generate candidate text descriptions based on the sample object features and sample topics using a pre-trained language model; Processing the candidate text description based on preset rules to obtain a target text description; Determining a text generation loss according to the target text description and the candidate text description; The pre-trained language model is adjusted according to the text generation loss to obtain a text generation model.
4. The method according to claim 3, characterized in that Processing the candidate text description based on a preset rule includes at least one of the following: Based on preset screening rules, screening the candidate text descriptions; The candidate text description is adjusted based on preset adjustment rules.
5. The method according to claim 2, characterized in that: Determining the style classification label of the text description includes: The style classification label of the text description is determined by a text classifier; wherein the text classifier is constructed based on the sample text output by the text generation model and the sample label of the sample text.
6. The method according to claim 1, characterized in that Also includes: In the case where the style classification label fails to be determined, a preset universal model in each of the text-generated image models is used as a target model.
7. The method according to any one of claims 1 to 6, characterized in that: The target image includes an advertisement image; wherein the text description is generated based on the advertisement theme and the object features of the advertisement recipient.
8. The method according to claim 7, characterized in that Also includes: Acquire multimedia data of a target object corresponding to the advertisement theme; Accordingly, generating a target image corresponding to the text description through the target model includes: The multimedia data and the text description are input into the object model to generate an object image through the object model.
9. An image generating device, characterized in that: include: A text acquisition module is used to obtain text descriptions; A label determination module, used to determine a style classification label of the text description; A model selection module, used to determine a target model corresponding to the style classification label from each text-generated image model; wherein each text-generated image model has a different image generation style; An image generation module is used to generate a target image corresponding to the text description through the target model.
10. An electronic device, characterized in that: The electronic device comprises: one or more processors; a storage device for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the image generating method according to any one of claims 1 to 8.
11. A storage medium comprising computer executable instructions, wherein the computer executable instructions are used to perform the image generation method according to any one of claims 1 to 8 when executed by a computer processor.
Citation Information
Cited By
Image generation method and apparatus, electronic device and storage medium
EP4807679A1